DeepSeek V4.1-Flash cuts KV cache 4x despite doubling model size, targeting AI agents
DeepSeek released its V4.1-Flash model on September 10, 2026, specifically engineered to reduce the cost of running AI agents that repeatedly read large contexts rather than generate long outputs. The model introduces a Causal Encoder-Decoder (CED) architecture that splits its 552 billion parameters into encoder and decoder halves, sharply reducing active parameters during input processing to just 8 billion. Additional techniques — including Compressed Sparse Attention 2 (CSA2) and FP4 KV caching — bring total KV cache storage down to 890 bytes per token, roughly one-quarter of its predecessor V4-Flash and 437 times smaller than the original V1. This matters economically because agentic workloads spend most compute re-reading the same context in loops, meaning KV cache — not token generation — is the dominant cost driver. Benchmark scores for agent tasks such as DeepSWE and Terminal-Bench rose substantially, from 54.4 to 74.2 and 82.7 to 90.6 respectively, compared to V4-Flash.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.


Discussion (0)
Log in to join the discussion and vote.
Log in