10 Sept 2026
DeepSeek V4.1 Flash: 890 Bytes KV Cache per Token
(01)DeepSeek V4.1 Flash introduces groundbreaking KV cache compression, reducing the footprint to just 890 bytes per token. Featuring an asymmetric Causal Encoder-Decoder architecture, it rivals much larger models in agentic tasks while remaining highly cost-effective under an MIT license.
DeepSeek