AI Agents
DeepSeek V4.1 Flash: 890 Bytes KV Cache per Token
DeepSeek V4.1 Flash introduces groundbreaking KV cache compression, reducing the footprint to just 890 bytes per token. Featuring an asymmetric Causal Encoder-Decoder architecture, it rivals much larger models in agentic tasks while remaining highly cost-effective under an MIT license.
9 days ago