When to Choose Each Option
Clear guidance based on your specific situation and needs.
Our Recommendation
There is still no single winner, but the balance has shifted since this page was first written — and it shifted on both sides at once. DeepSeek is now the stronger default for almost every cost-sensitive production decision. On live OpenRouter prices deepseek-v4-flash-0731 runs $0.14/M input and $0.28/M output against $0.67/$3.40 for kimi-k2.7-code, roughly 12x cheaper on output, and it exposes a 1,310,720-token context window against K2.7's 262,144 — the 'tie' this page previously recorded on context was simply wrong. The 0731 build is also the newest open weights in the comparison (304B, 167 GB, card updated 2026-08-01) and it ships under a plain MIT licence, where Kimi K3's card declares license: other with license_name: kimi-k3. If you self-host, fine-tune or redistribute, that licence line is a decision by itself. What changed in Moonshot's favour is the validation story, but it did not change for the model this page names. ARC Prize verified Kimi K3 on July 31, 2026 at 94.5% on ARC-AGI-1 Semi-Private and 60.4% on ARC-AGI-2 — the highest open-weight scores it has recorded. Kimi K2.7 Code's own coding gains remain self-reported on Moonshot's Kimi Code Bench v2. So the honest reading is: the Kimi lineage has earned independent credibility, and the way to collect it is to move up to K3, not to stay on K2.7. K2.7 Code still holds a real lead where this page always said it did — MCP tool-use (76.0 MCP Atlas, 81.1 MCP Mark Verified), throughput (180-260 tokens/sec on the HighSpeed variant), and roughly 30% lower reasoning-token usage than K2.6. The most useful finding for anyone actually routing traffic is that the choice between these two families is not the biggest lever. ARC Prize's own numbers show Kimi K3 falling from 60.4% to 12.4% on ARC-AGI-2 purely by dropping reasoning effort from max to low, and Simon Willison reports the same shape on the other side: V4-Flash-0731 gave him a disappointing result at the default reasoning level and a much better one at reasoning_effort high. A five-fold accuracy swing inside one model is larger than the gap between the two models. Benchmark a candidate at the effort level you will actually pay for, or the headline price is fiction. Context Studios' pattern is unchanged in shape and sharper in detail: default high-volume bounded coding to DeepSeek V4-Flash-0731 for cost and context, escalate the hardest reasoning to V4-Pro, and route MCP-orchestration-heavy agent loops to the Kimi lineage — but budget for Kimi K3 rather than K2.7 now that K3 is the verified end of that path, and pin your reasoning-effort setting in the same config where you pin the model. Additionally, on 21 August 2026 DeepSeek shipped DeepSeek-V4-Flash-Vision-Exp — an experimental multimodal API model (text + image) with V4-Flash-level text performance, listed on OpenRouter at $0.22/M input and $0.66/M output with 1M context. That extends V4 Flash's production cases to visual agentic work (screenshots, UIs, diagrams) at a fraction of frontier-multimodal cost. Since the 2026-09-10 snapshot the DeepSeek side moved again: V4-Flash-0731 halved to $0.065/M input and $0.18/M output, widening the output gap to kimi-k2.7-code (~$0.71/$3.50) to ~19x and keeping V4-Flash the cheapest 1.3M-token route; a V4.1 Flash test build (native multimodal, ~400 tok/s, cutoff 2026-09-10) is reported but not yet on OpenRouter. The routing recipe is unchanged: V4-Flash for volume, V4-Pro for the hardest reasoning, the Kimi lineage for MCP-heavy loops, K3 as the verified end. Since 10.09.2026 the V4.1-Flash preview is confirmed (552B MoE, 8B/16B active, 1M context, $0.30/$1.20) — and with the 14.09 redirect DeepSeek collapses its line into one Flash-class backbone that also serves the old v4-pro id, while Kimi keeps the steadier id semantics. On the efficiency axis, V4.1's KV cache is reported at 890 bytes per token (~510 GB of files for local runs) — compression and lookup tables now matter as much as raw parameter counts.
- Choose Kimi K2.7 Code when...
- Your workload is MCP-heavy and tool-call accuracy is the real bottleneck
- You are already on the Kimi K2.x lineage and want a drop-in upgrade path that ends at ARC-verified K3
- Token throughput matters more to you than token price, and the HighSpeed variant's 180-260 tokens/sec pays for itself
- You can run at high or max reasoning effort and are buying peak capability rather than cost-per-task
- Choose DeepSeek V4 when...
- Cost per token is your primary constraint — output tokens run roughly 19x cheaper on V4-Flash-0731
- You need a million-token context window; K2.7 Code tops out at 262,144 tokens
- You self-host, fine-tune or redistribute weights and want a plain MIT licence with no legal review
- You want the newest weights on either side, refreshed 2026-07-31, on an already-broad multi-provider deployment surface
- You want one family spanning a cheap Flash tier and a frontier Pro tier for routing
- You want the widest documented price lead: V4-Flash-0731 halved to $0.065/$0.18 per million tokens in the 2026-09-10 snapshot
- You want the September 2026 consolidation: one 1M-token, vision-capable Flash line (8B/16B active, 1/4 KV-cache footprint) that also absorbs the v4-pro endpoint on 14.09.