Kimi K2.7 vs DeepSeek V4 (2026): Open-Weight Coding After Kimi K3 and V4-Flash-0731
Kimi K2.7 vs DeepSeek V4 in 2026: live OpenRouter pricing, ARC-verified Kimi K3 scores and MIT-licensed V4-Flash-0731.
There is still no single winner, but the balance has shifted since this page was first written — and it shifted on both sides at once. DeepSeek is now the stronger default for almost every cost-sensitive production decision. On live OpenRouter prices deepseek-v4-flash-0731 runs $0.09/M input and $0.18/M output against $0.73/$3.50 for kimi-k2.7-code, roughly 19x cheaper on output, and it exposes a 1,048,576-token context window against K2.7's 262,144 — the 'tie' this page previously recorded on context was simply wrong. The 0731 build is also the newest open weights in the comparison (304B, 167 GB, card updated 2026-08-01) and it ships under a plain MIT licence, where Kimi K3's card declares license: other with license_name: kimi-k3. If you self-host, fine-tune or redistribute, that licence line is a decision by itself. What changed in Moonshot's favour is the validation story, but it did not change for the model this page names. ARC Prize verified Kimi K3 on July 31, 2026 at 94.5% on ARC-AGI-1 Semi-Private and 60.4% on ARC-AGI-2 — the highest open-weight scores it has recorded. Kimi K2.7 Code's own coding gains remain self-reported on Moonshot's Kimi Code Bench v2. So the honest reading is: the Kimi lineage has earned independent credibility, and the way to collect it is to move up to K3, not to stay on K2.7. K2.7 Code still holds a real lead where this page always said it did — MCP tool-use (76.0 MCP Atlas, 81.1 MCP Mark Verified), throughput (180-260 tokens/sec on the HighSpeed variant), and roughly 30% lower reasoning-token usage than K2.6. The most useful finding for anyone actually routing traffic is that the choice between these two families is not the biggest lever. ARC Prize's own numbers show Kimi K3 falling from 60.4% to 12.4% on ARC-AGI-2 purely by dropping reasoning effort from max to low, and Simon Willison reports the same shape on the other side: V4-Flash-0731 gave him a disappointing result at the default reasoning level and a much better one at reasoning_effort high. A five-fold accuracy swing inside one model is larger than the gap between the two models. Benchmark a candidate at the effort level you will actually pay for, or the headline price is fiction. Context Studios' pattern is unchanged in shape and sharper in detail: default high-volume bounded coding to DeepSeek V4-Flash-0731 for cost and context, escalate the hardest reasoning to V4-Pro, and route MCP-orchestration-heavy agent loops to the Kimi lineage — but budget for Kimi K3 rather than K2.7 now that K3 is the verified end of that path, and pin your reasoning-effort setting in the same config where you pin the model.
Detailed Comparison
A side-by-side analysis of key factors to help you make the right choice.
| Factor | Kimi K2.7 CodeRecommended | DeepSeek V4 | Winner |
|---|---|---|---|
| Release recency | Kimi K2.7 Code shipped June 12, 2026, but the lineage has moved on — Kimi K3's model card was last updated 2026-07-27, so K2.7 is no longer Moonshot's newest weights. | DeepSeek V4 launched April 24, 2026 and was refreshed on July 31, 2026 with V4-Flash-0731 (304B, 167 GB, card updated 2026-08-01) — the newest open weights on either side of this page. | |
| Independent benchmark validation | K2.7's headline coding gains are still self-reported on Moonshot's own Kimi Code Bench v2. The lineage does now hold third-party verification — ARC Prize verified 94.5% ARC-AGI-1 and 60.4% ARC-AGI-2 — but that belongs to Kimi K3, not to K2.7. | DeepSeek appears on independent leaderboards, and Artificial Analysis places V4-Flash-0731 ahead of the 428B MiniMax M3 — an independent read on the current build rather than on a predecessor. | |
| API cost | OpenRouter lists kimi-k2.7-code at $0.73/M input and $3.50/M output as of 2026-08-02. The $0.95/$4.00 figure previously quoted here from LLM Stats is stale. | deepseek-v4-flash-0731 lists $0.09/M input and $0.18/M output on OpenRouter ($0.14/$0.27 first-party), with V4-Pro at $0.435/$0.87 — output tokens run roughly 19x cheaper than K2.7 Code. | |
| MCP & agentic tool-use | Leads MCP tool-use benchmarks at launch: 76.0 MCP Atlas and 81.1 MCP Mark Verified | Strong general agentic coding, but no comparable published MCP tool-use leadership | |
| Inference speed & throughput | HighSpeed variant pushes 180 tokens/sec, up to 260 in short-context scenarios | Solid latency, especially V4-Flash, but no published throughput edge at this level | |
| Context window | kimi-k2.7-code exposes a 262,144-token context window on OpenRouter — large, but not million-token. The previous 'tie' on this page was wrong. | Both deepseek-v4-flash-0731 and deepseek-v4-pro expose 1,048,576 tokens on the same routing layer — a 4x larger window for whole-repo work. | |
| Production track record & availability | Kimi K2.7-Code has 665,880 Hugging Face downloads; the lineage's attention has shifted to K3 (559,924 downloads, 9,502 likes) barely six weeks after K2.7 landed. | DeepSeek-V4-Flash has 2,814,414 downloads — over 4x K2.7-Code — across multiple providers, and the 0731 build inherits that deployment surface rather than restarting it. | |
| Reasoning-token efficiency | Cuts reasoning-token usage roughly 30% versus K2.6, lowering cost on long agentic loops | Efficient chain-of-thought, but no comparable published reduction figure | |
| Reasoning effort versus model choice | ARC Prize's verified Kimi K3 run swings from 60.4% to 12.4% on ARC-AGI-2 purely by dropping reasoning effort from max to low — a 5x accuracy collapse inside one model, and cost moves with it ($1.59 per task at max). | Simon Willison got a disappointing result from V4-Flash-0731 at the default reasoning level and a much better one after setting reasoning_effort high. Neither family is safe to judge on a headline price: the effort setting dominates both quality and true cost. | |
| Licence clarity | Kimi K3's Hugging Face card declares license: other with license_name: kimi-k3 — a bespoke Moonshot licence that needs a legal read before you ship on it. | DeepSeek-V4-Flash and V4-Flash-0731 both declare a plain MIT licence on their model cards. For anyone self-hosting or redistributing weights, that is the difference between a review cycle and no review at all. | |
| Total Score | 3/ 10 | 6/ 10 | 1 ties |
Key Statistics
Real data from verified industry sources to support your decision.
OpenRouter model API (live)
ARC Prize — Verified results, Kimi K3
Hugging Face model API — deepseek-ai/DeepSeek-V4-Flash-0731
Hugging Face model API — moonshotai/Kimi-K3
Simon Willison — deepseek-ai/DeepSeek-V4-Flash-0731
Hugging Face model API — adoption counters
All statistics come from verified third-party sources. Source, year, and direct link are shown on each metric.
When to Choose Each Option
Clear guidance based on your specific situation and needs.
Choose Kimi K2.7 Code when...
- Your workload is MCP-heavy and tool-call accuracy is the real bottleneck
- You are already on the Kimi K2.x lineage and want a drop-in upgrade path that ends at ARC-verified K3
- Token throughput matters more to you than token price, and the HighSpeed variant's 180-260 tokens/sec pays for itself
- You can run at high or max reasoning effort and are buying peak capability rather than cost-per-task
Choose DeepSeek V4 when...
- Cost per token is your primary constraint — output tokens run roughly 19x cheaper on V4-Flash-0731
- You need a million-token context window; K2.7 Code tops out at 262,144 tokens
- You self-host, fine-tune or redistribute weights and want a plain MIT licence with no legal review
- You want the newest weights on either side, refreshed 2026-07-31, on an already-broad multi-provider deployment surface
- You want one family spanning a cheap Flash tier and a frontier Pro tier for routing
Our Recommendation
There is still no single winner, but the balance has shifted since this page was first written — and it shifted on both sides at once. DeepSeek is now the stronger default for almost every cost-sensitive production decision. On live OpenRouter prices deepseek-v4-flash-0731 runs $0.09/M input and $0.18/M output against $0.73/$3.50 for kimi-k2.7-code, roughly 19x cheaper on output, and it exposes a 1,048,576-token context window against K2.7's 262,144 — the 'tie' this page previously recorded on context was simply wrong. The 0731 build is also the newest open weights in the comparison (304B, 167 GB, card updated 2026-08-01) and it ships under a plain MIT licence, where Kimi K3's card declares license: other with license_name: kimi-k3. If you self-host, fine-tune or redistribute, that licence line is a decision by itself. What changed in Moonshot's favour is the validation story, but it did not change for the model this page names. ARC Prize verified Kimi K3 on July 31, 2026 at 94.5% on ARC-AGI-1 Semi-Private and 60.4% on ARC-AGI-2 — the highest open-weight scores it has recorded. Kimi K2.7 Code's own coding gains remain self-reported on Moonshot's Kimi Code Bench v2. So the honest reading is: the Kimi lineage has earned independent credibility, and the way to collect it is to move up to K3, not to stay on K2.7. K2.7 Code still holds a real lead where this page always said it did — MCP tool-use (76.0 MCP Atlas, 81.1 MCP Mark Verified), throughput (180-260 tokens/sec on the HighSpeed variant), and roughly 30% lower reasoning-token usage than K2.6. The most useful finding for anyone actually routing traffic is that the choice between these two families is not the biggest lever. ARC Prize's own numbers show Kimi K3 falling from 60.4% to 12.4% on ARC-AGI-2 purely by dropping reasoning effort from max to low, and Simon Willison reports the same shape on the other side: V4-Flash-0731 gave him a disappointing result at the default reasoning level and a much better one at reasoning_effort high. A five-fold accuracy swing inside one model is larger than the gap between the two models. Benchmark a candidate at the effort level you will actually pay for, or the headline price is fiction. Context Studios' pattern is unchanged in shape and sharper in detail: default high-volume bounded coding to DeepSeek V4-Flash-0731 for cost and context, escalate the hardest reasoning to V4-Pro, and route MCP-orchestration-heavy agent loops to the Kimi lineage — but budget for Kimi K3 rather than K2.7 now that K3 is the verified end of that path, and pin your reasoning-effort setting in the same config where you pin the model.
Frequently Asked Questions
Common questions about this comparison answered.
Need help deciding?
Book a free 30-minute consultation and we'll help you determine the best approach for your specific project.