Technology

Kimi K2.7 vs DeepSeek V4 (2026): Open-Weight Coding After Kimi K3 and V4-Flash-0731

Kimi K2.7 vs DeepSeek V4 in 2026: live OpenRouter pricing, ARC-verified Kimi K3 scores and MIT-licensed V4-Flash-0731.

3
Kimi K2.7 Code
vs
6
DeepSeek V4
Quick Verdict

There is still no single winner, but the balance has shifted since this page was first written — and it shifted on both sides at once. DeepSeek is now the stronger default for almost every cost-sensitive production decision. On live OpenRouter prices deepseek-v4-flash-0731 runs $0.09/M input and $0.18/M output against $0.73/$3.50 for kimi-k2.7-code, roughly 19x cheaper on output, and it exposes a 1,048,576-token context window against K2.7's 262,144 — the 'tie' this page previously recorded on context was simply wrong. The 0731 build is also the newest open weights in the comparison (304B, 167 GB, card updated 2026-08-01) and it ships under a plain MIT licence, where Kimi K3's card declares license: other with license_name: kimi-k3. If you self-host, fine-tune or redistribute, that licence line is a decision by itself. What changed in Moonshot's favour is the validation story, but it did not change for the model this page names. ARC Prize verified Kimi K3 on July 31, 2026 at 94.5% on ARC-AGI-1 Semi-Private and 60.4% on ARC-AGI-2 — the highest open-weight scores it has recorded. Kimi K2.7 Code's own coding gains remain self-reported on Moonshot's Kimi Code Bench v2. So the honest reading is: the Kimi lineage has earned independent credibility, and the way to collect it is to move up to K3, not to stay on K2.7. K2.7 Code still holds a real lead where this page always said it did — MCP tool-use (76.0 MCP Atlas, 81.1 MCP Mark Verified), throughput (180-260 tokens/sec on the HighSpeed variant), and roughly 30% lower reasoning-token usage than K2.6. The most useful finding for anyone actually routing traffic is that the choice between these two families is not the biggest lever. ARC Prize's own numbers show Kimi K3 falling from 60.4% to 12.4% on ARC-AGI-2 purely by dropping reasoning effort from max to low, and Simon Willison reports the same shape on the other side: V4-Flash-0731 gave him a disappointing result at the default reasoning level and a much better one at reasoning_effort high. A five-fold accuracy swing inside one model is larger than the gap between the two models. Benchmark a candidate at the effort level you will actually pay for, or the headline price is fiction. Context Studios' pattern is unchanged in shape and sharper in detail: default high-volume bounded coding to DeepSeek V4-Flash-0731 for cost and context, escalate the hardest reasoning to V4-Pro, and route MCP-orchestration-heavy agent loops to the Kimi lineage — but budget for Kimi K3 rather than K2.7 now that K3 is the verified end of that path, and pin your reasoning-effort setting in the same config where you pin the model.

Detailed Comparison

A side-by-side analysis of key factors to help you make the right choice.

Factor
Kimi K2.7 CodeRecommended
DeepSeek V4Winner
Release recency
Kimi K2.7 Code shipped June 12, 2026, but the lineage has moved on — Kimi K3's model card was last updated 2026-07-27, so K2.7 is no longer Moonshot's newest weights.
DeepSeek V4 launched April 24, 2026 and was refreshed on July 31, 2026 with V4-Flash-0731 (304B, 167 GB, card updated 2026-08-01) — the newest open weights on either side of this page.
Independent benchmark validation
K2.7's headline coding gains are still self-reported on Moonshot's own Kimi Code Bench v2. The lineage does now hold third-party verification — ARC Prize verified 94.5% ARC-AGI-1 and 60.4% ARC-AGI-2 — but that belongs to Kimi K3, not to K2.7.
DeepSeek appears on independent leaderboards, and Artificial Analysis places V4-Flash-0731 ahead of the 428B MiniMax M3 — an independent read on the current build rather than on a predecessor.
API cost
OpenRouter lists kimi-k2.7-code at $0.73/M input and $3.50/M output as of 2026-08-02. The $0.95/$4.00 figure previously quoted here from LLM Stats is stale.
deepseek-v4-flash-0731 lists $0.09/M input and $0.18/M output on OpenRouter ($0.14/$0.27 first-party), with V4-Pro at $0.435/$0.87 — output tokens run roughly 19x cheaper than K2.7 Code.
MCP & agentic tool-use
Leads MCP tool-use benchmarks at launch: 76.0 MCP Atlas and 81.1 MCP Mark Verified
Strong general agentic coding, but no comparable published MCP tool-use leadership
Inference speed & throughput
HighSpeed variant pushes 180 tokens/sec, up to 260 in short-context scenarios
Solid latency, especially V4-Flash, but no published throughput edge at this level
Context window
kimi-k2.7-code exposes a 262,144-token context window on OpenRouter — large, but not million-token. The previous 'tie' on this page was wrong.
Both deepseek-v4-flash-0731 and deepseek-v4-pro expose 1,048,576 tokens on the same routing layer — a 4x larger window for whole-repo work.
Production track record & availability
Kimi K2.7-Code has 665,880 Hugging Face downloads; the lineage's attention has shifted to K3 (559,924 downloads, 9,502 likes) barely six weeks after K2.7 landed.
DeepSeek-V4-Flash has 2,814,414 downloads — over 4x K2.7-Code — across multiple providers, and the 0731 build inherits that deployment surface rather than restarting it.
Reasoning-token efficiency
Cuts reasoning-token usage roughly 30% versus K2.6, lowering cost on long agentic loops
Efficient chain-of-thought, but no comparable published reduction figure
Reasoning effort versus model choice
ARC Prize's verified Kimi K3 run swings from 60.4% to 12.4% on ARC-AGI-2 purely by dropping reasoning effort from max to low — a 5x accuracy collapse inside one model, and cost moves with it ($1.59 per task at max).
Simon Willison got a disappointing result from V4-Flash-0731 at the default reasoning level and a much better one after setting reasoning_effort high. Neither family is safe to judge on a headline price: the effort setting dominates both quality and true cost.
Licence clarity
Kimi K3's Hugging Face card declares license: other with license_name: kimi-k3 — a bespoke Moonshot licence that needs a legal read before you ship on it.
DeepSeek-V4-Flash and V4-Flash-0731 both declare a plain MIT licence on their model cards. For anyone self-hosting or redistributing weights, that is the difference between a review cycle and no review at all.
Total Score3/ 106/ 101 ties
Release recency
Kimi K2.7 Code
Kimi K2.7 Code shipped June 12, 2026, but the lineage has moved on — Kimi K3's model card was last updated 2026-07-27, so K2.7 is no longer Moonshot's newest weights.
DeepSeek V4
DeepSeek V4 launched April 24, 2026 and was refreshed on July 31, 2026 with V4-Flash-0731 (304B, 167 GB, card updated 2026-08-01) — the newest open weights on either side of this page.
Independent benchmark validation
Kimi K2.7 Code
K2.7's headline coding gains are still self-reported on Moonshot's own Kimi Code Bench v2. The lineage does now hold third-party verification — ARC Prize verified 94.5% ARC-AGI-1 and 60.4% ARC-AGI-2 — but that belongs to Kimi K3, not to K2.7.
DeepSeek V4
DeepSeek appears on independent leaderboards, and Artificial Analysis places V4-Flash-0731 ahead of the 428B MiniMax M3 — an independent read on the current build rather than on a predecessor.
API cost
Kimi K2.7 Code
OpenRouter lists kimi-k2.7-code at $0.73/M input and $3.50/M output as of 2026-08-02. The $0.95/$4.00 figure previously quoted here from LLM Stats is stale.
DeepSeek V4
deepseek-v4-flash-0731 lists $0.09/M input and $0.18/M output on OpenRouter ($0.14/$0.27 first-party), with V4-Pro at $0.435/$0.87 — output tokens run roughly 19x cheaper than K2.7 Code.
MCP & agentic tool-use
Kimi K2.7 Code
Leads MCP tool-use benchmarks at launch: 76.0 MCP Atlas and 81.1 MCP Mark Verified
DeepSeek V4
Strong general agentic coding, but no comparable published MCP tool-use leadership
Inference speed & throughput
Kimi K2.7 Code
HighSpeed variant pushes 180 tokens/sec, up to 260 in short-context scenarios
DeepSeek V4
Solid latency, especially V4-Flash, but no published throughput edge at this level
Context window
Kimi K2.7 Code
kimi-k2.7-code exposes a 262,144-token context window on OpenRouter — large, but not million-token. The previous 'tie' on this page was wrong.
DeepSeek V4
Both deepseek-v4-flash-0731 and deepseek-v4-pro expose 1,048,576 tokens on the same routing layer — a 4x larger window for whole-repo work.
Production track record & availability
Kimi K2.7 Code
Kimi K2.7-Code has 665,880 Hugging Face downloads; the lineage's attention has shifted to K3 (559,924 downloads, 9,502 likes) barely six weeks after K2.7 landed.
DeepSeek V4
DeepSeek-V4-Flash has 2,814,414 downloads — over 4x K2.7-Code — across multiple providers, and the 0731 build inherits that deployment surface rather than restarting it.
Reasoning-token efficiency
Kimi K2.7 Code
Cuts reasoning-token usage roughly 30% versus K2.6, lowering cost on long agentic loops
DeepSeek V4
Efficient chain-of-thought, but no comparable published reduction figure
Reasoning effort versus model choice
Kimi K2.7 Code
ARC Prize's verified Kimi K3 run swings from 60.4% to 12.4% on ARC-AGI-2 purely by dropping reasoning effort from max to low — a 5x accuracy collapse inside one model, and cost moves with it ($1.59 per task at max).
DeepSeek V4
Simon Willison got a disappointing result from V4-Flash-0731 at the default reasoning level and a much better one after setting reasoning_effort high. Neither family is safe to judge on a headline price: the effort setting dominates both quality and true cost.
Licence clarity
Kimi K2.7 Code
Kimi K3's Hugging Face card declares license: other with license_name: kimi-k3 — a bespoke Moonshot licence that needs a legal read before you ship on it.
DeepSeek V4
DeepSeek-V4-Flash and V4-Flash-0731 both declare a plain MIT licence on their model cards. For anyone self-hosting or redistributing weights, that is the difference between a review cycle and no review at all.

Key Statistics

Real data from verified industry sources to support your decision.

Live routing prices, 2026-08-02: kimi-k2.7-code $0.73/M input, $3.50/M output, 262,144-token context; deepseek-v4-flash-0731 $0.09/M input, $0.18/M output, 1,048,576-token context; deepseek-v4-pro $0.435/$0.87. Output tokens are ~19x cheaper on V4-Flash and the context window is 4x larger.

OpenRouter model API (live)

As of July 31, 2026 ARC Prize verified Kimi K3 as the new open-weight high score on both benchmarks: 94.5% on ARC-AGI-1 Semi-Private at $0.77 per task and 60.4% on ARC-AGI-2 Semi-Private at $1.59 per task at max reasoning effort. The same model scores 86.7%/55.0% at high effort and 65.7%/12.4% at low effort.

ARC Prize — Verified results, Kimi K3

DeepSeek-V4-Flash-0731 ships 304B parameters (167 GB of weights) under a plain MIT licence; the model card was last modified 2026-08-01. This is the newest open-weight release on either side of this comparison.

Hugging Face model API — deepseek-ai/DeepSeek-V4-Flash-0731

Kimi K3's model card carries license: other with license_name: kimi-k3 — a custom Moonshot licence, not a standard OSI licence. Card last modified 2026-07-27.

Hugging Face model API — moonshotai/Kimi-K3

Willison rates V4-Flash-0731 at $0.14/M input and $0.27/M output as possibly the best value-per-intelligence model available, noting Artificial Analysis ranks the 304B model ahead of the 428B MiniMax M3 — and that the default reasoning level produced a disappointing result until he raised reasoning_effort to high.

Simon Willison — deepseek-ai/DeepSeek-V4-Flash-0731

Adoption splits by type of user: DeepSeek-V4-Flash has 2,814,414 downloads against 1,947 likes, while Kimi-K2.7-Code has 665,880 downloads against 1,347 likes and Kimi K3 has 559,924 downloads against 9,502 likes — DeepSeek wins on deployment volume, Moonshot on community enthusiasm.

Hugging Face model API — adoption counters

All statistics come from verified third-party sources. Source, year, and direct link are shown on each metric.

When to Choose Each Option

Clear guidance based on your specific situation and needs.

Choose Kimi K2.7 Code when...

  • Your workload is MCP-heavy and tool-call accuracy is the real bottleneck
  • You are already on the Kimi K2.x lineage and want a drop-in upgrade path that ends at ARC-verified K3
  • Token throughput matters more to you than token price, and the HighSpeed variant's 180-260 tokens/sec pays for itself
  • You can run at high or max reasoning effort and are buying peak capability rather than cost-per-task

Choose DeepSeek V4 when...

  • Cost per token is your primary constraint — output tokens run roughly 19x cheaper on V4-Flash-0731
  • You need a million-token context window; K2.7 Code tops out at 262,144 tokens
  • You self-host, fine-tune or redistribute weights and want a plain MIT licence with no legal review
  • You want the newest weights on either side, refreshed 2026-07-31, on an already-broad multi-provider deployment surface
  • You want one family spanning a cheap Flash tier and a frontier Pro tier for routing

Our Recommendation

There is still no single winner, but the balance has shifted since this page was first written — and it shifted on both sides at once. DeepSeek is now the stronger default for almost every cost-sensitive production decision. On live OpenRouter prices deepseek-v4-flash-0731 runs $0.09/M input and $0.18/M output against $0.73/$3.50 for kimi-k2.7-code, roughly 19x cheaper on output, and it exposes a 1,048,576-token context window against K2.7's 262,144 — the 'tie' this page previously recorded on context was simply wrong. The 0731 build is also the newest open weights in the comparison (304B, 167 GB, card updated 2026-08-01) and it ships under a plain MIT licence, where Kimi K3's card declares license: other with license_name: kimi-k3. If you self-host, fine-tune or redistribute, that licence line is a decision by itself. What changed in Moonshot's favour is the validation story, but it did not change for the model this page names. ARC Prize verified Kimi K3 on July 31, 2026 at 94.5% on ARC-AGI-1 Semi-Private and 60.4% on ARC-AGI-2 — the highest open-weight scores it has recorded. Kimi K2.7 Code's own coding gains remain self-reported on Moonshot's Kimi Code Bench v2. So the honest reading is: the Kimi lineage has earned independent credibility, and the way to collect it is to move up to K3, not to stay on K2.7. K2.7 Code still holds a real lead where this page always said it did — MCP tool-use (76.0 MCP Atlas, 81.1 MCP Mark Verified), throughput (180-260 tokens/sec on the HighSpeed variant), and roughly 30% lower reasoning-token usage than K2.6. The most useful finding for anyone actually routing traffic is that the choice between these two families is not the biggest lever. ARC Prize's own numbers show Kimi K3 falling from 60.4% to 12.4% on ARC-AGI-2 purely by dropping reasoning effort from max to low, and Simon Willison reports the same shape on the other side: V4-Flash-0731 gave him a disappointing result at the default reasoning level and a much better one at reasoning_effort high. A five-fold accuracy swing inside one model is larger than the gap between the two models. Benchmark a candidate at the effort level you will actually pay for, or the headline price is fiction. Context Studios' pattern is unchanged in shape and sharper in detail: default high-volume bounded coding to DeepSeek V4-Flash-0731 for cost and context, escalate the hardest reasoning to V4-Pro, and route MCP-orchestration-heavy agent loops to the Kimi lineage — but budget for Kimi K3 rather than K2.7 now that K3 is the verified end of that path, and pin your reasoning-effort setting in the same config where you pin the model.

Frequently Asked Questions

Common questions about this comparison answered.

It depends on your constraint. DeepSeek V4 is the safer pick for cost-sensitive, high-volume work: it is independently benchmarked (reported 83.7% on SWE-bench Verified, #14 on BenchLM), has been in production since April 2026, and its V4-Flash tier is among the cheapest serious coding APIs. Kimi K2.7 Code is stronger for MCP-heavy agentic workflows, leading tool-use benchmarks (76.0 MCP Atlas, 81.1 MCP Mark Verified) with high throughput — but its headline coding gains are still largely self-reported, so validate on your own tasks first.
DeepSeek V4 is cheaper. V4-Flash lists around $0.28 per million output tokens and V4-Pro around $0.87, among the lowest for serious coding models. Kimi K2.7 Code is priced at $0.95 per million input and $4.00 per million output, with cache hits as low as $0.19 per million — competitive, but well above DeepSeek's Flash tier on output cost.
Not yet, mostly. At launch, Kimi K2.7's headline coding gains (+21.8% over K2.6) come from Moonshot's own Kimi Code Bench v2, and independent SWE-bench numbers are still thin. DeepSeek V4, by contrast, already appears on independent leaderboards like Vals AI and BenchLM. Treat Kimi's launch figures as promising but unconfirmed until third-party benchmarks are published.
Yes — both are open-weight models, so you can run them on your own infrastructure for data-residency or compliance reasons, in addition to using their hosted APIs. DeepSeek V4 is already available across multiple providers (Fireworks, DeepInfra, Novita, SiliconFlow). Note that the MoE architectures are large: Kimi K2.7 is 1T total parameters and DeepSeek V4-Pro is 1.6T, so self-hosting the top tiers needs substantial memory bandwidth.
Partly — and not on the model this page names. ARC Prize verified Kimi K3 on July 31, 2026 at 94.5% on ARC-AGI-1 Semi-Private ($0.77/task) and 60.4% on ARC-AGI-2 ($1.59/task), the highest open-weight scores it has recorded. Kimi K2.7 Code itself still rests on Moonshot's own Kimi Code Bench v2. If independent validation is what you are waiting for, the answer is to move up to K3, not to trust K2.7's launch numbers.
DeepSeek is the cleaner answer. Both DeepSeek-V4-Flash and the newer V4-Flash-0731 declare a plain MIT licence on their Hugging Face cards. Kimi K3's card declares license: other with license_name: kimi-k3 — a custom Moonshot licence. MIT needs no legal review before you self-host, fine-tune or redistribute; a bespoke licence does.

Need help deciding?

Book a free 30-minute consultation and we'll help you determine the best approach for your specific project.

Free consultation
No obligation
Response within 24h