Technology

Kimi K2.7 vs DeepSeek V4 (2026): Open-Weight Coding After Kimi K3 and V4.1-Flash

Kimi K2.7 vs DeepSeek V4 in 2026: live OpenRouter pricing, ARC-verified Kimi K3 scores and MIT-licensed V4-Flash-0731.

Reviewed by Michael Kerkhoff, as of

Definition
Two Chinese labs anchor the open-weight coding race, and since this comparison was first written both of them have shipped again. Moonshot AI released Kimi K2.7 Code on June 12, 2026 — a 1-trillion-parameter mixture-of-experts model tuned for agentic tool-use and throughput — and has since moved the lineage on to Kimi K3, which ARC Prize verified on July 31, 2026 as the new open-weight high score on ARC-AGI-1 and ARC-AGI-2. DeepSeek launched V4 on April 24, 2026 and refreshed it on July 31, 2026 with DeepSeek-V4-Flash-0731: 304 billion parameters, 167 GB of weights, plain MIT licence. That means the honest 2026 question is no longer only "K2.7 or V4" but "which lineage do I standardise on, and at what reasoning effort". On live OpenRouter prices the cost gap is now roughly 19x on output tokens in DeepSeek's favour, and the context gap 4x — but reasoning effort moves ARC-AGI-2 accuracy by a factor of five within a single model, which is a larger swing than the choice between these two families.
Category
Technology
Options
Kimi K2.7 CodeDeepSeek V4

Detailed Comparison

A side-by-side analysis of key factors to help you make the right choice.

Kimi K2.7 Code vs DeepSeek V4
FactorKimi K2.7 CodeDeepSeek V4
Release recencyKimi K2.7 Code shipped June 12, 2026, but the lineage has moved on — Kimi K3's model card was last updated 2026-07-27, so K2.7 is no longer Moonshot's newest weights.DeepSeek V4 launched April 24, 2026 and was refreshed on July 31, 2026 with V4-Flash-0731 (304B, 167 GB, card updated 2026-08-01) — the newest open weights on either side of this page. Winner
Independent benchmark validationK2.7's headline coding gains are still self-reported on Moonshot's own Kimi Code Bench v2. The lineage does now hold third-party verification — ARC Prize verified 94.5% ARC-AGI-1 and 60.4% ARC-AGI-2 — but that belongs to Kimi K3, not to K2.7.DeepSeek appears on independent leaderboards, and Artificial Analysis places V4-Flash-0731 ahead of the 428B MiniMax M3 — an independent read on the current build rather than on a predecessor. Winner
API costOpenRouter lists kimi-k2.7-code at $0.73/M input and $3.50/M output as of 2026-08-02. The $0.95/$4.00 figure previously quoted here from LLM Stats is stale.deepseek-v4-flash-0731 lists $0.09/M input and $0.18/M output on OpenRouter ($0.14/$0.27 first-party), with V4-Pro at $0.435/$0.87 — output tokens run roughly 19x cheaper than K2.7 Code. Winner
MCP & agentic tool-useLeads MCP tool-use benchmarks at launch: 76.0 MCP Atlas and 81.1 MCP Mark Verified WinnerStrong general agentic coding, but no comparable published MCP tool-use leadership
Inference speed & throughputHighSpeed variant pushes 180 tokens/sec, up to 260 in short-context scenarios WinnerSolid latency, especially V4-Flash, but no published throughput edge at this level
Context windowkimi-k2.7-code exposes a 262,144-token context window on OpenRouter — large, but not million-token. The previous 'tie' on this page was wrong.Both deepseek-v4-flash-0731 and deepseek-v4-pro expose 1,048,576 tokens on the same routing layer — a 4x larger window for whole-repo work. Winner
Production track record & availabilityKimi K2.7-Code has 665,880 Hugging Face downloads; the lineage's attention has shifted to K3 (559,924 downloads, 9,502 likes) barely six weeks after K2.7 landed.DeepSeek-V4-Flash has 2,814,414 downloads — over 4x K2.7-Code — across multiple providers, and the 0731 build inherits that deployment surface rather than restarting it. Winner
Reasoning-token efficiencyCuts reasoning-token usage roughly 30% versus K2.6, lowering cost on long agentic loops WinnerEfficient chain-of-thought, but no comparable published reduction figure
Reasoning effort versus model choiceARC Prize's verified Kimi K3 run swings from 60.4% to 12.4% on ARC-AGI-2 purely by dropping reasoning effort from max to low — a 5x accuracy collapse inside one model, and cost moves with it ($1.59 per task at max).Simon Willison got a disappointing result from V4-Flash-0731 at the default reasoning level and a much better one after setting reasoning_effort high. Neither family is safe to judge on a headline price: the effort setting dominates both quality and true cost.
Licence clarityKimi K3's Hugging Face card declares license: other with license_name: kimi-k3 — a bespoke Moonshot licence that needs a legal read before you ship on it.DeepSeek-V4-Flash and V4-Flash-0731 both declare a plain MIT licence on their model cards. For anyone self-hosting or redistributing weights, that is the difference between a review cycle and no review at all. Winner
Price snapshot September 2026kimi-k2.7-code now lists $0.71/$3.50 (2026-09-10 snapshot); Kimi K3 at $3/$15 buys the independently verified scores but at roughly 16x V4-Flash's output price.V4-Flash-0731 halved to $0.065/M input and $0.18/M output on 2026-09-10 — the widest cost lead in this comparison's history (~19x on output) and still the cheapest 1.3M-token route. Winner
V4.1-Flash era (September 2026)Kimi holds stable ids and no redirects: k2.7-code (262K context) and K3 (1M context, $3/$15) keep their documented behaviour through the 09/2026 wave.DeepSeek consolidates: V4.1-Flash brings 1M native context, vision, 8B/16B active parameters and a 1/4-HBM KV cache, and swallows the v4-pro endpoint on 14.09 — one fast line to track, but two id-semantics shifts inside a week. Winner
Total Score · 1 ties3 / 128 / 12

Key Statistics

Real data from verified industry sources to support your decision.

  • Live routing prices, 2026-08-25: kimi-k2.7-code $0.67/M input, $3.40/M output, 262,144-token context; deepseek-v4-flash-0731 $0.14/M input, $0.28/M output, 1,310,720-token context; deepseek-v4-pro-0813 $1.122/$3.366. Output tokens are ~12x cheaper on V4-Flash and the context window is 5x larger. — OpenRouter model API (live) (2026)
  • As of July 31, 2026 ARC Prize verified Kimi K3 as the new open-weight high score on both benchmarks: 94.5% on ARC-AGI-1 Semi-Private at $0.77 per task and 60.4% on ARC-AGI-2 Semi-Private at $1.59 per task at max reasoning effort. The same model scores 86.7%/55.0% at high effort and 65.7%/12.4% at low effort. — ARC Prize — Verified results, Kimi K3 (2026)
  • DeepSeek-V4-Flash-0731 ships 304B parameters (167 GB of weights) under a plain MIT licence; the model card was last modified 2026-08-01. This is the newest open-weight release on either side of this comparison. — Hugging Face model API — deepseek-ai/DeepSeek-V4-Flash-0731 (2026)
  • Kimi K3's model card carries license: other with license_name: kimi-k3 — a custom Moonshot licence, not a standard OSI licence. Card last modified 2026-07-27. — Hugging Face model API — moonshotai/Kimi-K3 (2026)
  • Willison rates V4-Flash-0731 at $0.14/M input and $0.27/M output as possibly the best value-per-intelligence model available, noting Artificial Analysis ranks the 304B model ahead of the 428B MiniMax M3 — and that the default reasoning level produced a disappointing result until he raised reasoning_effort to high. — Simon Willison — deepseek-ai/DeepSeek-V4-Flash-0731 (2026)
  • Adoption splits by type of user: DeepSeek-V4-Flash has 2,814,414 downloads against 1,947 likes, while Kimi-K2.7-Code has 665,880 downloads against 1,347 likes and Kimi K3 has 559,924 downloads against 9,502 likes — DeepSeek wins on deployment volume, Moonshot on community enthusiasm. — Hugging Face model API — adoption counters (2026)
  • DeepSeek-V4-Flash-Vision-Exp: DeepSeek's experimental multimodal API model (text + image understanding) went live on the official DeepSeek API platform on 2026-08-21. It matches V4-Flash text capability and is offered on OpenRouter at $0.22/M input, $0.66/M output with a 1,048,576-token context — making V4 Flash the cheapest agentic multimodal option in this comparison. — DeepSeek API Docs — official release announcement (news260821) (2026)
  • OpenRouter live snapshot 2026-09-10: deepseek-v4-flash-0731 now lists $0.065/M input and $0.18/M output with a 1,310,720-token context — roughly half the late-August price ($0.14/$0.28). kimi-k2.7-code lists $0.71/$3.50 (262,144 tokens) and kimi-k3 $3/$15 (1,048,576 tokens). The V4-Flash output-price gap to K2.7 widens from ~12x to ~19x, while the unversioned first-party line sits at $0.0886/$0.177. — OpenRouter model API (live) (2026)
  • DeepSeek-V4.1-Flash confirmed by the official release note of 2026-09-10: a 552B-parameter MoE on a new causal encoder-decoder architecture with 8B active parameters for input and 16B for output, native visual understanding, served under the model id deepseek-flash; V4-Flash and V4-Flash-Vision-Exp are retired. — DeepSeek API Docs — news260910 (2026)
  • OpenRouter live snapshot 2026-09-11: deepseek/deepseek-v4.1-flash lists $0.30/M input and $1.20/M output with 1,048,576-token context. DeepSeek's note claims the KV cache needs just 1/4 the HBM and 1/8 the SSD storage of the previous generation, and off-peak rates are 50% of peak — a direct cut on the cache-hit share of agent costs. — OpenRouter model API (live) + DeepSeek release note (2026)
  • Redirect mechanics: from 04:00 UTC on 2026-09-14 all deepseek-v4-pro requests route to V4.1-Flash at V4.1-Flash rates until V4.1-Pro launches, and the retired v4-flash ids route temporarily as well — pinned model ids keep working but change what they resolve to. — DeepSeek API Docs — news260910 (2026)
  • DeepSeek V4.1 compresses its KV cache to 890 bytes per token; the model still needs ~510 GB of files for a local run — KV compression and learned n-gram lookup tables are the two axes of the efficiency frontier, independent of raw parameter count. — Cloud Codes — efficiency-frontier analysis (YouTube, 11.09.2026) (2026)
  • Redirect effective 2026-09-14: since 04:00 UTC all deepseek-v4-pro requests route to V4.1-Flash at V4.1-Flash rates until V4.1-Pro launches; V4-Flash and V4-Flash-Vision-Exp are retired and temporarily route to the same model (deepseek-flash) — the whole V4 line converges on one endpoint. — DeepSeek API Docs — news260910 (2026)
  • OpenRouter live snapshot 2026-09-14: deepseek-v4.1-flash $0.30/M input, $1.20/M output, cache reads $0.006/M, with off-peak windows at $0.15/$0.60 — the KV-cache compression (1/4 the HBM, 1/8 the SSD vs. the previous generation) feeds directly into cache-hit pricing, the dominant cost in agent loops. — OpenRouter model API (live) + DeepSeek release note news260910 (2026)

All statistics come from verified third-party sources. Source, year, and direct link are shown on each metric.

When to Choose Each Option

Clear guidance based on your specific situation and needs.

Our Recommendation

There is still no single winner, but the balance has shifted since this page was first written — and it shifted on both sides at once. DeepSeek is now the stronger default for almost every cost-sensitive production decision. On live OpenRouter prices deepseek-v4-flash-0731 runs $0.14/M input and $0.28/M output against $0.67/$3.40 for kimi-k2.7-code, roughly 12x cheaper on output, and it exposes a 1,310,720-token context window against K2.7's 262,144 — the 'tie' this page previously recorded on context was simply wrong. The 0731 build is also the newest open weights in the comparison (304B, 167 GB, card updated 2026-08-01) and it ships under a plain MIT licence, where Kimi K3's card declares license: other with license_name: kimi-k3. If you self-host, fine-tune or redistribute, that licence line is a decision by itself. What changed in Moonshot's favour is the validation story, but it did not change for the model this page names. ARC Prize verified Kimi K3 on July 31, 2026 at 94.5% on ARC-AGI-1 Semi-Private and 60.4% on ARC-AGI-2 — the highest open-weight scores it has recorded. Kimi K2.7 Code's own coding gains remain self-reported on Moonshot's Kimi Code Bench v2. So the honest reading is: the Kimi lineage has earned independent credibility, and the way to collect it is to move up to K3, not to stay on K2.7. K2.7 Code still holds a real lead where this page always said it did — MCP tool-use (76.0 MCP Atlas, 81.1 MCP Mark Verified), throughput (180-260 tokens/sec on the HighSpeed variant), and roughly 30% lower reasoning-token usage than K2.6. The most useful finding for anyone actually routing traffic is that the choice between these two families is not the biggest lever. ARC Prize's own numbers show Kimi K3 falling from 60.4% to 12.4% on ARC-AGI-2 purely by dropping reasoning effort from max to low, and Simon Willison reports the same shape on the other side: V4-Flash-0731 gave him a disappointing result at the default reasoning level and a much better one at reasoning_effort high. A five-fold accuracy swing inside one model is larger than the gap between the two models. Benchmark a candidate at the effort level you will actually pay for, or the headline price is fiction. Context Studios' pattern is unchanged in shape and sharper in detail: default high-volume bounded coding to DeepSeek V4-Flash-0731 for cost and context, escalate the hardest reasoning to V4-Pro, and route MCP-orchestration-heavy agent loops to the Kimi lineage — but budget for Kimi K3 rather than K2.7 now that K3 is the verified end of that path, and pin your reasoning-effort setting in the same config where you pin the model. Additionally, on 21 August 2026 DeepSeek shipped DeepSeek-V4-Flash-Vision-Exp — an experimental multimodal API model (text + image) with V4-Flash-level text performance, listed on OpenRouter at $0.22/M input and $0.66/M output with 1M context. That extends V4 Flash's production cases to visual agentic work (screenshots, UIs, diagrams) at a fraction of frontier-multimodal cost. Since the 2026-09-10 snapshot the DeepSeek side moved again: V4-Flash-0731 halved to $0.065/M input and $0.18/M output, widening the output gap to kimi-k2.7-code (~$0.71/$3.50) to ~19x and keeping V4-Flash the cheapest 1.3M-token route; a V4.1 Flash test build (native multimodal, ~400 tok/s, cutoff 2026-09-10) is reported but not yet on OpenRouter. The routing recipe is unchanged: V4-Flash for volume, V4-Pro for the hardest reasoning, the Kimi lineage for MCP-heavy loops, K3 as the verified end. Since 10.09.2026 the V4.1-Flash preview is confirmed (552B MoE, 8B/16B active, 1M context, $0.30/$1.20) — and with the 14.09 redirect DeepSeek collapses its line into one Flash-class backbone that also serves the old v4-pro id, while Kimi keeps the steadier id semantics. On the efficiency axis, V4.1's KV cache is reported at 890 bytes per token (~510 GB of files for local runs) — compression and lookup tables now matter as much as raw parameter counts.

Choose Kimi K2.7 Code when...
  • Your workload is MCP-heavy and tool-call accuracy is the real bottleneck
  • You are already on the Kimi K2.x lineage and want a drop-in upgrade path that ends at ARC-verified K3
  • Token throughput matters more to you than token price, and the HighSpeed variant's 180-260 tokens/sec pays for itself
  • You can run at high or max reasoning effort and are buying peak capability rather than cost-per-task
Choose DeepSeek V4 when...
  • Cost per token is your primary constraint — output tokens run roughly 19x cheaper on V4-Flash-0731
  • You need a million-token context window; K2.7 Code tops out at 262,144 tokens
  • You self-host, fine-tune or redistribute weights and want a plain MIT licence with no legal review
  • You want the newest weights on either side, refreshed 2026-07-31, on an already-broad multi-provider deployment surface
  • You want one family spanning a cheap Flash tier and a frontier Pro tier for routing
  • You want the widest documented price lead: V4-Flash-0731 halved to $0.065/$0.18 per million tokens in the 2026-09-10 snapshot
  • You want the September 2026 consolidation: one 1M-token, vision-capable Flash line (8B/16B active, 1/4 KV-cache footprint) that also absorbs the v4-pro endpoint on 14.09.

Common questions about this comparison answered.

Frequently Asked Questions

(01)Is Kimi K2.7 or DeepSeek V4 better for coding?
It depends on your constraint. DeepSeek V4 is the safer pick for cost-sensitive, high-volume work: it is independently benchmarked (reported 83.7% on SWE-bench Verified, #14 on BenchLM), has been in production since April 2026, and its V4-Flash tier is among the cheapest serious coding APIs. Kimi K2.7 Code is stronger for MCP-heavy agentic workflows, leading tool-use benchmarks (76.0 MCP Atlas, 81.1 MCP Mark Verified) with high throughput — but its headline coding gains are still largely self-reported, so validate on your own tasks first.
(02)Which is cheaper, Kimi K2.7 or DeepSeek V4?
DeepSeek V4 is cheaper. V4-Flash lists around $0.28 per million output tokens and V4-Pro around $0.87, among the lowest for serious coding models. Kimi K2.7 Code is priced at $0.95 per million input and $4.00 per million output, with cache hits as low as $0.19 per million — competitive, but well above DeepSeek's Flash tier on output cost.
(03)Are Kimi K2.7's benchmark scores independently verified?
Not yet, mostly. At launch, Kimi K2.7's headline coding gains (+21.8% over K2.6) come from Moonshot's own Kimi Code Bench v2, and independent SWE-bench numbers are still thin. DeepSeek V4, by contrast, already appears on independent leaderboards like Vals AI and BenchLM. Treat Kimi's launch figures as promising but unconfirmed until third-party benchmarks are published.
(04)Can I self-host Kimi K2.7 and DeepSeek V4?
Yes — both are open-weight models, so you can run them on your own infrastructure for data-residency or compliance reasons, in addition to using their hosted APIs. DeepSeek V4 is already available across multiple providers (Fireworks, DeepInfra, Novita, SiliconFlow). Note that the MoE architectures are large: Kimi K2.7 is 1T total parameters and DeepSeek V4-Pro is 1.6T, so self-hosting the top tiers needs substantial memory bandwidth.
(05)Are Kimi's benchmark scores independently verified now?
Partly — and not on the model this page names. ARC Prize verified Kimi K3 on July 31, 2026 at 94.5% on ARC-AGI-1 Semi-Private ($0.77/task) and 60.4% on ARC-AGI-2 ($1.59/task), the highest open-weight scores it has recorded. Kimi K2.7 Code itself still rests on Moonshot's own Kimi Code Bench v2. If independent validation is what you are waiting for, the answer is to move up to K3, not to trust K2.7's launch numbers.
(06)Which licence can I actually ship on?
DeepSeek is the cleaner answer. Both DeepSeek-V4-Flash and the newer V4-Flash-0731 declare a plain MIT licence on their Hugging Face cards. Kimi K3's card declares license: other with license_name: kimi-k3 — a custom Moonshot licence. MIT needs no legal review before you self-host, fine-tune or redistribute; a bespoke licence does.
(07)What changed in the September 2026 price cut?
On the 2026-09-10 OpenRouter snapshot deepseek-v4-flash-0731 lists $0.065/M input and $0.18/M output, down from $0.14/$0.28 in late August — roughly a halving. kimi-k2.7-code sits at $0.71/$3.50 and kimi-k3 at $3/$15 (1M-token window), so the output gap widened to ~19x. DeepSeek also reported a V4.1 Flash test build (native multimodal, ~400 tok/s, cutoff 2026-09-10) that is not yet on OpenRouter. Check live prices at routing time.
(08)What did V4.1-Flash change on 10 September 2026?
It became the confirmed flagship-tier Flash: 552B MoE, 8B active parameters on input and 16B on output, 1M-token context and native vision, listed at $0.30/$1.20 per million tokens in the 11.09 OpenRouter snapshot. V4-Flash and the Vision-Exp are retired, and from 14.09 04:00 UTC the deepseek-v4-pro id itself routes to V4.1-Flash at Flash rates — re-run your evals after that switch.

Need help deciding?

Book a free 30-minute consultation and we'll help you determine the best approach for your specific project.

Free consultation · No obligation · Personal reply