Kimi K3 Puts a 2.8T Open Model on Your Shortlist

Moonshot's Kimi K3 is a 2.8-trillion-parameter open-weight model at $3/$15 per million tokens. Why the July 27 weights date matters more than the benchmarks.

Kimi K3 Puts a 2.8T Open Model on Your Shortlist
If you maintain a model-selection policy, Kimi K3 is the first open-weight release that belongs in it.
Moonshot AI shipped the 2.8-trillion-parameter model on July 16, 2026, priced it at $3 per million input tokens and $15 per million output tokens, and set the open-weights date for July 27, 2026 ([EqualOcean](https://equalocean.com/news/2026071722035-moonshot-ai-unveils-2-8-trillion-parameter-kimi-k3-open-weights-due-july-27), [Trilogy AI](https://trilogyai.substack.com/p/kimi-k3-is-live-pricing-benchmarks)). We route production traffic across several model vendors instead of one, and every new frontier-class entrant has to earn its share on cost per accepted task before it gets any of it. Here is what shipped, which claims are still vendor-reported, and what actually changes for your routing.

What Moonshot Actually Shipped

Kimi K3 is a 2.8-trillion-parameter Mixture-of-Experts model with a 1-million-token context window and native vision, activating 16 of 896 experts per token.
Moonshot describes the architecture as built on "Kimi Delta Attention and Attention Residuals" and calls it the world's first open 3T-class model ([The Hindu](https://www.thehindu.com/sci-tech/technology/what-is-kimi-k3-chinas-first-open-ai-model-to-reach-28-trillion-parameters/article71232767.ece)). The expert-activation detail comes from launch-day technical summaries ([Next Tool](https://buttondown.com/nexttool/archive/open-models-just-hit-28-trillion-parameters-and-4)), and the July 16, 2026 release date is corroborated independently ([DEV Community](https://dev.to/jamilxt/kimi-k3-chinas-28-trillion-parameter-open-model-just-raised-the-bar-4in1)).

K3 is reachable now through the Moonshot API, kimi.com, Kimi Work and Kimi Code, and routing platforms list it as an ultra-large open-weight multimodal reasoning model aimed at large-repository navigation, tool use and long-horizon agentic work (OpenRouter). The weights themselves are not out. Moonshot's Hugging Face organisation carried no public Kimi K3 repository on July 18, 2026 (huggingface.co/moonshotai), which matches the stated July 27 date. For scale, launch coverage frames K3 as a step change in size over previous open-weight releases (The Stack).

The Benchmark Claims Are Still Vendor-Reported

Moonshot itself says Kimi K3 trails Anthropic's Claude Fable 5 and OpenAI's GPT-5.6 Sol on overall performance, while beating Claude Opus 4.8 and GPT-5.5 on the benchmarks it ran.
That framing comes from the company, not from an independent reproduction ([CNBC](https://www.cnbc.com/2026/07/17/moonshot-ai-kimi-k3-model-openai-anthropic-china.html)). Moonshot's broader public claim is that Kimi K3 rivals the top American labs ([BBC](https://www.bbc.com/news/articles/cy9w4q8pgp0o)).

Independent tracking is thinner than the headlines suggest. K3 debuted in the top five of the Artificial Analysis intelligence ranking, behind Anthropic's Claude Fable 5 and GPT-5.6 Sol, and led Arena.ai's front-end web development board (Wikipedia, Artificial Analysis) — though write-ups dated July 17, 2026 place it third or fourth rather than at one agreed position. At least one tracker declined to publish scores at all because the underlying figures were not verifiable at release (BenchLM). Treat every number as reported until the weights are public and somebody else runs them.

Our Read: Open Did Not Mean Cheap

The reflex assumption about an open-weight release is that it undercuts the incumbents on price. K3 does not. At $3 and $15 per million tokens it sits in frontier-adjacent territory, and Artificial Analysis flags it as expensive relative to models in the same price band (Artificial Analysis). If you expected the cost-per-task argument we made for Grok 4.5 to repeat here, it does not. This is a different move.

The economics of Kimi K3 change on July 27, 2026, when the weights land — not on the API price it launched with.
Open weights buy three things an API does not: the option to self-host when a vendor's terms or availability shift, a deployment you control for data-residency reasons, and a model you can still run in two years regardless of anyone's roadmap. That is the same hedging logic behind our read on [what a public Anthropic changes about your Claude stack](https://www.contextstudios.ai/blog/anthropic-ipo-changes-the-ground-under-your-claude-stack) and on the [split between US access controls and Chinese behaviour rules](https://www.contextstudios.ai/blog/the-us-gates-ai-access-china-gates-ai-behavior). A Chinese-origin open model does not remove that second constraint. It relocates it, and your compliance review has to say so in writing.

What This Means for You

Three moves, in order.

Do not re-route on launch-day numbers. Add K3 to your evaluation set, not your production mix. Score it on your own accepted-task rate, at your own prompt lengths, against whatever you run in production. A 1-million-token context window is only worth paying for if your workloads actually reach it.

Put July 27, 2026 in the calendar as the real decision point. That is when self-hosting becomes testable and when independent benchmark reproductions start to appear. Everything before that date is vendor-reported.

Write the open-weight option into your routing policy now. If your gateway or routing layer cannot take a fourth vendor without a code change, the largest open model in the world is not usable to you anyway. The same holds for your coding-agent stack: a model you cannot swap in behind an existing interface is a model you will never evaluate honestly.

If you want that turned into an actual policy — evaluation harness, routing rules, a documented fallback per workload — that is the kind of work we do.

Frequently Asked Questions

What is Kimi K3? Kimi K3 is a 2.8-trillion-parameter Mixture-of-Experts model released by Moonshot AI on July 16, 2026, with a 1-million-token context window and native vision. Moonshot calls it the world's first open 3T-class model (The Hindu).

When do the Kimi K3 weights become available? Moonshot has stated the open weights are due by July 27, 2026. Until then the model is reachable only through the API and Moonshot's own products (EqualOcean).

Is Kimi K3 better than Claude Fable 5 or GPT-5.6 Sol? No. By Moonshot's own account K3 trails both on overall performance, while beating Claude Opus 4.8 and GPT-5.5 on the benchmarks Moonshot ran (CNBC).

How much does Kimi K3 cost? Reported API pricing is $3 per million input tokens and $15 per million output tokens (Trilogy AI). Moonshot's own pricing documentation is the authoritative reference (Moonshot pricing).

Should I move production traffic to Kimi K3? Not on launch-day figures. Add it to your evaluation set, measure cost per accepted task on your own workloads, and revisit after the July 27, 2026 weights release, when independent reproductions exist.

Sources

  1. Moonshot AI — API pricing documentation
  2. Moonshot AI — blog
  3. Artificial Analysis — intelligence, performance and price analysis
  4. OpenRouter — API pricing and providers
  5. Hugging Face — moonshotai organisation
  6. BBC — China's Moonshot AI claims Kimi K3 can rival OpenAI and Anthropic
  7. CNBC — China's Moonshot AI unveils a model it says rivals US labs
  8. Wikipedia — Kimi (chatbot)
  9. The Hindu — What is Kimi K3, China's first open AI model to reach 2.8 trillion parameters
  10. EqualOcean — Moonshot AI unveils the model, open weights due July 27
  11. The Stack — China's open-weight model snaps at the heels of US models
  12. Trilogy AI — it is live: pricing, benchmarks, and the wait for open source
  13. Next Tool — Open models just hit 2.8 trillion parameters
  14. DEV Community — China's 2.8 trillion parameter open model
  15. BenchLM — Kimi 3 arrived before the data

Share article

Share: