Kimi K3 vs Claude Opus 5 (2026): The Weights Shipped — Now Read the Licence and the VRAM Bill
Kimi K3's weights shipped 27 July 2026. What actually decides it now: licence thresholds, a 1.5TB VRAM bill, price and reproducibility.
The premise this page was built on is gone, and it is worth being precise about why. Kimi K3's weights were promised for 27 July 2026 and delivered on 27 July 2026, at 16:29 UTC. Our check ran at 03:50 UTC the same morning and found nothing, which was accurate at that moment and misleading by the afternoon. Treat that as the lesson it is: a claim about whether something exists yet has a half-life measured in hours, and it does not age into the stale bucket, it just quietly becomes false. With that corrected, the comparison is stronger for Kimi K3 than it was, but on different ground. The weights are real, heavily adopted and already re-quantised by Red Hat and Unsloth. The price advantage holds at $3/$15 against $5/$25. The context window is a genuine tie at roughly a million tokens each. And K3 does something Opus does not: it publishes its eval results as machine-readable YAML next to downloadable weights, which turns vendor claims into something a third party can falsify. Nobody has published that rerun yet, so we score it a tie rather than a win. What the open weights do not do is make self-hosting a normal decision. At 2.8 trillion parameters, K3 needs over 1.5TB of VRAM before a KV cache and will not fit an eight-GPU B200 node. The practical answer is eight 288GB cards, and the most interesting number we found is that the cheaper card wins: 8x AMD MI355X delivered 48 tokens per second per dollar of GPU-hour against 33 on a B300 that costs roughly 2.4x more per GPU. If you have been costing self-hosting at NVIDIA prices, that is a real gap in the model. So: choose Kimi K3 when your use is internal, cost per token is binding, and 288GB-class hardware is something you have or can rent. Choose Claude Opus 5 when you want EU data governance with a Compliance API, when you would rather not own a serving stack, or when you are a Model-as-a-Service business approaching $20M in revenue — because at that line the Kimi K3 Licence stops being permissive and starts being a negotiation. The open weights are not the end of the decision. They move it from availability to economics and legal terms, which is where it should have been all along.
Detailed Comparison
A side-by-side analysis of key factors to help you make the right choice.
| Factor | Kimi K3Recommended | Claude Opus 5 | Winner |
|---|---|---|---|
| Open weights & self-hosting | Shipped: moonshotai/Kimi-K3 published 27 Jul 2026, 16:29 UTC — 96 safetensors shards, 837k downloads and 9.6k likes in six days | Closed and API-only by design; no weights are offered at any tier | |
| Commercial licence terms | Kimi K3 Licence, not MIT: internal use unrestricted, but a Model-as-a-Service business above $20M revenue over any 12 months needs a separate Moonshot agreement | Standard commercial API terms: no revenue threshold, no attribution duty, no separate agreement to negotiate | |
| Weight-release delivery | Met its own announced date — weights published on 27 July 2026, the day promised | Publishes no weights by design, so there is no release commitment to keep or miss | |
| Price per 1M tokens | $3 in / $15 out (OpenRouter, 3 Aug 2026) — roughly 40% below the Opus tier | $5 in / $25 out (OpenRouter, 3 Aug 2026), unchanged from Opus 4.8 | |
| Real cost of self-hosting | 2.8T parameters need over 1.5TB of VRAM before any KV cache — it will not fit an 8x B200 node; you need 8x MI355X or B300 (288GB each), or 16 B200s across two nodes | No infrastructure at all: no GPU procurement, no serving stack, no kernel debugging | |
| Independent benchmark verification | Ships its eval results as YAML inside the repo (GPQA-Diamond 93.5, HLE 56, DeepSWE 67.5, APEX-Agents 41) — still vendor-produced, but the public weights make them reproducible | Vendor-graded tables with no weights to reproduce them against | |
| EU data governance | China-hosted API, or self-hosting in your own region if you can fund the hardware | EU regions via Bedrock and Vertex plus a Compliance API — the shorter path to a signed DPA | |
| Context window | 1,048,576 tokens (OpenRouter, verified 3 Aug 2026) | 1,000,000 tokens (OpenRouter, verified 3 Aug 2026) | |
| Native multimodality | Native vision built in, not bolted on — the repo ships a vision processor and registers as an image-text-to-text model | Multimodal input is supported across the Claude platform | |
| Total Score | 3/ 9 | 2/ 9 | 4 ties |
Key Statistics
Real data from verified industry sources to support your decision.
Hugging Face API (re-probed 3 August 2026)
Kimi K3 Licence (first-party, in-repo)
Kimi K3 model card and in-repo .eval_results
OpenRouter models API (checked 3 August 2026)
Wafer, Ian Ye (31 July 2026)
Kimi K3 model card
All statistics come from verified third-party sources. Source, year, and direct link are shown on each metric.
When to Choose Each Option
Clear guidance based on your specific situation and needs.
Choose Kimi K3 when...
- You want weights you can actually download, inspect and pin — they exist now, and over 800,000 downloads say the supply chain is real
- Your use is internal, which puts you outside both licence thresholds entirely
- API cost is the binding constraint and roughly 40% per token decides it
- You already run 288GB-class accelerators, or your provider does, so a 1.5TB deployment is a procurement question rather than a fantasy
Choose Claude Opus 5 when...
- You need the model to work without owning, renting or debugging a serving stack
- You operate under EU or enterprise data-governance rules and want EU regions plus a Compliance API rather than a China-hosted endpoint
- You are a Model-as-a-Service business near the $20M revenue line and do not want a licence negotiation on the critical path
- You want benchmark claims from a vendor that carries contractual support, and you are not in a position to rerun evals yourself
Our Recommendation
The premise this page was built on is gone, and it is worth being precise about why. Kimi K3's weights were promised for 27 July 2026 and delivered on 27 July 2026, at 16:29 UTC. Our check ran at 03:50 UTC the same morning and found nothing, which was accurate at that moment and misleading by the afternoon. Treat that as the lesson it is: a claim about whether something exists yet has a half-life measured in hours, and it does not age into the stale bucket, it just quietly becomes false. With that corrected, the comparison is stronger for Kimi K3 than it was, but on different ground. The weights are real, heavily adopted and already re-quantised by Red Hat and Unsloth. The price advantage holds at $3/$15 against $5/$25. The context window is a genuine tie at roughly a million tokens each. And K3 does something Opus does not: it publishes its eval results as machine-readable YAML next to downloadable weights, which turns vendor claims into something a third party can falsify. Nobody has published that rerun yet, so we score it a tie rather than a win. What the open weights do not do is make self-hosting a normal decision. At 2.8 trillion parameters, K3 needs over 1.5TB of VRAM before a KV cache and will not fit an eight-GPU B200 node. The practical answer is eight 288GB cards, and the most interesting number we found is that the cheaper card wins: 8x AMD MI355X delivered 48 tokens per second per dollar of GPU-hour against 33 on a B300 that costs roughly 2.4x more per GPU. If you have been costing self-hosting at NVIDIA prices, that is a real gap in the model. So: choose Kimi K3 when your use is internal, cost per token is binding, and 288GB-class hardware is something you have or can rent. Choose Claude Opus 5 when you want EU data governance with a Compliance API, when you would rather not own a serving stack, or when you are a Model-as-a-Service business approaching $20M in revenue — because at that line the Kimi K3 Licence stops being permissive and starts being a negotiation. The open weights are not the end of the decision. They move it from availability to economics and legal terms, which is where it should have been all along.
Frequently Asked Questions
Common questions about this comparison answered.
Need help deciding?
Book a free 30-minute consultation and we'll help you determine the best approach for your specific project.