Provider Comparison

Kimi K3 vs Claude Opus 5 (2026): The Weights Shipped — Now Read the Licence and the VRAM Bill

Kimi K3's weights shipped 27 July 2026. What actually decides it now: licence thresholds, a 1.5TB VRAM bill, price and reproducibility.

3
Kimi K3
vs
2
Claude Opus 5
Quick Verdict

The premise this page was built on is gone, and it is worth being precise about why. Kimi K3's weights were promised for 27 July 2026 and delivered on 27 July 2026, at 16:29 UTC. Our check ran at 03:50 UTC the same morning and found nothing, which was accurate at that moment and misleading by the afternoon. Treat that as the lesson it is: a claim about whether something exists yet has a half-life measured in hours, and it does not age into the stale bucket, it just quietly becomes false. With that corrected, the comparison is stronger for Kimi K3 than it was, but on different ground. The weights are real, heavily adopted and already re-quantised by Red Hat and Unsloth. The price advantage holds at $3/$15 against $5/$25. The context window is a genuine tie at roughly a million tokens each. And K3 does something Opus does not: it publishes its eval results as machine-readable YAML next to downloadable weights, which turns vendor claims into something a third party can falsify. Nobody has published that rerun yet, so we score it a tie rather than a win. What the open weights do not do is make self-hosting a normal decision. At 2.8 trillion parameters, K3 needs over 1.5TB of VRAM before a KV cache and will not fit an eight-GPU B200 node. The practical answer is eight 288GB cards, and the most interesting number we found is that the cheaper card wins: 8x AMD MI355X delivered 48 tokens per second per dollar of GPU-hour against 33 on a B300 that costs roughly 2.4x more per GPU. If you have been costing self-hosting at NVIDIA prices, that is a real gap in the model. So: choose Kimi K3 when your use is internal, cost per token is binding, and 288GB-class hardware is something you have or can rent. Choose Claude Opus 5 when you want EU data governance with a Compliance API, when you would rather not own a serving stack, or when you are a Model-as-a-Service business approaching $20M in revenue — because at that line the Kimi K3 Licence stops being permissive and starts being a negotiation. The open weights are not the end of the decision. They move it from availability to economics and legal terms, which is where it should have been all along.

Detailed Comparison

A side-by-side analysis of key factors to help you make the right choice.

Factor
Kimi K3Recommended
Claude Opus 5Winner
Open weights & self-hosting
Shipped: moonshotai/Kimi-K3 published 27 Jul 2026, 16:29 UTC — 96 safetensors shards, 837k downloads and 9.6k likes in six days
Closed and API-only by design; no weights are offered at any tier
Commercial licence terms
Kimi K3 Licence, not MIT: internal use unrestricted, but a Model-as-a-Service business above $20M revenue over any 12 months needs a separate Moonshot agreement
Standard commercial API terms: no revenue threshold, no attribution duty, no separate agreement to negotiate
Weight-release delivery
Met its own announced date — weights published on 27 July 2026, the day promised
Publishes no weights by design, so there is no release commitment to keep or miss
Price per 1M tokens
$3 in / $15 out (OpenRouter, 3 Aug 2026) — roughly 40% below the Opus tier
$5 in / $25 out (OpenRouter, 3 Aug 2026), unchanged from Opus 4.8
Real cost of self-hosting
2.8T parameters need over 1.5TB of VRAM before any KV cache — it will not fit an 8x B200 node; you need 8x MI355X or B300 (288GB each), or 16 B200s across two nodes
No infrastructure at all: no GPU procurement, no serving stack, no kernel debugging
Independent benchmark verification
Ships its eval results as YAML inside the repo (GPQA-Diamond 93.5, HLE 56, DeepSWE 67.5, APEX-Agents 41) — still vendor-produced, but the public weights make them reproducible
Vendor-graded tables with no weights to reproduce them against
EU data governance
China-hosted API, or self-hosting in your own region if you can fund the hardware
EU regions via Bedrock and Vertex plus a Compliance API — the shorter path to a signed DPA
Context window
1,048,576 tokens (OpenRouter, verified 3 Aug 2026)
1,000,000 tokens (OpenRouter, verified 3 Aug 2026)
Native multimodality
Native vision built in, not bolted on — the repo ships a vision processor and registers as an image-text-to-text model
Multimodal input is supported across the Claude platform
Total Score3/ 92/ 94 ties
Open weights & self-hosting
Kimi K3
Shipped: moonshotai/Kimi-K3 published 27 Jul 2026, 16:29 UTC — 96 safetensors shards, 837k downloads and 9.6k likes in six days
Claude Opus 5
Closed and API-only by design; no weights are offered at any tier
Commercial licence terms
Kimi K3
Kimi K3 Licence, not MIT: internal use unrestricted, but a Model-as-a-Service business above $20M revenue over any 12 months needs a separate Moonshot agreement
Claude Opus 5
Standard commercial API terms: no revenue threshold, no attribution duty, no separate agreement to negotiate
Weight-release delivery
Kimi K3
Met its own announced date — weights published on 27 July 2026, the day promised
Claude Opus 5
Publishes no weights by design, so there is no release commitment to keep or miss
Price per 1M tokens
Kimi K3
$3 in / $15 out (OpenRouter, 3 Aug 2026) — roughly 40% below the Opus tier
Claude Opus 5
$5 in / $25 out (OpenRouter, 3 Aug 2026), unchanged from Opus 4.8
Real cost of self-hosting
Kimi K3
2.8T parameters need over 1.5TB of VRAM before any KV cache — it will not fit an 8x B200 node; you need 8x MI355X or B300 (288GB each), or 16 B200s across two nodes
Claude Opus 5
No infrastructure at all: no GPU procurement, no serving stack, no kernel debugging
Independent benchmark verification
Kimi K3
Ships its eval results as YAML inside the repo (GPQA-Diamond 93.5, HLE 56, DeepSWE 67.5, APEX-Agents 41) — still vendor-produced, but the public weights make them reproducible
Claude Opus 5
Vendor-graded tables with no weights to reproduce them against
EU data governance
Kimi K3
China-hosted API, or self-hosting in your own region if you can fund the hardware
Claude Opus 5
EU regions via Bedrock and Vertex plus a Compliance API — the shorter path to a signed DPA
Context window
Kimi K3
1,048,576 tokens (OpenRouter, verified 3 Aug 2026)
Claude Opus 5
1,000,000 tokens (OpenRouter, verified 3 Aug 2026)
Native multimodality
Kimi K3
Native vision built in, not bolted on — the repo ships a vision processor and registers as an image-text-to-text model
Claude Opus 5
Multimodal input is supported across the Claude platform

Key Statistics

Real data from verified industry sources to support your decision.

Kimi K3's weights were published to Hugging Face on 27 July 2026 at 16:29 UTC — 96 safetensors shards across 118 files, drawing 837,202 downloads and 9,666 likes within six days

Hugging Face API (re-probed 3 August 2026)

The Kimi K3 Licence is not MIT: a Model-as-a-Service business whose aggregate revenue passes $20M over any consecutive 12 months must sign a separate agreement with Moonshot before commercial use, and any product above 100M monthly active users or $20M monthly revenue must display "Kimi K3" in its interface. Purely internal use is exempt from both.

Kimi K3 Licence (first-party, in-repo)

Moonshot publishes K3's eval results as machine-readable YAML inside the model repository: GPQA-Diamond 93.5, Humanity's Last Exam 56, DeepSWE 67.5 and APEX-Agents 41 — vendor-produced numbers, but backed by downloadable weights

Kimi K3 model card and in-repo .eval_results

Kimi K3 lists at $3 per million input and $15 per million output tokens with a 1,048,576-token context, against Claude Opus 5 at $5 / $25 and 1,000,000 tokens

OpenRouter models API (checked 3 August 2026)

K3's 2.8 trillion parameters need more than 1.5TB of VRAM before allocating a KV cache, so the model does not fit an eight-GPU B200 node at all. On 8x AMD MI355X it reached 952 tokens/second per node and 48 tokens/second per dollar of GPU-hour, versus 33 on a B300 that is roughly 2.4x more expensive per GPU

Wafer, Ian Ye (31 July 2026)

K3 activates 16 of 896 experts per token across 2.8T parameters using Kimi Delta Attention and Attention Residuals, which Moonshot reports as roughly 2.5x the scaling efficiency of K2; it is billed as the first open 3T-class model and ships native vision

Kimi K3 model card

All statistics come from verified third-party sources. Source, year, and direct link are shown on each metric.

When to Choose Each Option

Clear guidance based on your specific situation and needs.

Choose Kimi K3 when...

  • You want weights you can actually download, inspect and pin — they exist now, and over 800,000 downloads say the supply chain is real
  • Your use is internal, which puts you outside both licence thresholds entirely
  • API cost is the binding constraint and roughly 40% per token decides it
  • You already run 288GB-class accelerators, or your provider does, so a 1.5TB deployment is a procurement question rather than a fantasy

Choose Claude Opus 5 when...

  • You need the model to work without owning, renting or debugging a serving stack
  • You operate under EU or enterprise data-governance rules and want EU regions plus a Compliance API rather than a China-hosted endpoint
  • You are a Model-as-a-Service business near the $20M revenue line and do not want a licence negotiation on the critical path
  • You want benchmark claims from a vendor that carries contractual support, and you are not in a position to rerun evals yourself

Our Recommendation

The premise this page was built on is gone, and it is worth being precise about why. Kimi K3's weights were promised for 27 July 2026 and delivered on 27 July 2026, at 16:29 UTC. Our check ran at 03:50 UTC the same morning and found nothing, which was accurate at that moment and misleading by the afternoon. Treat that as the lesson it is: a claim about whether something exists yet has a half-life measured in hours, and it does not age into the stale bucket, it just quietly becomes false. With that corrected, the comparison is stronger for Kimi K3 than it was, but on different ground. The weights are real, heavily adopted and already re-quantised by Red Hat and Unsloth. The price advantage holds at $3/$15 against $5/$25. The context window is a genuine tie at roughly a million tokens each. And K3 does something Opus does not: it publishes its eval results as machine-readable YAML next to downloadable weights, which turns vendor claims into something a third party can falsify. Nobody has published that rerun yet, so we score it a tie rather than a win. What the open weights do not do is make self-hosting a normal decision. At 2.8 trillion parameters, K3 needs over 1.5TB of VRAM before a KV cache and will not fit an eight-GPU B200 node. The practical answer is eight 288GB cards, and the most interesting number we found is that the cheaper card wins: 8x AMD MI355X delivered 48 tokens per second per dollar of GPU-hour against 33 on a B300 that costs roughly 2.4x more per GPU. If you have been costing self-hosting at NVIDIA prices, that is a real gap in the model. So: choose Kimi K3 when your use is internal, cost per token is binding, and 288GB-class hardware is something you have or can rent. Choose Claude Opus 5 when you want EU data governance with a Compliance API, when you would rather not own a serving stack, or when you are a Model-as-a-Service business approaching $20M in revenue — because at that line the Kimi K3 Licence stops being permissive and starts being a negotiation. The open weights are not the end of the decision. They move it from availability to economics and legal terms, which is where it should have been all along.

Frequently Asked Questions

Common questions about this comparison answered.

Yes. moonshotai/Kimi-K3 went live on Hugging Face on 27 July 2026 at 16:29 UTC, on the date Moonshot had announced. This page previously said the deadline had passed with nothing published — that came from an automated check run at 03:50 UTC the same morning, about thirteen hours before the upload. The finding was true when it was made and wrong by that afternoon. The weights have since been downloaded over 800,000 times, and Red Hat, Unsloth and others have published quantised builds.
Only with frontier-class hardware. At 2.8 trillion parameters the weights alone exceed 1.5TB of VRAM before you allocate any KV cache for a million-token context, which means the model will not fit on a standard eight-GPU B200 node. Practical configurations are eight 288GB cards — AMD MI355X or NVIDIA B300 — or sixteen B200s spanning two nodes with the cross-node all-reduce penalty that implies. If that sentence does not describe hardware you already have, the open weights are a licence to negotiate with a hosting provider, not a plan to run it yourself.
For internal use, yes, without conditions. Two thresholds matter beyond that. If you operate a Model-as-a-Service business — giving third parties meaningful control over inputs, parameters or training data — and your aggregate revenue crosses $20M over any consecutive twelve months, you must sign a separate agreement with Moonshot before commercial use. Separately, any product above 100 million monthly active users or $20M in monthly revenue must display "Kimi K3" prominently in its interface. Note these are two different triggers, not one: the revenue threshold governs the agreement, the user threshold governs attribution.
On list price, about 40% less: $3/$15 per million tokens against $5/$25, both verified on OpenRouter on 3 August 2026. That gap is real for API use. It is not the number to use if you are pricing self-hosting, where the comparison is against a multi-hundred-thousand-dollar GPU deployment and the engineering time to keep it serving.
Not yet, as far as we can verify. The GPQA, HLE, DeepSWE and APEX-Agents figures are Moonshot's own, published as YAML inside the repository. What has changed is that they are now falsifiable: with public weights, a third party can rerun them. That is a materially stronger position than a vendor table nobody can check, but it is not the same as an independent result, and we would not treat these numbers as confirmed until someone publishes one.

Need help deciding?

Book a free 30-minute consultation and we'll help you determine the best approach for your specific project.

Free consultation
No obligation
Response within 24h