Qwen3.8-Max vs. Kimi K3 (2026): A Shipped 2.78T Against a Launched 2.4T
Qwen3.8-Max vs Kimi K3 (2026): Qwen went GA and is cheaper, but K3 is the only one of the two you can actually download and serve yourself.
Qwen3.8-Max is now the cheaper model and Kimi K3 is still the only deployable one, and which of those matters more depends entirely on whether you ever plan to run the thing yourself. The price case for Qwen is real and no longer speculative. Since going generally available on 3 August 2026 it trades on the open market at $2.00 per million input tokens and $6.00 per million output, against Kimi K3's $3.00 and $15.00 — 33 percent cheaper in, 60 percent cheaper out, with cached input at $0.25 per million. It also takes video as input, which K3 does not, and it carries a 1 million token context window. If you are buying inference through an API and never intend to leave it, Qwen3.8-Max is the better commercial deal today. The deployment case for Kimi K3 hardened while nobody was looking. Moonshot's weights went live on 27 July 2026 at 16:29 UTC and have since been downloaded more than 960,000 times; the repository reports 2,779,931,837,184 parameters, which is not a rounded marketing figure but a summed tensor count anyone can re-derive. Ten independent providers now serve it — Together, Fireworks, BaseTen, Modal, DigitalOcean, Chutes, Morph, Wafer and Moonshot itself — which is why its price is contestable at all. Qwen3.8-Max is served by exactly one company: Alibaba. A single-vendor model has a list price, not a market price. The honest caveat on both sides: Alibaba's headline figures — 2.4 trillion total, 95 billion active, fifth in Text Arena — are vendor-reported and were announced alongside a weights release that has not happened. As of 4 August 2026 there is no Qwen3.8-Max repository on Hugging Face and none on ModelScope, Alibaba's own hub. Kimi K3's licence is likewise not an OSI licence; Hugging Face records it as "other". Neither model is open in the sense a compliance team means. Our recommendation: if the workload lives behind an API and cost per token is the constraint, take Qwen3.8-Max. If you have to self-host, negotiate against more than one supplier, or defend a number to someone who will check it, take Kimi K3 — and re-check the Qwen weights claim before you commit, because it is the kind of statement that inverts without warning.
Detailed Comparison
A side-by-side analysis of key factors to help you make the right choice.
| Factor | Qwen3.8-MaxRecommended | Kimi K3 | Winner |
|---|---|---|---|
| Weights you can actually download | None as of 4 August 2026 — no Hugging Face repository and no entry on Alibaba's own ModelScope hub; release announced for "next week" | Live since 27 July 2026, 16:29 UTC on Hugging Face; over 960,000 downloads and 9,800+ likes in the first week | |
| How the parameter count can be checked | 2.4 trillion total / 95 billion active — vendor-reported in the launch announcement, not independently re-derivable | 2,779,931,837,184 parameters summed from the published tensor index — anyone with the repository can recompute it | |
| Open-market price per million tokens | $2.00 input / $6.00 output, cached input $0.25 — the cheaper model on both axes | $3.00 input / $15.00 output, cached input $0.30 | |
| Number of suppliers you can buy from | One — Alibaba. A single-vendor model has a list price, not a market price | Ten independent providers including Together, Fireworks, BaseTen, Modal, DigitalOcean and Wafer, so the price is contestable | |
| Input modalities | Text, images and video in one model — video input is the clearest capability Kimi K3 does not have | Text and images; strong on coding and long-horizon agentic work, no video input | |
| Context window and output ceiling | 1,000,000 tokens of context, capped at 131,072 tokens of completion per request | 1,048,576 tokens of context, with no published per-request completion cap | |
| Licence you would actually be bound by | Not published — there is no weights release yet, so there is no licence to read | Published but not OSI-approved; Hugging Face records the licence as "other", with commercial thresholds in the text | |
| How fast the headline claims decay | Announcement-shaped: an unshipped weights release and vendor arena placements, both of which can change within days | Artifact-shaped: a dated repository, a download counter and a tensor index that do not move retroactively | |
| Total Score | 2/ 8 | 6/ 8 | 0 ties |
Key Statistics
Real data from verified industry sources to support your decision.
Hugging Face Model API — moonshotai/Kimi-K3
Hugging Face Model API — moonshotai/Kimi-K3
OpenRouter Models API
OpenRouter Endpoints API
ModelScope Model API (HTTP 404) and Hugging Face model search
Alibaba Cloud press room
All statistics come from verified third-party sources. Source, year, and direct link are shown on each metric.
When to Choose Each Option
Clear guidance based on your specific situation and needs.
Choose Qwen3.8-Max when...
- Your workload lives behind an API for the foreseeable future and cost per token is the binding constraint — Qwen is 33 percent cheaper on input and 60 percent on output.
- You need video as an input modality, which Kimi K3 does not accept at all.
- You are building for the Chinese market or already work inside Alibaba Cloud, Qoder or QoderWork.
- You are running exploratory work where vendor-reported benchmark placements are good enough and nobody downstream will ask you to source them.
Choose Kimi K3 when...
- You intend to self-host: K3's weights are the only ones of the two that exist as a downloadable artifact today.
- You want more than one supplier to negotiate against — ten providers serve K3, one serves Qwen3.8-Max.
- You have to defend every figure to a procurement or compliance reviewer, and need a parameter count that can be recomputed rather than quoted.
- Your evaluation horizon is longer than a news cycle and you would rather build on a dated artifact than on an announced release date.
Our Recommendation
Qwen3.8-Max is now the cheaper model and Kimi K3 is still the only deployable one, and which of those matters more depends entirely on whether you ever plan to run the thing yourself. The price case for Qwen is real and no longer speculative. Since going generally available on 3 August 2026 it trades on the open market at $2.00 per million input tokens and $6.00 per million output, against Kimi K3's $3.00 and $15.00 — 33 percent cheaper in, 60 percent cheaper out, with cached input at $0.25 per million. It also takes video as input, which K3 does not, and it carries a 1 million token context window. If you are buying inference through an API and never intend to leave it, Qwen3.8-Max is the better commercial deal today. The deployment case for Kimi K3 hardened while nobody was looking. Moonshot's weights went live on 27 July 2026 at 16:29 UTC and have since been downloaded more than 960,000 times; the repository reports 2,779,931,837,184 parameters, which is not a rounded marketing figure but a summed tensor count anyone can re-derive. Ten independent providers now serve it — Together, Fireworks, BaseTen, Modal, DigitalOcean, Chutes, Morph, Wafer and Moonshot itself — which is why its price is contestable at all. Qwen3.8-Max is served by exactly one company: Alibaba. A single-vendor model has a list price, not a market price. The honest caveat on both sides: Alibaba's headline figures — 2.4 trillion total, 95 billion active, fifth in Text Arena — are vendor-reported and were announced alongside a weights release that has not happened. As of 4 August 2026 there is no Qwen3.8-Max repository on Hugging Face and none on ModelScope, Alibaba's own hub. Kimi K3's licence is likewise not an OSI licence; Hugging Face records it as "other". Neither model is open in the sense a compliance team means. Our recommendation: if the workload lives behind an API and cost per token is the constraint, take Qwen3.8-Max. If you have to self-host, negotiate against more than one supplier, or defend a number to someone who will check it, take Kimi K3 — and re-check the Qwen weights claim before you commit, because it is the kind of statement that inverts without warning.
Frequently Asked Questions
Common questions about this comparison answered.
Need help deciding?
Book a free 30-minute consultation and we'll help you determine the best approach for your specific project.