Tencent Hy-4 Preview: 770B MoE Open-Weight Under $1 per Million Tokens
TL;DR
Hy-4 Preview is Tencent's new open-weight flagship: 770 billion total parameters, but only 49 billion active parameters per token, 1M token context, Apache 2.0. At $0.834 per million input tokens and $2.501 per million output tokens, it undercuts most Western frontier models in price. In a blind internal evaluation, it was slightly ahead of Kimi K3 and GLM-5.3. For builder stacks, this means: Flash-class for volume, Hy-4 for the mid-tier, Astra for precision jobs.
The Spec Sheet in Numbers
| Field | Hy-4 Preview |
|---|---|
| Release | August 28, 2026 |
| Architecture | Mixture-of-Experts, 78 layers (1 dense, 77 MoE with 256 routed + 1 shared expert) |
| Parameters | 770B total / 49B active per token |
| Context | 1M tokens, up to 64K completion tokens |
| License | Apache 2.0, weights on Hugging Face and ModelScope |
| Attention | Gated DeepSeek Sparse Attention with IndexCache, FP8 variant shipped |
The reported benchmark scores: GPQA Diamond 92.3 and SWE-bench Pro 65.7, plus Terminal-Bench 85.4. These figures come from Tencent's release materials and third-party evaluations — they should be read as "reported," not as independently verified values.
Pricing per Million Tokens
The three models in the current pricing lineup compared directly (Input/Output per million tokens, standard rates):
| Model | Input | Output | Cache Read | Context |
|---|---|---|---|---|
| GLM-5.3 Flash | $0.15 | $0.50 | $0.015 | ~1M |
| Hy-4 Preview | $0.834 | $2.501 | $0.042 | 1M |
| GPT-6 Astra | $10.00 | $50.00 | $1.00 | ~1M |
Hy-4 sits perfectly in the gap: about five times the cost of the Flash class, but one-twelfth to one-twentieth of the Astra rate. Output prices are the main lever here, as agent loops write more output than input.
A concrete calculation example — an agent loop with 40 calls, each ~8,000 input and 600 output tokens:
320K Input + 24K Output, per loop:
GLM-5.3 Flash: 0.32 × 0.15 + 0.024 × 0.50 = $0.06
Hy-4 Preview : 0.32 × 0.834 + 0.024 × 2.501 = $0.33
GPT-6 Astra : 0.32 × 10.00 + 0.024 × 50.00 = $4.40
At 1,000 loops daily, that comes to 60 vs. 330 vs. 4,400 dollars — the same ratio, just scaled.
What "Self-Optimized Training" Means Here
Tencent reports that the model helped optimize its own training pipeline, achieving a +31.8 percent throughput increase compared to the baseline run. We recognize this pattern from OpenAI's research post on September 6: the model isn't just trained through the pipeline; it suggests changes to the pipeline itself.
For builders, the practical value is small but real: you get more training and inference volume per dollar on the same hardware. Anyone comparing prices per million tokens shouldn't just look at the list price, but at the combination of price, active parameters, and context length — the active parameter count (49B) explains why Hy-4 can be cheaper per token than densely built models of a similar total size.
Which Flagship Gets Which Slot
The decision rule for the model pricing lineup, acting as a ladder from cheap to expensive:
- Volume Jobs (classification, extraction, first pass): GLM-5.3 Flash — multimodal, MIT license, 18B active parameters, the price anchor at the lower end.
- Mid-Tier with Long Context (document synthesis, large repos, office work): Hy-4 Preview — 1M context, 49B active parameters, Apache 2.0.
- Precision Jobs (heavy reasoning chains, acceptance tests): GPT-6 Astra — only when the measurable difference actually appears in your own evaluation runs.
Escalation ladder: Always start with the cheaper model and only upgrade if there is a proven loss of quality. This keeps the math verifiable — the three numbers above are on every provider's pricing sheet.
Local or API?
The weights are open: 770B in BF16 takes up about 1.8 TB; the FP8 variant takes about 900 GB (according to Tencent). This is below what four-SSD setups with layer streaming (see the Deltafin/K3 line) can handle, but it only makes sense with high-end RAM configurations.
Practical assessment: For most builders, the API route via Tencent Cloud TokenHub or OpenRouter is the right first step — $0.834/$2.501 per million is a small expense, even in series production. Going local is worthwhile if you already have a 96–128 GB setup (Mac Studio M5 Max or similar) and run constant volume; the ROI calculation is covered in the separate Apple hardware article.
FAQ
Why does the model have the "Preview" tag? Tencent released Hy-4 Preview as an early version and internally calls it the largest generational leap they have measured. With preview releases, you should expect rate and performance adjustments. Therefore: verify prices on provider pricing sheets rather than relying on this article. This is fine for prototypes, but for contractually fixed SLAs, you should wait for the final version.
Is Hy-4 worth it over GLM-5.3 Flash, or is Flash just cheaper? Both are true; it depends on the job. Flash is the price floor for volume tasks with 18B active parameters and multimodal input. Hy-4 costs about five times as much per token but offers higher single-token quality and, according to Tencent, a stronger showing on long engineering tasks — making it the right choice for the mid-tier between Flash and Astra.
What does Apache 2.0 mean specifically for commercial use? Apache 2.0 allows use, modification, and distribution in closed products without royalties, alongside a patent protection clause. For builders, this means the model can be embedded in SaaS products without fearing licensing chains. The same applies to fine-tuning and creating custom variants — the weights are available on Hugging Face and ModelScope.
How should the benchmark numbers be interpreted? GPQA Diamond 92.3 and SWE-bench Pro 65.7 come from Tencent's own materials; Terminal-Bench 85.4 and DeepSWE 64.3 are from third-party evaluations. The blind internal evaluation against Kimi K3 and GLM-5.3 was only published by Tencent itself — this is an indicator, not independent proof. The easiest figures to verify are the price-per-token numbers, as they are publicly listed on provider websites.
What impact does the self-optimized training pipeline have on the end user? The direct benefit is a cheaper token price: increased training throughput on the same hardware lowers the cost baseline. The second effect is model quality per active parameter, as 49B active parameters out of 770B total keep inference costs per token low. Both effects add up in the pricing sheet — which is why Hy-4 sits between Flash and Astra in this calculation.
Sources
- developersdigest.tech — Hy-4 Preview: 770B Open MoE, Pricing and Benchmarks
- progressiverobot.com — Release Context and Pricing Table
- aiweekly.co — Architectural Details (Layers, Experts), GPQA/SWE Scores
- mindstudio.ai — Blind Evaluation, Gated DSA/iHC, FP8
- testingcatalog.com — Pricing List, Access Methods (TokenHub, OpenRouter)
- openrouter.ai — GLM-5.3 Flash Pricing and Context
- openrouter.ai — GPT-6 Astra Pricing
- finout.io — Astra Cache Rates Compared
- startupfortune.com — Market Positioning
- huggingface.co/Tencent-Hunyuan — Official Weights and Model Cards