---
type: "BlogPosting"
title: "One Model, Six Hardware Tiers — What We Measured"
description: "Qwen3.8-Flash-Next from RTX 3090 to RTX PRO 6000: all measurements side by side. Why the jumps aren't linear and prefill is the secret reward."
resource: "https://www.contextstudios.ai/blog/one-model-six-hardware-tiers-what-we-measured"
language: "en"
tags: ["Local AI", "Self-Hosting", "LLM", "Benchmark", "Hardware"]
generated:
  by: "process:contextstudios-md/1"
  at: "2026-10-05T10:10:31.046Z"
status: "stable"
---

# One Model, Six Hardware Tiers — What We Measured

Published: 2026-10-01
Tags: Local AI, Self-Hosting, LLM, Benchmark, Hardware

![One Model, Six Hardware Tiers — What We Measured](https://wary-platypus-754.convex.cloud/api/storage/89f3f6b2-3feb-40cd-9cae-9cb9f7e591e8)

How fast does one and the same model run on different hardware? And what do you lose along the way? We put all the Qwen3.8-Flash-Next recipes from our database side by side — from the RTX 3090 to the RTX PRO 6000 node. The result is the most honest possible answer to the question "which hardware do I need": a table of measurements, not a marketing slide.

## The model in 30 seconds

Qwen3.8-Flash-Next is the open preview checkpoint toward Qwen4: a 125B main model plus a 51B n-gram embedding layer (partly kept on SSD), only ~6B active parameters per token, native image and video input, 262K context (1M via YaRN). The small active parameter count makes it the ideal test candidate for hardware tiers: it runs on almost anything — but how fast is the difference.

## The measurement series: six tiers, one model

All figures: best single-stream decode per recipe, prose/real workload preferred, from the recipe database (creator named):

| Tier | Setup | Quant | tok/s (typ) | Recipe by |
|---|---|---|---|---|
| RTX 3090 (1×) | EXL3 2.5 bpw | 41.5 GiB | ~38.6* | r0b0tlab |
| RTX 3090 (2×) | GGUF Q4, CPU-MoE | UD-Q4_K_XL | 36.8 | club3090 |
| RTX 5090 (1×) | GGUF Q2, CPU-MoE | UD-Q2_K_XL | ~59.7 | notwitcheer |
| DGX Spark (1×) | NVFP4 + FP8 KV | 6.78 bpw | 48.7 | MiaAI-Lab |
| DGX Spark (2×) | NVFP4 TP2 | 8.49 bpw | 54.4 | MiaAI-Lab |
| Mac Studio M4 Max | oMLX, MTP6 | oQ4e | ~83** | weschera |
| RTX PRO 6000 (1×) | NVFP4, SGLang | ModelOpt | ~207 | jpezzulli |

\* peak value. \** peak with MTP speculation; the real prose figure is lower.

## What the table really says

**1. The jumps are not linear.** From the 3090 (36.8) to the 5090 (~60) is a factor of 1.6 — but from the 5090 to the PRO 6000 (207) it's 3.5. The reason: the PRO 6000 recipes run native NVFP4 with full KV throughput, while the consumer cards work with CPU-MoE offloading (experts move into RAM). Same model family, fundamentally different operating modes.

**2. Unified memory is its own competition.** The Mac and Spark figures aren't directly comparable to the RTX figures: MLX with MTP speculation (weschera, ~83 tok/s peak) wins through speculation, not bandwidth. In the sober comparison, prose without drafter, Spark (48.7) and M5 Max (48.4, antirez/ds4) are practically tied — different ecosystems, identical reality.

**3. Quantization decides the tier boundary.** On the 3090, Flash-Next only runs at 2.5–3 bpw (41.5 GiB in 24 GB VRAM leaves no alternative). On the Spark it's NVFP4/6.78 bpw, on the PRO 6000 native FP4 class. The same verification benchmark against the BF16 original shows: the higher-bpw tiers gain not just speed but measurably more correctness. More VRAM buys you better bits.

**4. Prefill is the secret tier reward.** Decode tok/s dominate the discussion, but the prefill figures show the real hierarchy: ~664 (3090) → ~900 (5090) → ~1,492 (Spark/M5 Max) → ~12,073–14,842 (PRO 6000). Whoever feeds long documents or agent contexts doesn't experience the PRO 6000 as "somewhat faster" but as a different class.

## Which tier for whom?

- **Occasional chat and experiments:** 1× RTX 3090 — 37–39 tok/s is plenty for chatting, and you might already own the card.
- **Serious single-user use:** DGX Spark or Mac Studio — ~49 tok/s decode with a fuller quantization feels like "talking fluently."
- **RAG, agents, long contexts:** RTX PRO 6000 — the prefill argument decides, not the decode figure.
- **Multi-user:** 2× Spark or PRO 6000 with vLLM/SGLang — concurrency needs KV pool, not a single-stream peak.

Every row of this table is a real, reproducible recipe from our [Local AI database](/local-ai) — with serve flags, original repo, and measurement methodology. Pick your tier, copy the recipe, verify by measuring.

## Frequently Asked Questions

**Is the upgrade from 3090 to 5090 worth it for this model?**
For decode: +60% (36.8 → ~60 tok/s with a similar quant). For the experience: noticeable, but not stunning. The big jump is in prefill (664 → 900) and in the fact that the 5090 handles Q4 quants comfortably. Whoever already has a 3090 and only chats: probably not.

**Why is the Mac figure (83 tok/s) with MTP so much higher than the Spark figure?**
MTP (multi-token prediction) is speculation: guessing and verifying several tokens per step. weschera's oMLX implementation uses it aggressively (MTP6). Without speculation, Mac M5 Max and DGX Spark are tied (~48). Speculation is a software feature, not a hardware advantage.

**Is 2.5 bpw on the 3090 still the same model?**
It's the same model with measurably more deviation from the original. Fine for chat, critical for code agents — and exactly that is documented by the recipes' KLD figures. If you want code-agent quality, you need at least the 5090 class (Q4) or the Spark (NVFP4).

**What does 2× Spark buy over 1×?**
With TP2: full NVFP4 quality (8.49 bpw effective), 54.4 tok/s instead of 48.7 — and above all double the KV pool for longer contexts and more concurrency. The step is a quality and capacity jump at once, not just speed.

**Is it true that the PRO 6000 is "a different class"?**
For prefill: unambiguously yes — 12,000–14,800 tok/s vs. 664–1,492 on consumer/unified. That's the difference between "upload document, go get coffee" and "document is already read." Whoever runs agent workloads with long contexts buys for prefill, not decode.

## Sources

- [Local AI — Context Studios (recipe database, Qwen3.8-Flash-Next recipes)](https://www.contextstudios.ai/local-ai)
- [MiaAI-Lab — Qwen3.8-Flash-Next NVFP4 on 1×/2× DGX Spark (48.7/54.4 tok/s)](https://github.com/MiaAI-Lab)
- [vcruz305 — Qwen3.8-Flash-Next EXL3 3.05bpw on 1× Spark (53.0 tok/s)](https://www.contextstudios.ai/local-ai)
- [weschera — Qwen3.8-Flash-Next oMLX MTP6 on Mac Studio M4 Max](https://www.contextstudios.ai/local-ai)
- [antirez/ds4 — Qwen3.8-Flash-Next Q2/Q4 on M5 Max (SSD streaming of the n-gram tables)](https://dwarfstar.sh/blog/ds4-september-v41-qwen38-and-a-bigger-runtime)
- [jpezzulli — Qwen3.8-Flash-Next NVFP4 SGLang on 1× RTX PRO 6000 (207 tok/s, prefill 14,842)](https://www.contextstudios.ai/local-ai)
- [club3090 — Qwen3.8-Flash-Next CPU-MoE on 2× RTX 3090 (36.8 tok/s)](https://github.com/noonghunna/club-3090/blob/master/BENCHMARKS.md)


## Related

- [LLM Development](https://www.contextstudios.ai/llm-development.md)
- [AI Development](https://www.contextstudios.ai/ai-development.md)
- [LLM Integration](https://www.contextstudios.ai/llm-integration.md)
- [AI Consulting](https://www.contextstudios.ai/ai-consulting.md)
