Local AI

One Model, Six Hardware Tiers — What We Measured

One Model, Six Hardware Tiers — What We Measured

How fast does one and the same model run on different hardware? And what do you lose along the way? We put all the Qwen3.8-Flash-Next recipes from our database side by side — from the RTX 3090 to the RTX PRO 6000 node. The result is the most honest possible answer to the question "which hardware do I need": a table of measurements, not a marketing slide.

The model in 30 seconds

Qwen3.8-Flash-Next is the open preview checkpoint toward Qwen4: a 125B main model plus a 51B n-gram embedding layer (partly kept on SSD), only ~6B active parameters per token, native image and video input, 262K context (1M via YaRN). The small active parameter count makes it the ideal test candidate for hardware tiers: it runs on almost anything — but how fast is the difference.

The measurement series: six tiers, one model

All figures: best single-stream decode per recipe, prose/real workload preferred, from the recipe database (creator named):

TierSetupQuanttok/s (typ)Recipe by
RTX 3090 (1×)EXL3 2.5 bpw41.5 GiB~38.6*r0b0tlab
RTX 3090 (2×)GGUF Q4, CPU-MoEUD-Q4_K_XL36.8club3090
RTX 5090 (1×)GGUF Q2, CPU-MoEUD-Q2_K_XL~59.7notwitcheer
DGX Spark (1×)NVFP4 + FP8 KV6.78 bpw48.7MiaAI-Lab
DGX Spark (2×)NVFP4 TP28.49 bpw54.4MiaAI-Lab
Mac Studio M4 MaxoMLX, MTP6oQ4e~83**weschera
RTX PRO 6000 (1×)NVFP4, SGLangModelOpt~207jpezzulli

* peak value. ** peak with MTP speculation; the real prose figure is lower.

What the table really says

1. The jumps are not linear. From the 3090 (36.8) to the 5090 (~60) is a factor of 1.6 — but from the 5090 to the PRO 6000 (207) it's 3.5. The reason: the PRO 6000 recipes run native NVFP4 with full KV throughput, while the consumer cards work with CPU-MoE offloading (experts move into RAM). Same model family, fundamentally different operating modes.

2. Unified memory is its own competition. The Mac and Spark figures aren't directly comparable to the RTX figures: MLX with MTP speculation (weschera, ~83 tok/s peak) wins through speculation, not bandwidth. In the sober comparison, prose without drafter, Spark (48.7) and M5 Max (48.4, antirez/ds4) are practically tied — different ecosystems, identical reality.

3. Quantization decides the tier boundary. On the 3090, Flash-Next only runs at 2.5–3 bpw (41.5 GiB in 24 GB VRAM leaves no alternative). On the Spark it's NVFP4/6.78 bpw, on the PRO 6000 native FP4 class. The same verification benchmark against the BF16 original shows: the higher-bpw tiers gain not just speed but measurably more correctness. More VRAM buys you better bits.

4. Prefill is the secret tier reward. Decode tok/s dominate the discussion, but the prefill figures show the real hierarchy: ~664 (3090) → ~900 (5090) → ~1,492 (Spark/M5 Max) → ~12,073–14,842 (PRO 6000). Whoever feeds long documents or agent contexts doesn't experience the PRO 6000 as "somewhat faster" but as a different class.

Which tier for whom?

  • Occasional chat and experiments: 1× RTX 3090 — 37–39 tok/s is plenty for chatting, and you might already own the card.
  • Serious single-user use: DGX Spark or Mac Studio — ~49 tok/s decode with a fuller quantization feels like "talking fluently."
  • RAG, agents, long contexts: RTX PRO 6000 — the prefill argument decides, not the decode figure.
  • Multi-user: 2× Spark or PRO 6000 with vLLM/SGLang — concurrency needs KV pool, not a single-stream peak.

Every row of this table is a real, reproducible recipe from our Local AI database — with serve flags, original repo, and measurement methodology. Pick your tier, copy the recipe, verify by measuring.

Frequently Asked Questions

Is the upgrade from 3090 to 5090 worth it for this model? For decode: +60% (36.8 → ~60 tok/s with a similar quant). For the experience: noticeable, but not stunning. The big jump is in prefill (664 → 900) and in the fact that the 5090 handles Q4 quants comfortably. Whoever already has a 3090 and only chats: probably not.

Why is the Mac figure (83 tok/s) with MTP so much higher than the Spark figure? MTP (multi-token prediction) is speculation: guessing and verifying several tokens per step. weschera's oMLX implementation uses it aggressively (MTP6). Without speculation, Mac M5 Max and DGX Spark are tied (~48). Speculation is a software feature, not a hardware advantage.

Is 2.5 bpw on the 3090 still the same model? It's the same model with measurably more deviation from the original. Fine for chat, critical for code agents — and exactly that is documented by the recipes' KLD figures. If you want code-agent quality, you need at least the 5090 class (Q4) or the Spark (NVFP4).

What does 2× Spark buy over 1×? With TP2: full NVFP4 quality (8.49 bpw effective), 54.4 tok/s instead of 48.7 — and above all double the KV pool for longer contexts and more concurrency. The step is a quality and capacity jump at once, not just speed.

Is it true that the PRO 6000 is "a different class"? For prefill: unambiguously yes — 12,000–14,800 tok/s vs. 664–1,492 on consumer/unified. That's the difference between "upload document, go get coffee" and "document is already read." Whoever runs agent workloads with long contexts buys for prefill, not decode.

Sources

Relevant for your team? Let's talk for 30 minutes.

We sort out what of this actually works in your company — concrete, no slide marathon.

No commitment · 30 minutes · Proposal within 48 h