---
type: "BlogPosting"
title: "tok/s Is Not tok/s — How to Read Local LLM Benchmarks"
description: "Decode vs. prefill, mean wall TPS, quant fidelity: 4 traps in benchmark numbers — and the 5-minute self-check. With real measurements from 300 community recipes."
resource: "https://www.contextstudios.ai/blog/tok-s-is-not-tok-s-how-to-read-local-llm-benchmarks"
language: "en"
tags: ["Local AI", "Self-Hosting", "LLM", "Benchmark", "Hardware"]
generated:
  by: "process:contextstudios-md/1"
  at: "2026-10-05T08:50:21.478Z"
status: "stable"
---

# tok/s Is Not tok/s — How to Read Local LLM Benchmarks

Published: 2026-10-01
Tags: Local AI, Self-Hosting, LLM, Benchmark, Hardware

![tok/s Is Not tok/s — How to Read Local LLM Benchmarks](https://wary-platypus-754.convex.cloud/api/storage/2803a13c-917d-4aee-b378-353fd982fa86)

"That model runs at 250 tokens per second!" — you'll find that number in almost every benchmark post in the local AI scene. And it is almost never the whole truth. We analyzed the 300 community recipes in our local AI database and show which numbers you can trust, which ones need context — and how to measure cleanly yourself in five minutes.

## Trap 1: Decode, prefill, and the "mean wall TPS"

The biggest mistake happens at reading time: decode tok/s (how fast the text appears while being written) and prefill tok/s (how fast the prompt is processed) are different worlds. One recipe on an RTX PRO 6000 measured 14,842 tok/s prefill — at 207 tok/s decode. Whoever posts "2,000 tok/s" probably grabbed the prompt-processing figure, which behaves dramatically differently with long contexts or repeated prompts (prefix cache!).

Then there is the "mean wall TPS" trick: one recipe reports 251.8 tok/s as mean wall TPS across multiple streams. That is aggregate throughput — no single user ever sees that speed. Single-stream on the same hardware realistically lands at 40–60 tok/s. Both numbers are correct. Only one of them answers your question.

## Trap 2: The benchmark setup determines the result

Two recipes for the same 35B-A3B class on two RTX 5090s: 249.4 tok/s with NVFP4 quantization, 251.8 tok/s with AutoRound INT4. Sounds identical? The second figure comes from an 800-word essay prompt, the first from short structured queries. The benchmark prompt decides the outcome:

- Short structured prompts (tables, counting): high drafter acceptance, artificially fast numbers
- 800-word essay: real wall-clock speed without the drafter bonus
- Endless "count to 300" tests: pure show numbers that never occur in daily use

A serious recipe therefore separates code, prose, and structured output — each with speculative decoding on and off. Whoever quotes a single number without saying what was measured has usually picked the prettiest one.

## Trap 3: Quantization without a fidelity receipt

2.05 bits per weight sounds like "unusable", 8.49 bits like "obviously better". Reality is more complicated. A 2-bit EXL3 recipe (GLM-5.3-Flash on a single DGX Spark) reaches 18.9 tok/s decode with a documented deviation from the original (KLD metric, top-1 agreement 78.8%). A 4-bit EXL3 recipe on two Sparks decodes at 36.1 tok/s — twice as fast and measurably closer to the original.

The rule: bit width alone doesn't decide, bit width plus documented fidelity does. A recipe without a KLD/PPL figure is a claim, not a measurement. And: 2-bit can be perfectly fine for chat yet measurably break for code agents — which is why serious recipes separate this by use case.

## Trap 4: Vendor numbers vs. standard harnesses

Vendor benchmarks often run on their own scaffolds. Community measurements regularly show a 15–20 point gap between vendor scaffolds and standard harnesses (Terminal-Bench, SWE-bench). That doesn't mean vendors lie — but that their test setup is different (more forgiving). For purchase decisions the certified numbers matter; for real deployment, community recipes do.

## The 5-minute self-check

How to verify a recipe before you adopt it:

1. **Which number, which mode?** Decode or prefill? Single-stream or aggregate? Is it stated?
2. **Which prompt?** Prose essay (realistic) or "count to 300" (drafter show)?
3. **Which quantization, which fidelity?** bpw plus KLD/PPL against the original documented?
4. **Which software?** Engine version stated? A single vLLM update can move 10%.
5. **Measure yourself:** Same prompt, same settings, run it three times, take the median. Your rig, your version, your number.

## Which numbers can you trust?

The most honest numbers in our database come from recipes that separate all three dimensions: use cases (prose/code/structured), operating modes (speculative decoding on/off), and correctness (KLD against the original). Those recipes have the longest notes and the least spectacular numbers. That is exactly the quality marker.

All measurements quoted in this article come from the public recipe database on our [Local AI page](/local-ai) — every recipe links to the original repository with methodology and reproduction instructions.

## Frequently Asked Questions

**Why do two recipes for the same model on the same hardware differ so much?**
Because "same hardware" is rarely the same hardware: power limit (300 W vs. 600 W on an RTX 3090), PCIe lanes, engine version, KV cache type, and the benchmark prompt itself shift results by factors. A serious recipe documents exactly these parameters — which is why they are part of our recipe data fields.

**Is speculative decoding a cheated number?**
No, but it is a different operating mode. Speculative decoding (e.g. with a draft model) delivers faster output at the same quality — but acceptance rate depends on the prompt. Structured output benefits massively, free prose barely. Serious measurements separate the two; "just one number" is a warning sign.

**What do KLD and top-1 agreement mean for quantizations?**
KLD (KL divergence) measures how far the quantized model's output distribution deviates from the original — the lower, the more faithful. Top-1 agreement states how often the quantized and original model pick the same token (78.8% at 2-bit, above 94% for good 4-bit packs). Without that number, a quantization is a gamble with extra steps.

**Are high prefill numbers relevant to me?**
Very much so — if you process long documents or repeated system prompts. Prefill decides the wait until the first token. With recurring prompts the prefix cache joins in: the same prompt prefix isn't recomputed. For short chat sessions decode matters more; for RAG and agent workloads it's often the prefill.

**Why trust community recipes over vendor benchmarks?**
Because community recipes are reproducible: repository, engine version, prompt, hardware configuration — all open. Vendor numbers are marketing statements with an unknown scaffold. Our recipe database links every recipe to its original so you can verify it on your own rig.

## Sources

- [Local AI — Context Studios (recipe database with 300 recipes)](https://www.contextstudios.ai/local-ai)
- [club3090 — Qwen3.6 35B-A3B NVFP4 on 1× RTX 5090 (251.8 tok/s mean wall TPS)](https://github.com/noonghunna/club-3090/blob/master/BENCHMARKS.md)
- [vcruz305 — GLM-5.3-Flash EXL3 K2 on 1× DGX Spark (KLD fidelity documentation)](https://huggingface.co/vcruz305/GLM-5.3-Flash-EXL3-K2)
- [MiaAI-Lab — GLM-5.3-Flash EXL3 4bpw on 2× DGX Sparks](https://github.com/MiaAI-Lab/GLM-5.3-Flash-EXL3-2x-DGX-Sparks)
- [jpezzulli — Qwen3.8-Flash-Next NVFP4 SGLang on 1× RTX PRO 6000 (prefill 14,842 tok/s)](https://www.contextstudios.ai/local-ai)
- [Regolo — DeepSeek V4 Flash vs Qwen3.8-Flash-Next vs GLM-5.3-Flash (vendor scaffold discussion)](https://regolo.ai/deepseek-v4-flash-vs-qwen3-8-flash-next-vs-glm-5-3-flash-the-real-leader-in-quality-to-price-in-2026/)


## Related

- [LLM Development](https://www.contextstudios.ai/llm-development.md)
- [AI Development](https://www.contextstudios.ai/ai-development.md)
- [LLM Integration](https://www.contextstudios.ai/llm-integration.md)
- [AI Consulting](https://www.contextstudios.ai/ai-consulting.md)
