In July, our guide Best Open-Weight LLMs 2026 covered the best open models of the moment: Qwen3.5, GLM-5.2, gpt-oss, DeepSeek-V4, Gemma 4 and Mistral. Much has happened since. Three new Flash models have reshuffled the self-hosting landscape, and a 2.5-billion-parameter model now beats rivals four times its size. This article summarizes where things stand in September 2026 — and answers the question the guide deliberately left open: what hardware does this actually run on?
The three Flash models of September 2026
| Model | Architecture | Active per token | Context | License | Released |
|---|---|---|---|---|---|
| GLM-5.3-Flash (Z.ai) | 320B MoE, hybrid linear+sparse attention | ~18B | 1M | MIT | Aug 26, 2026 |
| DeepSeek V4.1 Flash | 552B MoE + Engram n-gram | ~14B | 1M | MIT | Sep 10, 2026 |
| Qwen3.8-Flash-Next (Alibaba) | 125B MoE + 51B n-gram, Gated DeltaNet | ~6B | 262K native, 1M with YaRN | Qwen Community | Aug 28, 2026 |
All three are mixture-of-experts models that activate only a fraction of their parameters per token — that is the only reason they run on consumer hardware at all. The differences are in the details.
GLM-5.3-Flash is the strongest open-weight coder in the Flash class: Terminal-Bench 2.1 at 84.3, DeepSWE v1.1 at 63.4. Z.ai published the MIT weights on August 26 — the same day the previously anonymous model „Ox Alpha" was identified. The catch for self-hosters: the native FP8 weights weigh in around 306 GiB, so a single 128 GB machine is not enough.
DeepSeek V4.1 Flash (open since September 10) is not an iteration of V4 Flash but a rebuild: its own tokenizer, its own inference graph, 189 GiB of Engram tables that stay on SSD in the Q2 package while 152 GiB of main weights run resident. On price it is the bargain: $0.15 input / $0.60 output per million tokens (off-peak).
Qwen3.8-Flash-Next is the preview checkpoint of the Qwen4 architecture: a 125B main model plus a 51B n-gram embedding layer, only ~6B active parameters, with native image and video input. 262K context natively, extendable to 1M via YaRN — but that YaRN configuration is then your deployment problem, not the provider's.
The hardware matrix: what actually fits where
The honest answer to "which hardware?" depends on quantization. The community quants of September 2026 paint a clear picture:
Laptop and desktop (16–64 GB): This is where Qwen3.8-27B (Apache 2.0, ~19 GB at 4-bit, native image and video input) is the number one — it tops several September leaderboards. MiniCPM5-2B (Apache 2.0, ~2.4 GB as Q4 GGUF) is the sleeper hit: 131K context, 69.1 on LiveCodeBench v6 — at 2.52B parameters it beats Qwen3.5-4B by almost 13 points. If you need agentic tool-calling on a budget, this is the affordable local path.
128 GB (Mac Studio, workstation): Qwen3.8-Flash-Next as a 4-bit MLX conversion (112 GB) is the new favorite here — an estimated ~49 tok/s on an M5 Max, because only 6B parameters stream per token. GLM-5.3-Flash fits tightly as a 3-bit quant (120 GB). DeepSeek V4.1 Flash runs here via SSD streaming: the ds4 package (Metal) keeps 152 GiB of main weights resident and reads the Engram tables from SSD.
DGX Spark (1× 128 GB unified): Three community quants bring GLM-5.3-Flash onto a single Spark — the most complete recipe (EXL3 2.05 bpw + DFlash2 drafter) decodes structured output at 64.1 tok/s (5-run median), prose at ~30 tok/s. DeepSeek V4.1 Flash as EXL3 1.59 bpw also runs on a single Spark: 15–17 tok/s fresh, up to ~25 tok/s with prefix cache. One Spark is enough for both — with quality trade-offs documented in the respective KLD-vs-BF16 fidelity receipts.
2–4× DGX Spark: The sweet spot for serious work. GLM-5.3-Flash as EXL3 4 bpw on 2× Spark (TP2): 48–52 tok/s code, 19 tok/s prose, 1M context. DeepSeek V4.1 Flash on 4× Spark (TP4): 40.5 tok/s single-stream with the official 510 GB checkpoint, and community quants (EXL3 3.5 bpw) push that to ~61 tok/s with a 3.4M-token KV pool at 1M context.
Multi-GPU node (8× H200 etc.): From here up, all three run at native FP8/BF16 quality. GLM-5.3-Flash realistically needs an 8-GPU node as its floor; DeepSeek V4.1 Flash wants 2× H200 minimum; Qwen3.8-Flash-Next is the most flexible with its smaller footprint.
What changed since July — in three sentences
First: the Flash class (18B active and below) has reached the quality threshold where self-hosting for agent workloads becomes serious — GLM-5.3-Flash ranks at agentic coding ahead of much of what only API models could do before. Second: n-gram Engram tables (DeepSeek, Qwen) have established the new design pattern — huge lookup tables stay on SSD, only the hot weights are resident. Third: the licenses have diverged — MIT (GLM-5.3-Flash, DeepSeek), Apache 2.0 (Qwen3.8-27B, MiniCPM5-2B, gpt-oss), Qwen Community (Flash-Next) and a custom license on the GLM-5.3 flagship. Check before production deployments.
Frequently Asked Questions
Which model is the best entry point into local LLM hosting in September 2026? For most people: Qwen3.8-27B. Apache 2.0, ~19 GB at 4-bit, runs in Ollama, LM Studio, llama.cpp and vLLM, native image and video input, and it tops the open-source leaderboards of its class. There is no good reason to start more expensively or less conveniently.
Do I absolutely need 8 GPUs for GLM-5.3-Flash? No — but for full quality, yes. The native FP8 weights (~306 GiB) realistically require an 8-GPU node. On a 128 GB Mac or a DGX Spark, community quants (3-bit or EXL3 2.05 bpw) run with documented fidelity measurements against BF16. The quality loss is measurable but acceptable for chat and agent workloads.
What is the difference between Qwen3.8-Flash (API) and Qwen3.8-Flash-Next? qwen3.8-flash is the API service on QwenCloud with 1M context out of the box. Qwen3.8-Flash-Next is the open checkpoint behind it — with 262K native context. If you want YaRN at 1M you configure it yourself; if you want it easy, use the API.
Does DeepSeek V4.1 Flash run on a single 128 GB device? Yes, with compromises. On the Mac via ds4 with SSD streaming of the Engram tables (Q2: 152 GiB resident). On the DGX Spark as EXL3 1.59 bpw at 15–25 tok/s. Neither compares to a 4× Spark setup (TP4, 40–61 tok/s) — but it runs.
Which license is the most straightforward for commercial use? Apache 2.0 (Qwen3.8-27B, MiniCPM5-2B, gpt-oss) and MIT (GLM-5.3-Flash, DeepSeek V4.1 Flash) are practically unrestricted for commercial use. The Qwen Community license on the Flash-Next checkpoint and the GLM-5.3 custom license on the flagship carry conditions — read them before a production deployment.
Does the July guide remain valid? Yes, as an overview of the established families (Qwen3.5, GLM-5.2, gpt-oss, DeepSeek-V4, Gemma 4). This article adds the Flash generation of August/September 2026. For our approach to running local AI in production, see Context Studios Local AI.
Sources
- Yotta Labs: DeepSeek V4 Flash vs GLM 5.3 Flash vs Qwen Flash-Next (2026)
- Regolo: DeepSeek V4 Flash vs Qwen3.8-Flash-Next vs GLM-5.3-Flash — quality-to-price 2026
- LLMCheck: State of Open-Source Local LLMs, September 2026
- vcruz305: GLM-5.3-Flash-EXL3-K2 (Hugging Face, single-Spark measurements)
- Market Intelligence Research: GLM-5.3-Flash on a single DGX Spark — three community quantizations
- NVIDIA Developer Forums: DeepSeek-V4.1-Flash EXL3 2.9 bpw on 2× DGX Sparks
- NVIDIA Developer Forums: DeepSeek-V4.1-Flash is now open-sourced (TP4 measurements)
- vcruz305: DeepSeek-V4.1-Flash-EXL3-DGX-Spark-recipe (single Spark)
- bot-lab-21: DeepSeek-V4.1-Flash-EXL3-3.5bpw-Pollard (4× Spark TP4)
- dwarfstar.sh: ds4 September 2026 — DeepSeek V4.1 and Qwen3.8 Flash Next on Metal
- Dev.to: GLM-5.3-Flash vs Qwen3.8-Flash-Next vs DeepSeek V4 Flash — benchmark tier
- Fernando Nog: DeepSeek-V4.1-Flash vs GLM-5.3-Flash vs Qwen3.8-Flash — API pricing snapshot
- CourionAI: MiniCPM5-2B — a 2.5B model beats bigger rivals
- Local Inference Lab: RTX 6000 Pro daily summaries — DSV4.1 EXL3, Qwen 3.8 Flash Next vLLM configs