Local AI

Current Open-Weight LLMs & the Hardware They Run On (September 2026)

Current Open-Weight LLMs & the Hardware They Run On (September 2026)

In July, our guide Best Open-Weight LLMs 2026 covered the best open models of the moment: Qwen3.5, GLM-5.2, gpt-oss, DeepSeek-V4, Gemma 4 and Mistral. Much has happened since. Three new Flash models have reshuffled the self-hosting landscape, and a 2.5-billion-parameter model now beats rivals four times its size. This article summarizes where things stand in September 2026 — and answers the question the guide deliberately left open: what hardware does this actually run on?

The three Flash models of September 2026

ModelArchitectureActive per tokenContextLicenseReleased
GLM-5.3-Flash (Z.ai)320B MoE, hybrid linear+sparse attention~18B1MMITAug 26, 2026
DeepSeek V4.1 Flash552B MoE + Engram n-gram~14B1MMITSep 10, 2026
Qwen3.8-Flash-Next (Alibaba)125B MoE + 51B n-gram, Gated DeltaNet~6B262K native, 1M with YaRNQwen CommunityAug 28, 2026

All three are mixture-of-experts models that activate only a fraction of their parameters per token — that is the only reason they run on consumer hardware at all. The differences are in the details.

GLM-5.3-Flash is the strongest open-weight coder in the Flash class: Terminal-Bench 2.1 at 84.3, DeepSWE v1.1 at 63.4. Z.ai published the MIT weights on August 26 — the same day the previously anonymous model „Ox Alpha" was identified. The catch for self-hosters: the native FP8 weights weigh in around 306 GiB, so a single 128 GB machine is not enough.

DeepSeek V4.1 Flash (open since September 10) is not an iteration of V4 Flash but a rebuild: its own tokenizer, its own inference graph, 189 GiB of Engram tables that stay on SSD in the Q2 package while 152 GiB of main weights run resident. On price it is the bargain: $0.15 input / $0.60 output per million tokens (off-peak).

Qwen3.8-Flash-Next is the preview checkpoint of the Qwen4 architecture: a 125B main model plus a 51B n-gram embedding layer, only ~6B active parameters, with native image and video input. 262K context natively, extendable to 1M via YaRN — but that YaRN configuration is then your deployment problem, not the provider's.

The hardware matrix: what actually fits where

The honest answer to "which hardware?" depends on quantization. The community quants of September 2026 paint a clear picture:

Laptop and desktop (16–64 GB): This is where Qwen3.8-27B (Apache 2.0, ~19 GB at 4-bit, native image and video input) is the number one — it tops several September leaderboards. MiniCPM5-2B (Apache 2.0, ~2.4 GB as Q4 GGUF) is the sleeper hit: 131K context, 69.1 on LiveCodeBench v6 — at 2.52B parameters it beats Qwen3.5-4B by almost 13 points. If you need agentic tool-calling on a budget, this is the affordable local path.

128 GB (Mac Studio, workstation): Qwen3.8-Flash-Next as a 4-bit MLX conversion (112 GB) is the new favorite here — an estimated ~49 tok/s on an M5 Max, because only 6B parameters stream per token. GLM-5.3-Flash fits tightly as a 3-bit quant (120 GB). DeepSeek V4.1 Flash runs here via SSD streaming: the ds4 package (Metal) keeps 152 GiB of main weights resident and reads the Engram tables from SSD.

DGX Spark (1× 128 GB unified): Three community quants bring GLM-5.3-Flash onto a single Spark — the most complete recipe (EXL3 2.05 bpw + DFlash2 drafter) decodes structured output at 64.1 tok/s (5-run median), prose at ~30 tok/s. DeepSeek V4.1 Flash as EXL3 1.59 bpw also runs on a single Spark: 15–17 tok/s fresh, up to ~25 tok/s with prefix cache. One Spark is enough for both — with quality trade-offs documented in the respective KLD-vs-BF16 fidelity receipts.

2–4× DGX Spark: The sweet spot for serious work. GLM-5.3-Flash as EXL3 4 bpw on 2× Spark (TP2): 48–52 tok/s code, 19 tok/s prose, 1M context. DeepSeek V4.1 Flash on 4× Spark (TP4): 40.5 tok/s single-stream with the official 510 GB checkpoint, and community quants (EXL3 3.5 bpw) push that to ~61 tok/s with a 3.4M-token KV pool at 1M context.

Multi-GPU node (8× H200 etc.): From here up, all three run at native FP8/BF16 quality. GLM-5.3-Flash realistically needs an 8-GPU node as its floor; DeepSeek V4.1 Flash wants 2× H200 minimum; Qwen3.8-Flash-Next is the most flexible with its smaller footprint.

What changed since July — in three sentences

First: the Flash class (18B active and below) has reached the quality threshold where self-hosting for agent workloads becomes serious — GLM-5.3-Flash ranks at agentic coding ahead of much of what only API models could do before. Second: n-gram Engram tables (DeepSeek, Qwen) have established the new design pattern — huge lookup tables stay on SSD, only the hot weights are resident. Third: the licenses have diverged — MIT (GLM-5.3-Flash, DeepSeek), Apache 2.0 (Qwen3.8-27B, MiniCPM5-2B, gpt-oss), Qwen Community (Flash-Next) and a custom license on the GLM-5.3 flagship. Check before production deployments.

Frequently Asked Questions

Which model is the best entry point into local LLM hosting in September 2026? For most people: Qwen3.8-27B. Apache 2.0, ~19 GB at 4-bit, runs in Ollama, LM Studio, llama.cpp and vLLM, native image and video input, and it tops the open-source leaderboards of its class. There is no good reason to start more expensively or less conveniently.

Do I absolutely need 8 GPUs for GLM-5.3-Flash? No — but for full quality, yes. The native FP8 weights (~306 GiB) realistically require an 8-GPU node. On a 128 GB Mac or a DGX Spark, community quants (3-bit or EXL3 2.05 bpw) run with documented fidelity measurements against BF16. The quality loss is measurable but acceptable for chat and agent workloads.

What is the difference between Qwen3.8-Flash (API) and Qwen3.8-Flash-Next? qwen3.8-flash is the API service on QwenCloud with 1M context out of the box. Qwen3.8-Flash-Next is the open checkpoint behind it — with 262K native context. If you want YaRN at 1M you configure it yourself; if you want it easy, use the API.

Does DeepSeek V4.1 Flash run on a single 128 GB device? Yes, with compromises. On the Mac via ds4 with SSD streaming of the Engram tables (Q2: 152 GiB resident). On the DGX Spark as EXL3 1.59 bpw at 15–25 tok/s. Neither compares to a 4× Spark setup (TP4, 40–61 tok/s) — but it runs.

Which license is the most straightforward for commercial use? Apache 2.0 (Qwen3.8-27B, MiniCPM5-2B, gpt-oss) and MIT (GLM-5.3-Flash, DeepSeek V4.1 Flash) are practically unrestricted for commercial use. The Qwen Community license on the Flash-Next checkpoint and the GLM-5.3 custom license on the flagship carry conditions — read them before a production deployment.

Does the July guide remain valid? Yes, as an overview of the established families (Qwen3.5, GLM-5.2, gpt-oss, DeepSeek-V4, Gemma 4). This article adds the Flash generation of August/September 2026. For our approach to running local AI in production, see Context Studios Local AI.

Sources

Relevant for your team? Let's talk for 30 minutes.

We sort out what of this actually works in your company — concrete, no slide marathon.

No commitment · 30 minutes · Proposal within 48 h