Local AI · Recipe · 2× DGX Spark

GLM-5.3-Flash 320B NVFP4 with vLLM on 2× DGX Spark (vision)

GLM-5.3-Flash 320B NVFP4 with vLLM on 2× DGX Spark (vision): 21.2 tok/s according to github.com (dataset as of Sep 29, 2026).

by sfxnz

21.2tok/sEveryday

1 request, prose prompt, with DFlash2 (per recipe)

Source: github.com
Intelligence (original model) · Artificial Analysis42with thinking

Engine

Engine
vLLM
Quantization
NVFP4 weight-only auf routed experts (~181 GiB), fp8_e4m3 KV (4.14-GiB-Pin), Marlin MoE, DFlash2-7
Model family
GLM-5.3-Flash
Context
327,680
Parameters
320B total / 18B active (README); HF-API safetensors total 165,496,249,182
Creator
sfxnz
GitHub stars
17
Repo updated
Sep 15, 2026

Measurements

Every sourced value of this recipe, each with its condition and source. Bars relative to the largest value in the group.

Decode · 1 request

  • Everyday (decode)21.2 tok/sBench_decode.py, streamed greedy, thinking off, 200 Completion-Tokens, 3-run-Median; Rebench 2026-09-02: Median 21.179 (3 Läufe); TTFT p50 0.33 s Source 
  • Peak (decode)67.6 tok/sStructured (Count 1→200) aus evidence/rebench bench.txt SUMMARY; acceptance_len 7.844 Source 

Prefill by context

  • 8k (~10,271 Prompt-Tokens)1,425 tok/sUnique-salt 8k-word needle, TTFT 7.2 s; Repeat desselben Prompts: 2600 tok/s warm (prefix cache) Source 

What you need

Hardware
2 × NVIDIA DGX Spark (GB10)
Weights
Engine
vLLM
Context
327,680 tokens

Related recipes

2× DGX Spark
97.6tok/sEveryday

HumanEval 97.6% / GSM8K 98.0% with FP8 KV and 4-bit dense Source 

Intelligence42with thinking
  • ≤ 3 bit: quantization may cost quality
  • Custom kernel required
  • Non-commercial
  • License unclear

EXL3/TR3 4bpw routed experts + 4-bit dense, BF16 elsewhere, FP8 KVTensorFold v0.5.0 (52 patches: DFlash2+copy drafts, 4-bit dense, FP8 KV, RoCE one-shot all-gather, shared KV pool, vision, tool calling)Weights Repo Updated Sep 30, 2026

Details

← Back to overview

Local AI in your company?

In a workshop we work out which models and which hardware fit your tasks, and build the first agent on your infrastructure.