Local AI · Recipe · 1× DGX Spark

GLM-5.3-Flash 320B UD-IQ3_XXS with llama.cpp on 1× DGX Spark

GLM-5.3-Flash 320B UD-IQ3_XXS with llama.cpp on 1× DGX Spark: 20.8 tok/s according to github.com (dataset as of Sep 29, 2026).

by Weschera

20.8tok/sMixed

1 request, mixed prompt set, with MTP

Source: github.com
Intelligence (original model) · Artificial Analysisno independent value

≤ 3 bit: quantization may cost quality

Engine

Engine
llama.cpp
Quantization
Unsloth Dynamic UD-IQ3_XXS GGUF (120,4 GB, 4 Shards); KV q8_0
Model family
GLM-5.3-Flash
Context
65,536
Parameters
320B / A18B MoE (README)
Creator
Weschera
GitHub stars
5
Repo updated
Sep 4, 2026

Measurements

Every sourced value of this recipe, each with its condition and source. Bars relative to the largest value in the group.

Decode · 1 request

  • Mixed (decode)20.8 tok/sMedian p10–p90 11.6–25.1, 50 Requests / 26K Tokens, thinking OFF, aus der Benchmark-Run; MTP n=2 Source 
  • Peak (decode)20.8 tok/sMedian p10–p90 11.6–25.1, 50 Requests / 26K Tokens, thinking OFF, aus der Benchmark-Run; MTP n=2 Source 

Prefill by context

  • 7K270 tok/s6.981-token-Prompt in 26 s, README '~270 tok/s' → approx Source 

What you need

Hardware
1 × NVIDIA DGX Spark (GB10)
Engine
llama.cpp
Context
65,536 tokens

Notes

What matters before you rebuild it.

  • ≤ 3 bit: quantization may cost qualityAt 3 bit and below the model may answer noticeably worse than the original. The intelligence number refers to the original.

Sources

Related recipes

2× DGX Spark
97.6tok/sEveryday

HumanEval 97.6% / GSM8K 98.0% with FP8 KV and 4-bit dense Source 

Intelligence42with thinking
  • ≤ 3 bit: quantization may cost quality
  • Custom kernel required
  • Non-commercial
  • License unclear

EXL3/TR3 4bpw routed experts + 4-bit dense, BF16 elsewhere, FP8 KVTensorFold v0.5.0 (52 patches: DFlash2+copy drafts, 4-bit dense, FP8 KV, RoCE one-shot all-gather, shared KV pool, vision, tool calling)Weights Repo Updated Sep 30, 2026

Details

← Back to overview

Local AI in your company?

In a workshop we work out which models and which hardware fit your tasks, and build the first agent on your infrastructure.