Local AI · Recipe · 1× DGX Spark

GLM-5.3-Flash 320B UD-Q2_K_XL with llama.cpp on 1× DGX Spark

GLM-5.3-Flash 320B UD-Q2_K_XL with llama.cpp on 1× DGX Spark: 17.9 tok/s according to github.com (dataset as of Sep 29, 2026).

by Weschera

17.9tok/sEveryday

1 request, prose prompt, context 32k, no speculative decoding

Source: github.com
Intelligence (original model) · Artificial Analysisno independent value

≤ 3 bit: quantization may cost quality

Engine

Engine
llama.cpp
Quantization
Unsloth UD-Q2_K_XL (109 GB, ~2,4-bit Experts dynamisch; Attention/Router höher)
Model family
GLM-5.3-Flash
Context
262,144
Parameters
320B / A18B MoE (README)
Creator
Weschera
GitHub stars
12
Repo updated
Aug 27, 2026

Measurements

Every sourced value of this recipe, each with its condition and source. Bars relative to the largest value in the group.

Decode · 1 request

  • Everyday (decode)17.9 tok/scontext 32K · Freeform prose, Acceptance ~50 %; ohne MTP 17.5 Source 
  • Peak (decode)33.8 tok/scontext 32K · Structured (counting), Acceptance 96 %; temp 0, warm Source 

Prefill by context

  • 2.2K291 tok/sREADME '~291 tok/s' → approx Source 

What you need

Hardware
1 × NVIDIA DGX Spark (GB10)
Engine
llama.cpp
Context
262,144 tokens

Notes

What matters before you rebuild it.

  • ≤ 3 bit: quantization may cost qualityAt 3 bit and below the model may answer noticeably worse than the original. The intelligence number refers to the original.

Sources

Related recipes

2× DGX Spark
97.6tok/sEveryday

HumanEval 97.6% / GSM8K 98.0% with FP8 KV and 4-bit dense Source 

Intelligence42with thinking
  • ≤ 3 bit: quantization may cost quality
  • Custom kernel required
  • Non-commercial
  • License unclear

EXL3/TR3 4bpw routed experts + 4-bit dense, BF16 elsewhere, FP8 KVTensorFold v0.5.0 (52 patches: DFlash2+copy drafts, 4-bit dense, FP8 KV, RoCE one-shot all-gather, shared KV pool, vision, tool calling)Weights Repo Updated Sep 30, 2026

Details

← Back to overview

Local AI in your company?

In a workshop we work out which models and which hardware fit your tasks, and build the first agent on your infrastructure.