Local AI · Recipe · 4× DGX Spark

GLM-5.2 753B INT4/INT8 with vLLM on 4× DGX Spark

GLM-5.2 753B INT4/INT8 with vLLM on 4× DGX Spark: 32.5 tok/s according to github.com (dataset as of Sep 29, 2026).

by Weschera

32.5tok/sPeak

1 request, best value reported by source, context 200k

Source: github.com
Intelligence (original model) · Artificial Analysisno independent value

Engine

Engine
vLLM
Quantization
QuantTrio Int4-Int8Mix (W4A16 MoE-Experts / W8A16 dense+MTP, compressed-tensors, data-free; Layer 0 + Attention-Indexer BF16); KV nvfp4_ds_mla (400 B/Token)
Model family
GLM-5.2
Context
316,000
Parameters
753B MoE, alle 256 Experts intakt (unpruned, README)
Creator
Weschera
GitHub stars
1
Repo updated
Aug 20, 2026

Measurements

Every sourced value of this recipe, each with its condition and source. Bars relative to the largest value in the group.

Decode · 1 request

  • Peak (decode)32.5 tok/scontext 200K · Speed-Night 2–3: 512-tok temp-0-Completions, TTFT excluded; peak 36.0; fp8-KV-Profil Source 

What you need

Hardware
4 × NVIDIA DGX Spark (GB10)
Engine
vLLM
Context
316,000 tokens

Notes

What matters before you rebuild it.

  • Experimental
  • Custom kernel required

Sources

Related recipes

← Back to overview

Local AI in your company?

In a workshop we work out which models and which hardware fit your tasks, and build the first agent on your infrastructure.