Local AI · Recipe · 3× DGX Spark

GLM-5.2 753B NVFP4 with vLLM on 3× DGX Spark

GLM-5.2 753B NVFP4 with vLLM on 3× DGX Spark: 13.6 tok/s according to github.com (dataset as of Sep 29, 2026).

by MiaAI-Lab

13.6tok/sMixed

1 request, mixed prompt set, with MTP (per recipe)

Source: github.com
Intelligence (original model) · Artificial Analysisno independent value

Engine

Engine
vLLM custom :k12l1-vision
Quantization
NVFP4+AQLM hybrid, nvfp4_ds_mla KV
Model family
GLM-5.2
Context
380,928
Parameters
~753B
Creator
MiaAI-Lab
GitHub stars
73
Repo updated
Sep 22, 2026

Measurements

Every sourced value of this recipe, each with its condition and source. Bars relative to the largest value in the group.

Decode · 1 request

  • Mixed (decode)13.6 tok/scontext kurz (Pfad A, nvfp4_ds_mla KV, Vision-Config) · „mixed“ = Prosa/Chat laut README; boot-matched, warm≥5, r1 13.7 / r2 13.6 (niedrigerer Wert); thinking off; README-Kopf: „~13.6–19“ Source 
  • Peak (decode)20.9 tok/scontext kurz (Pfad A, Vision-Config) · „structured“ = coding/tight format laut README; r1 21.0 / r2 20.9; README-Kopf „~21“ Source 

What you need

Hardware
3 × NVIDIA DGX Spark (GB10)
Engine
vLLM custom :k12l1-vision
Context
380,928 tokens

Related recipes

← Back to overview

Local AI in your company?

In a workshop we work out which models and which hardware fit your tasks, and build the first agent on your infrastructure.