Local AI · Recipe · 4× DGX Spark

GLM-5.3-Flash 320B NVFP4 with vLLM on 4× DGX Spark

GLM-5.3-Flash 320B NVFP4 with vLLM on 4× DGX Spark: 45 tok/s according to github.com (dataset as of Sep 29, 2026).

by MiaAI-Lab

45tok/sPeak

1 request, best value reported by source

Source: github.com
Intelligence (original model) · Artificial Analysis42with thinking

Engine

Engine
vLLM
Quantization
NVFP4
Model family
GLM-5.3-Flash
Context
262,144
Parameters
320B-A18B
Creator
MiaAI-Lab
GitHub stars
9
Repo updated
Sep 25, 2026

Measurements

Every sourced value of this recipe, each with its condition and source. Bars relative to the largest value in the group.

Decode · 1 request

  • Peak (decode)45 tok/s'~45 tok/s typical' auf realem Agent-Traffic; Wert ungefähr („~45 tok/s typical“), Methodik nicht tabelliert Source 

What you need

Hardware
4 × NVIDIA DGX Spark (GB10)
Weights
Engine
vLLM
Context
262,144 tokens

Related recipes

2× DGX Spark
97.6tok/sEveryday

HumanEval 97.6% / GSM8K 98.0% with FP8 KV and 4-bit dense Source 

Intelligence42with thinking
  • ≤ 3 bit: quantization may cost quality
  • Custom kernel required
  • Non-commercial
  • License unclear

EXL3/TR3 4bpw routed experts + 4-bit dense, BF16 elsewhere, FP8 KVTensorFold v0.5.0 (52 patches: DFlash2+copy drafts, 4-bit dense, FP8 KV, RoCE one-shot all-gather, shared KV pool, vision, tool calling)Weights Repo Updated Sep 30, 2026

Details

← Back to overview

Local AI in your company?

In a workshop we work out which models and which hardware fit your tasks, and build the first agent on your infrastructure.