Local AI · Recipe · 1× DGX Spark

Spark-X2.5-4B BF16 with SGLang on 1× DGX Spark

Spark-X2.5-4B BF16 with SGLang on 1× DGX Spark: 21.8 tok/s according to github.com (dataset as of Sep 29, 2026).

by Weschera

21.8tok/sEveryday

1 request, prose prompt, no speculative decoding

Source: github.com
Intelligence (original model) · Artificial Analysisno independent value

Engine

Engine
SGLang
Quantization
BF16 (unquantisiert)
Model family
Spark-X2.5-4B
Context
131,072
Parameters
4B (Modellname)
Creator
Weschera
GitHub stars
0
Repo updated
Sep 7, 2026

Measurements

Every sourced value of this recipe, each with its condition and source. Bars relative to the largest value in the group.

Decode · 1 request

  • Everyday (decode)21.8 tok/sThinking-off short prose, 2 Runs: 21.77 / 21.89 (411 tokens in 18.99 s im einen) → Mittel als approx; pagoda whole-run 19.62 output tok/s über 58,330 tokens (49 min 33 s) Source 
  • Peak (decode)21.8 tok/sThinking-off short prose, 2 Runs: 21.77 / 21.89 (411 tokens in 18.99 s im einen) → Mittel als approx; pagoda whole-run 19.62 output tok/s über 58,330 tokens (49 min 33 s) Source 

What you need

Hardware
1 × NVIDIA DGX Spark (GB10)
Weights
Engine
SGLang
Context
131,072 tokens

Related recipes

← Back to overview

Local AI in your company?

In a workshop we work out which models and which hardware fit your tasks, and build the first agent on your infrastructure.