Local AI · Recipe · 1× DGX Spark

Qwen3.8-Flash-Next 125B UD-Q4_K_XL with llama.cpp on 1× DGX Spark

Qwen3.8-Flash-Next 125B UD-Q4_K_XL with llama.cpp on 1× DGX Spark: 24.2 tok/s according to github.com (dataset as of Sep 29, 2026).

by Weschera

24.2tok/sEveryday

1 request, prose prompt, with Draft

Source: github.com
Intelligence (original model) · Artificial Analysisno independent value

Engine

Engine
llama.cpp
Quantization
Unsloth UD-Q4_K_XL GGUF (103,7 GiB, 4 Shards; PLE ≥4-bit) + MTP-Draft shared-Q8_0 (2,6 GiB)
Model family
Qwen3.8-Flash-Next
Context
262,144
Parameters
125B MoE, 6B aktiv, +51B n-gram-Embedding, +4B MTP (README)
Creator
Weschera
GitHub stars
8
Repo updated
Sep 2, 2026

Measurements

Every sourced value of this recipe, each with its condition and source. Bars relative to the largest value in the group.

Decode · 1 request

  • Everyday (decode)24.2 tok/sDraft-N 4, Acceptance 38 % — MTP bricht auf Prose gerade mal eben (n=8: 14.6, schlechter als off) Source 
  • Peak (decode)47.1 tok/sDraft-N 4 (empfohlen), Acceptance 90 %; server-side timings.predicted_per_second, temp 0, fixed seed; Sweep n=5: 48.3 (88 %) Source 

What you need

Hardware
1 × NVIDIA DGX Spark (GB10)
Engine
llama.cpp
Context
262,144 tokens

Related recipes

← Back to overview

Local AI in your company?

In a workshop we work out which models and which hardware fit your tasks, and build the first agent on your infrastructure.