Local AI · Recipe · 2× DGX Spark

Qwen3.8-Flash-Next NVFP4 with vLLM on 2× DGX Spark (Weschera)

Qwen3.8-Flash-Next NVFP4 with vLLM on 2× DGX Spark (Weschera): 33 tok/s according to github.com (dataset as of Sep 29, 2026).

by Weschera

33tok/sEveryday

1 request, prose prompt, with MTP

Source: github.com
Intelligence (original model) · Artificial Analysisno independent value

Engine

Engine
vLLM
Quantization
NVFP4 (ModelOpt MIXED_PRECISION: NVFP4 Experts, FP8 dense/PLE/MTP); BF16 KV (Pflicht — QSA lehnt fp8-KV ab)
Model family
Qwen3.8-Flash-Next
Context
131,072
Creator
Weschera
GitHub stars
0
Repo updated
Sep 5, 2026

Measurements

Every sourced value of this recipe, each with its condition and source. Bars relative to the largest value in the group.

Decode · 1 request

  • Everyday (decode)33 tok/sREADME '~33 tok/s (MTP-3)' → approx (~-Marker) Source 
  • Peak (decode)58 tok/sStructured JSON '~58 tok/s' → approx Source 

What you need

Hardware
2 × NVIDIA DGX Spark (GB10)
Weights
Engine
vLLM
Context
131,072 tokens

Related recipes

← Back to overview

Local AI in your company?

In a workshop we work out which models and which hardware fit your tasks, and build the first agent on your infrastructure.