Local AI · Recipe · 2× DGX Spark

Qwen3.8-Flash-Next NVFP4 with vLLM on 2× DGX Spark (vision)

Qwen3.8-Flash-Next NVFP4 with vLLM on 2× DGX Spark (vision): 43.2 tok/s according to github.com (dataset as of Sep 29, 2026).

by sfxnz

43.2tok/sEveryday

1 request, prose prompt, with MTP (per recipe)

Source: github.com
Intelligence (original model) · Artificial Analysis40with thinking

Engine

Engine
vLLM
Quantization
ModelOpt NVFP4 W4A4 routed experts, PLE-n-gram-Tabelle per-Tensor FP8, MTP-Experten FP8_BLOCK_SCALES (gs 128), KV auto, Marlin/FlashInfer via moe-backend auto
Model family
Qwen3.8-Flash-Next
Context
262,144
Parameters
HF-API safetensors total 119,602,003,859
Creator
sfxnz
GitHub stars
0
Repo updated
Sep 15, 2026

Measurements

Every sourced value of this recipe, each with its condition and source. Bars relative to the largest value in the group.

Decode · 1 request

  • Everyday (decode)43.2 tok/sBench_decode.py, streamed greedy, thinking off, 200 Completion-Tokens, 3-run-Median; evidence opt-c4-native-ctx-262144: 43.16; TTFT p50 0.17 s Source 
  • Peak (decode)55.5 tok/sStructured (Count 1→200); evidence: 55.47; TTFT p50 0.15 s Source 

What you need

Hardware
2 × NVIDIA DGX Spark (GB10)
Weights
Engine
vLLM
Context
262,144 tokens

Related recipes

← Back to overview

Local AI in your company?

In a workshop we work out which models and which hardware fit your tasks, and build the first agent on your infrastructure.