Local AI · Recipe · 2× DGX Spark

Qwen3.8-Flash-Next NVFP4 with SGLang on 2× DGX Spark

Qwen3.8-Flash-Next NVFP4 with SGLang on 2× DGX Spark: 41.7 tok/s according to github.com (dataset as of Sep 29, 2026).

by Weschera

41.7tok/sMixed

1 request, mixed prompt set, with NEXTN (MTP) (per recipe)

Source: github.com
Intelligence (original model) · Artificial Analysisno independent value

Engine

Engine
SGLang
Quantization
NVFP4 (modelopt_fp4, RadixArk-Checkpoint); BF16 SSM; NEXTN-Draft in-checkpoint
Model family
Qwen3.8-Flash-Next
Context
262,144
Creator
Weschera
GitHub stars
1
Repo updated
Aug 27, 2026

Measurements

Every sourced value of this recipe, each with its condition and source. Bars relative to the largest value in the group.

Decode · 1 request

  • Mixed (decode)41.7 tok/sV2-Recipe (SM121-Gate-Bypass + Dispatch-Kernel), temp 0, 3-Run-Median; gemessen bei mem-fraction 0.70 — nach Korrektur auf 0.80 wahrscheinlich konservativ Source 
  • Peak (decode)49.8 tok/sThinking, 6000-token-Generierungen (49.8–50.5 im Text); faster als non-thinking Source 

What you need

Hardware
2 × NVIDIA DGX Spark (GB10)
Engine
SGLang
Context
262,144 tokens

Notes

What matters before you rebuild it.

  • Custom kernel required
  • License unclear

Sources

Related recipes

← Back to overview

Local AI in your company?

In a workshop we work out which models and which hardware fit your tasks, and build the first agent on your infrastructure.