Local AI · Recipe · 1× DGX Spark

Qwen3.8-27B NVFP4 with SGLang on 1× DGX Spark (Weschera)

Qwen3.8-27B NVFP4 with SGLang on 1× DGX Spark (Weschera): 25.2 tok/s according to github.com (dataset as of Sep 29, 2026).

by Weschera

25.2tok/sMixed

1 request, mixed prompt set, with DFlash2 (per recipe)

Source: github.com
Intelligence (original model) · Artificial Analysisno independent value

Engine

Engine
SGLang
Quantization
NVFP4 (RadixArk) + BF16-Draft; quantized lm_head-Adapter (TP=1, 256-token selector chunks, FP32 top-k — r0b0tlab-Ansatz)
Model family
Qwen3.8-27B
Context
262,144
Parameters
27B (dicht, hybrid Gated-DeltaNet VLM lt. Schwester-Repo)
Creator
Weschera
GitHub stars
29
Repo updated
Aug 21, 2026

Measurements

Every sourced value of this recipe, each with its condition and source. Bars relative to the largest value in the group.

Decode · 1 request

  • Mixed (decode)25.2 tok/scontext 1024 in / 256 out · Ladder C1, Median aus 3 Source 
  • Peak (decode)42 tok/scontext 512 in / 2048 out · Dedicated C1, Median aus 5 Runs, gate-qualified boot (Boot-Lottery-Gate MIN_TPS=39); synthetic random tokens ~ worst case für den Drafter; Block-8-Vorgänger 34.08, r0b0tlab 28.38 Source 

What you need

Hardware
1 × NVIDIA DGX Spark (GB10)
Weights
Engine
SGLang
Context
262,144 tokens

Notes

What matters before you rebuild it.

  • Custom kernel required

Sources

Related recipes

← Back to overview

Local AI in your company?

In a workshop we work out which models and which hardware fit your tasks, and build the first agent on your infrastructure.