Local AI · Recipe · 1× DGX Spark

Qwen3.6-35B NVFP4 with SGLang on 1× DGX Spark

Qwen3.6-35B NVFP4 with SGLang on 1× DGX Spark: 92.2 tok/s according to github.com (dataset as of Sep 29, 2026).

by Weschera

92.2tok/sPeak

1 request, best value reported by source, context 64k

Source: github.com
Intelligence (original model) · Artificial Analysisno independent value

Engine

Engine
SGLang
Quantization
NVFP4 (qwen36-35b-nvfp4 Checkpoint)
Model family
Qwen3.6-35B
Context
262,144
Creator
Weschera
GitHub stars
1
Repo updated
Jul 12, 2026

Measurements

Every sourced value of this recipe, each with its condition and source. Bars relative to the largest value in the group.

Decode · 1 request

  • Peak (decode)92.2 tok/scontext 64K · Single stream (cache-off); vLLM 93.7 — innerhalb noise; 256k: 65.4/70.9 Source 

What you need

Hardware
1 × NVIDIA DGX Spark (GB10)
Weights
Engine
SGLang
Context
262,144 tokens

Related recipes

← Back to overview

Local AI in your company?

In a workshop we work out which models and which hardware fit your tasks, and build the first agent on your infrastructure.