no documented speed
Local AI · Recipe · 1× DGX Spark
Qwen3.6-35B NVFP4 with SGLang on 1× DGX Spark
Qwen3.6-35B NVFP4 with SGLang on 1× DGX Spark: 92.2 tok/s according to github.com (dataset as of Sep 29, 2026).
by Weschera
Engine
- Engine
- SGLang
- Quantization
- NVFP4 (qwen36-35b-nvfp4 Checkpoint)
- Model family
- Qwen3.6-35B
- Context
- 262,144
- Creator
- Weschera
- GitHub stars
- 1
- Repo updated
- Jul 12, 2026
Measurements
Every sourced value of this recipe, each with its condition and source. Bars relative to the largest value in the group.
Decode · 1 request
What you need
- Hardware
- 1 × NVIDIA DGX Spark (GB10)
- Weights
- Main weightsRedHatAI/Qwen3.6-35B-A3B-NVFP4 Apache-2.0
- Engine
- SGLang
- Context
- 262,144 tokens
Related recipes
1 request, prose prompt, with MTP (per recipe) Source
- Custom kernel required
- License unclear
1 request, prose prompt, with MTP (per recipe) Source
- Access on request
- License unclear
1 request, prose prompt, with MTP Source
1 request, prose prompt, prompt 50 tokens, with DSpark (per recipe) Source
- License unclear
1 request, prose prompt, with Drafting aus dem Prompt (Edit-Setup, lt. README 'drafts its guesses from your prompt') (per recipe) Source
- License unclear
Local AI in your company?
In a workshop we work out which models and which hardware fit your tasks, and build the first agent on your infrastructure.