1 request, best value reported by source, context 64k Source
Local AI · Recipe · 1× DGX Spark
Qwen3.6-35B NVFP4 with vLLM on 1× DGX Spark
by Weschera
No speed value is documented for this recipe, so this page is not indexed.
Engine
- Engine
- vLLM
- Quantization
- NVFP4 (RedHatAI) ; BF16/auto KV
- Model family
- Qwen3.6-35B
- Context
- 262,144
- Creator
- Weschera
- GitHub stars
- 0
- Repo updated
- Aug 9, 2026
Measurements
Every sourced value of this recipe, each with its condition and source. Bars relative to the largest value in the group.
No sourced measurements for this recipe.
What you need
- Hardware
- 1 × NVIDIA DGX Spark (GB10)
- Weights
- Main weightsRedHatAI/Qwen3.6-35B-A3B-NVFP4 Apache-2.0
- Draft modelRedHatAI/Qwen3.6-35B-A3B-speculator.dspark Apache-2.0
- Engine
- vLLM
- Context
- 262,144 tokens
Related recipes
1 request, prose prompt, with MTP (per recipe) Source
- Custom kernel required
- License unclear
1 request, prose prompt, with MTP (per recipe) Source
- Access on request
- License unclear
1 request, prose prompt, with MTP Source
1 request, prose prompt, prompt 50 tokens, with DSpark (per recipe) Source
- License unclear
1 request, prose prompt, with Drafting aus dem Prompt (Edit-Setup, lt. README 'drafts its guesses from your prompt') (per recipe) Source
- License unclear
Local AI in your company?
In a workshop we work out which models and which hardware fit your tasks, and build the first agent on your infrastructure.