1 request, prose prompt, with DSpark (per recipe) Source
Local AI · Recipe · 3× DGX Spark
DeepSeek-V4.1-Flash 763B EXL3 3.5 bpw with vLLM on 3× DGX Spark
DeepSeek-V4.1-Flash 763B EXL3 3.5 bpw with vLLM on 3× DGX Spark: 46 tok/s according to github.com (dataset as of Sep 29, 2026).
by tonyd2wild
Engine
- Engine
- vLLM
- Quantization
- EXL3 3.5 bpw routed experts (Pollard) + FP8 Rest
- Model family
- DeepSeek-V4.1-Flash
- Context
- 300,000
- Parameters
- ~763B MoE
- Creator
- tonyd2wild
- GitHub stars
- 86
- Repo updated
- Sep 19, 2026
Measurements
Every sourced value of this recipe, each with its condition and source. Bars relative to the largest value in the group.
Decode · 1 request
Prefill by context
What you need
- Hardware
- 3 × NVIDIA DGX Spark (GB10)
- Weights
- Main weightsbot-lab-21/DeepSeek-V4.1-Flash-EXL3-3.5bpw-Pollard MIT
- Basedeepseek-ai/DeepSeek-V4.1-Flash MIT
- Engine
- vLLM
- Context
- 300,000 tokens
Notes
What matters before you rebuild it.
- Experimental
- Custom kernel required
Sources
Related recipes
1 request, prose prompt, context 32788, with DSPARK Source
1 request, prose prompt, with DSpark (per recipe) Source
1 request, prose prompt, with MTP/EAGLE (spekulativ, ~59 ms/step) (per recipe) Source
- Custom kernel required
1 request, prose prompt, with DSpark (per recipe) Source
- ≤ 3 bit: quantization may cost quality
- Experimental
- Custom kernel required
1 request, prose prompt, with DSpark (per recipe) Source
- ≤ 3 bit: quantization may cost quality
- Custom kernel required
Local AI in your company?
In a workshop we work out which models and which hardware fit your tasks, and build the first agent on your infrastructure.