1 request, prose prompt, with DSpark (per recipe) Source
Local AI · Recipe · 2× DGX Spark
DeepSeek-V4.1-Flash EXL3 2.0 bpw with vLLM on 2× DGX Spark
DeepSeek-V4.1-Flash EXL3 2.0 bpw with vLLM on 2× DGX Spark: 42.6 tok/s according to github.com (dataset as of Sep 29, 2026).
by sfxnz
≤ 3 bit: quantization may cost quality
Engine
- Engine
- vLLM
- Quantization
- EXL3 2.0 bpw MCG (routed experts, tail-biting Viterbi re-encoded, scale refit), lm_head MXFP8, KV fp8 (8-GiB-Pin), DSpark-3
- Model family
- DeepSeek-V4.1-Flash
- Context
- 1,048,576
- Parameters
- Basis deepseek-ai/DeepSeek-V4.1-Flash: 763,205,315,794 Parameter (HF-API safetensors total); EXL3-2.0bpw-Pack ~334 GB auf der Platte
- Creator
- sfxnz
- GitHub stars
- 23
- Repo updated
- Sep 27, 2026
Measurements
Every sourced value of this recipe, each with its condition and source. Bars relative to the largest value in the group.
Decode · 1 request
Prefill by context
What you need
- Hardware
- 2 × NVIDIA DGX Spark (GB10)
- Weights
- Main weightssfxnz/DeepSeek-V4.1-Flash-EXL3 MIT
- Basedeepseek-ai/DeepSeek-V4.1-Flash MIT
- Engine
- vLLM
- Context
- 1,048,576 tokens
Notes
What matters before you rebuild it.
- ≤ 3 bit: quantization may cost qualityAt 3 bit and below the model may answer noticeably worse than the original. The intelligence number refers to the original.
- Experimental
- Custom kernel required
Sources
Related recipes
1 request, prose prompt, context 32788, with DSPARK Source
1 request, prose prompt, with DSpark (per recipe) Source
1 request, prose prompt, with MTP/EAGLE (spekulativ, ~59 ms/step) (per recipe) Source
- Custom kernel required
1 request, prose prompt, with DSpark (per recipe) Source
- ≤ 3 bit: quantization may cost quality
- Custom kernel required
1 request, prose prompt, with DSpark (per recipe) Source
- ≤ 3 bit: quantization may cost quality
- Experimental
- Custom kernel required
Local AI in your company?
In a workshop we work out which models and which hardware fit your tasks, and build the first agent on your infrastructure.