1 request, prose prompt, with DSpark (per recipe) Source
Local AI · Recipe · 4× DGX Spark
DeepSeek-V4-Flash-0731 304.6B FP8 with vLLM on 4× DGX Spark
DeepSeek-V4-Flash-0731 304.6B FP8 with vLLM on 4× DGX Spark: 68.8 tok/s according to github.com (dataset as of Sep 29, 2026).
Engine
- Engine
- vLLM
- Quantization
- Stock-Checkpoint (FP8 block-quantisiert, fp8_ds_mla KV)
- Model family
- DeepSeek-V4-Flash-0731
- Context
- 1,048,576
- Parameters
- 304.6B MoE (Summe der Safetensors-Tensoren lt. HF-API; 296.4B INT8-Experten)
- Creator
- fujitsupolycom
- GitHub stars
- 82
- Repo updated
- Sep 27, 2026
Measurements
Every sourced value of this recipe, each with its condition and source. Bars relative to the largest value in the group.
Decode · 1 request
Prefill by context
What you need
- Hardware
- 4 × NVIDIA DGX Spark (GB10)
- Weights
- Main weightsdeepseek-ai/DeepSeek-V4-Flash-0731 MIT
- Engine
- vLLM
- Context
- 1,048,576 tokens
Notes
What matters before you rebuild it.
- Custom kernel required
Sources
Related recipes
1 request, prose prompt, with DSpark (per recipe) Source
1 request, prose prompt, with DSpark Source
1 request, prose prompt, with DSpark (per recipe) Source
- Custom kernel required
1 request, prose prompt, with DFlash2 (per recipe) Source
- Custom kernel required
- Non-commercial
1 request, realistic prompt, with Qwen MTP3 probabilistic mit standard rejection (per recipe) Source
- Custom kernel required
Local AI in your company?
In a workshop we work out which models and which hardware fit your tasks, and build the first agent on your infrastructure.