1 request, JSON prompt Source
Local AI · Recipe · 2× DGX Spark
MiMo-V2.5 310B NVFP4 with vLLM on 2× DGX Spark (160k context, with MTP, tonyd2wild)
by tonyd2wild
Engine
- Engine
- vLLM 0.21.1rc1.dev85+gd87ee1893
- Quantization
- NVFP4 experts + MXFP8 dense, fp8_e4m3 KV
- Model family
- MiMo-V2.5
- Context
- 163,840
- Parameters
- 310B MoE / ~15B active (lt. README)
- Creator
- tonyd2wild
- GitHub stars
- 9
- Repo updated
- Jul 10, 2026
Measurements
Every sourced value of this recipe, each with its condition and source. Bars relative to the largest value in the group.
No sourced measurements for this recipe.
What you need
- Hardware
- 2 × NVIDIA DGX Spark (GB10)
- Weights
- Main weightslukealonso/MiMo-V2.5-NVFP4 MIT
- Engine
- vLLM 0.21.1rc1.dev85+gd87ee1893
- Context
- 163,840 tokens
Notes
What matters before you rebuild it.
- Experimental
- Custom kernel required
- Superseded
Sources
Related recipes
1 request, best value reported by source Source
- Superseded
1 request, mixed prompt set Source
- Experimental
- Custom kernel required
1 request, JSON prompt Source
- Custom kernel required
1 request, best value reported by source Source
- Custom kernel required
1 request, best value reported by source Source
Local AI in your company?
In a workshop we work out which models and which hardware fit your tasks, and build the first agent on your infrastructure.