1 request, prose prompt, with MTP (per recipe) Source
Local AI · Recipe · 1× DGX Spark
Gemma 4 26B-A4B NVFP4 with vLLM on 1× DGX Spark
by MiaAI-Lab
Engine
- Engine
- vLLM vllm/vllm-openai:nightly (ungepinnt)
- Quantization
- NVFP4 weights + fp8 KV
- Model family
- Gemma 4
- Context
- 262,144
- Parameters
- 26B MoE / 4B active
- Creator
- MiaAI-Lab
- GitHub stars
- 10
- Repo updated
- Sep 14, 2026
Measurements
Every sourced value of this recipe, each with its condition and source. Bars relative to the largest value in the group.
No sourced measurements for this recipe.
What you need
- Hardware
- 1 × NVIDIA DGX Spark (GB10)
- Weights
- Main weightsnvidia/Gemma-4-26B-A4B-NVFP4 Apache-2.0
- MTPgoogle/gemma-4-26B-A4B-it-assistant Apache-2.0
- Engine
- vLLM vllm/vllm-openai:nightly (ungepinnt)
- Context
- 262,144 tokens
Related recipes
1 request, prose prompt, with DFlash (per recipe) Source
1 request, synthetic prompt Source
1 request, prose prompt, with MTP (per recipe) Source
- Custom kernel required
- License unclear
1 request, prose prompt, with MTP (per recipe) Source
- Access on request
- License unclear
1 request, prose prompt, prompt 50 tokens, with DSpark (per recipe) Source
- License unclear
Local AI in your company?
In a workshop we work out which models and which hardware fit your tasks, and build the first agent on your infrastructure.