1 request, best value reported by source Source
Local AI · Recipe · 1× DGX Spark
Ornith-1.5 35B FP8 with vLLM on 1× DGX Spark
by sfxnz
No speed value is documented for this recipe, so this page is not indexed.
Engine
- Engine
- vLLM
- Quantization
- ModelOpt W4A16_NVFP4-Experten + FP8-Attention, FP8-KV, Marlin MoE, FlashInfer, MTP-3 (triton)
- Model family
- Ornith-1.5
- Context
- 262,144
- Parameters
- 35B-A3B (Repo-Titel); HF-API safetensors total 19,528,501,104
- Creator
- sfxnz
- GitHub stars
- 1
- Repo updated
- Sep 4, 2026
Measurements
Every sourced value of this recipe, each with its condition and source. Bars relative to the largest value in the group.
Time to first token
Shorter is better.
What you need
- Hardware
- 1 × NVIDIA DGX Spark (GB10)
- Weights
- Main weightsornith-ai/Ornith-1.5-35B-A3B-NVFP4 MIT
- Engine
- vLLM
- Context
- 262,144 tokens
Related recipes
1 request, best value reported by source Source
- License unclear
1 request, prose prompt, with MTP (per recipe) Source
- Custom kernel required
- License unclear
1 request, prose prompt, with MTP (per recipe) Source
- Access on request
- License unclear
1 request, prose prompt, with MTP Source
1 request, prose prompt, prompt 50 tokens, with DSpark (per recipe) Source
- License unclear
Local AI in your company?
In a workshop we work out which models and which hardware fit your tasks, and build the first agent on your infrastructure.