Local AI · Recipe · 1× DGX Spark

Ornith-1.5 35B FP8 with vLLM on 1× DGX Spark

by sfxnz

No speed value is documented for this recipe, so this page is not indexed.

Intelligence (original model) · Artificial Analysisno independent value

Engine

Engine
vLLM
Quantization
ModelOpt W4A16_NVFP4-Experten + FP8-Attention, FP8-KV, Marlin MoE, FlashInfer, MTP-3 (triton)
Model family
Ornith-1.5
Context
262,144
Parameters
35B-A3B (Repo-Titel); HF-API safetensors total 19,528,501,104
Creator
sfxnz
GitHub stars
1
Repo updated
Sep 4, 2026

Measurements

Every sourced value of this recipe, each with its condition and source. Bars relative to the largest value in the group.

Time to first token

Shorter is better.

  • Prompt not stated121 mscontext 256 in / 128 out (random) · Mean TTFT der completions-256/128 max-concurrency-1-Tabelle Source 

What you need

Hardware
1 × NVIDIA DGX Spark (GB10)
Engine
vLLM
Context
262,144 tokens

Related recipes

← Back to overview

Local AI in your company?

In a workshop we work out which models and which hardware fit your tasks, and build the first agent on your infrastructure.