---
type: "Article"
title: "Ornith-1.5 35B FP8 with vLLM on 1× DGX Spark"
description: "Ornith-1.5 35B FP8 with vLLM on 1× DGX Spark: recipe by sfxnz with speed, condition and source for every number, weights and license."
resource: "https://www.contextstudios.ai/local-ai/sfxnz--ornith-15-35b-a3b-nvfp4-1x-dgx-spark"
language: "en"
generated:
  by: "process:contextstudios-md/1"
  at: "2026-10-02T21:57:25.408Z"
status: "stable"
---

# Ornith-1.5 35B FP8 with vLLM on 1× DGX Spark

- Creator: [sfxnz](https://www.contextstudios.ai/local-ai/creators/sfxnz.md)
- Engine: vLLM
- Quantization: ModelOpt W4A16_NVFP4-Experten + FP8-Attention, FP8-KV, Marlin MoE, FlashInfer, MTP-3 (triton)
- Model family: Ornith-1.5
- Hardware: dgx-spark x1
- Context: 262144

## All numbers

- no measured speed on record
- ttft_ms: 120.92 (1 request) — [source](https://github.com/sfxnz/Ornith-1.5-35B-A3B-NVFP4-DGX-Spark/blob/main/README.md)

## Weights

- [ornith-ai/Ornith-1.5-35B-A3B-NVFP4](https://huggingface.co/ornith-ai/Ornith-1.5-35B-A3B-NVFP4) (main, mit)

[https://github.com/sfxnz/Ornith-1.5-35B-A3B-NVFP4-DGX-Spark](https://github.com/sfxnz/Ornith-1.5-35B-A3B-NVFP4-DGX-Spark)

## Related

- [Local AI on your hardware.](https://www.contextstudios.ai/local-ai.md)
