Local AI · Recipe · 1× DGX Spark

Nemotron-3.5 30B NVFP4 with vLLM on 1× DGX Spark

Nemotron-3.5 30B NVFP4 with vLLM on 1× DGX Spark: 93 tok/s according to github.com (dataset as of Sep 29, 2026).

by sfxnz

93tok/sPeak

1 request, best value reported by source

Source: github.com
Intelligence (original model) · Artificial Analysis13with thinking

Engine

Engine
vLLM
Quantization
NVFP4 (modelopt_mixed), FP8-KV, Marlin MoE, DSpark-3
Model family
Nemotron-3.5
Context
262,144
Parameters
30B total / ~3B active (README); HF-API safetensors total 17,820,210,764
Creator
sfxnz
GitHub stars
1
Repo updated
Sep 4, 2026

Measurements

Every sourced value of this recipe, each with its condition and source. Bars relative to the largest value in the group.

Decode · 1 request

  • Peak (decode)93 tok/sL.A.I.L-Lab-Bench, Run id 20260812T131332Z_d0f7ae; Quelle '~93 tok/s' → approx; Messmethodik/Prompt im Repo nicht dokumentiert Source 

Prefill by context

  • Prompt not stated678 tok/sQuelle '~678 tok/s' → approx; L.A.I.L-Lab-Bench Source 

Time to first token

Shorter is better.

  • Prompt not stated246 msQuelle '~246 ms' (TTFT p50) → approx Source 

What you need

Hardware
1 × NVIDIA DGX Spark (GB10)
Engine
vLLM
Context
262,144 tokens

Related recipes

← Back to overview

Local AI in your company?

In a workshop we work out which models and which hardware fit your tasks, and build the first agent on your infrastructure.