Local AI · Recipe · 2× DGX Spark

DeepSeek-V4-Flash NVFP4 with vLLM on 2× DGX Spark

DeepSeek-V4-Flash NVFP4 with vLLM on 2× DGX Spark: 65.4 tok/s according to github.com (dataset as of Sep 29, 2026).

by Weschera

65.4tok/sPeak

1 request, best value reported by source

Source: github.com
Intelligence (original model) · Artificial Analysisno independent value

Engine

Engine
vLLM
Quantization
NVFP4-KV-Pfad (Stage C: 584-byte padded envelope via nvfp4_ds_mla; NICHT der 416-byte true-layout Fix — dokumentiert gescheitert >411 prompt tokens)
Model family
DeepSeek-V4-Flash
Context
1,048,576
Creator
Weschera
GitHub stars
8
Repo updated
Jul 2, 2026

Measurements

Every sourced value of this recipe, each with its condition and source. Bars relative to the largest value in the group.

Decode · 1 request

  • Peak (decode)65.4 tok/sP256/g256, bester 1M-Profil-Wert; acceptance 0.718, accepted/draft 3.59; TTFC 0.324 s; >50 tok/s in allen 5 Probes Source 

What you need

Hardware
2 × NVIDIA DGX Spark (GB10)
Engine
vLLM
Context
1,048,576 tokens

Related recipes

← Back to overview

Local AI in your company?

In a workshop we work out which models and which hardware fit your tasks, and build the first agent on your infrastructure.