Local AI · Recipe · 4× DGX Spark

DeepSeek-V4-Flash 305B NVFP4 with vLLM on 4× DGX Spark

DeepSeek-V4-Flash 305B NVFP4 with vLLM on 4× DGX Spark: 42 tok/s according to github.com (dataset as of Sep 29, 2026).

by tonyd2wild

42tok/sEveryday

1 request, prose prompt, with DSpark (per recipe)

Source: github.com
Intelligence (original model) · Artificial Analysisno independent value

Engine

Engine
vLLM
Quantization
FP8/FP4 Original-Checkpoint + NVFP4 KV (nvfp4_ds_mla)
Model family
DeepSeek-V4-Flash
Context
1,048,576
Parameters
~305B MoE
Creator
tonyd2wild
GitHub stars
486
Repo updated
Sep 10, 2026

Measurements

Every sourced value of this recipe, each with its condition and source. Bars relative to the largest value in the group.

Decode · 1 request

  • Everyday (decode)42 tok/scontext kurz (5 Prosa-Prompts, 24–41 Prompt-Tokens) · C1 Median Kategorie prose, thinking off, 2026-09-02; CURRENT.md: „prose 42 tok/s“ Source 
  • Mixed (decode)74.2 tok/scontext kurz, 40 Prompts / 8 Kategorien · C1 decode Median über alle 40 Prompts Source 
  • Peak (decode)98.3 tok/scontext kurz (5 Coding-Prompts) · C1 Median Kategorie coding; CURRENT.md: „code 98 tok/s“ Source 

Prefill by context

  • ~1.5K4,868 tok/sCold, c1 (README SGLang-Repo Z.68); CURRENT.md: „~4.6K tok/s, flat from 14K to 182K“ Source 

What you need

Hardware
4 × NVIDIA DGX Spark (GB10)
Engine
vLLM
Context
1,048,576 tokens

Related recipes

← Back to overview

Local AI in your company?

In a workshop we work out which models and which hardware fit your tasks, and build the first agent on your infrastructure.