Local AI · Recipe · 4× DGX Spark

DeepSeek-V4-Flash 291B FP8 with vLLM on 4× DGX Spark

DeepSeek-V4-Flash 291B FP8 with vLLM on 4× DGX Spark: 46 tok/s according to github.com (dataset as of Sep 29, 2026).

by tonyd2wild

46tok/sEveryday

1 request, prose prompt, with DSpark

Source: github.com
Intelligence (original model) · Artificial Analysisno independent value

Engine

Engine
vLLM
Quantization
FP8
Model family
DeepSeek-V4-Flash
Context
—
Parameters
~291B
Creator
tonyd2wild
GitHub stars
7
Repo updated
Sep 2, 2026

Measurements

Every sourced value of this recipe, each with its condition and source. Bars relative to the largest value in the group.

Decode · 1 request

  • Everyday (decode)46 tok/sPrivater Produktions-Stack des Autors (DSpark-k=5-Drafter + fused-Markov-argmax), nicht der öffentliche MTP-k2-Launcher (README „That stack rides a private image“); non-streaming, temp 0, max_tokens 200, Nonce am Anfang jedes Requests; prose = realistischer Worst case (niedrigste Draft-Acceptance); TP2-Vergleich 42 -> 46 (+9 %) Source 
  • Peak (decode)120 tok/sPrivater Produktions-Stack des Autors (DSpark-k=5-Drafter + fused-Markov-argmax), nicht der öffentliche MTP-k2-Launcher (README „That stack rides a private image“); non-streaming, temp 0, max_tokens 200, Nonce am Anfang jedes Requests; Count-to-300-Headline-Stream peakte bei ~127 t/s Source 

What you need

Hardware
4 × NVIDIA DGX Spark (GB10)
Engine
vLLM

Related recipes

← Back to overview

Local AI in your company?

In a workshop we work out which models and which hardware fit your tasks, and build the first agent on your infrastructure.