Local AI · Recipe · 2× DGX Spark

DeepSeek-V4-Flash FP8 with vLLM on 2× DGX Spark (1024k context)

DeepSeek-V4-Flash FP8 with vLLM on 2× DGX Spark (1024k context): 83.8 tok/s according to github.com (dataset as of Sep 29, 2026).

by Weschera

83.8tok/sMixed

1 request, mixed prompt set, with Draft

Source: github.com
Intelligence (original model) · Artificial Analysisno independent value

Engine

Engine
vLLM
Quantization
Official BF16/FP8-Checkpoint mit nvfp4_ds_mla-KV (block 256); MoE FlashInfer B12X
Model family
DeepSeek-V4-Flash
Context
1,048,576
Creator
Weschera
GitHub stars
9
Repo updated
Jul 31, 2026

Measurements

Every sourced value of this recipe, each with its condition and source. Bars relative to the largest value in the group.

Decode · 1 request

  • Mixed (decode)83.8 tok/scontext 1k prompt / 512 output · 40K-Speed-Profil (config.40k-speed.example.env): Median aus 3 Reps, 1,024 prompt tokens forced 512 output, K7 greedy, Output-SHA-256 identisch zu spec-off; +3.09× vs 27.1030; 1,770/2,093 Draft-Tokens accepted Source 
  • Peak (decode)83.8 tok/scontext 1k prompt / 512 output · 40K-Speed-Profil (config.40k-speed.example.env): Median aus 3 Reps, 1,024 prompt tokens forced 512 output, K7 greedy, Output-SHA-256 identisch zu spec-off; +3.09× vs 27.1030; 1,770/2,093 Draft-Tokens accepted Source 

What you need

Hardware
2 × NVIDIA DGX Spark (GB10)
Engine
vLLM
Context
1,048,576 tokens

Related recipes

← Back to overview

Local AI in your company?

In a workshop we work out which models and which hardware fit your tasks, and build the first agent on your infrastructure.