Lokale KI · Rezept · 2× DGX Spark

DeepSeek-V4-Flash FP8 mit vLLM auf 2× DGX Spark (1024k Kontext)

DeepSeek-V4-Flash FP8 mit vLLM auf 2× DGX Spark (1024k Kontext): 83,8 tok/s laut github.com (Datensatz-Stand: 29. Sept. 2026).

von Weschera

83,8tok/sGemischt

1 Anfrage, gemischtes Prompt-Set, mit Draft

Quelle: github.com
Intelligenz (Originalmodell) · Artificial Analysiskein unabhängiger Wert

Engine

Engine
vLLM
Quantisierung
Official BF16/FP8-Checkpoint mit nvfp4_ds_mla-KV (block 256); MoE FlashInfer B12X
Modellfamilie
DeepSeek-V4-Flash
Kontext
1.048.576
Kreator
Weschera
GitHub-Sterne
9
Repo-Stand
31. Juli 2026

Messwerte

Alle belegten Werte dieses Rezepts, je mit Bedingung und Quelle. Balken im Verhältnis zum größten Wert der Gruppe.

Decode · 1 Anfrage

  • Gemischt (Decode)83,8 tok/sKontext 1k prompt / 512 output · 40K-Speed-Profil (config.40k-speed.example.env): Median aus 3 Reps, 1,024 prompt tokens forced 512 output, K7 greedy, Output-SHA-256 identisch zu spec-off; +3.09× vs 27.1030; 1,770/2,093 Draft-Tokens accepted Quelle 
  • Spitze (Decode)83,8 tok/sKontext 1k prompt / 512 output · 40K-Speed-Profil (config.40k-speed.example.env): Median aus 3 Reps, 1,024 prompt tokens forced 512 output, K7 greedy, Output-SHA-256 identisch zu spec-off; +3.09× vs 27.1030; 1,770/2,093 Draft-Tokens accepted Quelle 

Was brauche ich

Hardware
2 × NVIDIA DGX Spark (GB10)
Engine
vLLM
Kontext
1.048.576 Token

Verwandte Rezepte

← Zur Übersicht

Lokale KI im eigenen Unternehmen?

Im Workshop klären wir, welche Modelle und welche Hardware zu Ihren Aufgaben passen, und bauen den ersten Agenten auf Ihrer Infrastruktur.