Local AI · Recipe · 2× DGX Spark

DeepSeek-V4-Flash-Vision-Exp 284B MXFP4 with vLLM on 2× DGX Spark

DeepSeek-V4-Flash-Vision-Exp 284B MXFP4 with vLLM on 2× DGX Spark: 26.2 tok/s according to github.com (dataset as of Sep 29, 2026).

by sfxnz

26.2tok/sEveryday

1 request, prose prompt, with DSpark (per recipe)

Source: github.com
Intelligence (original model) · Artificial Analysisno independent value

Engine

Engine
vLLM
Quantization
MXFP4-Experten, FP8-Attention, fp8-KV 12-GiB-Pin, DSpark-6
Model family
DeepSeek-V4-Flash-Vision-Exp
Context
1,048,576
Parameters
284B total / 13B active MoE + native 32-layer ViT (~0.47B BF16) (README); HF-API safetensors total 304,646,824,126
Creator
sfxnz
GitHub stars
1
Repo updated
Sep 4, 2026

Measurements

Every sourced value of this recipe, each with its condition and source. Bars relative to the largest value in the group.

Decode · 1 request

  • Everyday (decode)26.2 tok/sBench_decode.py, streamed greedy, thinking off, 200 Completion-Tokens, 3-run-Median; Rebench 2026-09-02: 26.175; TTFT p50 0.39 s Source 
  • Peak (decode)26.2 tok/sBench_decode.py, streamed greedy, thinking off, 200 Completion-Tokens, 3-run-Median; Rebench 2026-09-02: 26.175; TTFT p50 0.39 s Source 

Prefill by context

  • 13,349 Prompt-Tokens2,134 tok/sUnique-salt needle, KV-hit Source 

What you need

Hardware
2 × NVIDIA DGX Spark (GB10)
Engine
vLLM
Context
1,048,576 tokens

Related recipes

2× DGX Spark
42.6tok/sEveryday

1 request, prose prompt, with DSpark (per recipe) Source 

Intelligence40with thinkingno thinking 25Original model's value. Heavily compressed (~2 bit), intelligence is likely much lower.
  • ≤ 3 bit: quantization may cost quality
  • Experimental
  • Custom kernel required

EXL3 2.0 bpw MCG (routed experts, tail-biting Viterbi re-encoded, scale refit), lm_head MXFP8, KV fp8 (8-GiB-Pin), DSpark-3vLLMWeights Repo Updated Sep 27, 2026

Details
2× DGX Spark
97.6tok/sEveryday

HumanEval 97.6% / GSM8K 98.0% with FP8 KV and 4-bit dense Source 

Intelligence42with thinking
  • ≤ 3 bit: quantization may cost quality
  • Custom kernel required
  • Non-commercial
  • License unclear

EXL3/TR3 4bpw routed experts + 4-bit dense, BF16 elsewhere, FP8 KVTensorFold v0.5.0 (52 patches: DFlash2+copy drafts, 4-bit dense, FP8 KV, RoCE one-shot all-gather, shared KV pool, vision, tool calling)Weights Repo Updated Sep 30, 2026

Details

← Back to overview

Local AI in your company?

In a workshop we work out which models and which hardware fit your tasks, and build the first agent on your infrastructure.