Local AI · Recipe · 3× DGX Spark

DeepSeek-V4.1-Flash 763B EXL3 3.5 bpw with vLLM on 3× DGX Spark

DeepSeek-V4.1-Flash 763B EXL3 3.5 bpw with vLLM on 3× DGX Spark: 46 tok/s according to github.com (dataset as of Sep 29, 2026).

by tonyd2wild

46tok/sMixed

1 request, mixed prompt set, with DSpark (per recipe)

Source: github.com
Intelligence (original model) · Artificial Analysis40with thinkingno thinking 25

Engine

Engine
vLLM
Quantization
EXL3 3.5 bpw routed experts (Pollard) + FP8 Rest
Model family
DeepSeek-V4.1-Flash
Context
300,000
Parameters
~763B MoE
Creator
tonyd2wild
GitHub stars
86
Repo updated
Sep 19, 2026

Measurements

Every sourced value of this recipe, each with its condition and source. Bars relative to the largest value in the group.

Decode · 1 request

  • Mixed (decode)46 tok/sC1 aggregate über 8 Kategorien (Counting ausgenommen); per-stream Mittel 51.46 Source 
  • Peak (decode)70.1 tok/sKategorie coding, exl3tp3a11 Source 

Prefill by context

  • 93.3351,199 tok/sCold, unique prefix, TTFT 77.865 s Source 

What you need

Hardware
3 × NVIDIA DGX Spark (GB10)
Engine
vLLM
Context
300,000 tokens

Related recipes

2× DGX Spark
42.6tok/sEveryday

1 request, prose prompt, with DSpark (per recipe) Source 

Intelligence40with thinkingno thinking 25Original model's value. Heavily compressed (~2 bit), intelligence is likely much lower.
  • ≤ 3 bit: quantization may cost quality
  • Experimental
  • Custom kernel required

EXL3 2.0 bpw MCG (routed experts, tail-biting Viterbi re-encoded, scale refit), lm_head MXFP8, KV fp8 (8-GiB-Pin), DSpark-3vLLMWeights Repo Updated Sep 27, 2026

Details

← Back to overview

Local AI in your company?

In a workshop we work out which models and which hardware fit your tasks, and build the first agent on your infrastructure.