Local AI · Recipe · 4× DGX Spark

Qwen3.8-27B EXL3 with vLLM on 4× DGX Spark

Qwen3.8-27B EXL3 with vLLM on 4× DGX Spark: 35.1 tok/s according to github.com (dataset as of Sep 29, 2026).

by fujitsupolycom

35.1tok/sEveryday

1 request, realistic prompt, with Qwen MTP3 probabilistic mit standard rejection (per recipe)

Source: github.com
Intelligence (original model) · Artificial Analysis34with thinkingno thinking 20

Engine

Engine
vLLM
Quantization
EXL3 K5/K6 gemischt, 21.59 GB (hf hydrated); FP8 KV (kv_cache_dtype fp8); EXL3-Prefill FP8, Tile 256
Model family
Qwen3.8-27B
Context
1,048,576
Parameters
27.8B dense (Qwen/Qwen3.8-27B, HF-API)
Creator
fujitsupolycom
GitHub stars
82
Repo updated
Sep 27, 2026

Measurements

Every sourced value of this recipe, each with its condition and source. Bars relative to the largest value in the group.

Decode · 1 request

  • Everyday (decode)35.1 tok/scontext 16K decode · N=5 Mittel; unique cold prompts, temp 1.0, top-p 0.95/top-k 20 aus der generation_config der Quelle; Workload-Typ in der Quelle nicht klassifiziert Source 
  • Peak (decode)48.5 tok/sCoding Peak, 15/15 requests, Mittel (Median 47.89, Range 44.57–52.13) Source 

Prefill by context

What you need

Hardware
4 × NVIDIA DGX Spark (GB10)
Engine
vLLM
Context
1,048,576 tokens

Related recipes

← Back to overview

Local AI in your company?

In a workshop we work out which models and which hardware fit your tasks, and build the first agent on your infrastructure.