Local AI · Recipe · 1× DGX Spark

Qwen3.8-27B NVFP4 with vLLM on 1× DGX Spark (Weschera)

Qwen3.8-27B NVFP4 with vLLM on 1× DGX Spark (Weschera): 45 tok/s according to github.com (dataset as of Sep 29, 2026).

by Weschera

45tok/sEveryday

1 request, prose prompt, with MTP

Source: github.com
Intelligence (original model) · Artificial Analysisno independent value

Engine

Engine
vLLM
Quantization
NVFP4 (unsloth; compressed-tensors) + fp8 KV + triton_attn (fp8-KV auf GB10 nur mit triton, FA kann kein fp8-KV auf sm_121)
Model family
Qwen3.8-27B
Context
262,144
Parameters
27B (dicht, hybrid Gated-DeltaNet VLM, 262K nativ, MTP-Head)
Creator
Weschera
GitHub stars
1
Repo updated
Aug 15, 2026

Measurements

Every sourced value of this recipe, each with its condition and source. Bars relative to the largest value in the group.

Decode · 1 request

  • Everyday (decode)45 tok/sApple-Silicon-Lane: mlx-community 4-bit + MTP-4bit-Sidecar via mlx-vlm, M4 Max 128 GB, ~45 tok/s (55 % acceptance, temp 0.7, single session) → approx; plain MLX 29.5 Source 
  • Peak (decode)45 tok/sApple-Silicon-Lane: mlx-community 4-bit + MTP-4bit-Sidecar via mlx-vlm, M4 Max 128 GB, ~45 tok/s (55 % acceptance, temp 0.7, single session) → approx; plain MLX 29.5 Source 

What you need

Hardware
1 × NVIDIA DGX Spark (GB10)
Weights
Engine
vLLM
Context
262,144 tokens

Related recipes

← Back to overview

Local AI in your company?

In a workshop we work out which models and which hardware fit your tasks, and build the first agent on your infrastructure.