Local AI · Recipe · 1× DGX Spark

Qwen3.8-27B NVFP4 with Atlas on 1× DGX Spark

Qwen3.8-27B NVFP4 with Atlas on 1× DGX Spark: 19 tok/s according to github.com (dataset as of Sep 29, 2026).

by Weschera

19tok/sEveryday

1 request, prose prompt, with MTP (per recipe)

Source: github.com
Intelligence (original model) · Artificial Analysisno independent value

Engine

Engine
Atlas
Quantization
NVFP4 (unsloth); MTP-Head bf16 (--mtp-quantization bf16); fp8 KV (Engine-Default)
Model family
Qwen3.8-27B
Context
32,768
Parameters
27B (dicht)
Creator
Weschera
GitHub stars
0
Repo updated
Aug 15, 2026

Measurements

Every sourced value of this recipe, each with its condition and source. Bars relative to the largest value in the group.

Decode · 1 request

  • Everyday (decode)19 tok/sProse temp 0.6 ('~19') → approx Source 
  • Peak (decode)31.6 tok/sMaintainer-Fingerprint-Protokoll: Atlas-README gibt Bereich '30-32'; 31.6 wörtlich im Quant-Ladder-README (gleiche Messkampagne, inkl. deterministic-twin 964-token run); eigener Lab 32.0/29.8/30.5 → approx Source 

What you need

Hardware
1 × NVIDIA DGX Spark (GB10)
Weights
Engine
Atlas
Context
32,768 tokens

Related recipes

← Back to overview

Local AI in your company?

In a workshop we work out which models and which hardware fit your tasks, and build the first agent on your infrastructure.