Local AI · Recipe · 4× DGX Spark

Qwen3.8-2.4T-A95B UD-Q1_0 with llama.cpp on 4× DGX Spark

Qwen3.8-2.4T-A95B UD-Q1_0 with llama.cpp on 4× DGX Spark: 7.76 tok/s according to github.com (dataset as of Sep 29, 2026).

by Weschera

7.76tok/sPeak

1 request, best value reported by source, prompt 121 tokens

Source: github.com
Intelligence (original model) · Artificial Analysisno independent value

≤ 3 bit: quantization may cost quality

Engine

Engine
llama.cpp
Quantization
Unsloth UD-Q1_0 GGUF (1-bit dynamic, 397,256,393,248 bytes, 10 Shards); MODEL_SHA256SUMS im Repo
Model family
Qwen3.8-2.4T-A95B
Context
65,536
Parameters
2.4T total / 95B aktiv (Modellname/README)
Creator
Weschera
GitHub stars
0
Repo updated
Aug 13, 2026

Measurements

Every sourced value of this recipe, each with its condition and source. Bars relative to the largest value in the group.

Decode · 1 request

  • Peak (decode)7.76 tok/scontext 121-token prompt · Qualified pagoda run 2026-08-12: 53,448 Completion-Tokens in 6,883.90 s, natural stop; API-Completion-Throughput genau dieses Long-Runs; first token 5.78 s; 64.87 % MTP-Acceptance Source 

What you need

Hardware
4 × NVIDIA DGX Spark (GB10)
Weights
Engine
llama.cpp
Context
65,536 tokens

Notes

What matters before you rebuild it.

  • ≤ 3 bit: quantization may cost qualityAt 3 bit and below the model may answer noticeably worse than the original. The intelligence number refers to the original.
  • Experimental
  • License unclear

Sources

Related recipes

← Back to overview

Local AI in your company?

In a workshop we work out which models and which hardware fit your tasks, and build the first agent on your infrastructure.