Local AI · Recipe · 1× DGX Spark

Qwen3.8-Flash-Next 80B UD-IQ4_XS with llama.cpp on 1× DGX Spark

Qwen3.8-Flash-Next 80B UD-IQ4_XS with llama.cpp on 1× DGX Spark: 26.9 tok/s according to github.com (dataset as of Sep 29, 2026).

by Weschera

26.9tok/sMixed

1 request, mixed prompt set, no speculative decoding

Source: github.com
Intelligence (original model) · Artificial Analysisno independent value

Engine

Engine
llama.cpp
Quantization
Unsloth UD-IQ4_XS GGUF (94 GB)
Model family
Qwen3.8-Flash-Next
Context
262,144
Parameters
80B-Klasse hybrid SSM + sparse-attention MoE (README; ohne Params-Auflösung)
Creator
Weschera
GitHub stars
9
Repo updated
Aug 27, 2026

Measurements

Every sourced value of this recipe, each with its condition and source. Bars relative to the largest value in the group.

Decode · 1 request

  • Mixed (decode)26.9 tok/s3-Run-Median, temp 0; 65K- und 262K-Config je 26.9 (262K: kurze Prompts) Source 
  • Peak (decode)26.9 tok/s3-Run-Median, temp 0; 65K- und 262K-Config je 26.9 (262K: kurze Prompts) Source 

Prefill by context

  • 185K181 tok/s@185K-Prompt; 65K-Config ~193 tok/s (~-Marker) Source 

What you need

Hardware
1 × NVIDIA DGX Spark (GB10)
Weights
Engine
llama.cpp
Context
262,144 tokens

Related recipes

← Back to overview

Local AI in your company?

In a workshop we work out which models and which hardware fit your tasks, and build the first agent on your infrastructure.