Local AI · Recipe · 1× DGX Spark

Qwen3.8-Flash-Next EXL3 3.05 bpw with ExLlamaV3 on 1× DGX Spark

by vcruz305

53tok/sEveryday

1 request, prose prompt, with MTP (per recipe)

Source: github.com
Intelligence (original model) · Artificial Analysis40with thinking

Engine

Engine
ExLlamaV3 + TabbyAPI Fork-Pin 94ba01d (gemessen auf 329e051)
Quantization
EXL3 3.05bpw (turboderp-Pack, h5, ngram5), 8-bit KV
Model family
Qwen3.8-Flash-Next
Context
262,144
Parameters
48 Layer, 512 Experten (MoE); Gesamtparameter in README nicht genannt
Creator
vcruz305
GitHub stars
57
Repo updated
Sep 24, 2026

Measurements

Every sourced value of this recipe, each with its condition and source. Bars relative to the largest value in the group.

No sourced measurements for this recipe.

What you need

Hardware
1 × NVIDIA DGX Spark (GB10)
Engine
ExLlamaV3 + TabbyAPI Fork-Pin 94ba01d (gemessen auf 329e051)
Context
262,144 tokens

Related recipes

← Back to overview

Local AI in your company?

In a workshop we work out which models and which hardware fit your tasks, and build the first agent on your infrastructure.