Qwen3.8-Flash-Next EXL3 3.05 bpw with ExLlamaV3 on 1× DGX Spark
1 request, prose prompt, with MTP (per recipe) Source
Good quality at a fluid pace for everyday use.
Open recipeVotre objectif
Être visible là où l'IA répond
Automatiser les processus
Construire un produit
Déployer des agents IA
Connecter et moderniser vos systèmes
Savoir où nous en sommes
Cas d'Usage
CRMRenforcer les relations clientsPopulaireE-CommerceAugmenter les ventes en ligneSystème de réservationRéservations 24h/24Gestion de projetCoordonner les équipesFacturationÊtre payé plus viteAnalytiqueDécisions basées sur donnéesResources · Local AI
Which open models run on DGX Spark, Mac, Strix Halo or RTX, how fast and how smart? Recipes from their creators, every number with its condition and source, every weight linked.
Quick pick for your hardware
Pick platform and count: the three recommendations for your hardware appear here.
282 recipes24 creators6 platforms217 with intelligence scoreas of Sep 26, 2026
Pick platform and count. Recommendations and the recipe list follow your choice.
Memory classes (rule of thumb)
A Context Studios rule of thumb, not a measurement. What actually fits is stated in each recipe.
No selection yet: the balanced pick per platform for its smallest setup.
1 request, prose prompt, with MTP (per recipe) Source
Good quality at a fluid pace for everyday use.
Open recipe1 request, prose prompt, short prompt, with MTP (DSpark built-in) (per recipe) Source
Good quality at a fluid pace for everyday use.
Open recipe1 request, mixed prompt set, with ndt=3 Source
Good quality at a fluid pace for everyday use.
Open recipe1 request, prose prompt, with MTP (per recipe) Source
Good quality at a fluid pace for everyday use.
Open recipe1 request, mixed prompt set, with MTP Source
Good quality at a fluid pace for everyday use.
Open recipe1 request, mixed prompt set, with FR-Spec + NEXTN (per recipe) Source
Good quality at a fluid pace for everyday use.
Open recipeEvery dot is a recipe: further right is faster, higher up is smarter. The shape shows how speed was measured.
209 of 282 recipes shown, only recipes with both a speed and an intelligence value.
The speed categories are not directly comparable: “peak” is a best case, mostly code or speculative decoding, “everyday” is a single request with a prose prompt.
Intelligence = original model per Artificial Analysis Intelligence Index v4.3.2, as of Sep 26, 2026. Quantization and engine can change the value in a recipe. Source: Artificial Analysis
Tap shows details, a second tap opens the recipe.
Filter by hardware, engine, quantization, model family and creator. Cards or table, same data.
1 request, prose prompt, no speculative decoding Source
1 request, prose prompt, no speculative decoding Source
1 request, prose prompt, context 19, with DSpark (per recipe) Source
1 request, prose prompt, no speculative decoding Source
1 request, prose prompt, prompt 450 tokens, no speculative decoding Source
1 request, prose prompt, no speculative decoding Source
1 request, prose prompt, with MTP (per recipe) Source
1 request, prose prompt, with MTP Source
1 request, prose prompt, with MTP (per recipe) Source
1 request, prose prompt, no speculative decoding Source
1 request, prose prompt, no speculative decoding Source
1 request, prose prompt, with DFlash (per recipe) Source
1 request, prose prompt, no speculative decoding Source
1 request, prose prompt, with MTP (per recipe) Source
1 request, prose prompt, with MTP (per recipe) Source
1 request, prose prompt, with Draft Source
1 request, prose prompt, with MTP (per recipe) Source
1 request, prose prompt, with MTP (per recipe) Source
1 request, prose prompt, with MTP (per recipe) Source
1 request, prose prompt, with DFlash2 (per recipe) Source
1 request, prose prompt, with MTP (per recipe) Source
1 request, prose prompt, with MTP (per recipe) Source
1 request, prose prompt, with DSpark (per recipe) Source
1 request, prose prompt, with DFlash2 (per recipe) Source
For Strix Halo we show quality instead of speed: agent tasks solved in Terminal-Bench-Mini (19 tasks in the terminal). Plus a few selected speed values with attribution; the full series live at the author’s site.
| Model | Quant | Setup | Engine | Solved | First try | Run |
|---|---|---|---|---|---|---|
| DeepSeek-V4.1-Flash | Q2 | 2× Strix Halo | DwarfStar pr-16-09-2026 / rocm 10.0 | 18/19 | Run ↗ | |
| Qwen3.8-Flash-Next | W4B | 1× Strix Halo | halogen-flash-server / unknown | 18/19 | Run ↗ | |
| DeepSeek-V4-Flash-0731 | MXFP4 | 2× Strix Halo | DwarfStar / rocm 7.14 | 15/19 | Run ↗ | |
| DeepSeek-V4-Flash-0731 | UD-IQ2_XXS | 1× Strix Halo | llama.cpp vulkan-performance / vulkan | 18/19 | Run ↗ | |
| DeepSeek-V4-Flash-0731 | IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8 | 1× Strix Halo | DwarfStar / rocm 10.0 | 18/19 | Run ↗ | |
| DeepSeek-V4-Flash-0731 | IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8 | 1× Strix Halo | DwarfStar / rocm 10.0 | 17/19 | Run ↗ | |
| DeepSeek-V4.1-Flash | Q2 | 2× Strix Halo | DwarfStar pr-22-09-2026 / rocm 10.0 | 16/19 | Run ↗ | |
| Qwen3.8-27B | UD-Q4_K_XL | 1× Strix Halo | llama.cpp / rocm 7.14 | 15/19 | Run ↗ | |
| DeepSeek-V4-Flash-0731 | UD-IQ3_XXS | 1× Strix Halo | llama.cpp / vulkan | 17/19 | Run ↗ | |
| Qwen3.8-27B | Q4_0_ROCMI4 | 1× Strix Halo | llama.cpp rocmfpx / rocm 7.14 | 16/19 | Run ↗ | |
| Qwen3.8-Flash-Next | UD-IQ4_XS | 1× Strix Halo | EngramHalo / rocm 10.0 | 16/19 | Run ↗ | |
| Qwen3.8-27B | Q4_0_ROCMFP4_STRIX | 1× Strix Halo | llama.cpp rocmfpx / rocm 7.14 | 15/19 | Run ↗ | |
| GLM 5.3 Flash · run incomplete | Q2 | 1× Strix Halo | DwarfStar / rocm 10.0 | 14/19 | Run ↗ | |
| Qwen3.6-27B | UD-Q8_K_XL | 1× Strix Halo | llama.cpp / rocm 7.14 | 12/19 | Run ↗ | |
| Muse-Glimmer-30B | UD-Q4_K_XL | 1× Strix Halo | llama.cpp / rocm 7.14 | 11/19 | Run ↗ | |
| Qwen3.6-35B-A3B · run incomplete | UD-Q4_K_XL | 1× Strix Halo | llama.cpp / rocm 7.14 | 11/19 | Run ↗ |
Results from Terminal-Bench-Mini by kyuz0, Apache-2.0 (based on Terminal-Bench 2.1, Apache-2.0). Terminal-Bench-Mini (core19) · kyuz0 · Apache-2.0
Summarised from Donato Capitella's videos and kyuz0's toolbox READMEs. Each source carries its date.
Avoid linux-firmware 20251125; the fix landed from 20260110.1 kyuz0's firmware notes (8 Jan 2026) give a downgrade to 20251111 as a stopgap.2 Run kernel 6.18.4 or newer only with a matching ROCm (7.2+ or nightly); older ROCm images are not stable on new kernels.3 kyuz0 currently tests on kernel 6.18.9.4
Keep the fixed VRAM carve-out small in the BIOS and hand about 124 GB to the GPU via kernel parameters: amd_iommu=off amdgpu.gttsize=126976 ttm.pages_limit=32505856. On the Ryzen AI Halo, amd-ttm --set 124 does the same. About 4 GB stay with the system.4
Headless, with tuned set to accelerator-performance and the IOMMU off. The often quoted “10–15% faster” is mostly prefill (+14 to +18%); decode gains +3 to +7%. Without the IOMMU, the NPU and DMA protection are off.5
kyuz0's toolboxes bundle engine and drivers. vulkan-radv is the most stable choice and recommended for all models. MTP is now part of standard llama.cpp; the old -mtp toolboxes are deprecated.6
MTP adds +22 to +47% decode speed on MoE models and +72 to +93% on dense 27B models. Prefill gets 1 to 27% slower.7,8 For Qwen3.8-Flash-Next, Donato currently recommends Halogen ahead of EngramHalo, noting this may change; EngramHalo needs -lm mmap --lazy-mode on.9,6
A model with few active parameters is fast on Strix Halo, and the software stack matters: in decode Vulkan is usually ahead of ROCm (in every pair only with the experimental RADV performance fork), MTP adds a clear decode step, and long prior context costs noticeable speed. Large MoE models run in low-bit quants, but much slower.
Decode in tok/s, one request, 2,048-token prompt + 128 output tokens
Measurements: Donato Capitella · local-llm-benchmarks.dev · retrieved Sep 26, 2026. Selected values quoted with attribution; full series at local-llm-benchmarks.dev.
1 request, prose prompt, no speculative decoding Source
1 request, mixed prompt set, with ndt=3 Source
1 request, synthetic prompt Source
1 request, synthetic prompt Source
1 request, code prompt Source
Datasheet per vendor, prices from German shops with retailer and retrieval date. Third-party prices, not a Context Studios offer.
| Platform | Memory | Bandwidth | Power | Native FP4 | Native FP8 | Price (DE) |
|---|---|---|---|---|---|---|
| NVIDIA DGX Spark (GB10)Grace Blackwell GB10nvidia.com ↗ developer.nvidia.com ↗ | 128 GBLPDDR5x unified (coherent) | 273 GB/s | n/a | yes | n/a | from €4,800NVIDIA Marketplace DE ↗ · as of Sep 25, 2026NVIDIA DGX Spark Founders Edition, 128 GB, 4 TB SSDAll 13 offers
|
| Mac / Apple SiliconUnified Memory; Chip/Speicher/Bandbreite je Rezept verschieden | n/aLPDDR (unified)M2 Max, M2 Ultra, M3 Ultra, M4 Max, M5 Max | n/a | n/a | n/a | n/a | from €5,859Apple Store DE ↗ · as of Sep 25, 2026Apple Mac Studio (M5 Max), M5 Max 18-Core CPU / 40-Core GPU, 128 GB, 512 GB SSDAll 4 offers
|
| AMD Strix Halo16× Zen 5 + Radeon 8060S (40 CU, RDNA 3.5, gfx1151)amd.com ↗ | 128 GBLPDDR5x-8000 (unified, max. 128 GB) | 256 GB/s | 55 (cTDP 45–120) W | no | no | from €3,799galaxus (via geizhals.de) ↗ · as of Sep 25, 2026GMKtec EVO-X2, Ryzen AI Max+ 395, 128 GB, 2 TB SSDAll 3 offers
|
| NVIDIA RTX 3090Ampere (GA102)images.nvidia.com ↗ nvidia.com ↗ | 24 GBGDDR6X | 936 GB/s | 350 W | no | no | from €2,290BC GmbH (via geizhals.de) ↗ · as of Sep 25, 2026NVIDIA GeForce RTX 3090 24 GB, neu – EVGA FTW3 Ultra Gamingused: €999–€1,500 · kleinanzeigen.de ↗ |
| NVIDIA RTX 5090Blackwell consumer (GB202)images.nvidia.com ↗ nvidia.com ↗ | 32 GBGDDR7 | 1792 GB/s | 575 W | yes | yes | from €5,555Future-X.de (via geizhals.de) ↗ · as of Sep 25, 2026NVIDIA GeForce RTX 5090 32 GB, Zotac Gaming AMP Extreme INFINITY |
| NVIDIA RTX PRO 6000 BlackwellBlackwell (GB202, 188 SM)nvidia.com ↗ nvidia.com ↗ | 96 GBGDDR7 ECC | 1792 GB/s | maxq: 300 W · server: bis 600 (konfigurierbar) W · workstation: 600 W | yes | yes | from €16,399computeruniverse.net (via geizhals.de) ↗ · as of Sep 25, 2026NVIDIA RTX PRO 6000 Blackwell 96 GB, Max-Q Workstation Edition (PNY, Smallbox VCNRTXPRO6000MQ-SB)All 2 offers
|
Prices in EUR incl. VAT, retailer and retrieval date per entry. per device.
What more Sparks buy you, why the box works despite modest bandwidth, and what matters when buying. Every number is sourced.
A Spark has 128 GB of unified memory,1 of which about 121 GiB are usable according to a recipe author.2 That suits mid-sized MoE models: in MiaAI Lab's recipe, Qwen3.6-35B does 95.1 tok/s for a single request and 317 tok/s in total across eight.3
Two Sparks connect with a single QSFP cable.4 Three form a switchless ring with three cables.5 For four, NVIDIA specifies a switch with at least four 200 Gbit/s ports.6 Community projects such as SparkRing also run four boxes as a switchless ring; NVIDIA does not document that path.7
More boxes mainly add memory. DeepSeek-V4.1-Flash takes 476 GiB on disk and only fits from three Sparks up; on four, the recipe reports 87.7 tok/s for a single prose request.8
Tensor parallelism scales speed well but not linearly: in NVIDIA's test (Llama 3.3 70B, 32K input), decode on two Sparks is 2.0× and on four 3.7× as fast as on one.9
273 GB/s of memory bandwidth is modest for a machine in this class.1 During decode the model reads its weights again for every token. Three techniques make up for that.
MoE models activate only a fraction of their parameters per token. Speculative decoding (MTP, DSpark, DFlash) lets a small helper model propose several tokens that the large model checks in one step. Tensor parallelism splits every weight matrix across Sparks: in NVIDIA's test, time per output token drops from 269 ms (one Spark) to 133 ms (two) and 72 ms (four).9
Prefill of long prompts gains less: 1.6× and 2.1× at 32K tokens (computed from NVIDIA's TTFT values). NVIDIA itself notes that inference scales sub-linearly when nodes have to synchronise often.9
Partner machines use the same GB10 chip (compute capability 12.1) as NVIDIA's Founders Edition, so the recipes are not tied to one model.10 The recipe's hardware note says which machine the author used.
The Founders Edition ships with 4 TB of NVMe.1 That fills up quickly: Qwen3.8-Flash-Next alone needs about 130 GiB.2
For stacking, NVIDIA qualifies the Amphenol NJAAKK-N911 (0.4 m) or NJAAKK0006 (0.5 m) and Luxshare LMTQF022-SD-R cables. Faster cables do not help, since each port tops out at 200 Gbit/s.11
GB10 machines differ: a software update cut the Founders Edition's idle draw clearly, while a Dell Pro Max GB10 stayed at 35–37 W.12
At idle, a Spark on current software draws 22–25 W.12 LLM inference mostly takes 60–90 W, full load just under 200 W; outside stress tests it stayed below 40 dBA.13
At the average German household rate of 37.0 ct/kWh, a Spark running around the clock costs roughly €16–24 a month under typical inference load and about €6 at idle (our calculation, 720 hours).14
Further reading: „DGX Spark: from one box to a cluster“ ↗ · 0xSero for EXO Labs, 25 Sep 2026. Some figures in the article are out of date by now; the checked values are above.
We curate and link. The work belongs to the people who published it.
Headline numbers in repos are often best cases. We separate three categories and show the best available one per card.
Decode for one request with plain text. The speed you feel in chat. Preferred as main number.
Mixed prompt set, often with speculative decoding. Sits between everyday and peak.
Best documented value, usually code or JSON with speculative decoding. Realistic for coding agents, not for prose.
Rank: category first (everyday before mixed before peak), then speed. Recipes without a documented main number come last.
Speed values from local-llm-benchmarks.dev and the Strix Halo toolboxes are not published as numbers while their license is unclear. We link them.
Artificial Analysis Intelligence Index v4.3.2 for the original model, retrieved on Sep 26, 2026 via the official API. The shown stage is the AA default stage, with the non-reasoning value in small type. Derived models (REAP, pruning) get no value.
Results from Terminal-Bench-Mini by kyuz0, Apache-2.0 (based on Terminal-Bench 2.1, Apache-2.0).
At 3 bit and below the model may answer noticeably worse than the original. The intelligence number refers to the original.
Hardware prices from German shops in EUR incl. VAT, as of Sep 25, 2026. Retailer and retrieval date are listed per offer.
All recipes, recommendations and measurements are public via the Context Studios MCP server, no key needed.
Public endpoint
https://mcp.contextstudios.ai/api/public/mcp
Tools · Public · no key
Example
claude mcp add contextstudios \
--transport http https://mcp.contextstudios.ai/api/public/mcp
find_local_ai_recipes({
platform: "dgx-spark",
count: 2,
goal: "balance"
})In a workshop we work out which models and which hardware fit your tasks, and build the first agent on your infrastructure.