Local AI

Local AI Hardware Guide 2026: DGX Spark vs. Mac Studio M5 Ultra vs. RTX 5090 vs. Strix Halo

Local AI Hardware Guide 2026: DGX Spark vs. Mac Studio M5 Ultra vs. RTX 5090 vs. Strix Halo

Update log

  • 24 Sep 2026: First version. Covers DGX Spark/GB10, AMD Strix Halo, RTX 3090/4090/5090/PRO 6000, MacBook Pro M5 Max and the new Mac Studio M5 Ultra (shipping since 22 Sep 2026). Includes our own benchmark run from 24 Sep 2026 (GLM-5.3-Flash and Qwen3.8-Flash-Next on 2× DGX Spark-class nodes). Prices as of September 2026.
  • Next refresh: 1 Oct 2026 (monthly: new models, new prices, fresh measurements).

Short answer: In September 2026, memory bandwidth sets how fast one conversation runs, and memory size decides which model fits at all. A used RTX 3090 is still the cheapest way to get fast tokens from a 27B model. The Mac Studio M5 Ultra (1.2 TB/s, up to 512 GB) is the fastest single-user box for large MoE models. The DGX Spark wins when you need CUDA, clustering and several parallel streams. Strix Halo gets you 128 GB for the least money, but prompt processing is slow.

We run DGX Spark-class GB10 nodes (ASUS Ascent GX10) with vLLM in production at Context Studios, with a Mac mini as agent host. This guide combines our own measurements with the best publicly documented numbers, and every figure is tagged by source type.

The one rule that explains every benchmark

Decode speed (the tokens you read) is roughly memory bandwidth ÷ bytes read per token. Prefill (how fast the model reads your prompt) is compute-bound. Memory size decides whether the model fits at all.

Mixture-of-experts (MoE) models such as Qwen3.8-Flash-Next (~6B active parameters) or GLM-5.3-Flash (320B total, 18B active) read only a small slice of their weights per token. That is why a 128 GB box with 273 GB/s can run them at usable speed, while a 70B dense model crawls on the same machine.

Comparison table (September 2026)

HardwareMemoryBandwidthPrice (Sep 2026)Power under loadBest stack
NVIDIA DGX Spark / ASUS GX10 (GB10)128 GB unified273 GB/s$4,699 list (US), from €5,038 net (EU)~160 WvLLM, SGLang, EXL3, llama.cpp
AMD Ryzen AI Max+ 395 (Framework, EVO-X2)128 GB unified (~96–100 GB for GPU)~256 GB/s$3,449–3,650 (128 GB)~112–120 Wllama.cpp Vulkan/ROCm
RTX 3090 (used)24 GB936 GB/s~$1,050–1,264250–350 WvLLM, llama.cpp, EXL3
RTX 4090 (used)24 GB1,008 GB/s~$2,268–2,362~450 WvLLM, llama.cpp
RTX 509032 GB1,792 GB/s$1,999 MSRP, ~$4,700 street~575 WvLLM, llama.cpp, NVFP4
RTX PRO 6000 Blackwell96 GB~1.8 TB/s$16,000 (Sep 2026)up to 600 WvLLM, SGLang, NVFP4
MacBook Pro M5 Maxup to 128 GB614 GB/snot verifiedlaptopMLX, mlx-serve, llama.cpp
Mac Studio M5 Ultra96/256/512 GB1.2 TB/sfrom $5,499; 512 GB from Octobernot yet measured under LLM loadMLX, LM Studio, llama.cpp

Prices moved a lot in 2026 because of the DRAM and GDDR7 shortage. The RTX PRO 6000 now lists about 87 % above its launch MSRP, and the GMKtec EVO-X2 128 GB went from $1,999 to $3,499–3,650.

Model × hardware × tokens per second

Single-stream decode unless noted. [R] = review or vendor, [C] = community report on X or GitHub, [CS] = measured by Context Studios.

ModelHardwareDecodeNotes
Qwen3.8-27B (dense) Q4RTX 3090 / 4090 / 509040 / 46 / 75 t/sllama.cpp stock, 4k ctx [R]
Qwen3.8-27BRTX 5090 + MTP77 → 155 t/sllama.cpp Q4_K_M [C]
Qwen3.8-27BRTX 3090, tuned vLLM~114 t/s, ~1,000 t/s at 64 streamssyv-ai recipe [C]
Qwen3.8-27BDGX Spark12 t/s stock → 31–58 t/s tunedNVFP4 + MTP [R]/[C]
Qwen3.8-27BStrix Halo11 → 24–30 t/sVulkan + MTP [R]/[C]
Qwen3.8-27BM5 Max 128 GB31 → 65–69 t/sMLX + MTP/DFlash2 [R]/[C]
Qwen3.8-Flash-Next (MoE)Mac Studio M5 Ultra100+ t/s short, 60–85 t/s at 64k–256kMacStories review [R]
Qwen3.8-Flash-NextMacBook Pro M5 Max56–101 t/smlx-serve, MTP [C]
Qwen3.8-Flash-Next1× DGX Spark65–81 t/svLLM/EXL3 + MTP [C]
Qwen3.8-Flash-Next2× DGX Spark, NVFP432.9 t/s, 722 t/s prefill (1k prompt)our run 24 Sep, no speculative decoding [CS]
Qwen3.8-Flash-NextStrix Halo22–38 t/sllama.cpp / Halogen [C]
Qwen3.8-Flash-NextRTX PRO 6000227 t/s, prefill 9,900 t/s[C]
Qwen3.8-Flash-Next2× RTX PRO 60001,610 t/s over 32 streamsstock vLLM [C]
GLM-5.3-Flash EXL3 4bpw2× DGX Spark (TP=2)32.9 t/s, 63.7 t/s aggregate at 4 streamsour run 24 Sep [CS]
GLM-5.3-Flash IQ2Strix Halo17–18 t/s[C]
GLM-5.2 (large MoE)4× GB10 cluster27 t/s, 76 t/s aggregate at 6 streamsvLLM TP4 + MTP, Jul/Aug [CS]
DeepSeek-V4.1-Flash2× DGX Spark23 → 33 t/sEXL3, 1M ctx [C]
DeepSeek-V4.1-Flash4× RTX PRO 6000~120–125 t/s flat to 131k[C]

Prefill matters more than people think. In the MacStories test, the M5 Ultra read a 15.5k-token prompt at 2,887 t/s, compared with 1,143 t/s on the M3 Ultra. An RTX 5090 reached about 3,000 t/s on a 6k prompt. On Strix Halo, a 1M-token prompt took 18 minutes to process. For agents that send 20–60k tokens of system prompt, tools and memory on every turn, prefill often decides whether a machine feels fast.

Community measurements from X (September 2026)

Single-stream decode unless noted. These are reports by builders on X, not our own runs. Setups differ (quantization, speculative decoding, engine patches), so treat each row as what is possible with that exact recipe, not as a guaranteed baseline. The gap between the 62.9 t/s reported for GLM-5.3-Flash on two Sparks and our 32.9 t/s shows how much a tuned recipe matters.

HardwareModelSetupDecodeNotesSource
1× DGX SparkQwen3.8-Flash-NextNVFP4, vLLM + MTP k=338 t/s53.8 t/s on code, 180 t/s aggregate at 8 streams, ~2,000 t/s prefill@jvr0x
1× DGX SparkQwen3.8-Flash-NextEXL3 vs. NVFP4, vLLM78.3 vs. 31.0 t/ssame machine; ~600 vs. ~150 t/s aggregate at 8 requests@ViC305
2× DGX Spark (TP=2)GLM-5.3-FlashEXL3 4bpw, vllm-exl362.9 t/s146.5 t/s aggregate at 4 streams@yume_arasaki
2× DGX Spark (TP=2)GLM-5.3-FlashNVFP4, patched vLLM46.3 t/ssingle stream, reasoning@mr_r0b0t
8× DGX SparkGLM-5.3-FlashNVFP4, vLLM96 t/s110 t/s at 4 streams@hypermac6502
16× DGX SparkKimi K3 (2.8T)vLLM + GB10 patch29.8 t/s87 t/s aggregate at 8 streams, ~20 t/s at 300k context@ciprianveg
RTX PRO 6000 (96 GB)Qwen3.8-Flash-NextSGLang, FP8 KV + MTP252.9 t/s819 t/s at 8 streams, 1,190 t/s at 16@wei_wang
4× RTX PRO 6000GLM-5.3-FlashvLLM240 t/s~12,000 t/s prefill, 550 t/s aggregate at 4 streams@hiheyhowareya
Strix Halo (Ryzen AI Max+ 395)Qwen3.8-Flash-NextLucebox ROCm38 t/s41.9 t/s at 8k context@luceboxai
RTX 4090 (24 GB)Qwen3.8-27BQ4_K_M, llama.cpp46 t/sflash attention, quantized KV cache@yume_arasaki
RTX 3090 / 4090 (24 GB)Qwen3.8-27BEXL3 3.5bpw + DFlash94–105 t/ssustained; 130–150 t/s on short answers@yume_arasaki
MacBook Pro M4 Max (128 GB)Qwen3.8-27BMLX 4-bit + MTP43 t/srange 42–44 t/s@yume_arasaki

Our measurement (Context Studios, 24 Sep 2026)

We ran a short, reproducible benchmark against the endpoints we use every day. Both models run on pairs of ASUS Ascent GX10 nodes (NVIDIA GB10, 128 GB each) with tensor parallel 2 over a 100G RoCE fabric. The Qwen endpoint is reached through a LiteLLM proxy on our Mac mini, so its numbers include one extra network hop.

ModelSetupSingle-stream decodePrefill (~1k prompt)4 parallel streams
GLM-5.3-Flash (320B MoE, 18B active)EXL3 4bpw, vLLM, 2× GX10 TP=232.9 t/s1,262 t/s63.7 t/s aggregate (~16.5 t/s per stream)
Qwen3.8-Flash-NextNVFP4, vLLM via LiteLLM, 2× GX1032.9 t/s722 t/s21.8 t/s aggregate (one stream dropped to 7 t/s)

Method. Streaming /v1/chat/completions, temperature 0, thinking disabled, fixed technical prompt. One warm-up run is discarded, then three runs with 400 output tokens; we report the median. Decode = (completion tokens − 1) ÷ time from first to last token. Prefill is estimated as prompt tokens ÷ time to first token for a ~1,040-token prompt, which includes network latency and is therefore a lower bound. The parallel test fires four identical 400-token requests at once and divides all generated tokens by wall time. Token counts come from the server's usage field. Our cluster serves production agents at the same time, so the parallel numbers in particular are a snapshot, not a lab result.

You can reproduce the run with our script against any OpenAI-compatible endpoint:

bash
python3 bench_local.py http://YOUR-HOST:8000 your-model-id 3 4

What we read from these numbers:

  1. Two Sparks give you a 320B-parameter model at interactive speed. 33 t/s single-stream feels fluid in chat and in coding agents.
  2. The community's higher Qwen numbers come from speculative decoding. Setups with MTP report 65–81 t/s on one Spark; our NVFP4 endpoint runs without MTP and lands at 33 t/s. Speculative decoding is the biggest single lever, it roughly doubled GLM-5.2 on our 4-node cluster (14.5 → 26.6 t/s).
  3. Parallel throughput depends on everything else running on the box. GLM scaled almost linearly to four streams; the Qwen pair was busy with agent traffic, and one stream stalled.

Earlier: 4× GB10 cluster with GLM-5.2 (July/August 2026)

What we measuredResult
GLM-5.2 decode, 1 / 2 / 3 / 4 / 6 parallel streams (aggregate)27.2 / 42.6 / 52.1 / 62.7 / 75.8 t/s
Prefill 8k / 32k prompt853 / 840 t/s
MTP effect (no draft → k=2)14.5 → 26.6 t/s
Fixing cudagraph capture sizes+38 % at 2 streams
Long conversation (22 turns, 104,846 tokens)no degradation
Cold start / full cluster restart520 s / 13–14 min

Lessons from running it: if your 2-stream speed is close to your 1-stream speed, your CUDA-graph capture sizes are missing (one config line gave us +38 %). Unified memory needs care, because an OOM killer can take the engine down and force-removed containers leaked memory until reboot. Download a checkpoint once and copy it over the fabric: ~15 minutes instead of ~48 hours for three more nodes.

Buying decision by budget, use case and model size

You want to…BudgetBuyWhy
Run 27–32B dense models fast for one person~€1,200–1,600Used RTX 3090Most tokens per euro; 24 GB fits Qwen3.8-27B Q4
Serve a small team (5–60 parallel requests) with a 27B model~€2,500–3,5002× RTX 3090 or one RTX 5090Batched vLLM throughput
Run 100–300B MoE models quietly on your desk~€3,500–4,000Strix Halo 128 GBMost memory per euro; accept slow prefill
Build on CUDA, prototype for data-center GPUs, or cluster~€5,000–6,000 per nodeDGX Spark / GX10NVFP4, vLLM recipes, 200G clustering
Best single-user speed on large MoE models and long contextfrom $5,499Mac Studio M5 Ultra1.2 TB/s and up to 512 GB
Work local-first on the roadhighMacBook Pro M5 Max 128 GB614 GB/s in a laptop
Serve 700B+ models at production speed$64k+4× RTX PRO 6000384 GB VRAM, ~120 t/s DeepSeek V4.1 Flash

Rules of thumb by model size (4-bit):

  • Up to 32B dense: fits in 24–32 GB VRAM. A single RTX card is the fastest option.
  • Around 70B dense: needs about 48 GB and runs slowly on unified memory (4–6 t/s on Strix Halo).
  • 100–320B MoE: needs 64–128 GB. Spark, Strix Halo, M5 Max or Mac Studio. Two Sparks run GLM-5.3-Flash at 4 bpw with room for long context.
  • 500B+ MoE: needs a cluster of Sparks, 256–512 GB Macs, or several RTX PRO 6000.

Cost per token: local vs. cloud

Assumptions: German electricity at €0.37/kWh (BDEW 2026 average), 3-year write-off (26,280 hours), running 24/7 at full load. That is the best case for local hardware.

SetupThroughputElectricity onlyIncl. hardwareCloud reference
2× GX10, GLM-5.3-Flash, 4 streams [CS]63.7 t/s~€0.52 / 1M tokens~€2.03 / 1M tokensGLM-5.3-Flash API: $0.50 / 1M output
4× GX10, GLM-5.2, 6 streams [CS]75.8 t/s~€0.87 / 1M tokens~€3.41 / 1M tokenssame
Used RTX 3090, batched vLLM, Qwen3.8-27B [C]~1,000 t/s~€0.04 / 1M tokens~€0.05 / 1M tokensQwen3.8-Flash API: $0.47 / 1M output

How we calculated the 2× GX10 row: 2 × ~160 W = 0.32 kW, or €0.118 per hour. Hardware 2 × €4,558 (German retail price, 28 Aug 2026) over 26,280 hours adds €0.347 per hour. 63.7 t/s is 229,320 tokens per hour, so €0.465 per hour works out to about €2.03 per million tokens.

  • The cloud is roughly 4× cheaper per token than our two-node GLM setup, even at full load.
  • A fully used RTX 3090 undercuts API prices for 27B models, but only with constant batched traffic.
  • A typical single user runs the hardware at under 10 % utilization, which multiplies the local cost by 10 or more.

Our conclusion: local AI rarely wins on price per token. It wins on data sovereignty, fixed costs, no rate limits and full control over models and context. If you only need tokens, the API is cheaper. If data must not leave the building, or you want agents running 24/7 without a metered bill, own the hardware.

Typical pitfalls by platform

  • DGX Spark: single-stream speed is limited by bandwidth; long cold prefills stall parallel decode; community recipes change weekly.
  • Strix Halo: the GTT kernel parameter is required; ROCm and kernel versions are fragile; long prompts take minutes.
  • RTX cards: VRAM runs out quickly; offloading to system RAM drops speed to single digits; power draw and noise are high; the 5090 street price is about double its MSRP.
  • Macs: prefill is still behind NVIDIA; many MLX conversions drop vision or MTP; the newest models often need beta builds.

FAQ

What is the best hardware for local LLMs in 2026?

It depends on model size and the number of users. For one person running 27–32B models, a used RTX 3090 offers the most speed per euro in September 2026. For 100B+ MoE models with long context, the Mac Studio M5 Ultra is the fastest single-user box, while the DGX Spark is the better choice if you need CUDA, clustering or many parallel requests.

Is the DGX Spark worth it compared with the Mac Studio M5 Ultra?

The M5 Ultra has about 4.4× the memory bandwidth (1.2 TB/s vs 273 GB/s), so it generates tokens faster for a single user. The Spark's strengths are CUDA, NVFP4, fast community vLLM recipes and cheap clustering over 200G networking. We chose Sparks because we serve several agents in parallel, and aggregate throughput matters more to us than single-stream speed.

How fast is GLM-5.3-Flash on two DGX Sparks?

In our run on 24 September 2026, GLM-5.3-Flash in EXL3 4bpw on two ASUS GX10 nodes with tensor parallel 2 reached 32.9 tokens per second for a single stream. With four parallel requests, the pair delivered 63.7 tokens per second in total, or about 16.5 per stream. Prefill for a ~1k-token prompt ran at more than 1,200 tokens per second, measured including network latency.

Can AMD Strix Halo run large models like GLM-5.3-Flash?

Yes. With 128 GB of unified memory, of which roughly 96–100 GB can go to the GPU, it runs 2- to 4-bit quants of 200–300B MoE models at about 13–38 t/s. The catch is prompt processing: long contexts can take minutes, and a 1M-token prompt took 18 minutes in one report. It is a strong capacity box for patient single users, not for agent loops with huge prompts.

How many tokens per second do I need for coding agents?

For interactive work, around 30 t/s decode feels fluid and 60+ t/s feels fast. For agents, prefill speed is just as important, because every turn sends tens of thousands of tokens of context. A machine with 50 t/s decode but 300 t/s prefill will feel slower in an agent harness than one with 40 t/s decode and 2,500 t/s prefill.

Is running local AI cheaper than using an API?

Usually not per token. Even at full 24/7 load, our two-node GLM-5.3-Flash setup costs about €2.03 per million tokens including hardware, while the GLM-5.3-Flash API charges $0.50 per million output tokens. Local wins on privacy, predictable fixed costs, no rate limits and full control, and a fully utilized used RTX 3090 can undercut API prices for 27B models.

Why are hardware prices so high in 2026?

A global DRAM and GDDR7 shortage has pushed up the price of everything memory-heavy. NVIDIA raised the DGX Spark from $3,999 to $4,699 in February 2026, the RTX PRO 6000 now lists at $16,000, and 128 GB Strix Halo boxes rose by around 75 %. Used cards with older GDDR6X memory, such as the RTX 3090, are the least affected.

Sources


Set up local AI for your team?

Context Studios advises on hardware, setup and operations for local AI: from choosing between a DGX Spark, a Mac Studio or a GPU server to running inference with vLLM in production. We run this stack ourselves every day, and the measurements in this guide come from it.

Talk to us about your local AI setup → · Our services

Relevant for your team? Let's talk for 30 minutes.

We sort out what of this actually works in your company — concrete, no slide marathon.

No commitment · 30 minutes · Proposal within 48 h