local-ai

Mac Studio M5 Ultra for Local AI: Benchmarks, Prices and Our Comparison with 4× DGX Spark (2026)

Mac Studio M5 Ultra for Local AI: Benchmarks, Prices and Our Comparison with 4× DGX Spark (2026)

Updated: September 24, 2026. Apple announced the Mac Studio with M5 Ultra on August 25, 2026. According to the Apple Store, the 512 GB configuration arrives "late October". We will update this guide as new measurements come in.

Short answer: In September 2026, the M5 Ultra is the fastest quiet single-box machine for large mixture-of-experts models such as Qwen3.8-Flash-Next or GLM-5.3-Flash. It delivers 1.2 TB/s of memory bandwidth and up to 256 GB (soon 512 GB) of unified memory. For dense models that fit into 32 GB, an RTX 5090 is still faster. If you need CUDA, clustering or many concurrent users, DGX Spark is the better fit. And if cost per token is all you care about, the cloud wins: by our math, local inference on the M5 Ultra costs roughly €1.70–2.70 per million output tokens, while the Qwen3.8-Flash API charges $0.47.

At Context Studios we have been running a cluster of four ASUS Ascent GX10 (NVIDIA GB10, the same hardware as DGX Spark) since summer 2026. So we asked ourselves the obvious question: should we have bought Mac Studios instead of Sparks? This guide is our honest answer, built on verified specs, independent reviews, numbers from the X community and our own measurements from September 24.

The verified specs in 60 seconds

SpecM5 Ultra (base)M5 Ultra (top)
CPU30 cores (10 super, 20 performance)36 cores
GPU64 cores with Neural Accelerators80 cores with Neural Accelerators
Neural Engine32 cores32 cores
Memory bandwidth1.2 TB/s1.2 TB/s
Unified memory96 GB256 GB or 512 GB (512 GB from late October, 80-core GPU only)
SSD1 TBup to 16 TB
Ports4× Thunderbolt 5 (120 Gb/s), 10Gb Ethernet, HDMI 2.1same
Max. continuous power (PSU)480 W480 W

Source: Apple tech specs and Apple Newsroom, August 25, 2026. The big change is a quad-die design: Apple uses UltraFusion to join four dies into one chip. Apple says bandwidth is 50 percent higher than on the M3 Ultra (819 GB/s). That number matters most for local LLMs, because generating each token means reading the active weights from memory.

Prices in the US and Germany

We read these prices directly from the Apple Store configurator (US and Germany) on September 24, 2026:

ConfigurationUSAGermany
M5 Ultra 30C/64G, 96 GB, 1 TB$5,499€6,599
M5 Ultra 36C/80G, 96 GB, 1 TB$6,799€8,029
M5 Ultra 30C/64G, 256 GB, 1 TB$9,499€10,999
M5 Ultra 36C/80G, 256 GB, 1 TB$10,799€12,429
Upgrade 96 → 256 GB+$4,000+€4,400
M5 Ultra 512 GBno price yetno price yet
For reference: M5 Max (entry)$2,499€2,999

Tom's Hardware reports 16 to 18 weeks of backorder for its review unit (80-core GPU, 256 GB, 4 TB, $12,299). Ars Technica expects the 512 GB model to be "a $20,000-and-up computer". That is an estimate, not an Apple price.

Benchmarks: what the M5 Ultra actually does

We don't own an M5 Ultra. Every Mac number here comes from three independent reviews with published methodology, plus community measurements that we label separately below. Results depend heavily on runtime, quantization and context length, so only compare numbers within the same table.

Qwen3.8-Flash-Next (MoE) with MLX

MacStories tested the M5 Ultra (256 GB) against the M3 Ultra (512 GB) using oMLX 0.7.0.dev2:

Context lengthM3 Ultra decodeM5 Ultra decodeM5 Ultra time to first token
4K59.0 t/s90.7 t/s2.5 s
16K53.5 t/s87.8 t/s6.2 s
64K45.3 t/s83.8 t/s23.7 s
128K37.0 t/s60.6 t/s49.2 s
256K38.6 t/s74.7 t/s101.5 s

For prompt processing (prefill), the M5 Ultra reads a 16K prompt at 2,887 t/s (M3 Ultra: 1,143 t/s) and still manages about 2,550 t/s at 256K. That is roughly 2.5 times the previous generation, thanks to the new Neural Accelerators in every GPU core. On short prompts, Flash-Next reaches a median of 108 t/s across three runs, 54% more than the M3 Ultra (70 t/s).

GLM-5.3-Flash with MLX

For GLM-5.3-Flash (MLX, mixed 4/8-bit), MacStories measures 41 t/s decode and 1,107 t/s prefill on a 16K prompt on the M5 Ultra. The M3 Ultra manages 26 t/s and 428 t/s. Note that GLM's writing rate here includes its reasoning tokens.

Qwen 3.8 27B (dense) with llama.cpp

Tom's Hardware measures Qwen 3.8 27B Q4_K_M with llama.cpp, without multi-token prediction:

ContextMac Studio M5 UltraMac Studio M4 MaxDGX Spark
046.3 t/s22.7 t/s11.8 t/s
16K43.4 t/s21.1 t/s11.0 t/s
64K32.4 t/s14.6 t/s9.3 t/s
128K30.1 t/s9.0 t/s7.7 t/s

Using LM Studio, Ars Technica gets "just over 50" tokens per second for the same 4-bit model. The M5 Max Studio manages about 31 t/s, while a Framework Desktop (Strix Halo) and an M2 Max Studio each land at roughly 18 to 20 t/s.

How it stacks up against the RTX 5090

MacStories also pitted Qwen 3.8 27B on the Mac against a gaming PC with an RTX 5090 (32 GB). On a 6,091-token prompt, the 5090 reads at 3,031 t/s, the M5 Ultra at 1,701 t/s and the M3 Ultra at 414 t/s. For writing, the 5090 hits 59 t/s and the M5 Ultra 48 t/s. As long as the model fits into 32 GB of VRAM, the 5090 is about 25% faster at decode and nearly twice as fast at prefill. Once it no longer fits, the picture flips: offloading to system RAM dropped the 5090 to 4.6 to 1.5 t/s in the same test.

What the X community is measuring

Since the first units shipped, new M5 Ultra numbers have been appearing on X every day. On September 24 we used X search to collect the most-viewed measurements. These are one-off runs with different runtimes and settings, not controlled benchmarks. They do show how fast MLX software is catching up.

WhoModel / setupResultLink
@mweinbach (163,000 views)Qwen3.8-Flash-Next, batched3,740 t/s prefill, 149 t/s decode ("nearly double what it was yesterday")Post
@viticci (MacStories)Flash-Next 4/8-bit, MLX-Serve v26.9.5over 3,500 t/s prefill, 110 t/s on short prosePost
@wei_wangFlash-Next, multi-turn agents with tool calls60–85 t/s, prefill about 2.5× faster than M3 Ultra on averagePost
@MiaAI_lab (144,000 views)GLM-5.3-Flash, M5 Ultra vs. 2× DGX SparkM5 Ultra: 28 t/s decode, 1,016 t/s prefill; 2× Spark: 39 t/s, 1,700 t/sPost
@MiaAI_labQwen3.8-Flash on 2× DGX Spark56.8 t/s single stream, 146 t/s with 4 streams, 3,500 t/s prefillPost

Two takeaways. First, Mac numbers are sometimes doubling within a day, because kernels for the new Neural Accelerators are only now being written. Second, the community's Spark recipes are well tuned. Whether "M5 Ultra vs. two Sparks" goes one way or the other depends on the model and on the day.

Our own measurements: 2× GX10 on September 24

To put the community numbers into context, we ran measurements on two nodes of our cluster on September 24 (ASUS Ascent GX10, tensor parallel over 100G RoCE). Method: streaming chat completions, temperature 0, thinking off, one warm-up, median of three runs, 400 output tokens.

Model on 2× GX10Quantization / engineSingle-stream decode4 concurrent streams (total)
GLM-5.3-FlashEXL3 4 bpw, vLLM32.9 t/s63.7 t/s (about 16–17 t/s per stream)
Qwen3.8-Flash-NextNVFP4, vLLM behind a LiteLLM proxy32.9 t/snot reliable (one stream stalled)

With a short prompt of roughly 1,000 tokens, we measured 1,262 t/s prefill for GLM-5.3-Flash, including network overhead. That is a lower bound, not a long-context figure. How this compares: in a single stream, the M5 Ultra is ahead of our two nodes on GLM-5.3-Flash with 41 t/s (MacStories). With four concurrent streams, the two GX10 deliver more in total. For Qwen3.8-Flash-Next, our current NVFP4 setup at 32.9 t/s is well behind the 56.8 t/s Mia AI Lab shows on two Sparks, and clearly behind the M5 Ultra. Our setup isn't at the community's level here, and we'd rather say so.

From the summer we also have numbers for the larger GLM-5.2 on all four nodes: 27.2 t/s with one stream, 75.8 t/s with six streams, and prefill of 853 t/s (8K) and 840 t/s (32K).

The prompt processing question: caught up, not ahead

Slow prompt ingestion has been Apple Silicon's best-known weakness. For agentic workflows with 50,000 tokens of context or more, it matters a lot. The M5 Ultra makes its biggest leap here, but doesn't fully close the gap to NVIDIA:

  • Against the RTX 5090, it reads about half as fast (1,701 vs. 3,031 t/s).
  • Against the DGX Spark, it is faster at short and medium context. Tom's Hardware measures a shorter time to first token up to 32K (42.2 s vs. 49.1 s at 32K). At 64K the two are even (about 104 s), and at 128K the Spark is ahead (236.8 s vs. 295.8 s).
  • Against the M3 Ultra, it is about 2.5 times faster.

That matches our experience with GB10: decode is held back by 273 GB/s of bandwidth, but compute-heavy prefill is strong.

Buying advice: which model fits which memory tier?

Rule of thumb: 4-bit weights need about 0.55–0.6 GB per billion parameters. Add KV cache for context and roughly 15–20% headroom for macOS. For MoE models, total parameters determine memory, while active parameters determine speed.

Memory tierFits wellTight or doesn't fitRecommendation
96 GB (from $5,499)Qwen 3.8 27B in 8-bit or BF16, gpt-oss-120b, Qwen3.8-Flash-Next in 4-bit with moderate context (one community report cites about 68 GB)GLM-5.3-Flash, long 256K contexts with Flash-NextA solid start for one model at medium context. A 128 GB M5 Max Studio is often the better buy here at half the price.
256 GB (from $9,499)Flash-Next oQ4e/oQ5e with 256K context (peak 155–179 GB per MacStories), GLM-5.3-Flash in MLX 4/8-bit, three parallel Flash-Next sessionsFlash-Next in 6 or 8-bit only with embedding tables offloaded to SSD; oMLX allows about 200 GB at mostThe sweet spot for agentic work with large MoE models.
512 GB (late October, price TBA)Candidates: DeepSeek-V4.1-Flash, GLM-5.3 at higher quantizations, several large models at onceNo measurements yetOnly buy if a specific model breaks 256 GB.

One detail from MacStories shows how hard that ceiling is: on the 256 GB machine, macOS killed Flash-Next in 6-bit for memory pressure at 176.6 GB. oMLX refused to load the 8-bit version at all. Plan for at least 20% headroom above the raw model size.

Cost per token: our math

Assumptions (adjust to your situation): electricity at €0.37/kWh (German BDEW average for 2026), three-year depreciation at 24/7 operation, full utilization. There is no power measurement for the M5 Ultra under LLM load yet. Ars Technica measured 75 W during video encoding, and Apple lists 480 W maximum continuous power. We calculate with that range. For one GX10 node we use 160 W (Tom's Hardware, DGX Spark under GPU load) and a purchase price of €4,558 (German retail, August 2026).

SystemModel / loadThroughputHardwareCost per 1M output tokens
Mac Studio M5 Ultra 80C/256 GBFlash-Next, 3 concurrent requests (MacStories)81.5 t/s€12,429€1.71–2.22
Mac Studio M5 Ultra 80C/256 GBFlash-Next, one request with a 6.5K prompt (MacStories)66.2 t/s€12,429€2.10–2.73
Mac Studio M5 Ultra 80C/256 GBGLM-5.3-Flash, one request (MacStories)41 t/s€12,429€3.39–4.41
2× ASUS GX10 (our measurement)GLM-5.3-Flash, 4 streams63.7 t/s€9,116~€2.03
2× ASUS GX10 (our measurement)GLM-5.3-Flash, one stream32.9 t/s€9,116~€3.93
4× ASUS GX10 (our cluster)GLM-5.2, 6 streams75.8 t/s€18,232~€3.41
Cloud APIQwen3.8-Flash (Alibaba Cloud)––$0.47
Cloud APIGLM-5.3-Flash (Z.ai)––$0.50

The verdict is clear: none of these machines pays off on token price. Even at full utilization, the API is several times cheaper. In practice, a single workstation often runs below 10% utilization, which pushes the cost per token up tenfold. You buy local inference for data privacy, predictable fixed costs, independence from rate limits and free choice of models. On energy, the Mac is strong: even at an assumed 480 W, it needs about 1.6 kWh per million tokens for Flash-Next with three requests. Our four-node cluster needs about 2.4 kWh for GLM-5.2.

Mac Studio M5 Ultra or DGX Spark: when to pick which?

We deliberately chose GB10 hardware in the summer, and for our use case we'd do it again. The reasons just aren't the ones most benchmark tables suggest.

What speaks for the Spark (from running it ourselves):

  • CUDA and community recipes. New models often get optimized vLLM or SGLang recipes with NVFP4 and speculative decoding on the Spark within days.
  • Scaling across nodes. With 200G ConnectX and a 100G RoCE switch, we run GLM-5.2 with tensor parallelism across four nodes. We download big checkpoints once and distribute them to the other nodes over RoCE in about 15 minutes instead of 48 hours.
  • Concurrency. Our cluster scales GLM-5.2 from 27.2 t/s on one stream to 75.8 t/s on six. Two nodes nearly double GLM-5.3-Flash throughput from one to four streams (32.9 → 63.7 t/s). Multiple agents and users share the hardware.
  • Entry price per node. A GX10 cost €4,558 in Germany in August. Prices have climbed sharply since (Cyberport: from €4,000 to €5,199 in six weeks).

What speaks for the Mac Studio:

  • Single-stream speed. On Qwen 3.8 27B, Tom's Hardware measures almost four times the decode rate of a single Spark (46.3 vs. 11.8 t/s). With 1.2 TB/s against 273 GB/s, that's simple physics.
  • One box, not a cluster. 256 GB in a single box replaces two Sparks, with no network fabric, no tensor-parallel config and none of the 13 to 14 minutes a restart of our cluster takes.
  • Quiet and efficient. The Mac is quiet; MacStories describes it as noticeably quieter than their own 5090 PC.
  • MLX momentum. The MLX community (oMLX, mlx-lm, MLX-Serve, MTP variants) is catching up fast, and new MoE models often show up as MLX quantizations on day one.

Our decision guide:

You want to…Pick
run a large MoE model fast and quietly as a single userMac Studio M5 Ultra, 256 GB
serve a team or several agents in parallelDGX Spark / GX10 (2+ nodes)
fine-tune models or use CUDA toolingDGX Spark
run a dense model up to 32 GB as fast as possibleRTX 5090
get 128 GB of unified memory at the lowest priceAMD Strix Halo (prefill is its weak spot)
test models with more than 400 GB of weights locallywait for the M5 Ultra 512 GB or use a Spark cluster with 3–4 nodes

How it compares with Strix Halo and the RTX 5090

PlatformMemoryBandwidthQwen 3.8 27B (decode)StrengthsWeaknesses
Mac Studio M5 Ultra96–512 GB1.2 TB/s46–50 t/slarge MoE models, quietprice, prefill below NVIDIA
DGX Spark / GX10128 GB273 GB/s11.8 t/s (llama.cpp), 33–35 t/s (SGLang + MTP, community)CUDA, clustering, concurrencysingle-stream decode
AMD Strix Haloup to 128 GB~256 GB/s18–20 t/s (Ars, LM Studio)price per GBprefill, software maturity
RTX 509032 GB1.79 TB/s59 t/s (MacStories), 75 t/s (llama.cpp)speed, prefillonly 32 GB, 575 W

For more on Qwen 3.8 27B across all platforms, see our Qwen 3.8 27B hardware guide. For the full cost picture from Mac mini to M5 Ultra, read What local AI on Apple hardware costs in 2026.

Quick start: Flash-Next on the Mac Studio

mlx-lm is enough for a first test. The commands below are a minimal example. Check the exact name of the quantization you want on Hugging Face first.

bash
pip install -U mlx-lm
mlx_lm.server --model mlx-community/<flash-next-4bit-repo> --port 8080

This gives you an OpenAI-compatible API on port 8080 that you can plug into Hermes Agent, Codex or Open WebUI. MacStories uses oMLX for multi-token prediction, which is still under development.

Bottom line

The Mac Studio M5 Ultra is the best single-box machine for large open models you can buy in September 2026. With 1.2 TB/s and 256 GB, models like Qwen3.8-Flash-Next become usable for everyday agentic work at 60 to 90 t/s with long context. But it isn't an all-rounder. NVIDIA still leads on prompt processing, a Spark cluster scales better with many concurrent users, and for dense 27B models an RTX 5090 is faster and cheaper.

For us at Context Studios, the GB10 cluster remains the right choice, because we serve agents in parallel and benefit from the community's CUDA recipes. If we were buying a single quiet workstation for large models today, it would be an M5 Ultra with 256 GB. Hold off on the 512 GB model until there's a price and benchmarks.

FAQ

How fast is the Mac Studio M5 Ultra for local LLMs?

It depends heavily on the model. With the MoE model Qwen3.8-Flash-Next, MacStories measures about 108 t/s on short prompts and still 83.8 t/s at 64K context. For the dense Qwen 3.8 27B, Tom's Hardware measures 46.3 t/s at zero context and 30.1 t/s at 128K with llama.cpp. That makes it roughly twice as fast as an M4 Max, and 54 to 94% faster than the M3 Ultra on Flash-Next depending on context.

Is 96 GB enough, or do I need 256 GB?

96 GB is enough for dense models up to about 70 billion parameters in 4-bit, and for Qwen3.8-Flash-Next in 4-bit with moderate context. If you want large MoE models such as GLM-5.3-Flash, or Flash-Next with 256K context, you need 256 GB. Keep in mind that macOS and the runtime won't hand over all of the memory. On the 256 GB model, oMLX allows only about 200 GB for a single model.

Is the Mac Studio M5 Ultra better than a DGX Spark?

For a single user, the M5 Ultra is much faster than a single Spark. On Qwen 3.8 27B, Tom's Hardware measures almost four times the decode rate. The DGX Spark wins on CUDA, fast community recipes, clustering over ConnectX and better scaling with many concurrent requests. Against two Sparks, it's an open race: depending on the model and the state of the software, sometimes the Mac leads and sometimes the Spark pair does.

How much does the Mac Studio M5 Ultra cost?

In the US, the base model with a 30-core CPU, 64-core GPU, 96 GB and 1 TB costs $5,499; in Germany it's €6,599. The 80-core GPU adds $1,300, and going from 96 to 256 GB adds $4,000. The fully loaded 256 GB model with a 1 TB SSD comes to $10,799, and there's no price for the 512 GB version yet.

Does local AI on the M5 Ultra pay off financially?

Not on token price alone. Our math comes to roughly €1.70 to €2.70 per million output tokens for Flash-Next at full utilization, while comparable cloud APIs sit around $0.50. Local inference pays off through data privacy, predictable costs, independence from vendors and rate limits, and free choice of models. If you process sensitive data or run agents around the clock, you get real value in return.

How much power does the Mac Studio M5 Ultra draw?

Apple lists a maximum continuous power of 480 W. That's the limit of the power supply, not a typical figure. Ars Technica measured about 75 W during video encoding, and we haven't seen an independent measurement under LLM load yet. For comparison, a DGX Spark draws about 160 W at the wall under GPU load, and an RTX 5090 alone can pull up to 575 W.


Setting up local AI for your team? Context Studios advises on hardware, setup and operations, from a single Mac Studio to a multi-node cluster. Get in touch

Sources

Relevant for your team? Let's talk for 30 minutes.

We sort out what of this actually works in your company — concrete, no slide marathon.

No commitment · 30 minutes · Proposal within 48 h