Local AI

AMD's hybrid wave: Lucebox, ACEMAGIC & Minisforum compared to Mac Studio and DGX Spark

AMD's hybrid wave: Lucebox, ACEMAGIC & Minisforum compared to Mac Studio and DGX Spark

The numbers are brutal: 224 GB of AI memory for around €8,700 on one side, 512 GB for an estimated €15,500 on the other — and in between, a whole wave of new AMD systems all promising the same thing: run large language models locally without an NVIDIA-server budget. While Apple's Mac Studio M5 Ultra with 512 GB still waits for its launch ("late October"), the AMD front is already here: Lucebox is shipping hybrid workstations with a discrete Radeon in series, ACEMAGIC sells 192-GB Strix Halo minis (EU store ships from a German warehouse), and Minisforum delivers the build-it-yourself kit with the MS-S1 Max plus eGPU docks. We took the entire wave apart — architecture, bandwidth, price per gigabyte, the software trap almost nobody mentions — and ran the numbers against our own recipe database with 428 measured community configurations: which of the 65 catalogued model families actually runs on which system?

The pattern behind it: three paths to fluid memory

If you infer large models, you're not fighting for FLOPS, you're fighting for memory: capacity (does the model fit?) and bandwidth (how fast do the weights arrive per token?). In 2026 there are exactly three answers:

Path 1 — Unified memory only. AMD's Strix Halo APUs (Ryzen AI Max 385/390, Max+ 395, Max+ PRO 495) solder LPDDR5X directly onto the chip: up to 192 GB shared between the CPU and the integrated Radeon. Bandwidth: ~256–273 GB/s. No VRAM bottleneck, but no faster shortcut either — every token reads from the same pool.

Path 2 — Hybrid split. Here a discrete graphics card (AMD Radeon AI PRO R9700, 32 GB GDDR6 at 640 GB/s) goes into the same machine, and specialized software splits the model: the "hot" parts — dense layers, the speculative-decoding drafter, frequently routed mixture-of-experts — run in fast VRAM, the "long tail" of experts in the large unified memory. That's exactly what Lucebox does with its engine; the principle also works with standard tools, but takes manual work.

Path 3 — Apple's monolith. The Mac Studio M5 Ultra with 512 GB of unified memory and ~1.2 TB/s of bandwidth (per Apple, +50% over the M3 Ultra) combines capacity and speed in one system — without split logic, but at a premium price.

The interesting question isn't which path is "right," but what each path costs per euro — and which model even fits. There, the math looks surprising.

The players in detail

Lucebox Zero 495 — the hybrid pioneer from Milan

Lucebox (founded 2025, roughly 2–10 people, official AMD partnership) pairs the Ryzen AI MAX+ PRO 495 with up to 192 GB of LPDDR5X-8533 (273 GB/s) with a Radeon AI PRO R9700 (32 GB GDDR6, 640 GB/s) in an 11.97-litre aluminium chassis with a 1,000 W Platinum power supply. Ubuntu Server preinstalled, Lucebox Engine and models set up, OpenAI-/Anthropic-compatible API from the first boot.

Launch pricing (first 200 units, €1,000 under list price): €6,499 for the Developer configuration with 128 GB + 2 TB; the 192 GB tier adds €2,199 — so around €8,700 for the full 224 GB total-memory configuration. An optional 100 GbE cluster kit (Intel E810, factory-installed) costs €869 per machine. By its own account, DeepSeek V4 Flash (284B) decodes at 86 tok/s across both chips — a vendor claim we won't treat as a measurement until an independent reproduction.

Two things are critical for the buying decision: Batch 3 is sold out, Batch 4 ships in November 2026 per the company — and co-founder Alessandro Puppo has already announced a price increase on X. For the 495 models the company names January 2027 as the delivery window. Buying here means buying from a young startup on a prepay equivalent — with EU consumer protection (ships from Italy) but without the backing of a corporate service network.

ACEMAGIC F9A — 192 GB Strix Halo in a 2-litre case

ACEMAGIC sells the Ryzen AI Max+ PRO 495 with the full 192 GB of LPDDR5X-8533 and integrated Radeon 8065S (40 CUs) for €6,399 (192 GB + 2 TB) — no discrete VRAM, pure unified-memory design. In return: a 2-litre chassis, 120 W TDP, 2-year warranty — and on the EU store (acemagic.eu), shipping from a German warehouse; other regional stores ship from their own locations. The catch is printed small on the product page: "Pre-sale: Late October" — so the box is exactly as unavailable as Apple's 512 GB Studio right now. For reference: without a discrete GPU, decode stays bound to the APU's ~256–273 GB/s; the F9A is a clean Strix Halo system with European support via the EU store, not a miracle box.

MSI PRO MAX EDGE AI+ — the "virtual VRAM" concept

The MSI competitor (Ryzen AI Max+ 395, 128 GB, Radeon 8060S) advertises "up to 96 GB VRAM" at first glance. Realistically: that's not a separate graphics card, it's a BIOS allocation — up to 96 GB of the 128 GB unified memory is assigned exclusively to the iGPU, leaving 32 GB for OS and applications. Bandwidth stays ~256 GB/s (LPDDR5X-8000). We have not verified the ~€4,500 price ourselves; the architectural principle is the important part, because it's often misunderstood: there is no second, faster memory pool here — only a partition of the same pool.

Minisforum — the kit vendor: MS-S1 Max, DEG1/DEG2 and the DIY path

Minisforum is the only vendor covering the whole spectrum — from turnkey Strix Halo workstation to open docks:

MS-S1 Max (Ryzen AI Max+ 395, 128 GB LPDDR5X-8000, Radeon 8060S, 130 W TDP, dual 10 GbE, USB4 v2 at 80 Gbps): the review standard was ~$2,639 (128 GB + 2 TB), the "Max AI Compute Edition" in the store €3,639–3,799 equivalent. Its differentiator versus all other Strix Halo minis: an internal PCIe x16 slot — electrically only 4.0 x4, though, the same bandwidth as an OCuLink link. It is effectively a built-in version of the eGPU dock.

DEG1 and DEG2 eGPU docks: the DEG1 is an OCuLink dock (PCIe 4.0 x4, 64 GT/s) with an x16 slot and ATX PSU input for $109–139 — officially rated up to an RTX 4090. The DEG2 adds dual OCuLink plus Thunderbolt 5 (up to 80 Gbps) and a built-in M.2 SSD slot. Together with a Radeon AI PRO R9700 (retail price per market overviews ~€1,750–1,850) and an SFX PSU, that's the "poor man's Lucebox".

And it works — with documented friction. A thoroughly documented community build (mini PC + DEG1 + R9700 via an M.2 adapter) reaches about 50–80 tok/s with Qwen 3.8 27B (Q4) on the R9700 — well above what the APU alone manages. Models up to ~160 GB (3-bit quantization of large MoEs) only fit in the machine at all with 32 GB VRAM plus 128/192 GB unified memory. But: the M.2-to-OCuLink adapter proved unreliable at Gen4 speeds (crashes, PCIe dropouts) — the dock route is clearly preferable to the adapter route. That's the honest state of the DIY path: technically functional, but tinkerer-grade rather than production-grade.

RTX 3090 — the used-market classic with the best bandwidth deal

The RTX 3090 (Ampere, 2020) has vanished from retail, but it's everywhere on the used market — and it holds an ace: 24 GB of GDDR6X at 936 GB/s, faster than any discrete card in this comparison (the R9700 manages 640). Used, it costs roughly €700–900 — the best bandwidth per euro a single card can buy. The catches: 24 GB per card (large models need clusters), 350 W per card, no new-purchase warranty — and Ampere only partially receives the latest quantization and inference features (NVFP4 infrastructure, modern FP8 paths).

Our recipe database holds 65 configurations on the 3090 — from single cards to 4-way clusters. The highlights: a 4×3090 cluster decodes Qwen3.8-27B with DFlash2 at up to 244.8 tok/s — the fastest 27B measurement in the entire database, faster than anything Spark or Strix Halo achieve. And the CPU-MoE trick (experts in system RAM, hot layers on the 3090s) brings the 305B class (DeepSeek-V4-Flash, GLM-5.3-Flash, Qwen3.8-Flash-Next) onto 2×3090 — the DIY brother of the Lucebox concept, though at 18–37 tok/s instead of the promised hybrid speed. If you already own a PC, a single used 3090 is the fastest upgrade available.

The reference systems: Mac Studio M5 Ultra and NVIDIA DGX Spark

For context, the two systems our recipe database knows best: the Mac Studio M5 Ultra with 512 GB (36/80) combines the largest capacity in the field with ~1.2 TB/s of bandwidth — EDU price estimated at ~€15,400–15,500 (derived from the 96 GB delta of €3,960; the official 512 GB price is still pending, availability "late October"). The NVIDIA DGX Spark (GB10, 128 GB, ~273 GB/s, CUDA ecosystem) remains the most comfortable entry into the NVIDIA world, but at ~€45/GB it's the most expensive math in this overview.

Which models really fit? The math against 65 model families

Spec sheets say nothing about models. So we ran the math against our recipe database — 428 measured community configurations, 65 model families from 3B to 763B parameters. Rule of thumb: memory need ≈ parameters × bits/weight ÷ 8, plus ~6% overhead. Usable capacity: total RAM minus ~10% for OS and KV-cache headroom.

The key takeaways of that math:

  • The 128 GB class (DGX Spark, MS-S1 Max, MSI) ends at ~180B parameters in 4-bit (Qwen3.8-Flash-Next 180B-A4B fits; DeepSeek-V4-Flash 305B no longer). In 8-bit, 99B is the limit (Qwen3.5-122B-A10B just misses). The stars of this class: Qwen3.8-Flash-Next, gpt-oss-120B, Nemotron-3-Super-120B — and the 35B-A3B family (Qwen3.6, Agents-A1) at near-full precision.
  • The 192 GB class (ACEMAGIC F9A) fits DeepSeek-V4-Flash (305B) entirely in 4-bit, and MiniMax-M3 (428B), Qwen3.5 (397B) in 3-bit. 8-bit runs up to ~124B (Ling-3.0-flash, gpt-oss-120B).
  • Lucebox 224 GB (192+32 hybrid) adds both worlds: 4-bit up to 305B like the F9A, but with the hot layers additionally on 640 GB/s — and in 3-bit up to GLM-5.1 (478B). It's the only path past 500B models (Kimi-K2.6 519B) in 2-bit.
  • RTX 3090 clusters (48/96 GB VRAM) natively reach the same 180B limit in 4-bit as the 128 GB UMA class — but with 3.4× the bandwidth per card (936 GB/s). The 305B class only goes via the CPU-MoE trick (18–37 tok/s documented). In return, the database holds the fastest 27B measurement of all: 244.8 tok/s on 4×3090. If you just want the fastest single card, buy a used 3090: 24 GB at 936 GB/s for under €1,000.
  • The Mac Studio 512 GB takes everything: all 65 families in the database — up to DeepSeek-V4.1-Flash (763B), GLM-5.2 (753B) and GLM-5.3 (743B) — in 4-bit. Even Kimi-K2.6 (519B) in 8-bit. In 2-bit, 519B+ fits with a huge KV-cache remainder. If you want to run the current 300–760B flash generations (GLM-5.3-Flash, DeepSeek-V4-Flash, Qwen3.8-Flash-Next) unquantized or in 8-bit, you need the 512 GB class — full stop.

A second look behind it: this math only says "fits". Whether it's fast is decided by bandwidth — and there the 512 GB Mac wins again, because 1.2 TB/s outruns any 4-bit decode of the 273 GB/s class by a factor of four. The hybrid systems buy their speed in VRAM, but only for the model part that lives there.

The honest math: euros per gigabyte of AI memory

SystemTotal AI memoryReal-world bandwidthPrice (approx.)€/GB
Minisforum MS-S1 Max128 GB (UMA)~256 GB/s~$3,600–3,800 (128 GB)~$29
MSI PRO MAX EDGE AI+128 GB (UMA, up to 96 GB to iGPU)~256 GB/s~€4,500 (unverified)~€35
ACEMAGIC F9A192 GB (UMA)256–273 GB/s€6,399~€33
Lucebox Zero 495 (192)224 GB (192 UMA + 32 VRAM)640 GB/s (VRAM) + 273 GB/s (UMA)~€8,700~€39
RTX 3090 (used, 1×)24 GB (VRAM)936 GB/s~€800 (card only)~€33 (card only)
RTX 3090 cluster (2× / 4×, used)48 / 96 GB (VRAM)936 GB/s per card~€3,100 / ~€5,700 (complete)~€65 / ~€59 (complete)
NVIDIA DGX Spark128 GB (UMA, CUDA)~273 GB/s~$4,000–5,000 entry~$45
Mac Studio M5 Ultra 512512 GB (UMA)~1,200 GB/s~€15,400–15,500 EDU (estimated)~€30

Three readings of this table. First — the hybrid concepts pay for their discrete 640 GB/s card with the worst €/GB value, but get the only way to use both memory types in parallel. Second — Apple's 512 GB Studio is, per gigabyte, cheaper than any AMD box with less memory, despite being the most expensive system overall; the bandwidth widens the gap further. Third — Minisforum is the price-performance crown in the Strix Halo class, but without discrete VRAM and (on the x16 slot) without more than x4 linkage.

The software trap almost everyone overlooks

Hardware is half the way. Standard software — llama.cpp, PyTorch, classic ROCm — sees two separate memory pools in a hybrid system and has no idea what to do with them. Lucebox solves this with a proprietary engine whose hand-tuned kernels are calibrated exactly to its own hardware (the core repository is on GitHub, the fine-tuning is not). On the DIY path you configure the split yourself — the community docs show it's possible, but it's engineering, not plug-and-play.

On the Apple side the stack is more mature: MLX and TensorFold deliver byte-exact, optimized inference on unified memory without any user effort — our own stress test of the MiaAI TensorFold recipe on 2× DGX Spark reproduced the recipe numbers within ±5% (and the recipe database carries the fresh numbers: GSM8K 98.8%, HumanEval 95.7% on the current v1.5 engine). ROCm on RDNA4 and Strix Halo is much better in 2026 than in 2025, but remains the youngest ecosystem in the field.

Risk and buying honesty

  • Startup risk is real: Lucebox is young (Batch 3 sold out, Batch 4 only in November, price increase announced, 495 delivery January 2027). If you test it, use the "test box" strategy: one machine, wait, then scale — and pay with a credit card that has chargeback protection.
  • Pre-sale means pre-sale: the ACEMAGIC F9A is "late October" — exactly the same waiting window as Apple's 512 GB Studio. None of these boxes will be on a shelf tomorrow.
  • The April episode is the warning: Apple pulled the previous Studio's 512 GB option from the market entirely in April 2026 (memory shortage) and followed up in September with "months of tight supply". For the new 512 GB configuration that means: whoever wants it must be fast when it unlocks. We run a four-channel monitoring system with an alarm in ≤5 minutes for exactly this — the launch is a "now" signal, not a "this quarter" signal.
  • ToS and limits: everything here is reading, comparing and buying yourself. No auto-checkout, no bot — the alarm chain of Telegram, ntfy and email pushes with a tap-to-order link is enough.

Who is what for?

  • One model, maximum offline, Apple ecosystem — or the complete 300B+ flash generation: Mac Studio M5 Ultra 512. No other system delivers 512 GB at 1.2 TB/s; per gigabyte it's even the cheapest option in the table, and it's the only one that swallows every model family in the database in 4-bit.
  • x86/ROCm recipes, Linux, cluster kit, DeepSeek-V4-Flash in 4-bit: Lucebox Zero 495 (192 GB tier) — the only series hybrid with discrete VRAM. But as a single test box after independent reviews first, not as a cluster prepay.
  • Best €/GB in Strix Halo, European support, 305B/4-bit is enough: ACEMAGIC F9A — if you can live with 273 GB/s and no discrete VRAM.
  • Existing PC + maximum bandwidth for under €1,000: a used RTX 3090 (24 GB, 936 GB/s) — and if you want a cluster, the database holds the fastest 27B measurement of all (244.8 tok/s on 4×3090). The price: power (350 W/card), noise and used-market risk.
  • Build it yourself, experiment, budget Strix Halo: Minisforum MS-S1 Max (turnkey) or MS-S1/DEG1/DEG2 + R9700 (DIY) — the path with the most learning effect and the second-worst comfort.
  • CUDA ecosystem, agent stack, recipe variety: DGX Spark — our daily driver; software maturity beats the bad €/GB math.

FAQ

Which models run concretely on the 128 GB class (DGX Spark, MS-S1 Max)? By the capacity math: in 4-bit (EXL3/NVFP4/MLX) everything up to Qwen3.8-Flash-Next (180B-A4B) — meaning the entire 35B-A3B family (Qwen3.6, Agents-A1, Qwen-AgentWorld), Gemma 4 (26B/31B), Qwen3.8-27B, gpt-oss-120B, Nemotron-3-Super-120B-A12B and Qwen3.5-122B-A10B. The 305B class (DeepSeek-V4-Flash, GLM-5.3-Flash, MiMo-V2.6) only runs in 2-bit quantizations with a significant quality loss. In 8-bit it ends at Qwen3.5-122B-A10B (99B) — just short. This matches our recipe database: 45 of 65 families have documented 4-bit recipes on this class.

Is the Lucebox number (86 tok/s on DeepSeek V4 Flash 284B) credible? The number comes from the vendor and is technically plausible — a 284B MoE has a small active core that fits the 640 GB/s card well — but it hasn't been independently reproduced. The architecture math (hot layers in fast VRAM, long tail in unified memory) is internally consistent, and a community build of the same combination reaches 50–80 tok/s on a 27B dense model. We treat the 86 tok/s as a vendor claim, not a measurement, until the first independent test.

Why can't I just build a Strix Halo APU with a normal graphics card on a standard motherboard? Because the decisive part of the Strix Halo concept is the soldered LPDDR5X: the APU has no AM5 socket, and the full bandwidth only materializes when the memory chips sit millimetres from the die. Socketed DDR5 sticks reach a fraction of that. The only real paths are turnkey systems (Lucebox, ACEMAGIC, MSI, Minisforum MS-S1 Max with an x4 slot) or the dock approach (Minisforum DEG1/DEG2 with a separate PSU) — a classic DIY motherboard with both worlds doesn't exist.

Is Apple's "notify me" feature enough to get the 512 GB version? No, not for this scenario. Apple's watchers work at product-category level and with delay; if you want one specific configuration that will be scarce at launch — and where Apple showed in April 2026 that it can pull it from the market entirely — you need an alarm at URL and configurator level. A self-built watcher (a cron job checking the order URLs every few minutes, plus configurator-JSON analysis) beats every standard tool because it monitors exactly one SKU instead of a catalogue.

What does the internal PCIe x16 slot on the Minisforum MS-S1 Max really mean? Physically x16, electrically PCIe 4.0 x4 — the same bandwidth as an OCuLink dock. For inference that's acceptable: weights move into VRAM once, then decode runs from the card's local memory; only model loading and split synchronization use the link. For gaming with 200 W cards it's also usable, but the slot is dimensioned spatially and thermally for compact cards — the limits are in the tests.

The RTX 3090 has "only" 24 GB — why is it in this comparison? Because per euro it's the fastest card in the field: 936 GB/s of bandwidth — more than the R9700 (640) and more than triple the Strix Halo APUs — on the used market for roughly €700–900. For anyone who already owns a PC, it's the fastest upgrade available; in clusters (2×/4×) the database holds the fastest 27B measurement (244.8 tok/s on 4×3090). The limits are stated honestly: 24 GB per card, 350 W draw, used market with no warranty, and the 305B class needs the CPU-MoE trick (18–37 tok/s) instead of native capacity.

So which system is the best for local AI? There's no "best", only "best fitting": capacity and speed together are only for the Mac Studio 512 GB crowd; the AMD hybrids buy flexibility (x86, ROCm, Linux) at the price of software work and startup risk; Minisforum is the budget and tinkerer corner; the DGX Spark remains the CUDA comfort path. The decision should rest on measured tok/s for your own model spectrum, not on the marketing table — our recipe database already delivers such measurements for Apple and NVIDIA systems; the AMD newcomers of this article are still open there.

Sources

Relevant for your team? Let's talk for 30 minutes.

We sort out what of this actually works in your company — concrete, no slide marathon.

No commitment · 30 minutes · Proposal within 48 h