Hardware · Review

AMD's Ryzen AI Halo puts 128GB of unified memory on your desk — and runs headfirst into a bandwidth wall

AMD's $3,999 answer to Nvidia's DGX Spark pairs a 16-core Zen 5 CPU with 128GB of unified LPDDR5X. It can hold enormous models; its 256 GB/s bus decides how fast it serves them.

By Flopper Hardware Desk 6 min read 12 sources
AMDStrix HaloLocal AIDGX Spark
AMD Ryzen AI Halo — Ryzen AI Max+ 395, 128 GB unified memory, 256 GB/s, 126 TOPS
Price
$3,999
CPU
16C / 32T
Zen 5, 5.1 GHz
iGPU
40 CU
Radeon 8060S
NPU
50 TOPS
XDNA 2, INT8
Memory
128 GB
LPDDR5X-8000
Bandwidth
256 GB/s
The ceiling

AMD's answer to Nvidia's DGX Spark is here. The $3,999 Ryzen AI Halo — a hardcover-book-sized desktop built on the "Strix Halo" Ryzen AI Max+ 395 — pairs 128 GB of unified LPDDR5X with a 16-core Zen 5 CPU, a 40-CU Radeon 8060S iGPU and an XDNA 2 NPU. It's a memory-capacity machine, not a memory-bandwidth one: it can hold enormous models, but its 256 GB/s bus is the ceiling on how fast it serves them.

01
The launch

AMD builds a DGX Spark of its own

AMD is shipping the Ryzen AI Halo Developer Platform, a turnkey mini-PC aimed squarely at engineers who want to run and prototype large local models without renting cloud GPUs. It went up for pre-order in June exclusively through Micro Center and hits shelves July 10 at $3,999 for the 128 GB / 2 TB configuration, per StorageReview. The pitch is explicitly comparative: Tom's Hardware titled its review "AMD builds a DGX Spark of its own," and AMD undercuts Nvidia's DGX Spark — recently repriced to roughly $4,699 amid the LPDDR5X/NAND crunch — by about $700, while adding a Windows 11 option the Linux-only Spark doesn't offer.

02
Silicon

What's actually inside

Cross-checked against AMD's product pages and independent trackers — not the marketing sheet.

ComponentSpecNotes / confidence
CPU16 cores / 32 threads, full Zen 5, up to 5.1 GHz (TSMC N4)Confirmed (AMD). Base 3.0 GHz; some trackers show a 2.0 GHz floor.
iGPURadeon 8060S, 40 CUs (2560 shaders), RDNA 3.5, up to 2.9 GHzConfirmed. ~14.85 TFLOPS FP32 / ~29.7 TFLOPS FP16 (dense). A ~59 TFLOPS FP16 figure is the dual-issue marketing peak.
NPUXDNA 2, 50 TOPS INT8 (dense)Confirmed; also 50 TOPS Block FP16. Sparse can be ~2× but AMD doesn't headline it.
"Combined AI"126 TOPS (NPU + GPU + CPU)AMD's quoted figure, but a marketing composite across precisions — no single engine hits it.
Memory128 GB unified LPDDR5X-8000, 256-bit256 GB/s theoretical; ~210–215 GB/s attainable. Up to 96 GB assignable as VRAM.
Power120 W sustained, ~140 W burstsPer MicroCenter / ServeTheHome. Die is configurable 45–120 W.
Chassis2 TB M.2 SSD; ~150 × 150 × 45 mm; < 1.2 kg; 10 GbEConfirmed (ServeTheHome).
256 GB/s is roughly one-twentieth of an H100's bandwidth. Strix Halo trades bandwidth for capacity — that's the whole story.
The core trade-off
03
Inference

The bandwidth wall — and where LTT Labs pinned it

The most rigorous teardown comes from LTT Labs, and most of the review pack builds on it. Using llama-bench (from llama.cpp) with the default pp512/tg128 profile, LTT Labs ran Qwen 3.6 35B, Gemma 4 31B and GLM 4.7 Flash. Their load-bearing finding: on the dense Gemma 4 model, Apple's Mac Studio generated text 2–3× faster — attributed to bandwidth, since the Mac's unified memory runs at ~800 GB/s versus the Halo's 256 GB/s.

Fig. 1 Memory bandwidth, GB/s. The Halo holds bigger models than a Mac Studio but feeds them over a far narrower pipe; a datacenter H100 (HBM3) is another order up again. Sources: LTT Labs (Halo, Mac), NVIDIA (H100).

The bright spot is efficiency. Running a model purely on the XDNA 2 NPU — CPU and GPU near idle — total chip power peaked at just 35 W while sustaining ~20 tokens/second on a ~20-billion-parameter model: a genuinely strong perf-per-watt story for always-on local inference. LTT Labs' broader thesis, echoed by Phoronix, is that the real value is the AMD Ryzen AI software stack and its pre-validated "Best Known Configurations," not the raw silicon.

04
Throughput

Reconciling the token numbers

Different outlets quote different tokens/second because they test different models on different engines.

Fig. 2 Decode throughput on a 120B-parameter model, tokens/second. On decode both boxes are bandwidth-bound and land within spitting distance. Source: Tom's Hardware.

The pattern holds across the wider data: Mixture-of-Experts models fly (community labs measured Qwen3 30B-A3B MoE at ~86 t/s, GPT-OSS 120B at ~53 t/s decode) because only ~3 B parameters activate per token, while same-size dense models get throttled by the 256 GB/s wall. StorageReview's vLLM serving tests had the Halo trailing the DGX Spark 2×–4× at higher concurrency — stretching to 8.8× on prefill-heavy GPT-OSS 120B work, the Halo's weakest, compute-bound axis. The "up to 200 billion parameters" line is a capacity claim, not a throughput promise: it fits; it won't fly.

05
The decision

Halo vs DGX Spark — and vs renting

Ryzen AI HaloDGX Spark
Price$3,999≈ $4,699
Unified memory128 GB128 GB
Operating systemWin 11 or LinuxLinux (DGX-OS) only
Networking10 GbE200 GbE ConnectX
Decode, 120B~34 t/s~39 t/s
vLLM servingbaseline2–4× faster

Against the Spark, the Halo is cheaper and dual-boots Windows 11 or AMD's Linux image, and XDA notes it does much of what the Spark does at a fraction of the power — but raw inference "trails DGX Spark and GB10 boxes." Against a Mac Studio, the Mac wins dense-model generation on bandwidth while the Halo undercuts it on price with an open ROCm/Linux stack. And against renting: a single H100 hour buys ~20× the bandwidth, so the Halo is for fitting 70B–120B-class models locally and iterating cheaply — not serving them at scale.

06
Provenance

How we verified this

Every number was routed through our hardware desk. Here's what's solid, what's marketing, and what we couldn't nail down.

Verified

Corroborated across independent sources

Price ($3,999), July 10 Micro Center availability, and the core specs (16C/32T Zen 5, Radeon 8060S 40 CU, XDNA 2 50 TOPS, 128 GB LPDDR5X-8000, 256 GB/s) check out against AMD's pages and multiple reviewers — and re-derive cleanly (8000 MT/s × 256-bit ÷ 8 = 256 GB/s). The DGX Spark framing and LTT Labs' bandwidth findings are consistent across the pack.
Claimed

Vendor composites and capacity claims

The 126 TOPS "combined AI" sums dissimilar engines across precisions; no single engine reaches it. "Up to 200 billion parameters" is a memory-capacity claim at aggressive quantization. The ~59 TFLOPS FP16 iGPU figure is the dual-issue best case — the honest dense number is ~29.7 TFLOPS.
Undisclosed

Unverified or disputed

MicroCenter's "45 tokens/second" headline lacks a stated model/quant. The DGX Spark price varies by source (~$4,699 vs $4,679). Real-world bandwidth (~210–215 GB/s) and community tokens/sec figures are forum-measured, not vendor-certified. The four primary anchors (LTT Labs, MicroCenter, Phoronix, Tom's) sit behind bot protection and were corroborated via secondary summaries — spot-check the live pages before republishing.
07
Bottom line

The verdict

Keep reading