AMD's Ryzen AI Halo puts 128GB of unified memory on your desk — and runs headfirst into a bandwidth wall
AMD's $3,999 answer to Nvidia's DGX Spark pairs a 16-core Zen 5 CPU with 128GB of unified LPDDR5X. It can hold enormous models; its 256 GB/s bus decides how fast it serves them.

AMD's answer to Nvidia's DGX Spark is here. The $3,999 Ryzen AI Halo — a hardcover-book-sized desktop built on the "Strix Halo" Ryzen AI Max+ 395 — pairs 128 GB of unified LPDDR5X with a 16-core Zen 5 CPU, a 40-CU Radeon 8060S iGPU and an XDNA 2 NPU. It's a memory-capacity machine, not a memory-bandwidth one: it can hold enormous models, but its 256 GB/s bus is the ceiling on how fast it serves them.
AMD builds a DGX Spark of its own
AMD is shipping the Ryzen AI Halo Developer Platform, a turnkey mini-PC aimed squarely at engineers who want to run and prototype large local models without renting cloud GPUs. It went up for pre-order in June exclusively through Micro Center and hits shelves July 10 at $3,999 for the 128 GB / 2 TB configuration, per StorageReview. The pitch is explicitly comparative: Tom's Hardware titled its review "AMD builds a DGX Spark of its own," and AMD undercuts Nvidia's DGX Spark — recently repriced to roughly $4,699 amid the LPDDR5X/NAND crunch — by about $700, while adding a Windows 11 option the Linux-only Spark doesn't offer.
What's actually inside
Cross-checked against AMD's product pages and independent trackers — not the marketing sheet.
| Component | Spec | Notes / confidence |
|---|---|---|
| CPU | 16 cores / 32 threads, full Zen 5, up to 5.1 GHz (TSMC N4) | Confirmed (AMD). Base 3.0 GHz; some trackers show a 2.0 GHz floor. |
| iGPU | Radeon 8060S, 40 CUs (2560 shaders), RDNA 3.5, up to 2.9 GHz | Confirmed. ~14.85 TFLOPS FP32 / ~29.7 TFLOPS FP16 (dense). A ~59 TFLOPS FP16 figure is the dual-issue marketing peak. |
| NPU | XDNA 2, 50 TOPS INT8 (dense) | Confirmed; also 50 TOPS Block FP16. Sparse can be ~2× but AMD doesn't headline it. |
| "Combined AI" | 126 TOPS (NPU + GPU + CPU) | AMD's quoted figure, but a marketing composite across precisions — no single engine hits it. |
| Memory | 128 GB unified LPDDR5X-8000, 256-bit | 256 GB/s theoretical; ~210–215 GB/s attainable. Up to 96 GB assignable as VRAM. |
| Power | 120 W sustained, ~140 W bursts | Per MicroCenter / ServeTheHome. Die is configurable 45–120 W. |
| Chassis | 2 TB M.2 SSD; ~150 × 150 × 45 mm; < 1.2 kg; 10 GbE | Confirmed (ServeTheHome). |
256 GB/s is roughly one-twentieth of an H100's bandwidth. Strix Halo trades bandwidth for capacity — that's the whole story.
The bandwidth wall — and where LTT Labs pinned it
The most rigorous teardown comes from LTT Labs,
and most of the review pack builds on it. Using llama-bench (from llama.cpp) with the default pp512/tg128 profile, LTT Labs ran
Qwen 3.6 35B, Gemma 4 31B and GLM 4.7 Flash. Their load-bearing finding: on the dense Gemma 4 model, Apple's Mac Studio generated text 2–3×
faster — attributed to bandwidth, since the Mac's unified memory runs at ~800 GB/s versus the Halo's 256 GB/s.
The bright spot is efficiency. Running a model purely on the XDNA 2 NPU — CPU and GPU near idle — total chip power peaked at just 35 W while sustaining ~20 tokens/second on a ~20-billion-parameter model: a genuinely strong perf-per-watt story for always-on local inference. LTT Labs' broader thesis, echoed by Phoronix, is that the real value is the AMD Ryzen AI software stack and its pre-validated "Best Known Configurations," not the raw silicon.
Reconciling the token numbers
Different outlets quote different tokens/second because they test different models on different engines.
The pattern holds across the wider data: Mixture-of-Experts models fly (community labs measured Qwen3 30B-A3B MoE at ~86 t/s, GPT-OSS 120B at ~53 t/s decode) because only ~3 B parameters activate per token, while same-size dense models get throttled by the 256 GB/s wall. StorageReview's vLLM serving tests had the Halo trailing the DGX Spark 2×–4× at higher concurrency — stretching to 8.8× on prefill-heavy GPT-OSS 120B work, the Halo's weakest, compute-bound axis. The "up to 200 billion parameters" line is a capacity claim, not a throughput promise: it fits; it won't fly.
Halo vs DGX Spark — and vs renting
| Ryzen AI Halo | DGX Spark | |
|---|---|---|
| Price | $3,999 | ≈ $4,699 |
| Unified memory | 128 GB | 128 GB |
| Operating system | Win 11 or Linux | Linux (DGX-OS) only |
| Networking | 10 GbE | 200 GbE ConnectX |
| Decode, 120B | ~34 t/s | ~39 t/s |
| vLLM serving | baseline | 2–4× faster |
Against the Spark, the Halo is cheaper and dual-boots Windows 11 or AMD's Linux image, and XDA notes it does much of what the Spark does at a fraction of the power — but raw inference "trails DGX Spark and GB10 boxes." Against a Mac Studio, the Mac wins dense-model generation on bandwidth while the Halo undercuts it on price with an open ROCm/Linux stack. And against renting: a single H100 hour buys ~20× the bandwidth, so the Halo is for fitting 70B–120B-class models locally and iterating cheaply — not serving them at scale.
How we verified this
Every number was routed through our hardware desk. Here's what's solid, what's marketing, and what we couldn't nail down.
Corroborated across independent sources
Vendor composites and capacity claims
Unverified or disputed
The verdict
Sources
- MicroCenter — Hands-On with the AMD Ryzen AI Halo
- LTT Labs — AI Dev Kit, Batteries Included
- Tom's Hardware — AMD builds a DGX Spark of its own
- Tom's Hardware — AMD challenges Nvidia's DGX Spark
- StorageReview — A Dual-OS, 200B-Parameter Desktop
- ServeTheHome — Developer System Review
- Phoronix — Open-source software review
- XDA — A fraction of the DGX Spark's power
- wccftech — AMD tackles Nvidia's DGX Spark
- linuxcompatible.org — Pocket-sized dev kit for local LLMs
- Level1Techs — Strix Halo LLM benchmark results
- AMD — Ryzen AI Max+ 395 product page
Hardware specs verified by Flopper.io's hardware desk against AMD's primary pages and independent trackers, not vendor marketing. Prepared for Flopper.io — "Compare the GPUs Powering AI." July 10, 2026.
Keep reading
- DatabaseAI desktops & personal dev kits, compared
DGX Spark, GB10 systems and Strix Halo mini-PCs side by side, with normalized specs.
- GuideApple Silicon explained
Why the Mac Studio's ~800 GB/s unified memory pulls ahead on dense-model inference.
- DatabaseCompare datacenter GPUs by precision
See how an H100 or H200 stacks up when throughput — not desk-side capacity — is the goal.