AMD opened IFA 2026 on September 4 with a deskside box aimed at people who already argue about HBM budgets: the Threadripper Halo Station. Jack Huynh called it the most powerful workstation in the world and said it can run models with more than a trillion parameters locally. I am not buying the slogan. I am buying the bill of materials — and the architectural bet against NVIDIA’s coherent DGX Station class.
This is a create-only read of what reporters on the floor documented today. AMD has not published a full SKU sheet, price, or availability window. Treat every “path to four” line as design intent until OEM configs land.
Spec sheet (show unit + platform class)
| Host CPU (assumed) | Threadripper PRO 9995WX — only 96-core TR PRO SKU; Huynh said “96 cores of Threadripper Pro” |
| CPU shape | 96C / 192T Zen 5 (“Shimada Peak”), boost ~5.4 GHz, up to 2 TB DDR5, 128 PCIe 5.0 lanes |
| Accelerators (show) | 2× Instinct MI350P (PCIe, CDNA 4 / TSMC N3) |
| Expand | Huynh: “path to four” MI350P → 576 GB HBM3E |
| HBM per card | 144 GB HBM3E at about 4 TB/s |
| Show HBM pool | 288 GB (2×144) |
| Compute (per STH) | ~2.3 PFLOPS MXFP8 per MI350P |
| Board power | up to ~600 W TBP per MI350P; CPU ~350 W class |
| Cooling | Liquid on CPU and accelerators (closed loops on the show unit) |
| Form | Tower workstation; show board reported by ServeTheHome as ASUS Pro WS WRX90E-SAGE SE |
| Price / ship / OEM | Not announced |
Primary coverage I used: ServeTheHome, StorageReview, TechPowerUp, Tom’s Hardware.
What the HBM math actually says
A trillion-parameter claim is a capacity claim until someone publishes tokens/sec and a harness. Rough weight floors:
| Precision | ~1T params | Fits 288 GB? | Fits 576 GB? |
|---|---|---|---|
| FP16 / BF16 | ~2 TB | No | No |
| FP8 / INT8 | ~1 TB | No | No |
| ~4-bit | ~500 GB | Borderline / no with KV | Yes for weights alone |
| More aggressive quant | lower | Maybe | Comfortable for weights |
So the four-card 576 GB story is the one that makes the “trillion-parameter local” line coherent at aggressive quantization. The two-card 288 GB show unit is a different product for most dense weights — closer to “very large MoE / heavy quant / multi-model deskside” than a full 1T FP8 resident set. StorageReview makes the same cut: ~500 GB for 1T at 4-bit, which wants the four-card pool, with the 2 TB host DDR5 and NVMe behind it for everything else.
Bandwidth is the other half. At ~4 TB/s HBM3E per MI350P, one card is in a different class from LPDDR5X “AI PC” unified memory. That is why this box sits above AMD’s Ryzen AI Halo / Max ladder and next to rack-adjacent Instinct iron, not next to a Strix/Kraken mini PC.
The DGX Station fight
NVIDIA’s recent deskside answer (GB300-class DGX Station / MSI XpertStation WS300 in StorageReview’s framing) bets on one big accelerator + Grace CPU sharing a large coherent memory pool over NVLink-C2C. AMD’s answer is the opposite:
- Discrete PCIe MI350P cards instead of a single coherent GPU+CPU package
- x86 Threadripper PRO host instead of Arm Grace (toolchain and driver story many of us already run)
- Scale by adding cards (2 → 4) instead of buying a denser single module
- Split pools: HBM on the Instincts, big DDR5 on the host — not one unified fabric
If your workload loves a single address space and NVLink-C2C bandwidth, the NVIDIA coherent pool still wins on paper. If your workload is ROCm + multi-GPU PCIe serving, or you refuse an Arm deskside toolchain, Halo Station is the first AMD-branded deskside that speaks that dialect with MI350-class HBM.
Nobody has published a public Halo Station tok/s table yet. Until then, the honest comparison is architectural, not a leaderboard.
Power, noise, and what was missing on stage
Two 600 W accelerators plus a ~350 W CPU is on the order of 1.5 kW of silicon before the rest of the box. Huynh’s liquid-cooling pitch (“stay super quiet”) is not marketing fluff at that TDP — it is the only way this stays a desk product instead of a lab cart. The show unit used per-device closed loops, not a shared custom loop, which ServeTheHome reads as COTS assembly more than a sealed appliance. Fine for a reference design. Awkward if you expected a sealed DGX-like brick with one service contract.
Also missing today:
- Named OEM / channel (direct AMD vs system builder)
- Interconnect detail between the four MI350Ps (PCIe topology, any Infinity Fabric / XGMI story)
- Storage layout, NIC (show unit had an SFP-class card — good, unspecified)
- Software image (ROCm version pin, Windows vs Linux, whether Project Zenith-class tooling shows up here or only on Ryzen AI Halo)
- Price. Street math on TR PRO 9995WX + 2–4× MI350P is “enterprise quote,” not Newegg.
How I would actually evaluate one
- Demand the four-card BOM in writing. If the ship SKU is two-card only, re-rate the product as a 288 GB HBM workstation, not a 1T weight box.
-
Measure multi-GPU serving, not a single-card demo.
vLLM/SGLang/llama.cpptensor-parallel across PCIe MI350Ps is the workload that decides whether “path to four” matters. - A/B against a coherent DGX Station class box on the same model and quant. Same tokenizer, same batch, same context. Record tok/s, TTFT, and PCIe traffic. Coherence vs discrete is the real question.
- Budget wall power and cooling. 1.5 kW-class silicon needs a circuit plan before a purchase order.
- Pin ROCm. Deskside Instinct only helps if your stack’s day-0 wheel matches the driver AMD ships on the image.
Where it sits in AMD’s ladder
| Rung | Product | Memory story |
|---|---|---|
| Edge / mini | Ryzen AI Halo / Max class | Unified LPDDR, tens–low hundreds of GB |
| Deskside HBM | Threadripper Halo Station | 288–576 GB HBM3E + up to 2 TB DDR5 |
| Rack | Instinct MI350X / platform | Full CDNA 4 rack SKUs |
That ladder is the strategic point. AMD is no longer answering NVIDIA only in the rack. It is answering Spark-class and Station-class deskside products with named HBM capacity.
Bottom line
Threadripper Halo Station is the first AMD deskside I would put on the same evaluation sheet as a GB300-class DGX Station — not because the keynote said “trillion parameters,” but because 144 GB HBM3E × N at MI350P bandwidth is a real capacity and bandwidth story for local and air-gapped inference. The open questions are boring and decisive: four-card ship config, multi-GPU interconnect behavior, ROCm image, and price. Until those land, treat today as a serious reference design announcement, not a buy.
I will revisit when an OEM posts a configure-to-order page with a wattage rating and a four-card option you can actually order.
Anpoo Sivanadi, Staff Software Engineer