AMD MI455X Ships: 432GB HBM4 Per GPU Rewrites the Memory Race

AMD’s Instinct MI455X is officially shipping, and the specification that will define this hardware generation is not compute — it is memory. Each accelerator carries 432GB of HBM4 across twelve 36GB stacks on a 2,048-bit memory interface, with peak bandwidth of roughly 23.3 TB/s. Against Nvidia’s Blackwell Ultra B300 at 288GB of HBM3e, AMD has opened a 50%+ memory-capacity advantage per GPU, and that advantage changes the mathematics of serving frontier-scale models.

The package itself is a study in modern chiplet engineering: 320 billion transistors spread across eight accelerator complex dies built on TSMC’s 2nm node, paired with two 3nm fabric-and-cache dies and two 3nm I/O dies. The eight compute die tile out 256 work-group processors executing Wave32 instructions — a datacenter-first design that has drifted deliberately away from AMD’s graphics-derived Radeon architecture. Peak engine clock is 2.4GHz, and the part is direct-liquid-cooled by necessity, not preference. Formally launched on July 23, 2026 as the flagship of the MI400 series, it is now the accelerator AMD’s entire rack-scale story hangs on.

Why memory capacity became the battleground

For most of the AI hardware era, marketing revolved around FLOPS. That made sense when models were small enough to fit comfortably on a handful of accelerators. It stopped making sense the moment frontier models crossed hundreds of billions of parameters. When a model, its KV-cache and its activations no longer fit in the memory attached to your compute, you pay a tax that no amount of raw FLOPS can remove: tensor-parallel sharding across more devices, interconnect traffic on every layer boundary, and idle compute bubbles while GPUs wait on each other.

Memory capacity attacks that tax directly. A GPU that holds more of the model locally needs fewer partners in the sharding group, generates less cross-device traffic, and delivers more usable throughput per watt. This is why AMD’s pitch for the MI455X is not ‘more FLOPS than Nvidia’ but ‘fewer GPUs per model, lower total cost per served token.’ On DeepSeek-V4-Flash — an open-weight model AMD uses as a public reference — the company claims up to 34x higher token throughput at high interactivity and up to 18x lower token cost versus the prior MI355X generation.

The generational math backs the ambition: versus the MI355X, AMD claims 1.5x the memory, 2.9x the memory bandwidth and 4x the MXFP4/MXFP8 throughput. Against Nvidia’s next-generation Rubin GPU — 288GB of HBM4 at a claimed 22 TB/s — AMD’s per-GPU figures lead on both capacity and bandwidth, which is precisely the comparison AMD’s own spec sheets now draw.

Inside the silicon

  • 432GB HBM4 across 12 stacks of 36GB on a 2,048-bit interface, ~23.3 TB/s peak bandwidth, with a 192MB L2 cache split into two 96MB sections
  • 40.26 PFLOPS peak compute at MXFP4 precision, roughly 20 PFLOPS at FP8 and MXFP6, 5 PFLOPS at FP16/BF16 matrix
  • 2nm + 3nm chiplet architecture — eight N2 compute die, two N3P fabric/cache die, two N3P I/O die, 320 billion transistors
  • Scale-up fabric — 3.6 TB/s per-GPU UALoE bi-directional bandwidth inside the tray; 600 GB/s UALink scale-out for multi-rack topologies
  • Reliability and virtualization — full-chip ECC with page retirement, SR-IOV for splitting one GPU across tenants, and an RAS feature set aimed at always-on serving
  • Three SKUs: MI455X (flagship, ships in Helios racks), MI450X (large-scale training/inference), MI430X (HPC and sovereign AI, up to 288 TFLOPS FP64)

One detail with real operational consequences is memory partitioning. NPS1 mode interleaves addresses across all twelve stacks for maximum bandwidth on shared workloads; NPS2 divides the GPU into two domains of six stacks each, keeping traffic local to each die group. Inference operators tuning per-request isolation will care which mode their serving stack selects.

The Helios rack: 72 GPUs, 31TB of pooled memory

Individual chips only tell half the story. AMD’s Helios rack bundles 72 MI455X accelerators into a single liquid-cooled system built on Meta’s Open Rack Wide standard, which AMD co-submitted to the Open Compute Project. Aggregate figures: up to 2.9 exaflops of FP4 compute, 31TB of total HBM4, and 1.7 petabytes per second of memory bandwidth. Nvidia’s comparable NVL72-class pod built on B300 parts carries roughly 20TB — the memory gap is architectural, not incremental.

The fabric numbers are where the design gets aggressive: roughly 260 TB/s of aggregate scale-up bandwidth per rack, with about 43 TB/s of scale-out connectivity handled by AMD’s Pensando networking silicon. The rack is built from repeatable four-GPU trays, which AMD argues is what makes exaflop-class results predictable and serviceable rather than a one-off benchmark configuration.

Oracle’s commitment to deploy 50,000 MI450-class GPUs in Helios racks — covered later this month — is the first hyperscale validation of this design at production scale.

The HBM4 supply picture behind the specs

None of these numbers exist without memory partners, and the supply side is scaling aggressively. Micron is ramping HBM output toward roughly 100,000 wafers per month by the end of 2026 — nearly double last year’s 40,000-50,000 — with 12-high HBM4 moving from 20-30% of its mix early in the year toward about 50% by year-end; its cumulative HBM4 revenue crossed $1 billion by June. Samsung has allocated roughly half of its 150,000 monthly HBM wafers to HBM4 after shipping the industry’s first parts in February, targeting $10 billion in HBM4 sales this year. SK hynix remains the incumbent leader and Nvidia’s primary supplier. The practical takeaway for buyers: at 432GB per GPU, every Helios rack consumes enormous HBM4 allocation, so supply commitments and delivery schedules deserve as much negotiation attention as unit pricing.

What buyers should actually do with this information

If your workloads are long-context inference, large-model serving or retrieval-heavy agents, the memory-capacity argument deserves a serious benchmark. Run your real serving traffic on both architectures if you can; per-token economics diverge sharply with context length. If your workloads are training-dominated with mature CUDA tooling, Nvidia’s ecosystem gravity remains real. The honest position for 2026: AMD has made itself benchmarkable, and that alone changes procurement conversations that were one-sided for two years.

Common mistakes when evaluating memory-rich accelerators

  • Comparing peak bandwidth only — realized throughput depends on partitioning mode, kernel maturity and how your serving stack shards; a paper 23.3 TB/s means little if utilization sits at 40%
  • Benchmarking short contexts — the 432GB advantage shows up in long-context KV-cache residency and large-batch serving, not in 2K-token single-stream tests
  • Ignoring the interconnect — a memory-rich GPU still bubbles if cross-rack traffic dominates; test at the topology you will actually deploy
  • Underestimating software porting — ROCm coverage of PyTorch, vLLM and SGLang is broad and improving, but migration effort belongs in the total-cost model

Build your own AI workstation — high-VRAM GPU options for local inference

View on Amazon

Leave a Reply

Your email address will not be published. Required fields are marked *