Blog

AI Accelerator Options in 2026: NVIDIA, AMD, Google, AWS, Cerebras, Tenstorrent

Twelve accelerators on memory per dollar, bandwidth per compute, MLPerf tokens per dollar, domain size, and software status.

Published Reading ...
AI Accelerator Options in 2026: NVIDIA, AMD, Google, AWS, Cerebras, Tenstorrent

Buying or renting AI compute in 2026 is a memory decision, not a FLOPS decision. This post puts twelve accelerators on the same axes using only vendor datasheets, public price lists, MLCommons result files, and SEC filings, checked on 9 September 2026. The derived metrics are ours. The inputs are not.

The chips

ChipVendorMemoryBandwidthDense FP8Scale-up domainStatus
H200NVIDIA141 GB HBM3e4.8 TB/s2.0 PFLOPS8Shipping, sold out in cloud per Q4 FY26 call
B200 (HGX)NVIDIA180 GB HBM3e8.0 TB/s4.5 PFLOPS8Shipping
B300 (HGX)NVIDIA270 GB HBM3e8.0 TB/s5.0 PFLOPS8Shipping
GB300 NVL72NVIDIA288 GB HBM3e8.0 TB/s5.0 PFLOPS72Shipping, most of NVIDIA volume
Vera Rubin NVL72NVIDIA288 GB HBM422 TB/sNot disclosed72Production shipments began August 2026
MI355XAMD288 GB HBM3e8.0 TB/s5.03 PFLOPS8Shipping
MI455X (Helios)AMD432 GB HBM423.3 TB/s20.1 PFLOPS72Ramp 2H 2026, no MLPerf, not in ROCm notes
Ironwood TPU7xGoogle192 GiB HBM7.38 TB/s4.61 PFLOPS9,216 podGA 31 March 2026, two zones
Trainium3AWS144 GB HBM3e4.9 TB/s2.52 PFLOPS64 to 144GA December 2025, no public price
Blackhole p150Tenstorrent32 GB GDDR6 + 180 MB SRAM512 GB/s0.66 PFLOPS (block FP8)Ethernet, 3.2 Tb/s per cardIn stock, $1,399
WSE-3 (CS-3)Cerebras44 GB SRAM21 PB/s125 PFLOPS (FP16, vendor)WaferShipping, CS-4 shipments from Q3 2026
Gaudi 3Intel128 GB HBM2e3.7 TB/s1.68 PFLOPS8Listed, written off in the FY2025 10-K

Table 1. Vendor-stated per-chip figures. NVIDIA and AMD dense figures are from spec tables. Rubin FP8 is not published. Cerebras does not state precision for its 125 PFLOPS. Tenstorrent publishes block FP8 only.

Groq is absent from the chip table on purpose. Its next chip ships as the NVIDIA Groq 3 LPX under a licence, and Groq’s own newsroom describes GroqCloud as running NVIDIA hardware alongside LPUs. It is a cloud provider now, not an accelerator option you can buy.

Bandwidth per unit of compute

Decode is bound by how fast weights and KV cache move, not by how fast the tensor cores multiply. The ratio below is HBM bandwidth in TB/s divided by dense FP8 PFLOPS. Lower means the chip is more compute-heavy relative to its memory, and spends more of its life waiting on HBM at inference batch sizes.

ChipTB/s per dense FP8 PFLOPS
H2002.40
Gaudi 32.20
Trainium31.94
B2001.78
Ironwood1.60
B3001.60
MI355X1.59
MI455X1.16

Table 2. Derived from Table 1. Every new generation lowers this ratio. FP8 and FP4 throughput grew faster than HBM bandwidth on every vendor’s roadmap.

The practical reading: a B300 has 2.5 times the FP8 of an H200 but only 1.67 times the bandwidth. For a 70B model at moderate batch, the bandwidth is what you are paying for. That makes memory per dollar the first sort key.

Memory per rented dollar

Public on-demand prices are rare. These are the ones that exist, from the provider’s own price page or price API on the day.

ChipProvider$/GPU-hourGB per $/hrTB/s per $/hr
MI355XVultr (bare metal, 8 GPU)2.591113.09
MI325XVultr (bare metal, 8 GPU)4.62551.30
Trainium2AWS Capacity Blocks2.24431.30
IronwoodGoogle, 3-year commit5.40361.37
B300Nebius7.85341.02
MI355XOracle OCI8.60330.93
H200Nebius4.50311.07
B200Lambda6.69271.20
GB200 NVL72CoreWeave10.50180.76
IronwoodGoogle, on demand12.00160.62

Table 3. Prices as listed on 9 September 2026. Vultr’s MI355X price sits below its own MI325X price and may be promotional. GB300 NVL72, Trainium3, and Rubin have no public hourly price anywhere.

Two things stand out. AMD’s MI355X on Vultr delivers two and a half times the memory bandwidth per dollar of the best NVIDIA part with a public price. And Google’s on-demand Ironwood is the most expensive memory on the list, while the same chip on a three-year commitment is competitive. The commitment, not the silicon, is the product.

MLPerf tokens per dollar

MLCommons publishes system totals. We divided each 8-GPU server result for Llama 2 70B by eight, then by the hourly price above.

Figure 1. Derived from MLPerf Inference v6.0 (B200, B300, MI355X) and v5.1 (H200, MI325X) server results divided by GPU count and price. Exact values in Table 4.

ChipMLPerf roundTokens/s per GPU$/hrTokens/s per $
B300v6.0, NVIDIA DGX B30013,4147.851,709
B200v6.0, HPE 8x B20012,9546.691,936
MI355Xv6.0, AMD 8x MI355X12,5358.60 (OCI) or 2.59 (Vultr)1,458 or 4,840
GB300 NVL72v6.0, NVIDIA, per GPU of 7212,059Not publicNot computable
H200v5.1, ASUS 8x H2004,2744.50950
MI325Xv5.1, AMD 8x MI325X4,0034.62867

Table 4. Per-GPU figures are our division of the published system total. Result IDs are in the MLCommons repositories for v5.1 and v6.0.

On this benchmark the MI355X is within 7 percent of the B300 per GPU. Price decides the rest. Ironwood, Trainium3, Gaudi 3, and Helios have no MLPerf inference submission in any round through v6.0, so they cannot be placed on this chart. Cerebras and Tenstorrent do not submit either.

Scale-up domain

The number of accelerators that share memory at full interconnect speed sets the largest model you can serve without crossing a network hop.

PlatformDomainPer-accelerator linkMemory in domain
HGX B300, MI355X UBB, Gaudi 381.8 TB/s, 1.08 TB/s, 24x200 GbE2.1 TB, 2.3 TB, 1 TB
GB300 NVL72721.8 TB/s NVLink 520 TB
Vera Rubin NVL72723.6 TB/s NVLink 620.7 TB
Helios (MI455X)723.6 TB/s UALink over Ethernet31 TB
Trainium3 UltraServer64 or 1442 TB/s NeuronLink-v49.2 TB or 20.7 TB
Ironwood pod9,2161.2 TB/s ICI, 3D torus1.77 PB
Galaxy Blackhole32 chips, Ethernet800 GbE ports1 TB GDDR6 + 6.2 GB SRAM
CS-3 clusterUp to 2,048 wafersFabricMemoryX up to 1.2 PB

Table 5. Vendor-stated. Helios matches Rubin on link bandwidth and exceeds it on memory, on paper. No customer shipment has been announced as of this writing.

Software status

StackState as of 9 September 2026Documented limitation
CUDA 13.3, TensorRT-LLM, DynamoMature. Every MLPerf NVIDIA entry uses itRubin toolchain not yet in release notes
ROCm 10.0.0 (26 August 2026)vLLM 0.27, SGLang 0.5.15, PyTorch 2.13 on MI355X9 to 25 percent training slowdown on gfx950 from an AOTriton kernel choice, workaround in notes. MI455X not in supported list
Neuron SDK 2.32Trainium3 supported since December 2025vLLM Neuron plugin is Beta. DeepSeek V3 and R1 are not production models
JAX and PyTorch on TPU7xSupportedvLLM TPU plugin has no GA label. TensorFlow not supported
tt-metal, tt-forge (Apache 2.0)Llama 3.1 8B “Complete” on BlackholeLlama 3.3 70B is “Functional”, not “Complete”, on Blackhole. 70B numbers published for Wormhole Galaxy only: 72.5 tokens/s per user
CerebrasInference as API. gpt-oss-120b at $0.35 in, $0.75 out per millionLlama 3.3 70B removed from the public price list. No 2026 independent speed measurement found
Intel Gaudi 1.24vLLM plugin v0.26 upstreamLazy mode deprecated. Intel’s 10-K calls the Gaudi effort “unsuccessful”

Table 6. From release notes, model support matrices, and filings. “Complete” and “Functional” are Tenstorrent’s own support tiers.

What we would pick

  • Renting for 70B-class inference today. MI355X on Oracle or Vultr, or B200 on Lambda. Run the per-dollar table with your provider’s quote. The silicon gap is 7 percent. The price gap is up to 3x.
  • Renting for models that need more than 2 TB in one domain. GB300 NVL72 is the only 72-GPU domain shipping in volume with a benchmark record. Rubin started shipping in August 2026 and has no MLPerf entry. Helios has neither shipments nor ROCm notes yet.
  • Committed capacity for a year or more. Ironwood at the three-year rate is competitive on memory per dollar and comes with the largest domain in the industry. On demand it is the worst value here.
  • Owning hardware under $10,000. Tenstorrent is the only option with an open stack and a public price. Treat 70B as experimental on Blackhole until the matrix says “Complete”.
  • Lowest latency per user. Cerebras and Groq as APIs, not as hardware purchases. The Groq silicon roadmap now belongs to NVIDIA’s catalogue.
  • Gaudi 3. No. Intel wrote off $1.3 billion of inventory across 2024 and 2025 and its 2026 earnings calls do not mention it.

Scope

We measured nothing ourselves. Every performance number is a vendor datasheet or an MLCommons result file, and every price is a public list price on one day. We did not test training, and the bandwidth ratio in Table 2 is a heuristic for decode, not a benchmark. Cloud prices for GB300, Rubin, Trainium3, MI455X, and Groq 3 LPX do not exist publicly, so those chips cannot be ranked on cost. Tenstorrent’s Galaxy Blackhole list price moved from $110,000 in April 2026 to $160,000 in September 2026 on the vendor’s own pages. We will rerun the tables when MLPerf v6.1 lands.

Choosing accelerators for a build? Book a free 30 minute consultation.