NVIDIA · Training & inference
B200 Blackwell
Perf
20 PFLOPS FP4
Memory
192 GB HBM3e
Power
1000 W
The default unit of frontier training. Multi-quarter backlog continues.
Frontier intelligence is now bottlenecked by megawatts and HBM stacks. Here is the board of the chips that matter and the labs they are pointed at.
The default unit of frontier training. Multi-quarter backlog continues.
Sampling Q3 ’26; the first chip designed end-to-end with AI co-pilots.
Gemini 3 trained here. Pods of 9,216 deliver 42 EFLOPS in one fabric.
Anthropic's Project Rainier — 400k chips committed.
The first credible second source. ROCm now usable for PyTorch out of the box.
Targets Copilot & Azure OpenAI inference, not training.
Token latency leader; 3,200 tok/s on a 70B class model.
Wafer-scale. 900,000 cores on a single piece of silicon.