NVIDIA Introduces the Hopper GPU Architecture
Published:
NVIDIA CEO Jensen Huang announced the Hopper GPU architecture at GTC (GPU Technology Conference) on March 22, 2022, during a keynote delivered in a chip fab-styled theatrical setting featuring a life-size H100 GPU prop. The flagship H100 (die GH100) was manufactured on TSMC N4 (4nm class) and contained 80 billion transistors — 2.6× the transistor count of the Ampere A100 (54 billion on Samsung 8nm). H100 SXM5 (the high-performance server module) shipped with 80 GB HBM3 at 3.35 TB/s memory bandwidth (vs A100’s 80 GB HBM2e at 2.0 TB/s), and delivered approximately 3.35 TFLOPS BF16/FP16 dense training throughput (vs 312 TFLOPS for A100) — roughly a 3× improvement in raw training arithmetic. The 4th-generation Tensor Cores added FP8 support, and the Transformer Engine (a hardware-software co-design) automatically selected between FP8 and FP16 precision on a per-layer, per-tensor basis during transformer training — using FP8 where precision loss was acceptable (typically activations and gradients after normalization) and FP16 where it was critical (typically weights and output projections), without requiring model developers to manually annotate precision choices.
NVLink 4.0 connected H100 GPUs at 900 GB/s bidirectional bandwidth per GPU in NVL configurations (vs 600 GB/s NVLink 3.0 in A100), and the 4th-generation NVSwitch (deployed in the DGX H100 SuperPOD rack) provided 57.6 TB/s all-reduce bandwidth across an 8-GPU node. Multi-Instance GPU (MIG) allowed one H100 to be partitioned into up to 7 isolated GPU instances (each with dedicated HBM, L2 cache, and SMs), enabling secure multi-tenant sharing of a single GPU — critical for cloud providers offering smaller GPU slices to inference workloads that didn’t fill an entire GPU. The H100 PCIe variant (for servers without SXM interconnect) shipped with 80 GB HBM2e at 350W; the SXM5 variant ran at 700W TDP. NVLink-C2C (chip-to-chip) debuted in the Grace Hopper superchip (GH200), combining an H100 GPU with a Grace Arm CPU via a high-bandwidth C2C interconnect providing 900 GB/s CPU-GPU bandwidth — the first tightly coupled CPU-GPU product from NVIDIA.
H100 was available in cloud deployments starting in Q3 2022, with Google Cloud, Microsoft Azure, and Oracle Cloud Infrastructure among the first to offer H100 instances. H100 became the standard compute unit for training frontier LLMs: the GPT-4 training run (early 2023) used approximately 25,000 A100s (H100 clusters had not yet reached scale for that run), while subsequent models from Anthropic, Google, and Meta used H100 clusters. The per-GPU retail price ranged from $25,000 to $40,000 MSRP, with spot prices during peak shortage (late 2023) reaching $50,000–$70,000 on secondary markets. NVIDIA’s data center revenue grew from $11.5 billion in FY2023 (ending January 2023) to $47.5 billion in FY2024 (ending January 2024), driven primarily by H100 demand from hyperscalers and AI startups building training and inference infrastructure for the generative AI wave that began in November 2022 with ChatGPT’s launch.
