NVIDIA Becomes Central to the AI Infrastructure Boom

2 minute read

Published:

By November 2024, NVIDIA had become the indispensable infrastructure provider for large-scale AI training and inference, with its market capitalization briefly reaching $3.4 trillion in June 2024 — briefly making it the world’s most valuable company and reflecting the AI infrastructure spending wave triggered by ChatGPT’s November 2022 launch. The H100 SXM5 (Hopper architecture, TSMC N4, 80 GB HBM3, 3.35 TFLOPS BF16/FP16 dense training throughput, 700W TDP) had become the standard unit for LLM training: OpenAI trained GPT-4 on an estimated 25,000 A100s (Hopper’s predecessor); Anthropic, Google, and Meta deployed H100 clusters of similar scale. H200 SXM5 (announced December 2023, HBM3e 141 GB at 4.8 TB/s bandwidth vs H100’s 80 GB at 3.35 TB/s) offered approximately 1.9× inference throughput on memory-bound LLM serving workloads. NVIDIA’s GB200 NVL72 rack (announced March 2024, available late 2024 in Blackwell generation): 36 Grace Arm CPUs + 72 Blackwell B200 GPUs in a 72-GPU NVLink domain, delivering 1.44 ExaFLOPS FP4 for AI inference — targeting trillion-parameter model serving at economics impractical with H100.

CUDA’s competitive advantage by November 2024 was not merely the CUDA API itself (which AMD’s ROCm and Intel’s oneAPI partially replicated) but the accumulated software ecosystem: PyTorch 2.x compiled to CUDA-specific operations via torch.compile, cuDNN (deep neural network primitives), cuBLAS (matrix math), NCCL (collective communication between GPUs across nodes using NVLink at 900 GB/s intra-node or InfiniBand NDR 400 Gbps inter-node), and TensorRT-LLM (continuous batching and quantization for LLM inference). NVSwitch 3.0 in H100 SXM5 nodes connected 8 GPUs with 900 GB/s all-reduce bandwidth — 36× higher than PCIe 4.0 ×16 (32 GB/s). NVLink 4.0 extended that fabric to 2-node configurations (NVL16). For inter-node connectivity, InfiniBand NDR 400 (400 Gbps per port) was the standard, with NVIDIA’s acquisition of Mellanox (2020, $6.9B) giving it control over both the GPU and the networking fabric used in AI clusters. NVIDIA Spectrum-X (Ethernet-based AI networking fabric announced 2023) provided an alternative to InfiniBand for hyperscalers building Ethernet-only fabrics.

The moat extended to power and cooling engineering: the GB200 NVL72 consumed approximately 120 kW per rack, requiring liquid cooling infrastructure (direct liquid cooling or rear-door heat exchangers) that air-cooled data center designs could not support at density. AMD’s MI300X (HBM3, 192 GB per GPU, announced December 2023) offered competitive performance on memory-bandwidth-bound inference workloads and was being deployed at Microsoft Azure and other hyperscalers, representing the strongest GPU alternative to H100/H200. Google’s TPU v5p (announced December 2023), Amazon Trainium2 (announced November 2023), and Microsoft Azure Maia 100 (announced November 2023) were custom ASIC alternatives that traded CUDA software compatibility for specific workload optimization. By Q3 2024, NVIDIA’s data center revenue had reached $26.3 billion in a single quarter (against $3.8 billion in Q3 2022 pre-ChatGPT), demonstrating the unprecedented capital reallocation toward GPU infrastructure that the LLM era had produced.