NVIDIA Announces Fermi GeForce GTX 480
Published:
NVIDIA launched the GeForce GTX 480 on March 26, 2010 at $499, introducing the Fermi architecture (GF100 chip) as its first GPU designed explicitly to serve dual roles as a graphics processor and a general-purpose parallel computing platform. NVIDIA had first revealed Fermi at its GPU Technology Conference in September 2009 with a detailed architecture paper, generating anticipation for what appeared to be a significant compute-capability advance. The GF100 die was manufactured on TSMC’s 40nm process and contained 3.0 billion transistors across a 529 mm² die — more than double the GT200 chip (1.4 billion transistors, 470 mm²) used in the GTX 280 (2008). The GTX 480 shipped with 480 CUDA cores (15 Streaming Multiprocessors × 32 cores each) out of GF100’s full complement of 512 cores (16 SMs), with one SM disabled — a common yield management practice for complex dies. Each SM in Fermi was substantially redesigned from the GT200: it gained a configurable L1 cache (16 or 48 KB, shared with shared memory in a 48/16 or 16/48 KB split), connected to a 768 KB unified L2 cache shared across the entire chip — a conventional memory hierarchy absent from earlier CUDA architectures that had instead relied on purely programmer-managed shared memory and no L2. The GTX 480 offered 1.35 TFLOPS of single-precision and 168 GFLOPS of double-precision floating-point performance, with the 1:8 ratio between DP and SP reflecting a consumer-grade configuration (the professional Tesla C2050/C2070 cards, launched in 2010, offered 1:2 DP:SP ratio for scientific computing).
The GF100 chip’s power consumption was Fermi’s most discussed characteristic. The GTX 480 carried a 250W TDP (Thermal Design Power) — up from the GTX 285’s 183W — requiring dual 6-pin PCIe power connectors and a 600W minimum system power supply. Under sustained load, GTX 480 cards routinely operated at junction temperatures approaching 90–105°C, with GPU fans spinning at maximum RPMs and producing noise levels compared unfavorably to vacuum cleaners in contemporary reviews. The performance implications of the thermal behavior were also apparent: graphics cards routinely throttled clock speeds to manage temperature, and the GTX 480’s high operating temperatures left less thermal headroom before throttling than competing cards. AMD’s Radeon HD 5870, launched in September 2009 at $379, had beaten NVIDIA to market by six months using ATI’s Cypress (RV870) chip on TSMC 40nm: the HD 5870 offered competitive rasterization performance with 188W TDP, DirectX 11 support, and a $120 lower launch price. In gaming benchmarks, the GTX 480 and HD 5870 were closely matched, with the GTX 480 ahead in some titles and the HD 5870 ahead in others — giving AMD a strong competitive position during the months before Fermi launched.
Fermi’s engineering significance was in what it established for GPU computing rather than consumer gaming. CUDA Compute Capability 2.0 introduced ECC memory support (critical for scientific computing where a single bit error in a simulation is unacceptable), concurrent kernel execution (multiple CUDA kernels running simultaneously on one GPU), unified virtual address space (simplifying CPU-GPU memory management), and atomic operations on global and shared memory that made it practical to implement complex parallel algorithms. The professional Fermi-based Tesla cards found rapid adoption in supercomputers: the Tianhe-1A supercomputer (November 2010, ranked #1 on TOP500) used 7,168 NVIDIA Tesla M2050 (Fermi) GPUs. Most consequentially for computing history, the GTX 580 (November 2010, a refined full-chip Fermi successor at $499 with all 512 CUDA cores enabled and better thermal management) became the GPU on which Alex Krizhevsky and Ilya Sutskever trained AlexNet in 2012 — two GTX 580 cards with 3 GB GDDR5 each running for about a week — the deep learning result that triggered the modern neural network era. Fermi’s L1/L2 cache hierarchy and CUDA improvements made that training computationally feasible, establishing the precedent of consumer GPU hardware as the primary training substrate for neural networks that would define the decade to follow.
