Microsoft Introduces the Maia 200 AI Accelerator

less than 1 minute read

Published:

Microsoft introduced Maia 200 in January 2026 as its second-generation in-house AI accelerator and designed it specifically around large-scale inference. The chip is manufactured on TSMC’s 3 nm process.

Maia 200 includes native FP8 and FP4 tensor computation, 216 GB of HBM3e memory with about 7 TB/s of bandwidth, and 272 MB of on-chip SRAM. Microsoft says the design delivers more than 10 petaFLOPS at FP4 and more than 5 petaFLOPS at FP8.

Those numbers reflect the economics of modern AI serving. Generating tokens for large models is often limited by memory movement and accelerator utilization, not only arithmetic. Microsoft’s decision to build custom silicon is also strategic: Azure can optimize hardware, networking, software, and pricing together instead of relying entirely on third-party GPUs.