Microsoft introduces Maia 200, a new in-house AI inference accelerator built on 3nm
Microsoft has announced Maia 200, its next-generation custom AI accelerator aimed at improving the economics of AI inference. The company says the chip is built on a 3nm process, features a redesigned memory system, and is already deployed in a US data centre, with more regions planned.
Microsoft has unveiled Maia 200, its next-generation in-house AI accelerator designed primarily for inference workloads. In a post on the Official Microsoft Blog dated January 26, 2026, the company said Maia 200 is engineered to improve the cost and speed of AI token generation, positioning it as part of Microsoft’s broader push to build first-party silicon for its cloud and product ecosystem.

Microsoft said Maia 200 is built on TSMC’s 3nm process and includes FP8/FP4 tensor capabilities alongside a redesigned memory system. The company framed the chip as a major step in making large-model serving more efficient, a critical issue as demand for AI applications rises across consumer and enterprise products.
According to Microsoft, Maia 200 has already been deployed in its US Central datacenter region near Des Moines, Iowa, with the US West 3 region near Phoenix, Arizona, planned next. The company also described a software stack and tooling approach, including an SDK preview, intended to help developers and internal teams build and optimise model execution for Maia hardware.
The announcement matters because hyperscalers are increasingly building custom chips to reduce dependence on a single supplier ecosystem and to tailor performance to their own workloads. This trend has already seen long-running internal silicon efforts at other major cloud providers, and Maia 200 adds to that competitive landscape by focusing on inference efficiency at scale.
Microsoft’s blog post also highlighted system-level design, describing an Ethernet-based scale-up networking approach and performance characteristics meant to keep large clusters utilised and cost-effective. While raw chip specifications are only one part of the story, the company is signalling that Maia 200 is designed as a full-stack platform play: silicon, networking and software integration.
For customers, the practical question is how quickly Maia-backed capacity becomes broadly available in Azure regions and whether it leads to more predictable pricing and performance for AI features. For the industry, Maia 200 is another marker that the economics of inference — not just training — is now a central battleground in the race to deploy AI at scale.
As Microsoft expands deployment to additional data centres, developers and enterprises will watch for clearer benchmarks in real-world model serving, ecosystem maturity of the SDK, and how quickly the chip’s advantages translate into improved responsiveness and lower costs for AI products and APIs.