Microsoft Unveils Maia 200 Inference Chip to Minimize AI Serving Prices
Microsoft just lately launched Maia 200, a custom-built accelerator geared toward reducing the price of operating synthetic intelligence workloads at cloud scale, as main suppliers look to curb hovering inference bills and reduce dependence on Nvidia graphics processors.
The chip is designed particularly for inference, the part during which educated fashions produce textual content, photographs and different outputs. As AI providers transition from pilots to on a regular basis manufacturing use, the price of producing tokens has grow to be an more and more vital share of total spending. Microsoft stated Maia 200 is meant to handle these economics by means of lower-precision compute, high-bandwidth reminiscence and networking optimized for giant AI clusters.
“As we speak, we’re proud to introduce Maia 200, a breakthrough inference accelerator engineered to dramatically enhance the economics of AI token technology,” Scott Guthrie, Microsoft’s government vice chairman for Cloud and AI, wrote in a weblog put up saying the chip.
Maia 200 is constructed on TSMC’s 3-nanometer course of and is designed round lower-precision math utilized in fashionable inference workloads. Microsoft stated every chip incorporates greater than 140 billion transistors and delivers greater than 10 petaFLOPS in 4-bit precision (FP4), and greater than 5 petaFLOPS in 8-bit precision (FP8), inside a 750-watt thermal envelope. The chip contains 216 gigabytes of HBM3e reminiscence with 7 terabytes per second of bandwidth, 272 megabytes of on-chip SRAM, and knowledge motion engines to scale back bottlenecks that may restrict real-world throughput even when uncooked compute is excessive.
“Crucially, FLOPS aren’t the one ingredient for sooner AI,” Guthrie wrote. “Feeding knowledge is equally vital.”
The launch comes as Microsoft, Google, and Amazon make investments closely in {custom} silicon alongside Nvidia GPUs. Google’s TPU household and Amazon’s Trainium chips provide options inside their cloud providers, and Microsoft has lengthy signaled that it needs higher management over prices and capability in its AI infrastructure. Maia 200 follows Maia 100, launched in 2023, and the corporate is positioning the brand new chip as an inference-focused workhorse for its AI merchandise.
Microsoft stated Maia 200 will help a number of fashions, together with “the most recent GPT-5.2 fashions from OpenAI,” and can be used to ship a performance-per-dollar benefit to Microsoft Foundry and Microsoft 365 Copilot. The corporate additionally stated its Microsoft Superintelligence workforce plans to make use of Maia 200 for artificial knowledge technology and reinforcement studying because it develops in-house fashions. Guthrie wrote that, for artificial knowledge pipelines, Maia 200’s design can speed up the technology and filtering of “high-quality, domain-specific knowledge.”
The chip can also be an effort to compete on headline efficiency with hyperscaler rivals. Guthrie wrote that Maia 200 is “essentially the most performant, first-party silicon from any hyperscaler,” including that it gives “thrice the FP4 efficiency of the third technology Amazon Trainium” and “FP8 efficiency above Google’s seventh technology TPU.” Reuters-style comparisons typically hinge on vendor-provided benchmarks, and Microsoft didn’t, in its put up, present full take a look at configurations for these claims.
Source link
#Microsoft #Unveils #Maia #Inference #Chip #Minimize #Serving #Prices #Campus #Expertise

