Microsoft Maia 300: What We Know About Microsoft's Next AI Chip
![]() |
| Image: Trazika / Pixabay |
Microsoft is preparing the next major step in its custom AI-silicon strategy. The company is reportedly planning to unveil Maia 300, the successor to the Maia 200 accelerator, as early as September 2026.
Unlike Nvidia’s general-purpose GPUs, Maia is designed specifically around Microsoft’s cloud infrastructure and AI workloads. The goal is straightforward: improve the economics of running large AI models at Azure scale while giving Microsoft greater control over its hardware supply chain.
But there is an important distinction between what is confirmed and what is still based on industry reporting. Microsoft has not yet publicly released the full specifications of Maia 300, so many of the technical details circulating online should not be treated as final.
Maia 300 Is Reportedly Coming in 2026
According to a recent report from Reuters, citing The Information, Microsoft plans to unveil Maia 300 in the fall of 2026, potentially as soon as September. Microsoft has confirmed that it continues to invest in custom silicon, but it has not publicly confirmed the Maia 300 launch date or its final specifications.
The timing would give Microsoft a relatively rapid progression from Maia 200 to its next-generation accelerator.
Microsoft introduced its original Maia accelerator in 2023 and followed it with Maia 200 in January 2026. Maia 200 is already deployed in Microsoft data centers, including facilities in Iowa and Arizona.
The Maia 300 therefore represents more than a routine chip refresh. Microsoft is trying to build a repeatable custom-silicon platform that can evolve alongside its rapidly growing AI infrastructure.
What Maia 200 Tells Us About Maia 300
The best indication of Microsoft's direction comes from Maia 200.
Maia 200 is a purpose-built inference accelerator rather than a conventional CPU or GPU. Microsoft designed it around the specific requirements of serving AI models, where memory capacity, memory bandwidth, networking and power efficiency can be just as important as raw compute performance.
The chip uses TSMC’s 3nm process, includes native FP8 and FP4 tensor computation, and has 216GB of HBM3e memory with 7TB/s of bandwidth alongside 272MB of on-chip SRAM. Microsoft rates it at more than 10 petaflops of FP4 performance and more than 5 petaflops at FP8.
Microsoft also says its Maia 200 system delivers more than 30% better performance per dollar than the latest-generation hardware previously deployed in its fleet.
That philosophy is likely to carry into Maia 300: optimize the entire system for the workloads Microsoft actually runs instead of simply maximizing a single benchmark number.
Is Maia 300 a 2nm Chip?
This is one of the most frequently reported specifications, but it should currently be treated as unconfirmed.
Several semiconductor-industry reports and research notes have associated Maia 300 with TSMC’s 2nm process. Some industry forecasts also place the chip in the second half of 2026. However, Microsoft has not publicly announced Maia 300's manufacturing process.
If the reports are accurate, moving from Maia 200's 3nm-class design to TSMC 2nm would give Microsoft access to a newer process generation with improvements in transistor density and efficiency.
But the process node alone does not determine the real-world performance of an AI accelerator. Memory architecture, packaging, interconnects, software and power delivery can have an equally important impact.
Memory Could Be the Real Story
Large AI models are increasingly constrained by memory capacity and bandwidth, particularly during inference.
Microsoft clearly recognized this with Maia 200. Its unusually large SRAM allocation and 216GB of HBM3e are designed to keep AI workloads supplied with data rather than allowing compute units to sit idle waiting for memory.
There are industry reports suggesting Maia 300 could use a newer generation of HBM and potentially provide substantially more memory than Maia 200. However, Microsoft has not confirmed Maia 300's HBM generation, capacity or bandwidth.
Claims circulating online about specific configurations such as 288GB of HBM4 or HBM4E should therefore be regarded as estimates rather than specifications.
That distinction matters. Memory configuration could ultimately be one of the most important differences between Maia 200 and Maia 300.
Microsoft Wants Hundreds of Thousands of Maia 300 Chips
The most significant recent development may not be the chip's architecture at all. It is Microsoft's reported production ambitions.
Reuters reported that Microsoft has been discussing manufacturing capacity with TSMC for more than 300,000 Maia 300 chips, with deliveries targeted for 2027. The same report said Microsoft ultimately wants capacity for more than one million chips, although supply constraints could affect those plans.
Microsoft's Azure Maia general manager Andrew Wall confirmed the company's continuing investment in custom silicon but said the reported production figures do not represent the full scale of the program.
That is an important signal.
Microsoft is no longer treating custom AI silicon as a small experimental project. The company is planning infrastructure deployment on a scale where even relatively small improvements in cost per token, power efficiency or performance can translate into enormous savings.
Maia 300 Is Not About Replacing Nvidia Overnight
It would be misleading to describe Maia 300 simply as Microsoft's "Nvidia killer."
Microsoft continues to deploy Nvidia and AMD accelerators throughout Azure. Its own Maia chips are intended to become another component of a heterogeneous infrastructure strategy.
The reason for developing custom silicon is control and optimization.
Buying GPUs from external suppliers gives Microsoft access to extremely capable hardware, but it also exposes the company to pricing, availability, product cycles and supply constraints. A custom accelerator lets Microsoft optimize silicon, software, networking and data-center design around its own workloads.
At Azure's scale, that can be economically valuable even if Maia 300 is not faster than every competing GPU.
Microsoft Is Building the Software Around Maia, Too
Hardware is only part of the strategy.
With Maia 200, Microsoft introduced the Maia Software Development Kit, including support for PyTorch and a Triton compiler, as well as tools such as a simulator and cost calculator. The company is attempting to make its custom accelerators usable within the broader AI software ecosystem rather than creating a completely isolated programming environment.
That will become increasingly important with Maia 300.
Nvidia's enormous advantage is not just its hardware. CUDA, libraries, frameworks and developer familiarity form a mature software ecosystem.
Microsoft therefore has to make Maia sufficiently easy to target and sufficiently compatible with existing AI workflows if it wants developers and internal teams to use the hardware at scale.
What Could Maia 300 Power?
Microsoft's Maia 200 is already intended for large-scale inference across Azure services and Microsoft's own AI products.
Microsoft has said Maia 200 will support workloads involving Microsoft Foundry, Microsoft 365 Copilot and OpenAI models, including GPT-5.2-era workloads.
Maia 300 is likely to follow the same general direction, particularly as inference demand grows.
The economics of inference are becoming increasingly important. Every AI-generated response requires compute, memory bandwidth and electricity. At the scale of Microsoft's cloud, improving the cost of generating each token can have a direct impact on operating expenses.
That makes custom inference accelerators particularly attractive.
The Anthropic Angle
Reuters also reported that Microsoft is seeking to persuade major cloud customers, including Anthropic, to use Maia 300.
If Microsoft can move workloads from Maia's internal use to broader Azure customer deployments, the business case for developing successive generations becomes considerably stronger.
It would effectively turn Maia from an internal optimization project into an important part of Azure's compute portfolio.
That is still a goal rather than a confirmed Maia 300 customer deployment.
What We Actually Know So Far
As of August 14, 2026, the picture looks like this:
| Feature | Maia 300 status |
|---|---|
| Product name | Maia 300 |
| Manufacturer/foundry | TSMC is reportedly involved |
| Expected unveiling | Fall 2026, potentially September |
| Manufacturing node | Reported as 2nm, not officially confirmed |
| AI focus | Expected to emphasize large-scale AI workloads/inference |
| HBM capacity | Not officially disclosed |
| HBM generation | Not officially disclosed |
| Compute performance | Not officially disclosed |
| Power consumption | Not officially disclosed |
| Production plans | Reports indicate 300,000+ units targeted for 2027 |
| Long-term production ambition | Reportedly more than 1 million units |
| Availability | Not officially announced |
The key point is that there is currently much more reliable information about Microsoft's strategy and production plans than about Maia 300's final silicon specifications.
Why Maia 300 Matters
Microsoft has spent billions building AI infrastructure, and the economics of that infrastructure are becoming as important as model performance.
The company cannot indefinitely optimize its AI business solely around buying the same general-purpose accelerators available to every other hyperscaler. Custom silicon provides another lever: Microsoft can tune hardware to its models, software stack and data-center architecture.
Maia 300 could therefore be important even if it does not dramatically outperform Nvidia's newest accelerator on a conventional benchmark.
The more meaningful question is whether Microsoft can achieve a better cost-per-token, performance-per-watt and total system efficiency for the workloads running across Azure.
That is the metric that matters when millions or billions of AI requests are processed.
The Bottom Line
Maia 300 is real, but many of the specifications circulating online are not yet official.
The strongest current reporting points to a fall 2026 unveiling, potentially in September, followed by a substantial production ramp in 2027. Microsoft is reportedly working with TSMC on capacity for hundreds of thousands of chips, with ambitions that could eventually exceed one million units.
The technical details remain the biggest unanswered question.
A 2nm process and next-generation HBM are widely associated with Maia 300 in industry reporting, but Microsoft has not yet confirmed those specifications. Until the company formally introduces the chip, claims about exact compute performance, memory capacity, bandwidth, power consumption or architectural features should be treated as estimates.
What is already clear is the strategic direction.
Microsoft wants Maia to become a serious pillar of Azure's AI infrastructure—not simply a side project. If Maia 300 delivers meaningful improvements in inference economics and Microsoft can manufacture it at the scale currently being reported, the chip could become one of the company's most important pieces of AI infrastructure over the next several years.

