Captive silicon. Maia 200 runs only inside Microsoft Azure and cannot be rented from any provider or bought as hardware. It is reachable only indirectly through Azure services built on it.

Microsoft Maia 200

Type: ASICArchitecture: Maia 200Form factor: CustomReleased: 2026Process: TSMC N3Spec confidence: Official
FP8 (dense)
5,050
TFLOPS
Memory
216 GB
HBM3e
Bandwidth
7.0 TB/s
memory
TDP
750 W

Overview

Microsoft Maia 200 is Microsoft's second-generation in-house AI accelerator and its first purpose-built for inference, announced and deployed on 26 January 2026. Each chip is fabricated on TSMC's 3nm process with over 140 billion transistors inside a 750W SoC TDP envelope, and pairs 216GB of HBM3e at 7 TB/s with 272MB of software-managed on-die SRAM split between cluster-level and tile-level pools. Microsoft publishes 10.1 PetaOPS of FP4 and states that FP4 throughput is twice FP8 and eight times BF16, which places FP8 at roughly 5.05 PFLOPS. Compute is arranged as tiles, each combining a Tile Tensor Unit for FP8, FP6 and FP4 matrix multiplication with a Tile Vector Processor covering FP8, BF16, FP16 and FP32, grouped into clusters around shared SRAM and a multi-level DMA subsystem. Scale-up runs over standard Ethernet rather than a proprietary fabric: an on-die NIC delivers 2.8 TB/s of bidirectional bandwidth, four accelerators per tray are directly connected in a switchless Fully Connected Quad, and a switched second tier reaches 6,144 accelerators. Microsoft claims 30 percent better performance per dollar than the newest hardware already in its fleet, three times the FP4 throughput of AWS Trainium3, and FP8 throughput above Google's TPU v7. Maia 200 is deployed first in the US Central region near Des Moines, Iowa, with US West 3 near Phoenix, Arizona to follow, serving OpenAI GPT-5.2 models, Microsoft Foundry and Microsoft 365 Copilot. It is captive Azure silicon and is not sold or rented as hardware. Microsoft states no dense or sparse basis for any of these figures and describes no structured-sparsity hardware, so the basis is not stated; the values are recorded as peak throughput with no sparse counterpart.

Performance

Peak theoretical throughput by precision type

PrecisionPeak
FP64
No verified data available
FP32
No verified data available
TF32
No verified data available
BF16
Brain Float 16
1,263TFLOPS
FP16
No verified data available
FP8
8-bit floating point
5,050TFLOPS
FP6
No verified data available
FP4
4-bit floating point
10,100TFLOPS
INT8
No verified data available

Specifications

Architecture

Maia 200

Form Factor

Custom

Launch Year

2026

Process Node

TSMC N3

Memory

216 GB HBM3e

Bandwidth

7,000 GB/s

TDP

750 W

Max power (Flopper estimate)

~862 W est. Flopper estimate: 750 W TDP x 1.15. The vendor publishes no maximum board power for this part.

Transistors

140.0 billion

Interconnect

2.8 TB/s Ethernet scale-up (Maia AI Transport Layer)

direction and scope not stated by vendor

Spec Confidence

Official

Full Specifications

Chip Design
Transistors 140.0 billion
Process Node TSMC N3
Memory
Memory 216 GB
Memory Type HBM3e
Bandwidth 7.0 TB/s
Interface Width No verified data available
On-Die SRAM 272 MB
Interconnect & I/O
GPU-to-GPU Ethernet scale-up (Maia AI Transport Layer)
Interconnect Bandwidth 2.8 TB/s direction and scope not stated by vendor
Power & Thermal
TDP 750 W
Max power (Flopper estimate) ~862 W est.
Cooling Air or liquid cooled, second-generation closed-loop liquid cooling heat exchanger unit
Enterprise Features
Compute APIs PyTorch, Triton, Maia SDK, NPL, MCCL
General
Form Factor Custom
Architecture Maia 200
Process Node TSMC N3
Launch Year 2026

Datasheet & Resources

Data Provenance

Every figure traced to a source

Primary Source

Publisher
Microsoft
Published
2026-01-26

Data Quality

Spec confidence
Official
Clock basis
Boost
Core precisions with figures
3 of 9
Normalization
All values in TFLOPS

Compare cloud providers offering on-demand GPU instances for AI training, inference, and HPC workloads.

Browse GPU Cloud Providers

Frequently Asked Questions

How many TFLOPS does the Microsoft Maia 200 have?

The Microsoft Maia 200 delivers 5,050 TFLOPS FP8 at peak. Flopper does not currently have verified FP32 and FP16 throughput figures for it.

What is the power consumption of the Microsoft Maia 200?

The Microsoft Maia 200 has a TDP (Thermal Design Power) rating of 750 watts.

How much memory does the Microsoft Maia 200 have?

The Microsoft Maia 200 is equipped with 216 GB of memory with 7,000 GB/s of memory bandwidth.

What architecture is the Microsoft Maia 200 based on?

The Microsoft Maia 200 is based on the Maia 200 architecture, launched in 2026.

Stay Updated on GPU Releases

Get notified when new GPUs are added or specifications are updated.

Loading verification...

No spam, unsubscribe anytime.

Back to GPUs

© 2026 Flopper.io - Compare the GPUs Powering AI