Google TPU 8i NEW
Overview
Google's eighth-generation inference TPU, announced alongside the TPU 8t in April 2026 and tuned for serving rather than training. Google publishes a peak of 10.1 PFLOPS of FP4 per chip, with 288 GB of HBM at 8,601 GB/s — about 1.3 times the TPU 8t's bandwidth — and 384 MB of on-chip Vmem SRAM, three times the 8t's, enough to hold a long-context KV cache on silicon. It uses a Boardfly topology rather than a torus, joining up to 1,152 chips with a maximum seven-hop ICI network diameter and up to 1,024 of them active. FP4 is the only precision Google states; nothing else is published or inferred. Google publishes no sparse figure for any TPU, so this is a dense rate.
Performance Metrics
Peak theoretical throughput by precision type
| Precision | Bits | Peak TFLOPS | |
|---|---|---|---|
| FP4 | 4 | 10100.0 |
Power Specifications
TDP
--
Max Power
--
Power Connector
PCIe Slot
Cooling
Air
Memory Specifications
Capacity
288 GB
Type
--
Bandwidth
8601 GB/s
Interface
--
Hardware & Design
Form Factor
--
Architecture
TPU 8
Process Node
--
Launch Year
2026
Variant
Standard
Market Segment
Professional
Documentation & Resources
Common Use Cases
The Google TPU 8i is optimized for high-performance computing tasks with TPU 8 architecture delivering high TFLOPS of compute power.
Where to Rent
Compare cloud providers offering on-demand GPU instances for AI training, inference, and HPC workloads.
Browse GPU Cloud ProvidersSimilar GPUs
Systems Using This GPU
Pre-configured systems featuring the Google TPU 8i
| System | GPU Count | Peak Performance | Total Power | |
|---|---|---|---|---|
Google TPU 8i Pod TPU · pod · 2026 | 1024x TPU 8i | 10342.40 PFLOPS | -- | View System |
Stay Updated on GPU Releases
Get notified when new GPUs are added or specifications are updated.
No spam, unsubscribe anytime.