Qualcomm Cloud AI 100 Ultra
Overview
Qualcomm Cloud AI 100 Ultra is the largest card in Qualcomm's first data center AI family, a 150 watt PCIe inference accelerator built on a 7 nm process and aimed squarely at generative AI serving economics rather than at peak throughput. It carries 64 AI cores on a single card and Qualcomm's product brief rates it at up to 870 INT8 TOPS, which is the only compute figure that brief publishes. An FP16 rating of 288 TFLOPS appears in our records, attributed to a Qualcomm product page spec table, but that page has since moved and now renders only through JavaScript, so the figure could not be re-verified and is deliberately not listed among the precisions below. Its most distinctive feature is memory hierarchy rather than raw compute: each AI core owns 9 megabytes of software-managed on-die SRAM for 576 megabytes in total, by far the largest local scratchpad of any conventional accelerator, backed by 128 GB of error-corrected LPDDR4x delivering 548 GB/s. The design accepts an order of magnitude less DRAM bandwidth than an HBM part in exchange for capacity, cost and a power envelope low enough that eight cards fit comfortably in a single server, and Qualcomm markets the result on performance per dollar, claiming a 100 billion parameter generative model runs on one card and that a single server holds models eight times larger than competing solutions allow. The card is a full-height three-quarter-length PCIe form factor on a Gen 4 sixteen-lane host interface, with no accelerator-to-accelerator fabric of any kind, so scaling is done over PCIe and the host network rather than over a proprietary link. Four sibling SKUs share the same core at lower core counts and clocks, from a 16-core 75 watt entry card upward, and Qualcomm's own footnote that each AI core holds 9 MB of SRAM reconciles every variant's memory figure exactly, confirming they are bins of one design. Qualcomm publishes throughput for only these two datatypes and no clock speed, die size or transistor count for any member of the family. Commercially this part matters more than its age suggests: Qualcomm's successor Dragonfly line, announced in October 2025 as the rack-scale AI200 and AI250, has published no performance figures at all and is not yet available, leaving the Cloud AI 100 Ultra as the Qualcomm accelerator that ships, that has a real product brief, and that can be rented today through Cirrascale.
Performance Metrics
Peak theoretical throughput by precision type
| Precision | Bits | Peak TFLOPS | Efficiency |
|---|---|---|---|
| INT8 | 8 | 870.0 | 5.800 TFLOPS/W |
Power Specifications
TDP
150 W
Max Power
173 W
Power Connector
PCIe Slot
Cooling
Air
Memory Specifications
Capacity
128 GB
Type
LPDDR4X
Bandwidth
548 GB/s
Interface
--
Hardware & Design
Form Factor
PCIe
Architecture
Cloud AI 100
Process Node
7nm
Launch Year
2023
Variant
Standard
Market Segment
Professional
Full Specifications
| Compute Engine | |
|---|---|
| AI Cores | 64 |
| Memory | |
| VRAM | 128 GB |
| Memory Type | LPDDR4X |
| Bandwidth | 548 GB/s |
| On-Die SRAM | 576 MB |
| Interconnect & I/O | |
| PCIe | Gen4 x16 |
| Power & Thermal | |
| TDP | 150 W |
| Enterprise Features | |
| ECC Memory | Yes |
| Physical & Media | |
| Card Length | 237.9 mm |
| General | |
| Form Factor | PCIe |
| Architecture | Cloud AI 100 |
| Launch Year | 2023 |
Documentation & Resources
Common Use Cases
The Qualcomm Cloud AI 100 Ultra is optimized for high-performance computing tasks with Cloud AI 100 architecture delivering high TFLOPS of compute power.
Where to Rent
Compare cloud providers offering on-demand GPU instances for AI training, inference, and HPC workloads.
Browse GPU Cloud ProvidersStay Updated on GPU Releases
Get notified when new GPUs are added or specifications are updated.
No spam, unsubscribe anytime.