Qualcomm Cloud AI 100 Ultra
Overview
Qualcomm Cloud AI 100 Ultra is the largest card in Qualcomm's first data center AI family, a 150 watt PCIe inference accelerator built on a 7 nm process and aimed squarely at generative AI serving economics rather than at peak throughput. It carries 64 AI cores on a single card and Qualcomm's product brief rates it at up to 870 INT8 TOPS, which is the only compute figure that brief publishes. An FP16 rating of 288 TFLOPS appears in our records, attributed to a Qualcomm product page spec table, but that page has since moved and now renders only through JavaScript, so the figure could not be re-verified and is deliberately not listed among the precisions below. Its most distinctive feature is memory hierarchy rather than raw compute: each AI core owns 9 megabytes of software-managed on-die SRAM for 576 megabytes in total, by far the largest local scratchpad of any conventional accelerator, backed by 128 GB of error-corrected LPDDR4x delivering 548 GB/s. The design accepts an order of magnitude less DRAM bandwidth than an HBM part in exchange for capacity, cost and a power envelope low enough that eight cards fit comfortably in a single server, and Qualcomm markets the result on performance per dollar, claiming a 100 billion parameter generative model runs on one card and that a single server holds models eight times larger than competing solutions allow. The card is a full-height three-quarter-length PCIe form factor on a Gen 4 sixteen-lane host interface, with no accelerator-to-accelerator fabric of any kind, so scaling is done over PCIe and the host network rather than over a proprietary link. Four sibling SKUs share the same core at lower core counts and clocks, from a 16-core 75 watt entry card upward, and Qualcomm's own footnote that each AI core holds 9 MB of SRAM reconciles every variant's memory figure exactly, confirming they are bins of one design. Qualcomm publishes throughput for only these two datatypes and no clock speed, die size or transistor count for any member of the family. Commercially this part matters more than its age suggests: Qualcomm's successor Dragonfly line, announced in October 2025 as the rack-scale AI200 and AI250, has published no performance figures at all and is not yet available, leaving the Cloud AI 100 Ultra as the Qualcomm accelerator that ships, that has a real product brief, and that can be rented today through Cirrascale.
Performance
Peak theoretical throughput by precision type
| Precision | Peak |
|---|---|
FP64 | No verified data available |
FP32 | No verified data available |
TF32 | No verified data available |
BF16 | No verified data available |
FP16 | No verified data available |
FP8 | No verified data available |
FP6 | No verified data available |
FP4 | No verified data available |
INT8 8-bit integer | 870TOPS |
Dense peak divided by accelerator TDP. Board power only: excludes host CPUs, networking, cooling and facility overhead. TDP for this part: 150 W.
Specifications
Architecture
Cloud AI 100
Form Factor
PCIe
Launch Year
2023
Process Node
7nm
Memory
128 GB LPDDR4X
Bandwidth
548 GB/s
TDP
150 W
Max power (Flopper estimate)
~173 W est. Flopper estimate: 150 W TDP x 1.15. The vendor publishes no maximum board power for this part.
AI Cores
64
Spec Confidence
Official
Full Specifications
| Compute Engine | |
|---|---|
| AI Cores | 64 |
| Memory | |
| Memory | 128 GB |
| Memory Type | LPDDR4X |
| Bandwidth | 548 GB/s |
| Interface Width | No verified data available |
| On-Die SRAM | 576 MB |
| Interconnect & I/O | |
| PCIe | Gen4 x16 |
| Power & Thermal | |
| TDP | 150 W |
| Max power (Flopper estimate) | ~173 W est. |
| Enterprise Features | |
| ECC Memory | Yes |
| Physical & Media | |
| Card Length | 237.9 mm |
| General | |
| Form Factor | PCIe |
| Architecture | Cloud AI 100 |
| Process Node | 7nm |
| Launch Year | 2023 |
Datasheet & Resources
Flopper Datasheet
Qualcomm Cloud AI 100 Ultra specifications, generated from our database. Printable.
Vendor product page
Qualcomm documentation for the Cloud AI 100 Ultra
Qualcomm Cloud AI 100 product page and specification table
Qualcomm · Latest version
Data Provenance
Every figure traced to a source
Primary Source
- Publisher
- Qualcomm
- Published
- No verified data available
Data Quality
- Spec confidence
- Official
- Clock basis
- Boost
- Core precisions with figures
- 1 of 9
- Normalization
- All values in TFLOPS
Compare cloud providers offering on-demand GPU instances for AI training, inference, and HPC workloads.
Browse GPU Cloud ProvidersFrequently Asked Questions
How many TFLOPS does the Qualcomm Cloud AI 100 Ultra have?
The verified peak figures Flopper holds for the Qualcomm Cloud AI 100 Ultra are 870 TOPS INT8. Flopper does not currently have verified FP32, FP16 and FP8 throughput figures for it.
What is the power consumption of the Qualcomm Cloud AI 100 Ultra?
The Qualcomm Cloud AI 100 Ultra has a TDP (Thermal Design Power) rating of 150 watts.
How much memory does the Qualcomm Cloud AI 100 Ultra have?
The Qualcomm Cloud AI 100 Ultra is equipped with 128 GB of memory with 548 GB/s of memory bandwidth.
What architecture is the Qualcomm Cloud AI 100 Ultra based on?
The Qualcomm Cloud AI 100 Ultra is based on the Cloud AI 100 architecture, launched in 2023.
Stay Updated on GPU Releases
Get notified when new GPUs are added or specifications are updated.
No spam, unsubscribe anytime.