Shipping and rentable today, which makes it the only Qualcomm data center accelerator you can actually obtain. That is worth stating plainly because Qualcomm's newer and far more widely covered Dragonfly parts are not: the AI200 has not shipped, the AI250 is not expected to begin commercial sampling before mid-2027, and Qualcomm has published no throughput figure of any kind for either. This older 7 nm card is the opposite on both counts, with a full Qualcomm product brief behind it and published INT8 and FP16 performance. Cirrascale rents it in eight-card servers on its AI Innovation Cloud with public monthly pricing, and also runs a separate Qualcomm inference service on the same silicon. Note that Qualcomm does not label these performance figures as dense or sparse anywhere; they are recorded here as dense, on the reasoning set out in this entry's data provenance.
QU

Qualcomm Cloud AI 100 Ultra

Cloud AI 100 PCIe 2023 7nm
VRAM
128 GB
LPDDR4X
TDP
150 W
Bandwidth
548 GB/s
memory
AI Cores
64

Overview

Qualcomm Cloud AI 100 Ultra is the largest card in Qualcomm's first data center AI family, a 150 watt PCIe inference accelerator built on a 7 nm process and aimed squarely at generative AI serving economics rather than at peak throughput. It carries 64 AI cores on a single card and Qualcomm's product brief rates it at up to 870 INT8 TOPS, which is the only compute figure that brief publishes. An FP16 rating of 288 TFLOPS appears in our records, attributed to a Qualcomm product page spec table, but that page has since moved and now renders only through JavaScript, so the figure could not be re-verified and is deliberately not listed among the precisions below. Its most distinctive feature is memory hierarchy rather than raw compute: each AI core owns 9 megabytes of software-managed on-die SRAM for 576 megabytes in total, by far the largest local scratchpad of any conventional accelerator, backed by 128 GB of error-corrected LPDDR4x delivering 548 GB/s. The design accepts an order of magnitude less DRAM bandwidth than an HBM part in exchange for capacity, cost and a power envelope low enough that eight cards fit comfortably in a single server, and Qualcomm markets the result on performance per dollar, claiming a 100 billion parameter generative model runs on one card and that a single server holds models eight times larger than competing solutions allow. The card is a full-height three-quarter-length PCIe form factor on a Gen 4 sixteen-lane host interface, with no accelerator-to-accelerator fabric of any kind, so scaling is done over PCIe and the host network rather than over a proprietary link. Four sibling SKUs share the same core at lower core counts and clocks, from a 16-core 75 watt entry card upward, and Qualcomm's own footnote that each AI core holds 9 MB of SRAM reconciles every variant's memory figure exactly, confirming they are bins of one design. Qualcomm publishes throughput for only these two datatypes and no clock speed, die size or transistor count for any member of the family. Commercially this part matters more than its age suggests: Qualcomm's successor Dragonfly line, announced in October 2025 as the rack-scale AI200 and AI250, has published no performance figures at all and is not yet available, leaving the Cloud AI 100 Ultra as the Qualcomm accelerator that ships, that has a real product brief, and that can be rented today through Cirrascale.

Performance Metrics

Peak theoretical throughput by precision type

PrecisionBitsPeak TFLOPSEfficiency
INT8 8 870.0 5.800 TFLOPS/W

Power Specifications

TDP

150 W

Max Power

173 W

Power Connector

PCIe Slot

Cooling

Air

Memory Specifications

Capacity

128 GB

Type

LPDDR4X

Bandwidth

548 GB/s

Interface

--

Hardware & Design

Form Factor

PCIe

Architecture

Cloud AI 100

Process Node

7nm

Launch Year

2023

Variant

Standard

Market Segment

Professional

Full Specifications

Compute Engine
AI Cores 64
Memory
VRAM 128 GB
Memory Type LPDDR4X
Bandwidth 548 GB/s
On-Die SRAM 576 MB
Interconnect & I/O
PCIe Gen4 x16
Power & Thermal
TDP 150 W
Enterprise Features
ECC Memory Yes
Physical & Media
Card Length 237.9 mm
General
Form Factor PCIe
Architecture Cloud AI 100
Launch Year 2023

Documentation & Resources

Common Use Cases

General Compute AI/ML Workloads Data Processing

The Qualcomm Cloud AI 100 Ultra is optimized for high-performance computing tasks with Cloud AI 100 architecture delivering high TFLOPS of compute power.

Where to Rent

Compare cloud providers offering on-demand GPU instances for AI training, inference, and HPC workloads.

Browse GPU Cloud Providers

Stay Updated on GPU Releases

Get notified when new GPUs are added or specifications are updated.

Loading verification...

No spam, unsubscribe anytime.

Back to GPUs