Qwen3.5-122B-A10B hardware requirements
Fits on one Instinct MI350X OAM at BF16. For 1,000 tokens/s at 32k context you need 4 x H100 SXM5 80GB at $7.16/hr or 4 x RTX PRO 6000 Blackwell Max-Q Workstation Edition at $2.00/hr.
Bare minimum
- Tokens/s
- 296
- Per hour
- $6.16
- Per million tokens
- $5.78
- Memory used
- 90%
- Tokens/s
- 239
- Per hour
- $2.00
- Per million tokens
- $2.33
- Memory used
- 68%
- Tokens/s
- 3.1k
- Per hour
- $2.00
- Per million tokens
- $0.18
- Memory used
- 77%
Configurations
Every datacenter GPU in the catalog, sized for these settings. Change precision, context and concurrency and the whole page follows. Each point is one replica: the smallest tensor-parallel group of that GPU that holds the model.
2 x GB300 NVL72 GPU 288GB: 7.2k tokens/s per replica at $17.24/hr.2 x GB200 NVL72 GPU 186GB: 7.2k tokens/s per replica at $21.00/hr.2 x Instinct MI350X OAM: 7.2k tokens/s per replica at $12.32/hr.2 x B300 SXM 262GB: 7.2k tokens/s per replica at $14.80/hr.2 x B200 SXM 180GB: 6.9k tokens/s per replica at $8.18/hr.2 x Instinct MI325X OAM: 5.4k tokens/s per replica at $7.60/hr.2 x Instinct MI300X 192GB: 4.8k tokens/s per replica at $4.78/hr.4 x H200 SXM 141GB: 8.2k tokens/s per replica at $11.96/hr.4 x H100 SXM5 80GB: 5.7k tokens/s per replica at $7.16/hr.4 x H200 NVL 141GB: 8.2k tokens/s per replica at $15.16/hr.4 x H100 NVL 94GB: 6.7k tokens/s per replica at $12.76/hr.4 x H100 PCIe 80GB: 3.4k tokens/s per replica at $10.00/hr.4 x RTX PRO 6000 Blackwell Workstation Edition: 3.1k tokens/s per replica at $7.20/hr.4 x RTX PRO 6000 Blackwell Server Edition: 2.7k tokens/s per replica at $2.36/hr.4 x RTX PRO 6000 Blackwell Max-Q Workstation Edition: 3.1k tokens/s per replica at $2.00/hr.16 x GeForce RTX 5090 32GB: 10k tokens/s per replica at $15.84/hr.8 x RTX 6000 Ada 48GB: 3.1k tokens/s per replica at $6.72/hr.8 x L40S 48GB: 2.8k tokens/s per replica at $6.96/hr.4 x A100 PCIe 80GB: 3.3k tokens/s per replica at $5.40/hr.4 x A100 SXM4 80GB: 3.5k tokens/s per replica at $5.60/hr.8 x A100 SXM4 40GB: 5.0k tokens/s per replica at $10.32/hr.16 x RTX 5000 Ada 32GB: 3.3k tokens/s per replica at $13.28/hr.8 x RTX PRO 5000 Blackwell 48GB: 4.3k tokens/s per replica at $7.68/hr.16 x RTX PRO 4500 Blackwell 32GB: 5.1k tokens/s per replica at $11.52/hr.8 x L40 48GB: 2.8k tokens/s per replica at $6.56/hr.16 x GeForce RTX 4090 24GB: 5.7k tokens/s per replica at $9.60/hr.16 x A30 24GB: 5.3k tokens/s per replica at $11.73/hr.8 x RTX A6000 48GB: 2.5k tokens/s per replica at $4.00/hr.8 x A40 48GB: 2.2k tokens/s per replica at $3.92/hr.16 x RTX PRO 4000 Blackwell 24GB: 3.8k tokens/s per replica at $9.12/hr.16 x V100S PCIe 32GB: 6.5k tokens/s per replica at $14.08/hr.16 x A10 24GB: 3.4k tokens/s per replica at $20.64/hr.16 x L4 24GB: 1.7k tokens/s per replica at $7.84/hr.16 x GeForce RTX 3090 24GB: 5.3k tokens/s per replica at $8.00/hr.16 x GeForce RTX 3090 Ti 24GB: 5.7k tokens/s per replica at $7.36/hr.16 x RTX A5000 24GB: 4.4k tokens/s per replica at $4.32/hr.
Priced configurations only; 37 more fit without a live price and appear in the table. Marker size grows with the tensor-parallel width. Roofline estimates.
| GPU | TP | Tokens/s | |
|---|---|---|---|
| 1 | ~11k | Specs | |
| 2 | 20k | Specs | |
| 2 | 7.2k | Specs | |
| 2 | 7.2k | Rent | |
| 2 | 7.2k | Rent | |
| 2 | 7.2k | Rent | |
| 2 | 7.2k | Rent | |
| 2 | 6.9k | Rent | |
| 2 | 7.2k | Specs | |
| 2 | 5.4k | Rent | |
| 2 | 4.8k | Rent | |
| 8 | 5.2k | Specs | |
| 4 | 8.2k | Rent | |
| 4 | 5.7k | Rent | |
| 4 | 5.6k | Specs | |
| 4 | ~8.2k | Rent | |
| 4 | ~6.7k | Rent | |
| 4 | 3.4k | Rent | |
| 4 | 3.4k | Specs | |
| 4 | 3.1k | Rent | |
| 4 | ~2.7k | Rent | |
| 4 | 3.1k | Rent | |
| 8 | 4.0k | Specs | |
GeForce RTX 5090 32GB 2 nodes | 16 | 10k | Rent |
| 4 | 5.6k | Specs | |
| 8 | ~3.1k | Rent | |
| 4 | 5.6k | Specs | |
| 8 | 2.8k | Rent | |
| 4 | 3.3k | Specs | |
| 4 | 3.3k | Rent | |
| 4 | 3.5k | Rent | |
| 8 | 5.0k | Specs | |
| 8 | 5.0k | Rent | |
GeForce RTX 5090 D V2 24GB 2 nodes | 16 | ~7.6k | Specs |
| 8 | ~3.1k | Specs | |
RTX 5000 Ada 32GB 2 nodes | 16 | ~3.3k | Rent |
| 8 | ~4.3k | Rent | |
| 8 | ~4.3k | Specs | |
RTX PRO 4500 Blackwell 32GB 2 nodes | 16 | ~5.1k | Rent |
Radeon AI PRO R9700 32GB 2 nodes | 16 | 3.6k | Specs |
Instinct MI100 32GB 2 nodes | 16 | 7.0k | Specs |
Arc Pro B70 32GB 2 nodes | 16 | ~3.5k | Specs |
| 8 | 2.8k | Rent | |
| 8 | 5.3k | Specs | |
GeForce RTX 4090 24GB 2 nodes | 16 | 5.7k | Rent |
A30 24GB 2 nodes | 16 | 5.3k | Rent |
RTX 4500 Ada 24GB 2 nodes | 16 | ~2.5k | Specs |
| 8 | 2.5k | Rent | |
| 8 | 2.2k | Rent | |
| 4 | 6.8k | Specs | |
| 4 | 6.8k | Specs | |
RTX PRO 4000 Blackwell 24GB 2 nodes | 16 | ~3.8k | Rent |
| 16 | 6.8k | Specs | |
RTX A5500 24GB 2 nodes | 16 | ~4.4k | Specs |
V100S PCIe 32GB 2 nodes | 16 | 6.5k | Rent |
A10 24GB 2 nodes | 16 | 3.4k | Rent |
| 8 | 2.8k | Specs | |
| 8 | 2.8k | Specs | |
Radeon RX 7900 XTX 24GB 2 nodes | 16 | 5.5k | Specs |
L4 24GB 2 nodes | 16 | 1.7k | Rent |
| 8 | 2.8k | Specs | |
Arc Pro B65 32GB 2 nodes | 16 | ~3.5k | Specs |
Arc Pro B60 24GB 2 nodes | 16 | ~2.6k | Specs |
L2 24GB 2 nodes | 16 | 1.7k | Specs |
RTX PRO 4000 Blackwell SFF 24GB 2 nodes | 16 | ~2.5k | Specs |
| 8 | 2.8k | Specs | |
Radeon PRO W7800 32GB 2 nodes | 16 | 3.3k | Specs |
GeForce RTX 3090 24GB 2 nodes | 16 | 5.3k | Rent |
A10G 24GB 2 nodes | 16 | 3.4k | Specs |
GeForce RTX 3090 Ti 24GB 2 nodes | 16 | 5.7k | Rent |
Radeon PRO W6800 32GB 2 nodes | 16 | 2.9k | Specs |
RTX A5000 24GB 2 nodes | 16 | 4.4k | Rent |
Radeon PRO V710 28GB 2 nodes | 16 | 2.5k | Specs |
AMD vs NVIDIA
At BF16 the NVIDIA pick is 4 x RTX PRO 6000 Blackwell Max-Q Workstation Edition at about 3.1k tokens/s for $2.00/hr; the AMD pick is 2 x Instinct MI300X 192GB at about 4.8k tokens/s for $4.78/hr. Per rental dollar NVIDIA delivers 1.53x the tokens of AMD here.
NVIDIA: 1.5k tokens/s per dollar, 2.5k tokens/s per kW (4 x RTX PRO 6000 Blackwell Max-Q Workstation Edition).AMD: 999 tokens/s per dollar, 3.2k tokens/s per kW (2 x Instinct MI300X 192GB).Intel: no live price, 2.3k tokens/s per kW (4 x Data Center GPU Max 1550 128GB).Other: no live price, 1.2k tokens/s per kW (8 x BR100).
Fleet what-if
Compare whole fleets at these settings: a thousand of one part against a hundred of another. Each fleet splits into replicas of its tensor-parallel width; GPUs left over sit idle.
| Fleet | Replicas | Tokens/s |
|---|---|---|
| A 1,000 x H200 SXM 141GB | 250 x TP4 | 2.05M |
| B 100 x GB200 NVL72 GPU 186GB | 50 x TP2 | 360k |
1,000 x H200 SXM 141GB: 2.05M tokens/s, $2,990/hr, 700 kW.100 x GB200 NVL72 GPU 186GB: 360k tokens/s, $1,050/hr, no power figure.
H200 SXM 141GB: 2.05M tokens/s at 1000 GPUs.GB200 NVL72 GPU 186GB: 3.60M tokens/s at 1000 GPUs.
Size Qwen3.5-122B-A10B for a tokens per second target
The calculator starts from your traffic instead of a fleet: set a target rate and read the replica count for every GPU.
Frequently asked questions
How much GPU memory does Qwen3.5-122B-A10B need?
The weights take 250 GB at the native BF16 precision (250 GB at BF16, 125 GB at FP8, 70.4 GB at INT4). Each concurrent stream adds 24 KB of KV cache per token: 0.1 GB at 4,096 tokens and 0.8 GB at 32,768 tokens.
What is the bare minimum to run Qwen3.5-122B-A10B?
1 x Instinct MI350X OAM at BF16 with 4,096 tokens of context and one stream. The cheapest live rental for that footprint is 4 x RTX PRO 6000 Blackwell Max-Q Workstation Edition at $2.00 per hour.
How many H100s do you need to run Qwen3.5-122B-A10B?
4 H100s at BF16 for 4,096 tokens of context and one stream. For 32,768 tokens and 32 concurrent streams the tensor-parallel group is 4 H100s, and 1,000 tokens per second takes 4 H100s across 1 replicas.
How much does Qwen3.5-122B-A10B cost per million tokens on an H100?
About $0.35 per million output tokens at BF16, 32,768 tokens of context and 32 streams, using the cheapest live on-demand price of $1.79 per GPU-hour. The cheapest part per token is 4 x RTX PRO 6000 Blackwell Max-Q Workstation Edition at $0.18.
What precision does Qwen3.5-122B-A10B ship in?
The published checkpoint is BF16. FP8 and INT4 figures describe post-training quantisations that halve and quarter the footprint at a small accuracy cost.
All figures are roofline estimates from the model's config.json and each GPU's published memory, bandwidth and tensor-core peak, with fixed efficiency factors. Prices are the cheapest live on-demand listing per GPU when the site has one. They are a planning floor; a tuned serving stack can do better. Only the 12 full-attention layers keep a per-token cache. The 36 linear-attention layers hold a fixed 151 MB state per sequence.

