Alibaba Qwen3

Qwen3-32B hardware requirements

Fits on one Instinct MI350X OAM at BF16. For 1,000 tokens/s at 32k context you need 8 x H100 SXM5 80GB at $14.32/hr or 12 x RTX PRO 6000 Blackwell Max-Q Workstation Edition at $6.00/hr.

Parameters
32.8B
Active per token
All
dense model
KV cache per token
256 KB
16-bit cache
Native precision
BF16

Bare minimum

Fewest GPUs
1 x Instinct MI350X OAM
BF16, 4,096 tokens, one stream
Tokens/s
90.1
Per hour
$6.16
Per million tokens
$18.99
Memory used
24%
Rent on DigitalOcean
Cheapest per hour
1 x RTX PRO 6000 Blackwell Max-Q Workstation Edition
BF16, 4,096 tokens, one stream, live price
Tokens/s
20.2
Per hour
$0.50
Per million tokens
$6.88
Memory used
73%
Rent on RunPod
Cheapest per million tokens
4 x RTX PRO 6000 Blackwell Max-Q Workstation Edition
BF16, 32,768 tokens, 32 streams
Tokens/s
455
Per hour
$2.00
Per million tokens
$1.22
Memory used
93%
Rent on RunPod

Configurations

Every datacenter GPU in the catalog, sized for these settings. Change precision, context and concurrency and the whole page follows. Each point is one replica: the smallest tensor-parallel group of that GPU that holds the model.

Vendor

2 x GB300 NVL72 GPU 288GB: 1.1k tokens/s per replica at $17.24/hr.2 x GB200 NVL72 GPU 186GB: 1.1k tokens/s per replica at $21.00/hr.2 x Instinct MI350X OAM: 1.1k tokens/s per replica at $12.32/hr.2 x B300 SXM 262GB: 1.1k tokens/s per replica at $14.80/hr.2 x B200 SXM 180GB: 1.0k tokens/s per replica at $8.18/hr.2 x Instinct MI325X OAM: 804 tokens/s per replica at $7.60/hr.2 x Instinct MI300X 192GB: 710 tokens/s per replica at $4.78/hr.4 x H200 SXM 141GB: 1.2k tokens/s per replica at $11.96/hr.8 x H100 SXM5 80GB: 1.6k tokens/s per replica at $14.32/hr.4 x H200 NVL 141GB: 1.2k tokens/s per replica at $15.16/hr.4 x H100 NVL 94GB: 990 tokens/s per replica at $12.76/hr.8 x H100 PCIe 80GB: 959 tokens/s per replica at $20.00/hr.4 x RTX PRO 6000 Blackwell Workstation Edition: 455 tokens/s per replica at $7.20/hr.4 x RTX PRO 6000 Blackwell Server Edition: 405 tokens/s per replica at $2.36/hr.4 x RTX PRO 6000 Blackwell Max-Q Workstation Edition: 455 tokens/s per replica at $2.00/hr.16 x GeForce RTX 5090 32GB: 1.5k tokens/s per replica at $15.84/hr.8 x RTX 6000 Ada 48GB: 460 tokens/s per replica at $6.72/hr.8 x L40S 48GB: 414 tokens/s per replica at $6.96/hr.8 x A100 PCIe 80GB: 928 tokens/s per replica at $10.80/hr.8 x A100 SXM4 80GB: 978 tokens/s per replica at $11.20/hr.16 x A100 SXM4 40GB: 1.3k tokens/s per replica at $20.64/hr.16 x RTX 5000 Ada 32GB: 487 tokens/s per replica at $13.28/hr.8 x RTX PRO 5000 Blackwell 48GB: 644 tokens/s per replica at $7.68/hr.16 x RTX PRO 4500 Blackwell 32GB: 758 tokens/s per replica at $11.52/hr.8 x L40 48GB: 414 tokens/s per replica at $6.56/hr.16 x GeForce RTX 4090 24GB: 853 tokens/s per replica at $9.60/hr.16 x A30 24GB: 789 tokens/s per replica at $11.73/hr.8 x RTX A6000 48GB: 368 tokens/s per replica at $4.00/hr.8 x A40 48GB: 334 tokens/s per replica at $3.92/hr.16 x RTX PRO 4000 Blackwell 24GB: 569 tokens/s per replica at $9.12/hr.16 x V100S PCIe 32GB: 959 tokens/s per replica at $14.08/hr.16 x A10 24GB: 508 tokens/s per replica at $20.64/hr.16 x L4 24GB: 254 tokens/s per replica at $7.84/hr.16 x GeForce RTX 3090 24GB: 792 tokens/s per replica at $8.00/hr.16 x GeForce RTX 3090 Ti 24GB: 853 tokens/s per replica at $7.36/hr.16 x RTX A5000 24GB: 650 tokens/s per replica at $4.32/hr.

Priced configurations only; 37 more fit without a live price and appear in the table. Marker size grows with the tensor-parallel width. Roofline estimates.

GPUTPTokens/s
1 ~1.6k Specs
2 2.9k Specs
2 1.1k Specs
2 1.1k Rent
2 1.1k Rent
2 1.1k Rent
2 1.1k Rent
2 1.0k Rent
2 1.1k Specs
2 804 Rent
2 710 Rent
8 767 Specs
4 1.2k Rent
8 1.6k Rent
4 832 Specs
4 ~1.2k Rent
4 ~990 Rent
8 959 Rent
8 959 Specs
4 455 Rent
4 ~405 Rent
4 455 Rent
8 589 Specs
16 1.5k Rent
4 832 Specs
8 ~460 Rent
4 832 Specs
8 414 Rent
8 928 Specs
8 928 Rent
8 978 Rent
16 1.3k Specs
16 1.3k Rent
16 ~1.1k Specs
8 ~460 Specs
16 ~487 Rent
8 ~644 Rent
8 ~644 Specs
16 ~758 Rent
16 541 Specs
16 1.0k Specs
16 ~514 Specs
8 414 Rent
8 786 Specs
16 853 Rent
A30 24GB 2 nodes
16 789 Rent
16 ~365 Specs
8 368 Rent
8 334 Rent
4 1.0k Specs
4 1.0k Specs
16 ~569 Rent
16 1.0k Specs
16 ~650 Specs
16 959 Rent
A10 24GB 2 nodes
16 508 Rent
8 414 Specs
8 414 Specs
16 812 Specs
L4 24GB 2 nodes
16 254 Rent
8 414 Specs
16 ~514 Specs
16 ~386 Specs
L2 24GB 2 nodes
16 254 Specs
16 ~365 Specs
8 414 Specs
16 487 Specs
16 792 Rent
A10G 24GB 2 nodes
16 508 Specs
16 853 Rent
16 433 Specs
16 650 Rent
16 379 Specs

AMD vs NVIDIA

At BF16 the NVIDIA pick is 4 x RTX PRO 6000 Blackwell Max-Q Workstation Edition at about 455 tokens/s for $2.00/hr; the AMD pick is 2 x Instinct MI300X 192GB at about 710 tokens/s for $4.78/hr. Per rental dollar NVIDIA delivers 1.53x the tokens of AMD here.

NVIDIA: 227 tokens/s per dollar, 379 tokens/s per kW (4 x RTX PRO 6000 Blackwell Max-Q Workstation Edition).AMD: 149 tokens/s per dollar, 473 tokens/s per kW (2 x Instinct MI300X 192GB).Intel: no live price, 347 tokens/s per kW (4 x Data Center GPU Max 1550 128GB).Other: no live price, 174 tokens/s per kW (8 x BR100).

NVIDIA 49 parts fit
Best value: 4 x RTX PRO 6000 Blackwell Max-Q Workstation Edition, 455 tok/s for $2.00/hr
Fewest GPUs: 2 x Rubin SXM, 2.9k tok/s
AMD 17 parts fit
Best value: 2 x Instinct MI300X 192GB, 710 tok/s for $4.78/hr
Fewest GPUs: 1 x Instinct MI455X OAM, 1.6k tok/s
Intel 5 parts fit
Fewest GPUs: 4 x Data Center GPU Max 1550 128GB, 832 tok/s
Other 2 parts fit
Fewest GPUs: 8 x BR100, 767 tok/s

Fleet what-if

Compare whole fleets at these settings: a thousand of one part against a hundred of another. Each fleet splits into replicas of its tensor-parallel width; GPUs left over sit idle.

A
B
FleetReplicasTokens/s
A 1,000 x H200 SXM 141GB 250 x TP4 305k
B 100 x GB200 NVL72 GPU 186GB 50 x TP2 54k

1,000 x H200 SXM 141GB: 305k tokens/s, $2,990/hr, 700 kW.100 x GB200 NVL72 GPU 186GB: 54k tokens/s, $1,050/hr, no power figure.

H200 SXM 141GB: 305k tokens/s at 1000 GPUs.GB200 NVL72 GPU 186GB: 536k tokens/s at 1000 GPUs.

Size Qwen3-32B for a tokens per second target

The calculator starts from your traffic instead of a fleet: set a target rate and read the replica count for every GPU.

Frequently asked questions

How much GPU memory does Qwen3-32B need?

The weights take 65.5 GB at the native BF16 precision (65.5 GB at BF16, 32.8 GB at FP8, 18.4 GB at INT4). Each concurrent stream adds 256 KB of KV cache per token: 1.1 GB at 4,096 tokens and 8.6 GB at 32,768 tokens.

What is the bare minimum to run Qwen3-32B?

1 x Instinct MI350X OAM at BF16 with 4,096 tokens of context and one stream. The cheapest live rental for that footprint is 1 x RTX PRO 6000 Blackwell Max-Q Workstation Edition at $0.50 per hour.

How many H100s do you need to run Qwen3-32B?

one H100 at BF16 for 4,096 tokens of context and one stream. For 32,768 tokens and 32 concurrent streams the tensor-parallel group is 8 H100s, and 1,000 tokens per second takes 8 H100s across 1 replicas.

How much does Qwen3-32B cost per million tokens on an H100?

About $2.48 per million output tokens at BF16, 32,768 tokens of context and 32 streams, using the cheapest live on-demand price of $1.79 per GPU-hour. The cheapest part per token is 4 x RTX PRO 6000 Blackwell Max-Q Workstation Edition at $1.22.

What precision does Qwen3-32B ship in?

The published checkpoint is BF16. FP8 and INT4 figures describe post-training quantisations that halve and quarter the footprint at a small accuracy cost.

All figures are roofline estimates from the model's config.json and each GPU's published memory, bandwidth and tensor-core peak, with fixed efficiency factors. Prices are the cheapest live on-demand listing per GPU when the site has one. They are a planning floor; a tuned serving stack can do better.