Z.ai GLM-5

GLM-5.2 hardware requirements

Needs at least 8 GPUs at BF16 (Instinct MI350X OAM). For 1,000 tokens/s at 32k context you need 16 x Instinct MI300X 192GB at $38.24/hr.

Parameters
753B
Active per token
41.7B
mixture of experts
KV cache per token
88 KB
16-bit cache
Native precision
BF16

Bare minimum

Fewest GPUs
8 x Instinct MI350X OAM
BF16, 4,096 tokens, one stream
Tokens/s
487
Per hour
$49.28
Per million tokens
$28.08
Memory used
68%
Rent on DigitalOcean
Cheapest per hour
8 x Instinct MI325X OAM
BF16, 4,096 tokens, one stream, live price
Tokens/s
366
Per hour
$30.40
Per million tokens
$23.10
Memory used
76%
Rent on DigitalOcean
Cheapest per million tokens
16 x Instinct MI300X 192GB
BF16, 32,768 tokens, 32 streams
Tokens/s
8.6k
Per hour
$38.24
Per million tokens
$1.24
Memory used
54%
Rent on RunPod

Configurations

Every datacenter GPU in the catalog, sized for these settings. Change precision, context and concurrency and the whole page follows. Each point is one replica: the smallest tensor-parallel group of that GPU that holds the model.

Vendor

8 x GB300 NVL72 GPU 288GB: 7.4k tokens/s per replica at $68.96/hr.16 x GB200 NVL72 GPU 186GB: 13k tokens/s per replica at $168/hr.8 x Instinct MI350X OAM: 7.4k tokens/s per replica at $49.28/hr.8 x B300 SXM 262GB: 7.4k tokens/s per replica at $59.20/hr.16 x B200 SXM 180GB: 12k tokens/s per replica at $65.44/hr.8 x Instinct MI325X OAM: 5.5k tokens/s per replica at $30.40/hr.16 x Instinct MI300X 192GB: 8.6k tokens/s per replica at $38.24/hr.16 x H200 SXM 141GB: 7.8k tokens/s per replica at $47.84/hr.16 x H200 NVL 141GB: 7.8k tokens/s per replica at $60.64/hr.

Priced configurations only; 8 more fit without a live price and appear in the table. Marker size grows with the tensor-parallel width. Roofline estimates.

GPUTPTokens/s
4 ~11k Specs
8 20k Specs
8 7.4k Specs
8 7.4k Rent
16 13k Rent
8 7.4k Rent
8 7.4k Rent
16 12k Rent
16 13k Specs
8 5.5k Rent
16 8.6k Rent
16 7.8k Rent
16 5.3k Specs
16 ~7.8k Rent
MI250X 128GB 2 nodes
16 5.3k Specs
16 5.3k Specs
16 6.5k Specs

AMD vs NVIDIA

At BF16 the NVIDIA pick is 16 x B200 SXM 180GB at about 12k tokens/s for $65.44/hr; the AMD pick is 16 x Instinct MI300X 192GB at about 8.6k tokens/s for $38.24/hr. Per rental dollar AMD delivers 1.18x the tokens of NVIDIA here.

NVIDIA: 191 tokens/s per dollar, 781 tokens/s per kW (16 x B200 SXM 180GB).AMD: 225 tokens/s per dollar, 716 tokens/s per kW (16 x Instinct MI300X 192GB).Intel: no live price, 554 tokens/s per kW (16 x Data Center GPU Max 1550 128GB).

NVIDIA 9 parts fit
Best value: 16 x B200 SXM 180GB, 12k tok/s for $65.44/hr
Fewest GPUs: 8 x Rubin SXM, 20k tok/s
AMD 7 parts fit
Best value: 16 x Instinct MI300X 192GB, 8.6k tok/s for $38.24/hr
Fewest GPUs: 4 x Instinct MI455X OAM, 11k tok/s
Intel 1 parts fit
Fewest GPUs: 16 x Data Center GPU Max 1550 128GB, 5.3k tok/s

Fleet what-if

Compare whole fleets at these settings: a thousand of one part against a hundred of another. Each fleet splits into replicas of its tensor-parallel width; GPUs left over sit idle.

A
B
FleetReplicasTokens/s
A 1,008 x H200 SXM 141GB 63 x TP16 490k
B 112 x GB200 NVL72 GPU 186GB 7 x TP16 91k

1,008 x H200 SXM 141GB: 490k tokens/s, $3,014/hr, 706 kW.112 x GB200 NVL72 GPU 186GB: 91k tokens/s, $1,176/hr, no power figure.

H200 SXM 141GB: 490k tokens/s at 1008 GPUs.GB200 NVL72 GPU 186GB: 817k tokens/s at 1008 GPUs.

Size GLM-5.2 for a tokens per second target

The calculator starts from your traffic instead of a fleet: set a target rate and read the replica count for every GPU.

Frequently asked questions

How much GPU memory does GLM-5.2 need?

The weights take 1,507 GB at the native BF16 precision (1,507 GB at BF16, 753 GB at FP8, 424 GB at INT4). Each concurrent stream adds 88 KB of KV cache per token: 0.4 GB at 4,096 tokens and 2.9 GB at 32,768 tokens.

What is the bare minimum to run GLM-5.2?

8 x Instinct MI350X OAM at BF16 with 4,096 tokens of context and one stream. The cheapest live rental for that footprint is 8 x Instinct MI325X OAM at $30.40 per hour.

How many H100s do you need to run GLM-5.2?

more than sixteen H100s at BF16 for 4,096 tokens of context and one stream. For 32,768 tokens and 32 concurrent streams the tensor-parallel group is more than sixteen H100s.

How much does GLM-5.2 cost per million tokens on an H100?

No live H100 rental price is on file right now, so the cost per token cannot be quoted. The configuration table lists every part that has one.

What precision does GLM-5.2 ship in?

The published checkpoint is BF16. FP8 and INT4 figures describe post-training quantisations that halve and quarter the footprint at a small accuracy cost.

All figures are roofline estimates from the model's config.json and each GPU's published memory, bandwidth and tensor-core peak, with fixed efficiency factors. Prices are the cheapest live on-demand listing per GPU when the site has one. They are a planning floor; a tuned serving stack can do better. Active parameters computed from the expert geometry, net of the multi-token-prediction layer. A sparse-attention indexer sits over the MLA cache; the cache figure ignores the indexer.