Infrastructure Tools
LLM GPU Calculator
How many GPUs do you need to run a model at your tokens per second target? Pick a model, a precision and a traffic profile to see which datacenter GPUs fit and how many of each you need.
Weights at BF16
141 GB
Active parameters
70.6B
dense model
KV cache per token
320 KB
KV cache per stream
10.7 GB
at 32,768 tokens
Memory needed
485 GB
32 concurrent streams
The Meta repo is gated; geometry read from the unsloth mirror, which ships the same config.
GPUs for 1,000 tokens/s 45 of 73 parts fit at these settings
Roofline estimates. Figures marked ~ use a derived tensor-core rate.
| GPU | Memory | Total GPUs | |
|---|---|---|---|
| 432 GB | 2 | Specs | |
| 288 GB | 2 | Specs | |
| 288 GB | 4 | Specs | |
| 288 GB | 4 | Rent | |
| 186 GB | 4 | Rent | |
| 288 GB | 4 | Rent | |
| 262.5 GB | 4 | Rent | |
| 180 GB | 4 | Rent | |
| 192 GB | 4 | Specs | |
| 256 GB | 4 | Rent | |
| 192 GB | 8 | Rent | |
| 64 GB | 16 | Specs | |
| 141 GB | 8 | Rent | |
| 80 GB | 8 | Rent | |
| 128 GB | 8 | Specs | |
| 141 GB | 8 | Rent | |
| 94 GB | 8 | Rent | |
| 80 GB | 16 | Rent | |
| 80 GB | 16 | Specs | |
| 96 GB | 16 | Rent | |
| 96 GB | 16 | Rent | |
| 96 GB | 16 | Rent | |
Data Center GPU Max 1100 48GB 2 nodes | 48 GB | 32 | Specs |
| 32 GB | |||
| 128 GB | 8 | Specs | |
RTX 6000 Ada 48GB 2 nodes | 48 GB | 32 | Rent |
| 128 GB | 8 | Specs | |
L40S 48GB 2 nodes | 48 GB | 32 | Rent |
| 80 GB | 16 | Specs | |
| 80 GB | 16 | Rent | |
| 80 GB | 16 | Rent | |
A100 PCIe 40GB 2 nodes | 40 GB | 32 | Specs |
A100 SXM4 40GB 2 nodes | 40 GB | 32 | Rent |
| 24 GB | |||
RTX 5880 Ada 48GB 2 nodes | 48 GB | 32 | Specs |
| 32 GB | |||
RTX PRO 5000 Blackwell 48GB 2 nodes | 48 GB | 32 | Rent |
| 72 GB | 24 | Specs | |
| 32 GB | |||
| 32 GB | |||
| 32 GB | |||
| 32 GB | |||
L40 48GB 2 nodes | 48 GB | 32 | Rent |
| 64 GB | 16 | Specs | |
| 24 GB | |||
| 24 GB | |||
| 24 GB | |||
RTX A6000 48GB 2 nodes | 48 GB | 48 | Rent |
A40 48GB 2 nodes | 48 GB | 48 | Rent |
| 141 GB | 8 | Specs | |
| 96 GB | 8 | Specs | |
| 24 GB | |||
| 32 GB | |||
| 24 GB | |||
| 32 GB | |||
| 24 GB | |||
Radeon PRO W7900 Dual Slot 48GB 2 nodes | 48 GB | 32 | Specs |
Radeon PRO W7900 48GB 2 nodes | 48 GB | 32 | Specs |
| 24 GB | |||
| 24 GB | |||
L20 48GB 2 nodes | 48 GB | 32 | Specs |
| 32 GB | |||
| 24 GB | |||
| 24 GB | |||
| 24 GB | |||
Radeon PRO W7800 48GB 2 nodes | 48 GB | 32 | Specs |
| 32 GB | |||
| 24 GB | |||
| 24 GB | |||
| 24 GB | |||
| 32 GB | |||
| 24 GB | |||
| 28 GB |
How this is calculated
- Weights are parameters times bytes per parameter: 2 at BF16, 1 at FP8, 0.5625 at INT4 or FP4 including block scales. The figure is the full checkpoint, not the active subset.
- KV cache follows the model's config.json: layers times KV heads times head dimension, K and V, in 16-bit. Sliding-window layers are charged for the window only, MLA models for their latent, and linear-attention layers for a fixed state per stream.
- A GPU offers 90 percent of its memory to weights and cache, less 1 GB of workspace. Tensor parallel widens across 1, 2, 4, 8 and 16 GPUs until the pool fits.
- Decode is bandwidth-bound: every step streams the active weights and every stream's cache at 75 percent of peak bandwidth, with a compute floor at 60 percent of the dense tensor peak. FP8 weights compute at FP8 where the part has it; 4-bit weights compute at 16-bit unless the part has FP4 tensor cores.
- Replicas are the target rate divided by one replica's rate, rounded up. Total GPUs is replicas times the tensor-parallel width. Measured throughput from a tuned serving stack can beat these numbers; they are a planning floor, not a benchmark.

