Infrastructure Tools

LLM GPU Calculator

How many GPUs do you need to run a model at your tokens per second target? Pick a model, a precision and a traffic profile to see which datacenter GPUs fit and how many of each you need.

Start from a model instead

Llama 3.3 70B Instruct

Meta Native BF16 Hardware requirements page
Weights at BF16
141 GB
Active parameters
70.6B
dense model
KV cache per token
320 KB
KV cache per stream
10.7 GB
at 32,768 tokens
Memory needed
485 GB
32 concurrent streams

The Meta repo is gated; geometry read from the unsloth mirror, which ships the same config.

GPUs for 1,000 tokens/s 45 of 73 parts fit at these settings

Roofline estimates. Figures marked ~ use a derived tensor-core rate.

GPUMemoryTotal GPUs
432 GB 2 Specs
288 GB 2 Specs
288 GB 4 Specs
288 GB 4 Rent
186 GB 4 Rent
288 GB 4 Rent
262.5 GB 4 Rent
180 GB 4 Rent
192 GB 4 Specs
256 GB 4 Rent
192 GB 8 Rent
64 GB 16 Specs
141 GB 8 Rent
80 GB 8 Rent
128 GB 8 Specs
141 GB 8 Rent
94 GB 8 Rent
80 GB 16 Rent
80 GB 16 Specs
96 GB 16 Rent
96 GB 16 Rent
96 GB 16 Rent
48 GB 32 Specs
32 GB
128 GB 8 Specs
48 GB 32 Rent
128 GB 8 Specs
L40S 48GB 2 nodes
48 GB 32 Rent
80 GB 16 Specs
80 GB 16 Rent
80 GB 16 Rent
40 GB 32 Specs
40 GB 32 Rent
24 GB
48 GB 32 Specs
32 GB
48 GB 32 Rent
72 GB 24 Specs
32 GB
32 GB
32 GB
32 GB
L40 48GB 2 nodes
48 GB 32 Rent
64 GB 16 Specs
24 GB
24 GB
24 GB
48 GB 48 Rent
A40 48GB 2 nodes
48 GB 48 Rent
141 GB 8 Specs
96 GB 8 Specs
24 GB
32 GB
24 GB
32 GB
24 GB
48 GB 32 Specs
48 GB 32 Specs
24 GB
24 GB
L20 48GB 2 nodes
48 GB 32 Specs
32 GB
24 GB
24 GB
24 GB
48 GB 32 Specs
32 GB
24 GB
24 GB
24 GB
32 GB
24 GB
28 GB

How this is calculated

  • Weights are parameters times bytes per parameter: 2 at BF16, 1 at FP8, 0.5625 at INT4 or FP4 including block scales. The figure is the full checkpoint, not the active subset.
  • KV cache follows the model's config.json: layers times KV heads times head dimension, K and V, in 16-bit. Sliding-window layers are charged for the window only, MLA models for their latent, and linear-attention layers for a fixed state per stream.
  • A GPU offers 90 percent of its memory to weights and cache, less 1 GB of workspace. Tensor parallel widens across 1, 2, 4, 8 and 16 GPUs until the pool fits.
  • Decode is bandwidth-bound: every step streams the active weights and every stream's cache at 75 percent of peak bandwidth, with a compute floor at 60 percent of the dense tensor peak. FP8 weights compute at FP8 where the part has it; 4-bit weights compute at 16-bit unless the part has FP4 tensor cores.
  • Replicas are the target rate divided by one replica's rate, rounded up. Total GPUs is replicas times the tensor-parallel width. Measured throughput from a tuned serving stack can beat these numbers; they are a planning floor, not a benchmark.