Meta Llama 3

Llama 3.3 70B Instruct hardware requirements

Fits on one Instinct MI350X OAM at BF16. For 1,000 tokens/s at 32k context you need 8 x H100 SXM5 80GB at $14.32/hr or 16 x RTX PRO 6000 Blackwell Max-Q Workstation Edition at $8.00/hr.

Parameters
70.6B
Active per token
All
dense model
KV cache per token
320 KB
16-bit cache
Native precision
BF16

Bare minimum

Fewest GPUs
1 x Instinct MI350X OAM
BF16, 4,096 tokens, one stream
Tokens/s
42.1
Per hour
$6.16
Per million tokens
$40.62
Memory used
51%
Rent on DigitalOcean
Cheapest per hour
2 x RTX PRO 6000 Blackwell Max-Q Workstation Edition
BF16, 4,096 tokens, one stream, live price
Tokens/s
17.9
Per hour
$1.00
Per million tokens
$15.50
Memory used
78%
Rent on RunPod
Cheapest per million tokens
8 x RTX PRO 6000 Blackwell Max-Q Workstation Edition
BF16, 32,768 tokens, 32 streams
Tokens/s
603
Per hour
$4.00
Per million tokens
$1.84
Memory used
66%
Rent on RunPod

Configurations

Every datacenter GPU in the catalog, sized for these settings. Change precision, context and concurrency and the whole page follows. Each point is one replica: the smallest tensor-parallel group of that GPU that holds the model.

Vendor

2 x GB300 NVL72 GPU 288GB: 753 tokens/s per replica at $17.24/hr.4 x GB200 NVL72 GPU 186GB: 1.4k tokens/s per replica at $42.00/hr.2 x Instinct MI350X OAM: 753 tokens/s per replica at $12.32/hr.2 x B300 SXM 262GB: 753 tokens/s per replica at $14.80/hr.4 x B200 SXM 180GB: 1.4k tokens/s per replica at $16.36/hr.2 x Instinct MI325X OAM: 564 tokens/s per replica at $7.60/hr.4 x Instinct MI300X 192GB: 945 tokens/s per replica at $9.56/hr.4 x H200 SXM 141GB: 856 tokens/s per replica at $11.96/hr.8 x H100 SXM5 80GB: 1.1k tokens/s per replica at $14.32/hr.4 x H200 NVL 141GB: 856 tokens/s per replica at $15.16/hr.8 x H100 NVL 94GB: 1.3k tokens/s per replica at $25.52/hr.8 x H100 PCIe 80GB: 673 tokens/s per replica at $20.00/hr.8 x RTX PRO 6000 Blackwell Workstation Edition: 603 tokens/s per replica at $14.40/hr.8 x RTX PRO 6000 Blackwell Server Edition: 538 tokens/s per replica at $4.72/hr.8 x RTX PRO 6000 Blackwell Max-Q Workstation Edition: 603 tokens/s per replica at $4.00/hr.16 x RTX 6000 Ada 48GB: 570 tokens/s per replica at $13.44/hr.16 x L40S 48GB: 513 tokens/s per replica at $13.92/hr.8 x A100 PCIe 80GB: 652 tokens/s per replica at $10.80/hr.8 x A100 SXM4 80GB: 687 tokens/s per replica at $11.20/hr.16 x A100 SXM4 40GB: 924 tokens/s per replica at $20.64/hr.16 x RTX PRO 5000 Blackwell 48GB: 799 tokens/s per replica at $15.36/hr.16 x L40 48GB: 513 tokens/s per replica at $13.12/hr.16 x RTX A6000 48GB: 456 tokens/s per replica at $8.00/hr.16 x A40 48GB: 414 tokens/s per replica at $7.84/hr.

Priced configurations only; 21 more fit without a live price and appear in the table. Marker size grows with the tensor-parallel width. Roofline estimates.

GPUTPTokens/s
2 ~2.2k Specs
2 2.1k Specs
2 753 Specs
2 753 Rent
4 1.4k Rent
2 753 Rent
2 753 Rent
4 1.4k Rent
4 1.4k Specs
2 564 Rent
4 945 Rent
8 539 Specs
4 856 Rent
8 1.1k Rent
4 584 Specs
4 ~856 Rent
8 ~1.3k Rent
8 673 Rent
8 673 Specs
8 603 Rent
8 ~538 Rent
8 603 Rent
16 730 Specs
4 584 Specs
16 ~570 Rent
4 584 Specs
L40S 48GB 2 nodes
16 513 Rent
8 652 Specs
8 652 Rent
8 687 Rent
16 924 Specs
16 924 Rent
16 ~570 Specs
16 ~799 Rent
8 ~453 Specs
L40 48GB 2 nodes
16 513 Rent
8 552 Specs
16 456 Rent
A40 48GB 2 nodes
16 414 Rent
4 713 Specs
8 1.3k Specs
16 513 Specs
16 513 Specs
L20 48GB 2 nodes
16 513 Specs
16 513 Specs

AMD vs NVIDIA

At BF16 the NVIDIA pick is 8 x RTX PRO 6000 Blackwell Max-Q Workstation Edition at about 603 tokens/s for $4.00/hr; the AMD pick is 4 x Instinct MI300X 192GB at about 945 tokens/s for $9.56/hr. Per rental dollar NVIDIA delivers 1.53x the tokens of AMD here.

NVIDIA: 151 tokens/s per dollar, 251 tokens/s per kW (8 x RTX PRO 6000 Blackwell Max-Q Workstation Edition).AMD: 98.8 tokens/s per dollar, 315 tokens/s per kW (4 x Instinct MI300X 192GB).Intel: no live price, 243 tokens/s per kW (4 x Data Center GPU Max 1550 128GB).Other: no live price, 122 tokens/s per kW (8 x BR100).

NVIDIA 31 parts fit
Best value: 8 x RTX PRO 6000 Blackwell Max-Q Workstation Edition, 603 tok/s for $4.00/hr
Fewest GPUs: 2 x Rubin SXM, 2.1k tok/s
AMD 11 parts fit
Best value: 4 x Instinct MI300X 192GB, 945 tok/s for $9.56/hr
Fewest GPUs: 2 x Instinct MI455X OAM, 2.2k tok/s
Intel 2 parts fit
Fewest GPUs: 4 x Data Center GPU Max 1550 128GB, 584 tok/s
Other 1 parts fit
Fewest GPUs: 8 x BR100, 539 tok/s

Fleet what-if

Compare whole fleets at these settings: a thousand of one part against a hundred of another. Each fleet splits into replicas of its tensor-parallel width; GPUs left over sit idle.

A
B
FleetReplicasTokens/s
A 1,000 x H200 SXM 141GB 250 x TP4 214k
B 100 x GB200 NVL72 GPU 186GB 25 x TP4 36k

1,000 x H200 SXM 141GB: 214k tokens/s, $2,990/hr, 700 kW.100 x GB200 NVL72 GPU 186GB: 36k tokens/s, $1,050/hr, no power figure.

H200 SXM 141GB: 214k tokens/s at 1000 GPUs.GB200 NVL72 GPU 186GB: 357k tokens/s at 1000 GPUs.

Size Llama 3.3 70B Instruct for a tokens per second target

The calculator starts from your traffic instead of a fleet: set a target rate and read the replica count for every GPU.

Frequently asked questions

How much GPU memory does Llama 3.3 70B Instruct need?

The weights take 141 GB at the native BF16 precision (141 GB at BF16, 70.6 GB at FP8, 39.7 GB at INT4). Each concurrent stream adds 320 KB of KV cache per token: 1.3 GB at 4,096 tokens and 10.7 GB at 32,768 tokens.

What is the bare minimum to run Llama 3.3 70B Instruct?

1 x Instinct MI350X OAM at BF16 with 4,096 tokens of context and one stream. The cheapest live rental for that footprint is 2 x RTX PRO 6000 Blackwell Max-Q Workstation Edition at $1.00 per hour.

How many H100s do you need to run Llama 3.3 70B Instruct?

2 H100s at BF16 for 4,096 tokens of context and one stream. For 32,768 tokens and 32 concurrent streams the tensor-parallel group is 8 H100s, and 1,000 tokens per second takes 8 H100s across 1 replicas.

How much does Llama 3.3 70B Instruct cost per million tokens on an H100?

About $3.53 per million output tokens at BF16, 32,768 tokens of context and 32 streams, using the cheapest live on-demand price of $1.79 per GPU-hour. The cheapest part per token is 8 x RTX PRO 6000 Blackwell Max-Q Workstation Edition at $1.84.

What precision does Llama 3.3 70B Instruct ship in?

The published checkpoint is BF16. FP8 and INT4 figures describe post-training quantisations that halve and quarter the footprint at a small accuracy cost.

All figures are roofline estimates from the model's config.json and each GPU's published memory, bandwidth and tensor-core peak, with fixed efficiency factors. Prices are the cheapest live on-demand listing per GPU when the site has one. They are a planning floor; a tuned serving stack can do better. The Meta repo is gated; geometry read from the unsloth mirror, which ships the same config.