OpenAI gpt-oss

gpt-oss-20b hardware requirements

Fits on one Instinct MI350X OAM at INT4 / FP4. For 1,000 tokens/s at 32k context you need 1 x H100 SXM5 80GB at $1.79/hr or 1 x RTX PRO 6000 Blackwell Max-Q Workstation Edition at $0.50/hr.

Parameters
20.9B
Active per token
3.6B
mixture of experts
KV cache per token
48 KB
16-bit cache
Native precision
MXFP4

Bare minimum

Fewest GPUs
1 x Instinct MI350X OAM
INT4 / FP4, 4,096 tokens, one stream
Tokens/s
2.8k
Per hour
$6.16
Per million tokens
$0.61
Memory used
4%
Rent on DigitalOcean
Cheapest per hour
1 x RTX A5000 24GB
INT4 / FP4, 4,096 tokens, one stream, live price
Tokens/s
271
Per hour
$0.27
Per million tokens
$0.28
Memory used
54%
Rent on RunPod
Cheapest per million tokens
1 x RTX PRO 6000 Blackwell Max-Q Workstation Edition
INT4 / FP4, 32,768 tokens, 32 streams
Tokens/s
1.5k
Per hour
$0.50
Per million tokens
$0.09
Memory used
41%
Rent on RunPod

Configurations

Every datacenter GPU in the catalog, sized for these settings. Change precision, context and concurrency and the whole page follows. Each point is one replica: the smallest tensor-parallel group of that GPU that holds the model.

Vendor

1 x GB300 NVL72 GPU 288GB: 6.9k tokens/s per replica at $8.62/hr.1 x GB200 NVL72 GPU 186GB: 6.9k tokens/s per replica at $10.50/hr.1 x Instinct MI350X OAM: 6.9k tokens/s per replica at $6.16/hr.1 x B300 SXM 262GB: 6.9k tokens/s per replica at $7.40/hr.1 x B200 SXM 180GB: 6.6k tokens/s per replica at $4.09/hr.1 x Instinct MI325X OAM: 5.2k tokens/s per replica at $3.80/hr.1 x Instinct MI300X 192GB: 4.6k tokens/s per replica at $2.39/hr.1 x H200 SXM 141GB: 4.1k tokens/s per replica at $2.99/hr.1 x H100 SXM5 80GB: 2.9k tokens/s per replica at $1.79/hr.1 x H200 NVL 141GB: 4.1k tokens/s per replica at $3.79/hr.1 x H100 NVL 94GB: 3.4k tokens/s per replica at $3.19/hr.1 x H100 PCIe 80GB: 1.7k tokens/s per replica at $2.50/hr.1 x RTX PRO 6000 Blackwell Workstation Edition: 1.5k tokens/s per replica at $1.80/hr.1 x RTX PRO 6000 Blackwell Server Edition: 1.4k tokens/s per replica at $0.59/hr.1 x RTX PRO 6000 Blackwell Max-Q Workstation Edition: 1.5k tokens/s per replica at $0.50/hr.2 x GeForce RTX 5090 32GB: 2.9k tokens/s per replica at $1.98/hr.1 x RTX 6000 Ada 48GB: 826 tokens/s per replica at $0.84/hr.1 x L40S 48GB: 743 tokens/s per replica at $0.87/hr.1 x A100 PCIe 80GB: 1.7k tokens/s per replica at $1.35/hr.1 x A100 SXM4 80GB: 1.8k tokens/s per replica at $1.40/hr.2 x A100 SXM4 40GB: 2.5k tokens/s per replica at $2.58/hr.2 x RTX 5000 Ada 32GB: 942 tokens/s per replica at $1.66/hr.1 x RTX PRO 5000 Blackwell 48GB: 1.2k tokens/s per replica at $0.96/hr.2 x RTX PRO 4500 Blackwell 32GB: 1.5k tokens/s per replica at $1.44/hr.1 x L40 48GB: 743 tokens/s per replica at $0.82/hr.2 x GeForce RTX 4090 24GB: 1.6k tokens/s per replica at $1.20/hr.2 x A30 24GB: 1.5k tokens/s per replica at $1.47/hr.1 x RTX A6000 48GB: 661 tokens/s per replica at $0.50/hr.1 x A40 48GB: 599 tokens/s per replica at $0.49/hr.2 x RTX PRO 4000 Blackwell 24GB: 1.1k tokens/s per replica at $1.14/hr.2 x V100S PCIe 32GB: 1.9k tokens/s per replica at $1.76/hr.2 x A10 24GB: 981 tokens/s per replica at $2.58/hr.2 x L4 24GB: 490 tokens/s per replica at $0.98/hr.2 x GeForce RTX 3090 24GB: 1.5k tokens/s per replica at $1.00/hr.2 x GeForce RTX 3090 Ti 24GB: 1.6k tokens/s per replica at $0.92/hr.2 x RTX A5000 24GB: 1.3k tokens/s per replica at $0.54/hr.

Priced configurations only; 37 more fit without a live price and appear in the table. Marker size grows with the tensor-parallel width. Roofline estimates.

GPUTPTokens/s
1 ~20k Specs
1 19k Specs
1 6.9k Specs
1 6.9k Rent
1 6.9k Rent
1 6.9k Rent
1 6.9k Rent
1 6.6k Rent
1 6.9k Specs
1 ~5.2k Rent
1 ~4.6k Rent
1 ~1.4k Specs
1 ~4.1k Rent
1 ~2.9k Rent
1 ~2.8k Specs
1 ~4.1k Rent
1 ~3.4k Rent
1 ~1.7k Rent
1 ~1.7k Specs
1 1.5k Rent
1 ~1.4k Rent
1 1.5k Rent
1 ~1.1k Specs
2 2.9k Rent
1 ~2.8k Specs
1 ~826 Rent
1 ~2.8k Specs
1 ~743 Rent
1 ~1.7k Specs
1 ~1.7k Rent
1 ~1.8k Rent
2 ~2.5k Specs
2 ~2.5k Rent
2 ~2.2k Specs
1 ~826 Specs
2 ~942 Rent
1 ~1.2k Rent
1 ~1.2k Specs
2 ~1.5k Rent
2 ~1.0k Specs
2 ~2.0k Specs
2 ~994 Specs
1 ~743 Rent
1 ~1.4k Specs
2 ~1.6k Rent
2 ~1.5k Rent
2 ~706 Specs
1 ~661 Rent
1 ~599 Rent
1 ~3.4k Specs
1 ~3.4k Specs
2 ~1.1k Rent
2 ~2.0k Specs
2 ~1.3k Specs
2 ~1.9k Rent
2 ~981 Rent
1 ~743 Specs
1 ~743 Specs
2 ~1.6k Specs
2 ~490 Rent
1 ~743 Specs
2 ~994 Specs
2 ~745 Specs
2 ~490 Specs
2 ~706 Specs
1 ~743 Specs
2 ~942 Specs
2 ~1.5k Rent
2 ~981 Specs
2 ~1.6k Rent
2 ~837 Specs
2 ~1.3k Rent
2 ~732 Specs

AMD vs NVIDIA

At INT4 / FP4 the NVIDIA pick is 1 x RTX PRO 6000 Blackwell Max-Q Workstation Edition at about 1.5k tokens/s for $0.50/hr; the AMD pick is 1 x Instinct MI300X 192GB at about 4.6k tokens/s for $2.39/hr. Per rental dollar NVIDIA delivers 1.62x the tokens of AMD here.

NVIDIA: 3.1k tokens/s per dollar, 5.1k tokens/s per kW (1 x RTX PRO 6000 Blackwell Max-Q Workstation Edition).AMD: 1.9k tokens/s per dollar, 6.1k tokens/s per kW (1 x Instinct MI300X 192GB).Intel: no live price, 4.7k tokens/s per kW (1 x Data Center GPU Max 1550 128GB).Other: no live price, 2.5k tokens/s per kW (1 x BR100).

NVIDIA 49 parts fit
Best value: 1 x RTX PRO 6000 Blackwell Max-Q Workstation Edition, 1.5k tok/s for $0.50/hr
Fewest GPUs: 1 x Rubin SXM, 19k tok/s
AMD 17 parts fit
Best value: 1 x Instinct MI300X 192GB, 4.6k tok/s for $2.39/hr
Fewest GPUs: 1 x Instinct MI455X OAM, 20k tok/s
Intel 5 parts fit
Fewest GPUs: 1 x Data Center GPU Max 1550 128GB, 2.8k tok/s
Other 2 parts fit
Fewest GPUs: 1 x BR100, 1.4k tok/s

Fleet what-if

Compare whole fleets at these settings: a thousand of one part against a hundred of another. Each fleet splits into replicas of its tensor-parallel width; GPUs left over sit idle.

A
B
FleetReplicasTokens/s
A 1,000 x H200 SXM 141GB 1,000 x TP1 4.13M
B 100 x GB200 NVL72 GPU 186GB 100 x TP1 688k

1,000 x H200 SXM 141GB: 4.13M tokens/s, $2,990/hr, 700 kW.100 x GB200 NVL72 GPU 186GB: 688k tokens/s, $1,050/hr, no power figure.

H200 SXM 141GB: 4.13M tokens/s at 1000 GPUs.GB200 NVL72 GPU 186GB: 6.88M tokens/s at 1000 GPUs.

Size gpt-oss-20b for a tokens per second target

The calculator starts from your traffic instead of a fleet: set a target rate and read the replica count for every GPU.

Frequently asked questions

How much GPU memory does gpt-oss-20b need?

The weights take 11.8 GB at the native INT4 / FP4 precision (41.8 GB at BF16, 20.9 GB at FP8, 11.8 GB at INT4). Each concurrent stream adds 48 KB of KV cache per token: 0.1 GB at 4,096 tokens and 0.8 GB at 32,768 tokens.

What is the bare minimum to run gpt-oss-20b?

1 x Instinct MI350X OAM at INT4 / FP4 with 4,096 tokens of context and one stream. The cheapest live rental for that footprint is 1 x RTX A5000 24GB at $0.27 per hour.

How many H100s do you need to run gpt-oss-20b?

one H100 at INT4 / FP4 for 4,096 tokens of context and one stream. For 32,768 tokens and 32 concurrent streams the tensor-parallel group is one H100, and 1,000 tokens per second takes 1 H100s across 1 replicas.

How much does gpt-oss-20b cost per million tokens on an H100?

About $0.17 per million output tokens at INT4 / FP4, 32,768 tokens of context and 32 streams, using the cheapest live on-demand price of $1.79 per GPU-hour. The cheapest part per token is 1 x RTX PRO 6000 Blackwell Max-Q Workstation Edition at $0.09.

What precision does gpt-oss-20b ship in?

The published checkpoint is MXFP4. That 4-bit format is what the lab validated, so the INT4 column is the native one; BF16 figures describe a dequantised copy nobody would deploy.

All figures are roofline estimates from the model's config.json and each GPU's published memory, bandwidth and tensor-core peak, with fixed efficiency factors. Prices are the cheapest live on-demand listing per GPU when the site has one. They are a planning floor; a tuned serving stack can do better. Active parameters per the OpenAI model card. Half the layers attend over a 128-token sliding window.