gpt-oss-120b hardware requirements
Fits on one Instinct MI350X OAM at INT4 / FP4. For 1,000 tokens/s at 32k context you need 2 x H100 SXM5 80GB at $3.58/hr or 2 x RTX PRO 6000 Blackwell Max-Q Workstation Edition at $1.00/hr.
Bare minimum
- Tokens/s
- 2.0k
- Per hour
- $6.16
- Per million tokens
- $0.86
- Memory used
- 24%
- Tokens/s
- 444
- Per hour
- $0.50
- Per million tokens
- $0.31
- Memory used
- 72%
- Tokens/s
- 2.0k
- Per hour
- $1.00
- Per million tokens
- $0.14
- Memory used
- 57%
Configurations
Every datacenter GPU in the catalog, sized for these settings. Change precision, context and concurrency and the whole page follows. Each point is one replica: the smallest tensor-parallel group of that GPU that holds the model.
1 x GB300 NVL72 GPU 288GB: 4.6k tokens/s per replica at $8.62/hr.1 x GB200 NVL72 GPU 186GB: 4.6k tokens/s per replica at $10.50/hr.1 x Instinct MI350X OAM: 4.6k tokens/s per replica at $6.16/hr.1 x B300 SXM 262GB: 4.6k tokens/s per replica at $7.40/hr.1 x B200 SXM 180GB: 4.4k tokens/s per replica at $4.09/hr.1 x Instinct MI325X OAM: 3.5k tokens/s per replica at $3.80/hr.1 x Instinct MI300X 192GB: 3.1k tokens/s per replica at $2.39/hr.1 x H200 SXM 141GB: 2.8k tokens/s per replica at $2.99/hr.2 x H100 SXM5 80GB: 3.7k tokens/s per replica at $3.58/hr.1 x H200 NVL 141GB: 2.8k tokens/s per replica at $3.79/hr.2 x H100 NVL 94GB: 4.3k tokens/s per replica at $6.38/hr.2 x H100 PCIe 80GB: 2.2k tokens/s per replica at $5.00/hr.2 x RTX PRO 6000 Blackwell Workstation Edition: 2.0k tokens/s per replica at $3.60/hr.2 x RTX PRO 6000 Blackwell Server Edition: 1.7k tokens/s per replica at $1.18/hr.2 x RTX PRO 6000 Blackwell Max-Q Workstation Edition: 2.0k tokens/s per replica at $1.00/hr.4 x GeForce RTX 5090 32GB: 3.7k tokens/s per replica at $3.96/hr.4 x RTX 6000 Ada 48GB: 2.0k tokens/s per replica at $3.36/hr.4 x L40S 48GB: 1.8k tokens/s per replica at $3.48/hr.2 x A100 PCIe 80GB: 2.1k tokens/s per replica at $2.70/hr.2 x A100 SXM4 80GB: 2.2k tokens/s per replica at $2.80/hr.4 x A100 SXM4 40GB: 3.2k tokens/s per replica at $5.16/hr.4 x RTX 5000 Ada 32GB: 1.2k tokens/s per replica at $3.32/hr.4 x RTX PRO 5000 Blackwell 48GB: 2.8k tokens/s per replica at $3.84/hr.4 x RTX PRO 4500 Blackwell 32GB: 1.9k tokens/s per replica at $2.88/hr.4 x L40 48GB: 1.8k tokens/s per replica at $3.28/hr.8 x GeForce RTX 4090 24GB: 3.9k tokens/s per replica at $4.80/hr.8 x A30 24GB: 3.7k tokens/s per replica at $5.86/hr.4 x RTX A6000 48GB: 1.6k tokens/s per replica at $2.00/hr.4 x A40 48GB: 1.4k tokens/s per replica at $1.96/hr.8 x RTX PRO 4000 Blackwell 24GB: 2.6k tokens/s per replica at $4.56/hr.4 x V100S PCIe 32GB: 2.4k tokens/s per replica at $3.52/hr.8 x A10 24GB: 2.3k tokens/s per replica at $10.32/hr.8 x L4 24GB: 1.2k tokens/s per replica at $3.92/hr.8 x GeForce RTX 3090 24GB: 3.7k tokens/s per replica at $4.00/hr.8 x GeForce RTX 3090 Ti 24GB: 3.9k tokens/s per replica at $3.68/hr.8 x RTX A5000 24GB: 3.0k tokens/s per replica at $2.16/hr.
Priced configurations only; 37 more fit without a live price and appear in the table. Marker size grows with the tensor-parallel width. Roofline estimates.
| GPU | TP | Tokens/s | |
|---|---|---|---|
| 1 | ~13k | Specs | |
| 1 | 13k | Specs | |
| 1 | 4.6k | Specs | |
| 1 | 4.6k | Rent | |
| 1 | 4.6k | Rent | |
| 1 | 4.6k | Rent | |
| 1 | 4.6k | Rent | |
| 1 | 4.4k | Rent | |
| 1 | 4.6k | Specs | |
| 1 | ~3.5k | Rent | |
| 1 | ~3.1k | Rent | |
| 2 | ~1.8k | Specs | |
| 1 | ~2.8k | Rent | |
| 2 | ~3.7k | Rent | |
| 1 | ~1.9k | Specs | |
| 1 | ~2.8k | Rent | |
| 2 | ~4.3k | Rent | |
| 2 | ~2.2k | Rent | |
| 2 | ~2.2k | Specs | |
| 2 | 2.0k | Rent | |
| 2 | ~1.7k | Rent | |
| 2 | 2.0k | Rent | |
| 4 | ~2.5k | Specs | |
| 4 | 3.7k | Rent | |
| 1 | ~1.9k | Specs | |
| 4 | ~2.0k | Rent | |
| 1 | ~1.9k | Specs | |
| 4 | ~1.8k | Rent | |
| 2 | ~2.1k | Specs | |
| 2 | ~2.1k | Rent | |
| 2 | ~2.2k | Rent | |
| 4 | ~3.2k | Specs | |
| 4 | ~3.2k | Rent | |
| 8 | ~5.3k | Specs | |
| 4 | ~2.0k | Specs | |
| 4 | ~1.2k | Rent | |
| 4 | ~2.8k | Rent | |
| 2 | ~1.5k | Specs | |
| 4 | ~1.9k | Rent | |
| 4 | ~1.3k | Specs | |
| 4 | ~2.5k | Specs | |
| 4 | ~1.3k | Specs | |
| 4 | ~1.8k | Rent | |
| 2 | ~1.8k | Specs | |
| 8 | ~3.9k | Rent | |
| 8 | ~3.7k | Rent | |
| 8 | ~1.7k | Specs | |
| 4 | ~1.6k | Rent | |
| 4 | ~1.4k | Rent | |
| 1 | ~2.3k | Specs | |
| 2 | ~4.4k | Specs | |
| 8 | ~2.6k | Rent | |
| 4 | ~2.5k | Specs | |
| 8 | ~3.0k | Specs | |
| 4 | ~2.4k | Rent | |
| 8 | ~2.3k | Rent | |
| 4 | ~1.8k | Specs | |
| 4 | ~1.8k | Specs | |
| 8 | ~3.8k | Specs | |
| 8 | ~1.2k | Rent | |
| 4 | ~1.8k | Specs | |
| 4 | ~1.3k | Specs | |
| 8 | ~1.8k | Specs | |
| 8 | ~1.2k | Specs | |
| 8 | ~1.7k | Specs | |
| 4 | ~1.8k | Specs | |
| 4 | ~1.2k | Specs | |
| 8 | ~3.7k | Rent | |
| 8 | ~2.3k | Specs | |
| 8 | ~3.9k | Rent | |
| 4 | ~1.1k | Specs | |
| 8 | ~3.0k | Rent | |
| 8 | ~1.8k | Specs |
AMD vs NVIDIA
At INT4 / FP4 the NVIDIA pick is 2 x RTX PRO 6000 Blackwell Max-Q Workstation Edition at about 2.0k tokens/s for $1.00/hr; the AMD pick is 1 x Instinct MI300X 192GB at about 3.1k tokens/s for $2.39/hr. Per rental dollar NVIDIA delivers 1.54x the tokens of AMD here.
NVIDIA: 2.0k tokens/s per dollar, 3.3k tokens/s per kW (2 x RTX PRO 6000 Blackwell Max-Q Workstation Edition).AMD: 1.3k tokens/s per dollar, 4.1k tokens/s per kW (1 x Instinct MI300X 192GB).Intel: no live price, 3.1k tokens/s per kW (1 x Data Center GPU Max 1550 128GB).Other: no live price, 1.6k tokens/s per kW (2 x BR100).
Fleet what-if
Compare whole fleets at these settings: a thousand of one part against a hundred of another. Each fleet splits into replicas of its tensor-parallel width; GPUs left over sit idle.
| Fleet | Replicas | Tokens/s |
|---|---|---|
| A 1,000 x H200 SXM 141GB | 1,000 x TP1 | 2.76M |
| B 100 x GB200 NVL72 GPU 186GB | 100 x TP1 | 461k |
1,000 x H200 SXM 141GB: 2.76M tokens/s, $2,990/hr, 700 kW.100 x GB200 NVL72 GPU 186GB: 461k tokens/s, $1,050/hr, no power figure.
H200 SXM 141GB: 2.76M tokens/s at 1000 GPUs.GB200 NVL72 GPU 186GB: 4.61M tokens/s at 1000 GPUs.
Size gpt-oss-120b for a tokens per second target
The calculator starts from your traffic instead of a fleet: set a target rate and read the replica count for every GPU.
Frequently asked questions
How much GPU memory does gpt-oss-120b need?
The weights take 65.7 GB at the native INT4 / FP4 precision (234 GB at BF16, 117 GB at FP8, 65.7 GB at INT4). Each concurrent stream adds 72 KB of KV cache per token: 0.2 GB at 4,096 tokens and 1.2 GB at 32,768 tokens.
What is the bare minimum to run gpt-oss-120b?
1 x Instinct MI350X OAM at INT4 / FP4 with 4,096 tokens of context and one stream. The cheapest live rental for that footprint is 1 x RTX PRO 6000 Blackwell Max-Q Workstation Edition at $0.50 per hour.
How many H100s do you need to run gpt-oss-120b?
one H100 at INT4 / FP4 for 4,096 tokens of context and one stream. For 32,768 tokens and 32 concurrent streams the tensor-parallel group is 2 H100s, and 1,000 tokens per second takes 2 H100s across 1 replicas.
How much does gpt-oss-120b cost per million tokens on an H100?
About $0.27 per million output tokens at INT4 / FP4, 32,768 tokens of context and 32 streams, using the cheapest live on-demand price of $1.79 per GPU-hour. The cheapest part per token is 2 x RTX PRO 6000 Blackwell Max-Q Workstation Edition at $0.14.
What precision does gpt-oss-120b ship in?
The published checkpoint is MXFP4. That 4-bit format is what the lab validated, so the INT4 column is the native one; BF16 figures describe a dequantised copy nobody would deploy.
All figures are roofline estimates from the model's config.json and each GPU's published memory, bandwidth and tensor-core peak, with fixed efficiency factors. Prices are the cheapest live on-demand listing per GPU when the site has one. They are a planning floor; a tuned serving stack can do better. Active parameters per the OpenAI model card. Half the layers attend over a 128-token sliding window, so the cache grows at half the rate the head count suggests.

