DeepSeek-V4-Pro hardware requirements
Needs at least 4 GPUs at INT4 / FP4 (Instinct MI350X OAM). For 1,000 tokens/s at 32k context you need 16 x H100 SXM5 80GB at $28.64/hr or 16 x RTX PRO 6000 Blackwell Max-Q Workstation Edition at $8.00/hr.
Bare minimum
- Tokens/s
- 770
- Per hour
- $24.64
- Per million tokens
- $8.89
- Memory used
- 81%
- Tokens/s
- 575
- Per hour
- $8.00
- Per million tokens
- $3.86
- Memory used
- 61%
- Tokens/s
- 16k
- Per hour
- $8.00
- Per million tokens
- $0.14
- Memory used
- 62%
Configurations
Every datacenter GPU in the catalog, sized for these settings. Change precision, context and concurrency and the whole page follows. Each point is one replica: the smallest tensor-parallel group of that GPU that holds the model.
4 x GB300 NVL72 GPU 288GB: 21k tokens/s per replica at $34.48/hr.8 x GB200 NVL72 GPU 186GB: 40k tokens/s per replica at $84.00/hr.4 x Instinct MI350X OAM: 21k tokens/s per replica at $24.64/hr.4 x B300 SXM 262GB: 21k tokens/s per replica at $29.60/hr.8 x B200 SXM 180GB: 39k tokens/s per replica at $32.72/hr.4 x Instinct MI325X OAM: 16k tokens/s per replica at $15.20/hr.8 x Instinct MI300X 192GB: 27k tokens/s per replica at $19.12/hr.8 x H200 SXM 141GB: 24k tokens/s per replica at $23.92/hr.16 x H100 SXM5 80GB: 30k tokens/s per replica at $28.64/hr.8 x H200 NVL 141GB: 24k tokens/s per replica at $30.32/hr.16 x H100 NVL 94GB: 35k tokens/s per replica at $51.04/hr.16 x H100 PCIe 80GB: 18k tokens/s per replica at $40.00/hr.16 x RTX PRO 6000 Blackwell Workstation Edition: 16k tokens/s per replica at $28.80/hr.16 x RTX PRO 6000 Blackwell Server Edition: 14k tokens/s per replica at $9.44/hr.16 x RTX PRO 6000 Blackwell Max-Q Workstation Edition: 16k tokens/s per replica at $8.00/hr.16 x A100 PCIe 80GB: 17k tokens/s per replica at $21.60/hr.16 x A100 SXM4 80GB: 18k tokens/s per replica at $22.40/hr.
Priced configurations only; 14 more fit without a live price and appear in the table. Marker size grows with the tensor-parallel width. Roofline estimates.
| GPU | TP | Tokens/s | |
|---|---|---|---|
| 4 | ~62k | Specs | |
| 4 | 59k | Specs | |
| 4 | 21k | Specs | |
| 4 | 21k | Rent | |
| 8 | 40k | Rent | |
| 4 | 21k | Rent | |
| 4 | 21k | Rent | |
| 8 | 39k | Rent | |
| 8 | 40k | Specs | |
| 4 | ~16k | Rent | |
| 8 | ~27k | Rent | |
| 16 | ~14k | Specs | |
| 8 | ~24k | Rent | |
H100 SXM5 80GB 2 nodes | 16 | ~30k | Rent |
| 8 | ~16k | Specs | |
| 8 | ~24k | Rent | |
H100 NVL 94GB 2 nodes | 16 | ~35k | Rent |
H100 PCIe 80GB 2 nodes | 16 | ~18k | Rent |
H800 PCIe 80GB 2 nodes | 16 | ~18k | Specs |
| 16 | 16k | Rent | |
| 16 | ~14k | Rent | |
| 16 | 16k | Rent | |
| 8 | ~16k | Specs | |
| 8 | ~15k | Specs | |
A800 PCIe 80GB 2 nodes | 16 | ~17k | Specs |
A100 PCIe 80GB 2 nodes | 16 | ~17k | Rent |
A100 SXM4 80GB 2 nodes | 16 | ~18k | Rent |
RTX PRO 5000 Blackwell 72GB 2 nodes | 16 | ~12k | Specs |
Instinct MI210 PCIe 2 nodes | 16 | ~13k | Specs |
| 8 | ~6.1k | Specs | |
H20 96GB 2 nodes | 16 | ~11k | Specs |
AMD vs NVIDIA
At INT4 / FP4 the NVIDIA pick is 16 x RTX PRO 6000 Blackwell Max-Q Workstation Edition at about 16k tokens/s for $8.00/hr; the AMD pick is 8 x Instinct MI300X 192GB at about 27k tokens/s for $19.12/hr. Per rental dollar NVIDIA delivers 1.43x the tokens of AMD here.
NVIDIA: 2.0k tokens/s per dollar, 3.3k tokens/s per kW (16 x RTX PRO 6000 Blackwell Max-Q Workstation Edition).AMD: 1.4k tokens/s per dollar, 4.4k tokens/s per kW (8 x Instinct MI300X 192GB).Intel: no live price, 3.4k tokens/s per kW (8 x Data Center GPU Max 1550 128GB).Other: no live price, 1.6k tokens/s per kW (16 x BR100).
Fleet what-if
Compare whole fleets at these settings: a thousand of one part against a hundred of another. Each fleet splits into replicas of its tensor-parallel width; GPUs left over sit idle.
| Fleet | Replicas | Tokens/s |
|---|---|---|
| A 1,000 x H200 SXM 141GB | 125 x TP8 | 3.02M |
| B 104 x GB200 NVL72 GPU 186GB | 13 x TP8 | 523k |
1,000 x H200 SXM 141GB: 3.02M tokens/s, $2,990/hr, 700 kW.104 x GB200 NVL72 GPU 186GB: 523k tokens/s, $1,092/hr, no power figure.
H200 SXM 141GB: 3.02M tokens/s at 1000 GPUs.GB200 NVL72 GPU 186GB: 5.03M tokens/s at 1000 GPUs.
Size DeepSeek-V4-Pro for a tokens per second target
The calculator starts from your traffic instead of a fleet: set a target rate and read the replica count for every GPU.
Frequently asked questions
How much GPU memory does DeepSeek-V4-Pro need?
The weights take 899 GB at the native INT4 / FP4 precision (3,198 GB at BF16, 1,599 GB at FP8, 899 GB at INT4). Each concurrent stream adds 4.1 KB of KV cache per token: 0.0 GB at 4,096 tokens and 0.1 GB at 32,768 tokens.
What is the bare minimum to run DeepSeek-V4-Pro?
4 x Instinct MI350X OAM at INT4 / FP4 with 4,096 tokens of context and one stream. The cheapest live rental for that footprint is 16 x RTX PRO 6000 Blackwell Max-Q Workstation Edition at $8.00 per hour.
How many H100s do you need to run DeepSeek-V4-Pro?
16 H100s at INT4 / FP4 for 4,096 tokens of context and one stream. For 32,768 tokens and 32 concurrent streams the tensor-parallel group is 16 H100s, and 1,000 tokens per second takes 16 H100s across 1 replicas.
How much does DeepSeek-V4-Pro cost per million tokens on an H100?
About $0.27 per million output tokens at INT4 / FP4, 32,768 tokens of context and 32 streams, using the cheapest live on-demand price of $1.79 per GPU-hour. The cheapest part per token is 16 x RTX PRO 6000 Blackwell Max-Q Workstation Edition at $0.14.
What precision does DeepSeek-V4-Pro ship in?
The published checkpoint is FP4. That 4-bit format is what the lab validated, so the INT4 column is the native one; BF16 figures describe a dequantised copy nobody would deploy.
All figures are roofline estimates from the model's config.json and each GPU's published memory, bandwidth and tensor-core peak, with fixed efficiency factors. Prices are the cheapest live on-demand listing per GPU when the site has one. They are a planning floor; a tuned serving stack can do better. Experts are stored in FP4 and attention in FP8. The KV cache is a single 512-wide latent per layer compressed 4x to 128x on most layers; 4.2 KB per token is an estimate from compress_ratios and the true figure depends on the serving stack. Active parameters computed from the expert geometry.
