DeepSeek DeepSeek V4 Estimated geometry

DeepSeek-V4-Flash hardware requirements

Fits on one Instinct MI350X OAM at INT4 / FP4. For 1,000 tokens/s at 32k context you need 4 x H100 SXM5 80GB at $7.16/hr or 2 x RTX PRO 6000 Blackwell Max-Q Workstation Edition at $1.00/hr.

Parameters
291B
Active per token
13.8B
mixture of experts
KV cache per token
8.1 KB
16-bit cache
Native precision
FP4

Bare minimum

Fewest GPUs
1 x Instinct MI350X OAM
INT4 / FP4, 4,096 tokens, one stream
Tokens/s
770
Per hour
$6.16
Per million tokens
$2.22
Memory used
59%
Rent on DigitalOcean
Cheapest per hour
2 x RTX PRO 6000 Blackwell Max-Q Workstation Edition
INT4 / FP4, 4,096 tokens, one stream, live price
Tokens/s
328
Per hour
$1.00
Per million tokens
$0.85
Memory used
89%
Rent on RunPod
Cheapest per million tokens
2 x RTX PRO 6000 Blackwell Max-Q Workstation Edition
INT4 / FP4, 32,768 tokens, 32 streams
Tokens/s
5.0k
Per hour
$1.00
Per million tokens
$0.06
Memory used
94%
Rent on RunPod

Configurations

Every datacenter GPU in the catalog, sized for these settings. Change precision, context and concurrency and the whole page follows. Each point is one replica: the smallest tensor-parallel group of that GPU that holds the model.

Vendor

1 x GB300 NVL72 GPU 288GB: 12k tokens/s per replica at $8.62/hr.1 x GB200 NVL72 GPU 186GB: 12k tokens/s per replica at $10.50/hr.1 x Instinct MI350X OAM: 12k tokens/s per replica at $6.16/hr.1 x B300 SXM 262GB: 12k tokens/s per replica at $7.40/hr.1 x B200 SXM 180GB: 11k tokens/s per replica at $4.09/hr.1 x Instinct MI325X OAM: 8.7k tokens/s per replica at $3.80/hr.1 x Instinct MI300X 192GB: 7.7k tokens/s per replica at $2.39/hr.2 x H200 SXM 141GB: 13k tokens/s per replica at $5.98/hr.4 x H100 SXM5 80GB: 18k tokens/s per replica at $7.16/hr.2 x H200 NVL 141GB: 13k tokens/s per replica at $7.58/hr.2 x H100 NVL 94GB: 11k tokens/s per replica at $6.38/hr.4 x H100 PCIe 80GB: 10k tokens/s per replica at $10.00/hr.2 x RTX PRO 6000 Blackwell Workstation Edition: 5.0k tokens/s per replica at $3.60/hr.2 x RTX PRO 6000 Blackwell Server Edition: 4.4k tokens/s per replica at $1.18/hr.2 x RTX PRO 6000 Blackwell Max-Q Workstation Edition: 5.0k tokens/s per replica at $1.00/hr.8 x GeForce RTX 5090 32GB: 18k tokens/s per replica at $7.92/hr.4 x RTX 6000 Ada 48GB: 5.0k tokens/s per replica at $3.36/hr.4 x L40S 48GB: 4.5k tokens/s per replica at $3.48/hr.4 x A100 PCIe 80GB: 10k tokens/s per replica at $5.40/hr.4 x A100 SXM4 80GB: 11k tokens/s per replica at $5.60/hr.8 x A100 SXM4 40GB: 15k tokens/s per replica at $10.32/hr.8 x RTX 5000 Ada 32GB: 5.7k tokens/s per replica at $6.64/hr.4 x RTX PRO 5000 Blackwell 48GB: 7.1k tokens/s per replica at $3.84/hr.8 x RTX PRO 4500 Blackwell 32GB: 8.9k tokens/s per replica at $5.76/hr.4 x L40 48GB: 4.5k tokens/s per replica at $3.28/hr.8 x GeForce RTX 4090 24GB: 10.0k tokens/s per replica at $4.80/hr.8 x A30 24GB: 9.2k tokens/s per replica at $5.86/hr.4 x RTX A6000 48GB: 4.0k tokens/s per replica at $2.00/hr.4 x A40 48GB: 3.7k tokens/s per replica at $1.96/hr.8 x RTX PRO 4000 Blackwell 24GB: 6.7k tokens/s per replica at $4.56/hr.8 x V100S PCIe 32GB: 11k tokens/s per replica at $7.04/hr.8 x A10 24GB: 5.9k tokens/s per replica at $10.32/hr.8 x L4 24GB: 3.0k tokens/s per replica at $3.92/hr.8 x GeForce RTX 3090 24GB: 9.3k tokens/s per replica at $4.00/hr.8 x GeForce RTX 3090 Ti 24GB: 5.9k tokens/s per replica at $3.68/hr.8 x RTX A5000 24GB: 4.1k tokens/s per replica at $2.16/hr.

Priced configurations only; 37 more fit without a live price and appear in the table. Marker size grows with the tensor-parallel width. Roofline estimates.

GPUTPTokens/s
1 ~34k Specs
1 32k Specs
1 12k Specs
1 12k Rent
1 12k Rent
1 12k Rent
1 12k Rent
1 11k Rent
1 12k Specs
1 ~8.7k Rent
1 ~7.7k Rent
4 ~8.4k Specs
2 ~13k Rent
4 ~18k Rent
2 ~9.1k Specs
2 ~13k Rent
2 ~11k Rent
4 ~10k Rent
4 ~10k Specs
2 5.0k Rent
2 ~4.4k Rent
2 5.0k Rent
4 ~6.4k Specs
8 18k Rent
2 ~9.1k Specs
4 ~5.0k Rent
2 ~9.1k Specs
4 ~4.5k Rent
4 ~10k Specs
4 ~10k Rent
4 ~11k Rent
8 ~15k Specs
8 ~15k Rent
8 ~13k Specs
4 ~5.0k Specs
8 ~5.7k Rent
4 ~7.1k Rent
4 ~7.1k Specs
8 ~8.9k Rent
8 ~6.3k Specs
8 ~12k Specs
8 ~6.0k Specs
4 ~4.5k Rent
4 ~8.6k Specs
8 ~10.0k Rent
8 ~9.2k Rent
8 ~4.3k Specs
4 ~4.0k Rent
4 ~3.7k Rent
2 ~6.1k Specs
2 ~6.1k Specs
8 ~6.7k Rent
8 ~12k Specs
8 ~7.6k Specs
8 ~11k Rent
8 ~5.9k Rent
4 ~4.5k Specs
4 ~4.5k Specs
8 ~9.5k Specs
8 ~3.0k Rent
4 ~4.5k Specs
8 ~6.0k Specs
8 ~4.5k Specs
8 ~3.0k Specs
8 ~4.3k Specs
4 ~4.5k Specs
8 ~5.7k Specs
8 ~9.3k Rent
8 ~5.9k Specs
8 ~5.9k Rent
8 ~5.1k Specs
8 ~4.1k Rent
8 ~4.1k Specs

AMD vs NVIDIA

At INT4 / FP4 the NVIDIA pick is 2 x RTX PRO 6000 Blackwell Max-Q Workstation Edition at about 5.0k tokens/s for $1.00/hr; the AMD pick is 1 x Instinct MI300X 192GB at about 7.7k tokens/s for $2.39/hr. Per rental dollar NVIDIA delivers 1.54x the tokens of AMD here.

NVIDIA: 5.0k tokens/s per dollar, 8.3k tokens/s per kW (2 x RTX PRO 6000 Blackwell Max-Q Workstation Edition).AMD: 3.2k tokens/s per dollar, 10k tokens/s per kW (1 x Instinct MI300X 192GB).Intel: no live price, 7.6k tokens/s per kW (2 x Data Center GPU Max 1550 128GB).Other: no live price, 3.8k tokens/s per kW (4 x BR100).

NVIDIA 49 parts fit
Best value: 2 x RTX PRO 6000 Blackwell Max-Q Workstation Edition, 5.0k tok/s for $1.00/hr
Fewest GPUs: 1 x Rubin SXM, 32k tok/s
AMD 17 parts fit
Best value: 1 x Instinct MI300X 192GB, 7.7k tok/s for $2.39/hr
Fewest GPUs: 1 x Instinct MI455X OAM, 34k tok/s
Intel 5 parts fit
Fewest GPUs: 2 x Data Center GPU Max 1550 128GB, 9.1k tok/s
Other 2 parts fit
Fewest GPUs: 4 x BR100, 8.4k tok/s

Fleet what-if

Compare whole fleets at these settings: a thousand of one part against a hundred of another. Each fleet splits into replicas of its tensor-parallel width; GPUs left over sit idle.

A
B
FleetReplicasTokens/s
A 1,000 x H200 SXM 141GB 500 x TP2 6.65M
B 100 x GB200 NVL72 GPU 186GB 100 x TP1 1.17M

1,000 x H200 SXM 141GB: 6.65M tokens/s, $2,990/hr, 700 kW.100 x GB200 NVL72 GPU 186GB: 1.17M tokens/s, $1,050/hr, no power figure.

H200 SXM 141GB: 6.65M tokens/s at 1000 GPUs.GB200 NVL72 GPU 186GB: 11.66M tokens/s at 1000 GPUs.

Size DeepSeek-V4-Flash for a tokens per second target

The calculator starts from your traffic instead of a fleet: set a target rate and read the replica count for every GPU.

Frequently asked questions

How much GPU memory does DeepSeek-V4-Flash need?

The weights take 164 GB at the native INT4 / FP4 precision (582 GB at BF16, 291 GB at FP8, 164 GB at INT4). Each concurrent stream adds 8.1 KB of KV cache per token: 0.0 GB at 4,096 tokens and 0.3 GB at 32,768 tokens.

What is the bare minimum to run DeepSeek-V4-Flash?

1 x Instinct MI350X OAM at INT4 / FP4 with 4,096 tokens of context and one stream. The cheapest live rental for that footprint is 2 x RTX PRO 6000 Blackwell Max-Q Workstation Edition at $1.00 per hour.

How many H100s do you need to run DeepSeek-V4-Flash?

4 H100s at INT4 / FP4 for 4,096 tokens of context and one stream. For 32,768 tokens and 32 concurrent streams the tensor-parallel group is 4 H100s, and 1,000 tokens per second takes 4 H100s across 1 replicas.

How much does DeepSeek-V4-Flash cost per million tokens on an H100?

About $0.11 per million output tokens at INT4 / FP4, 32,768 tokens of context and 32 streams, using the cheapest live on-demand price of $1.79 per GPU-hour. The cheapest part per token is 2 x RTX PRO 6000 Blackwell Max-Q Workstation Edition at $0.06.

What precision does DeepSeek-V4-Flash ship in?

The published checkpoint is FP4. That 4-bit format is what the lab validated, so the INT4 column is the native one; BF16 figures describe a dequantised copy nobody would deploy.

All figures are roofline estimates from the model's config.json and each GPU's published memory, bandwidth and tensor-core peak, with fixed efficiency factors. Prices are the cheapest live on-demand listing per GPU when the site has one. They are a planning floor; a tuned serving stack can do better. Experts are stored in FP4 and attention in FP8. The KV cache is a single 512-wide latent per layer compressed on most layers; 8.3 KB per token is an estimate from compress_ratios. Active parameters computed from the expert geometry.