Moonshot AI Kimi K2

Kimi K2 Instruct hardware requirements

Needs at least 4 GPUs at FP8 (Instinct MI350X OAM). For 1,000 tokens/s at 32k context you need 16 x H100 SXM5 80GB at $28.64/hr or 16 x RTX PRO 6000 Blackwell Max-Q Workstation Edition at $8.00/hr.

Parameters
1.03T
Active per token
32.0B
mixture of experts
KV cache per token
69 KB
16-bit cache
Native precision
FP8

Bare minimum

Fewest GPUs
4 x Instinct MI350X OAM
FP8, 4,096 tokens, one stream
Tokens/s
669
Per hour
$24.64
Per million tokens
$10.23
Memory used
93%
Rent on DigitalOcean
Cheapest per hour
16 x RTX PRO 6000 Blackwell Max-Q Workstation Edition
FP8, 4,096 tokens, one stream, live price
Tokens/s
500
Per hour
$8.00
Per million tokens
$4.45
Memory used
70%
Rent on RunPod
Cheapest per million tokens
16 x RTX PRO 6000 Blackwell Max-Q Workstation Edition
FP8, 32,768 tokens, 32 streams
Tokens/s
4.9k
Per hour
$8.00
Per million tokens
$0.46
Memory used
75%
Rent on RunPod

Configurations

Every datacenter GPU in the catalog, sized for these settings. Change precision, context and concurrency and the whole page follows. Each point is one replica: the smallest tensor-parallel group of that GPU that holds the model.

Vendor

4 x GB300 NVL72 GPU 288GB: 6.5k tokens/s per replica at $34.48/hr.8 x GB200 NVL72 GPU 186GB: 12k tokens/s per replica at $84.00/hr.4 x Instinct MI350X OAM: 6.5k tokens/s per replica at $24.64/hr.8 x B300 SXM 262GB: 12k tokens/s per replica at $59.20/hr.8 x B200 SXM 180GB: 12k tokens/s per replica at $32.72/hr.8 x Instinct MI325X OAM: 9.3k tokens/s per replica at $30.40/hr.8 x Instinct MI300X 192GB: 8.2k tokens/s per replica at $19.12/hr.16 x H200 SXM 141GB: 13k tokens/s per replica at $47.84/hr.16 x H100 SXM5 80GB: 9.1k tokens/s per replica at $28.64/hr.16 x H200 NVL 141GB: 13k tokens/s per replica at $60.64/hr.16 x H100 NVL 94GB: 11k tokens/s per replica at $51.04/hr.16 x H100 PCIe 80GB: 5.5k tokens/s per replica at $40.00/hr.16 x RTX PRO 6000 Blackwell Workstation Edition: 4.9k tokens/s per replica at $28.80/hr.16 x RTX PRO 6000 Blackwell Server Edition: 4.4k tokens/s per replica at $9.44/hr.16 x RTX PRO 6000 Blackwell Max-Q Workstation Edition: 4.9k tokens/s per replica at $8.00/hr.16 x A100 PCIe 80GB: 5.3k tokens/s per replica at $21.60/hr.16 x A100 SXM4 80GB: 5.6k tokens/s per replica at $22.40/hr.

Priced configurations only; 11 more fit without a live price and appear in the table. Marker size grows with the tensor-parallel width. Roofline estimates.

GPUTPTokens/s
4 ~19k Specs
4 18k Specs
4 6.5k Specs
4 6.5k Rent
8 12k Rent
4 6.5k Rent
8 12k Rent
8 12k Rent
8 12k Specs
8 9.3k Rent
8 8.2k Rent
16 13k Rent
16 9.1k Rent
16 ~8.9k Specs
16 ~13k Rent
H100 NVL 94GB 2 nodes
16 ~11k Rent
16 5.5k Rent
16 5.5k Specs
16 4.9k Rent
16 ~4.4k Rent
16 4.9k Rent
MI250X 128GB 2 nodes
16 ~8.9k Specs
16 ~8.9k Specs
16 ~5.3k Specs
16 ~5.3k Rent
16 ~5.6k Rent
16 11k Specs
H20 96GB 2 nodes
16 11k Specs

AMD vs NVIDIA

At FP8 the NVIDIA pick is 16 x RTX PRO 6000 Blackwell Max-Q Workstation Edition at about 4.9k tokens/s for $8.00/hr; the AMD pick is 8 x Instinct MI300X 192GB at about 8.2k tokens/s for $19.12/hr. Per rental dollar NVIDIA delivers 1.43x the tokens of AMD here.

NVIDIA: 610 tokens/s per dollar, 1.0k tokens/s per kW (16 x RTX PRO 6000 Blackwell Max-Q Workstation Edition).AMD: 428 tokens/s per dollar, 1.4k tokens/s per kW (8 x Instinct MI300X 192GB).Intel: no live price, 930 tokens/s per kW (16 x Data Center GPU Max 1550 128GB).

NVIDIA 20 parts fit
Best value: 16 x RTX PRO 6000 Blackwell Max-Q Workstation Edition, 4.9k tok/s for $8.00/hr
Fewest GPUs: 4 x Rubin SXM, 18k tok/s
AMD 7 parts fit
Best value: 8 x Instinct MI300X 192GB, 8.2k tok/s for $19.12/hr
Fewest GPUs: 4 x Instinct MI455X OAM, 19k tok/s
Intel 1 parts fit
Fewest GPUs: 16 x Data Center GPU Max 1550 128GB, 8.9k tok/s

Fleet what-if

Compare whole fleets at these settings: a thousand of one part against a hundred of another. Each fleet splits into replicas of its tensor-parallel width; GPUs left over sit idle.

A
B
FleetReplicasTokens/s
A 1,008 x H200 SXM 141GB 63 x TP16 824k
B 104 x GB200 NVL72 GPU 186GB 13 x TP8 161k

1,008 x H200 SXM 141GB: 824k tokens/s, $3,014/hr, 706 kW.104 x GB200 NVL72 GPU 186GB: 161k tokens/s, $1,092/hr, no power figure.

H200 SXM 141GB: 824k tokens/s at 1008 GPUs.GB200 NVL72 GPU 186GB: 1.56M tokens/s at 1008 GPUs.

Size Kimi K2 Instruct for a tokens per second target

The calculator starts from your traffic instead of a fleet: set a target rate and read the replica count for every GPU.

Frequently asked questions

How much GPU memory does Kimi K2 Instruct need?

The weights take 1,026 GB at the native FP8 precision (2,053 GB at BF16, 1,026 GB at FP8, 577 GB at INT4). Each concurrent stream adds 69 KB of KV cache per token: 0.3 GB at 4,096 tokens and 2.3 GB at 32,768 tokens.

What is the bare minimum to run Kimi K2 Instruct?

4 x Instinct MI350X OAM at FP8 with 4,096 tokens of context and one stream. The cheapest live rental for that footprint is 16 x RTX PRO 6000 Blackwell Max-Q Workstation Edition at $8.00 per hour.

How many H100s do you need to run Kimi K2 Instruct?

16 H100s at FP8 for 4,096 tokens of context and one stream. For 32,768 tokens and 32 concurrent streams the tensor-parallel group is 16 H100s, and 1,000 tokens per second takes 16 H100s across 1 replicas.

How much does Kimi K2 Instruct cost per million tokens on an H100?

About $0.87 per million output tokens at FP8, 32,768 tokens of context and 32 streams, using the cheapest live on-demand price of $1.79 per GPU-hour. The cheapest part per token is 16 x RTX PRO 6000 Blackwell Max-Q Workstation Edition at $0.46.

What precision does Kimi K2 Instruct ship in?

The published checkpoint is FP8. FP8 halves the BF16 footprint with the accuracy the lab validated; INT4 is a further community quantisation.

All figures are roofline estimates from the model's config.json and each GPU's published memory, bandwidth and tensor-core peak, with fixed efficiency factors. Prices are the cheapest live on-demand listing per GPU when the site has one. They are a planning floor; a tuned serving stack can do better. Active parameters per the Moonshot model card.