Moonshot AI Kimi K3

Kimi K3 hardware requirements

Needs at least 8 GPUs at INT4 / FP4 (B300 SXM 262GB). For 1,000 tokens/s at 32k context you need 16 x Instinct MI300X 192GB at $38.24/hr.

Parameters
2.78T
Active per token
104B
mixture of experts
KV cache per token
27 KB
16-bit cache
Native precision
MXFP4

Top options for Kimi K3

The strongest part each vendor offers for Kimi K3 at INT4 / FP4 with 32,768 tokens per user, rentable parts first. Each card is one replica, the fewest accelerators that hold the model, and the line it draws on the graph below. Select a card to plot it.

Tokens per second as users grow

Each line is one replica, the fewest accelerators that hold the model, serving more users at once. It climbs while memory bandwidth is the limit, flattens once compute is, and stops where memory runs out at 32,768 tokens per user. Hover a line to bring it forward.

8 x GB300 NVL72 GPU 288GB 8 x Instinct MI355X OAM 16 x Ascend 950DT 16 x Gaudi 3 128GB 16 x TPU v7 192GB

8 x GB300 NVL72 GPU 288GB and 8 x Instinct MI355X OAM draw the same line: identical memory and bandwidth give identical curves.

8 x GB300 NVL72 GPU 288GB: peaks at 28k tokens per second when memory fills at 488 users.8 x Instinct MI355X OAM: peaks at 28k tokens per second when memory fills at 488 users.16 x Ascend 950DT: peaks at 25k tokens per second when memory fills at 481 users.16 x Gaudi 3 128GB: peaks at 22k tokens per second when memory fills at 297 users.16 x TPU v7 192GB: peaks at 48k tokens per second and still has memory at 1024 users.

Memory fills at: 8 x GB300 NVL72 GPU 288GB 488 users, 8 x Instinct MI355X OAM 488 users, 16 x Ascend 950DT 481 users, 16 x Gaudi 3 128GB 297 users, 16 x TPU v7 192GB beyond 1,024 users. Roofline estimates at INT4 / FP4.

Bare minimum

Fewest GPUs
8 x B300 SXM 262GB
INT4 / FP4, 4,096 tokens, one stream
Tokens/s
691
Per hour
$39.92
Per million tokens
$16.05
Memory used
77%
Rent on Together AI
Cheapest per hour
8 x Instinct MI325X OAM
INT4 / FP4, 4,096 tokens, one stream, live price
Tokens/s
518
Per hour
$30.40
Per million tokens
$16.29
Memory used
79%
Rent on DigitalOcean
Cheapest per million tokens
16 x Instinct MI300X 192GB
INT4 / FP4, 32,768 tokens, 32 streams
Tokens/s
15k
Per hour
$38.24
Per million tokens
$0.71
Memory used
54%
Rent on RunPod

Configurations

Every datacenter GPU in the catalog, sized for these settings. Change precision, context and concurrency and the whole page follows. Each point is one replica: the smallest tensor-parallel group of that GPU that holds the model.

Vendor

8 x Instinct MI355X OAM: 13k tokens/s per replica at $47.92/hr.8 x GB300 NVL72 GPU 288GB: 13k tokens/s per replica at $76.22/hr.16 x GB200 NVL72 GPU 186GB: 23k tokens/s per replica at $168/hr.8 x Instinct MI350X OAM: 13k tokens/s per replica at $49.28/hr.8 x B300 SXM 262GB: 13k tokens/s per replica at $39.92/hr.16 x B200 SXM 180GB: 22k tokens/s per replica at $65.44/hr.8 x Instinct MI325X OAM: 9.7k tokens/s per replica at $30.40/hr.16 x Instinct MI300X 192GB: 15k tokens/s per replica at $38.24/hr.16 x H200 SXM 141GB: 14k tokens/s per replica at $39.68/hr.16 x H200 NVL 141GB: 14k tokens/s per replica at $60.64/hr.

Priced configurations only; 18 more fit without a live price and appear in the table. Marker size grows with the tensor-parallel width. Roofline estimates.

Rent
4 ~20k tok/s Specs →
8 35k tok/s Specs →
TPU 8t announced
8 ~11k tok/s Specs →
TPU 8i announced
8 ~14k tok/s Specs →
8 13k tok/s Rent →
8 13k tok/s Rent →
16 23k tok/s Rent →
8 13k tok/s Rent →
TPU v7 192GB 2 nodes
16 ~21k tok/s Specs →
8 13k tok/s Rent →
16 22k tok/s Rent →
Huawei logo
Ascend 970 announced
8 23k tok/s Specs →
16 23k tok/s Specs →
8 ~9.7k tok/s Rent →
16 ~15k tok/s Rent →
Huawei logo
Ascend 960 announced
8 15k tok/s Specs →
16 ~14k tok/s Rent →
16 ~9.3k tok/s Specs →
16 ~14k tok/s Rent →
Huawei logo
Ascend 910C 2 nodes
16 ~9.1k tok/s Specs →
Huawei logo
Ascend 950DT 2 nodes
16 11k tok/s Specs →
Huawei logo
Ascend 950PR 2 nodes
16 4.5k tok/s Specs →
Gaudi 3 128GB 2 nodes
16 ~11k tok/s Specs →
16 ~1.6k tok/s Specs →
Huawei logo
Atlas 350 2 nodes
16 4.0k tok/s Specs →
MI250X 128GB 2 nodes
16 ~9.3k tok/s Specs →
16 ~9.3k tok/s Specs →
16 ~5.1k tok/s Specs →

By vendor

At INT4 / FP4 the NVIDIA pick is 16 x H200 SXM 141GB at about 14k tokens/s for $39.68/hr; the AMD pick is 16 x Instinct MI300X 192GB at about 15k tokens/s for $38.24/hr. Per rental dollar AMD delivers 1.15x the tokens of NVIDIA here.

NVIDIA: 344 tokens/s per dollar, 1.2k tokens/s per kW (16 x H200 SXM 141GB).AMD: 394 tokens/s per dollar, 1.3k tokens/s per kW (16 x Instinct MI300X 192GB).Huawei: no live price, no power figure (16 x Ascend 950DT).Intel: no live price, 730 tokens/s per kW (16 x Gaudi 3 128GB).Google: no live price, no power figure (16 x TPU v7 192GB).Other: no live price, 649 tokens/s per kW (16 x Cloud AI 100 Ultra).

NVIDIA 9 parts fit
Best value: 16 x H200 SXM 141GB, 14k tok/s for $39.68/hr
Fewest GPUs: 8 x Rubin SXM, 35k tok/s
AMD 7 parts fit
Best value: 16 x Instinct MI300X 192GB, 15k tok/s for $38.24/hr
Fewest GPUs: 4 x Instinct MI455X OAM, 20k tok/s
Huawei 4 parts fit
Fewest GPUs: 16 x Ascend 950DT, 11k tok/s
Intel 2 parts fit
Fewest GPUs: 16 x Gaudi 3 128GB, 11k tok/s
Google 1 parts fit
Fewest GPUs: 16 x TPU v7 192GB, 21k tok/s
Other 1 parts fit
Fewest GPUs: 16 x Cloud AI 100 Ultra, 1.6k tok/s

Fleet what-if

Compare whole fleets at these settings: a thousand of one part against a hundred of another. Each fleet splits into replicas of its tensor-parallel width; GPUs left over sit idle.

A
B
FleetTokens/s
1,008 x H200 SXM 141GB
859k tok/s
112 x GB200 NVL72 GPU 186GB
159k tok/s

1,008 x H200 SXM 141GB: 859k tokens/s, $2,500/hr, 706 kW.112 x GB200 NVL72 GPU 186GB: 159k tokens/s, $1,176/hr, no power figure.

H200 SXM 141GB: 859k tokens/s at 1008 GPUs.GB200 NVL72 GPU 186GB: 1.43M tokens/s at 1008 GPUs.

Size Kimi K3 for a tokens per second target

The calculator starts from your traffic instead of a fleet: set a target rate and read the replica count for every GPU.

Frequently asked questions

How much GPU memory does Kimi K3 need?

The weights take 1,564 GB at the native INT4 / FP4 precision (5,560 GB at BF16, 2,780 GB at FP8, 1,564 GB at INT4). Each concurrent stream adds 27 KB of KV cache per token: 0.1 GB at 4,096 tokens and 0.9 GB at 32,768 tokens.

What is the bare minimum to run Kimi K3?

8 x B300 SXM 262GB at INT4 / FP4 with 4,096 tokens of context and one stream. The cheapest live rental for that footprint is 8 x Instinct MI325X OAM at $30.40 per hour.

How many H100s do you need to run Kimi K3?

more than sixteen H100s at INT4 / FP4 for 4,096 tokens of context and one stream. For 32,768 tokens and 32 concurrent streams the tensor-parallel group is more than sixteen H100s.

How much does Kimi K3 cost per million tokens on an H100?

No live H100 rental price is on file right now, so the cost per token cannot be quoted. The configuration table lists every part that has one.

What precision does Kimi K3 ship in?

The published checkpoint is MXFP4. That 4-bit format is what the lab validated, so the INT4 column is the native one; BF16 figures describe a dequantised copy nobody would deploy.

All figures are roofline estimates from the model's config.json and each GPU's published memory, bandwidth and tensor-core peak, with fixed efficiency factors. Prices are the cheapest live on-demand listing per GPU when the site has one. They are a planning floor; a tuned serving stack can do better. Kimi Linear hybrid: 69 KDA linear-attention layers hold a fixed 434 MB of FP32 state per sequence, and only 24 layers keep an MLA cache (27 KB per token). Ships in MXFP4; 896 routed experts at a reduced width, 16 per token. Active parameters per the Moonshot model card.

© 2026 Flopper.io - Compare the hardware powering AI