Kimi K3 hardware requirements
Needs at least 8 GPUs at INT4 / FP4 (B300 SXM 262GB). For 1,000 tokens/s at 32k context you need 16 x Instinct MI300X 192GB at $38.24/hr.
Top options for Kimi K3
The strongest part each vendor offers for Kimi K3 at INT4 / FP4 with 32,768 tokens per user, rentable parts first. Each card is one replica, the fewest accelerators that hold the model, and the line it draws on the graph below. Select a card to plot it.
Tokens per second as users grow
Each line is one replica, the fewest accelerators that hold the model, serving more users at once. It climbs while memory bandwidth is the limit, flattens once compute is, and stops where memory runs out at 32,768 tokens per user. Hover a line to bring it forward.
8 x GB300 NVL72 GPU 288GB and 8 x Instinct MI355X OAM draw the same line: identical memory and bandwidth give identical curves.
8 x GB300 NVL72 GPU 288GB: peaks at 28k tokens per second when memory fills at 488 users.8 x Instinct MI355X OAM: peaks at 28k tokens per second when memory fills at 488 users.16 x Ascend 950DT: peaks at 25k tokens per second when memory fills at 481 users.16 x Gaudi 3 128GB: peaks at 22k tokens per second when memory fills at 297 users.16 x TPU v7 192GB: peaks at 48k tokens per second and still has memory at 1024 users.
Memory fills at: 8 x GB300 NVL72 GPU 288GB 488 users, 8 x Instinct MI355X OAM 488 users, 16 x Ascend 950DT 481 users, 16 x Gaudi 3 128GB 297 users, 16 x TPU v7 192GB beyond 1,024 users. Roofline estimates at INT4 / FP4.
Bare minimum
- Tokens/s
- 691
- Per hour
- $39.92
- Per million tokens
- $16.05
- Memory used
- 77%
- Tokens/s
- 518
- Per hour
- $30.40
- Per million tokens
- $16.29
- Memory used
- 79%
- Tokens/s
- 15k
- Per hour
- $38.24
- Per million tokens
- $0.71
- Memory used
- 54%
Configurations
Every datacenter GPU in the catalog, sized for these settings. Change precision, context and concurrency and the whole page follows. Each point is one replica: the smallest tensor-parallel group of that GPU that holds the model.
8 x Instinct MI355X OAM: 13k tokens/s per replica at $47.92/hr.8 x GB300 NVL72 GPU 288GB: 13k tokens/s per replica at $76.22/hr.16 x GB200 NVL72 GPU 186GB: 23k tokens/s per replica at $168/hr.8 x Instinct MI350X OAM: 13k tokens/s per replica at $49.28/hr.8 x B300 SXM 262GB: 13k tokens/s per replica at $39.92/hr.16 x B200 SXM 180GB: 22k tokens/s per replica at $65.44/hr.8 x Instinct MI325X OAM: 9.7k tokens/s per replica at $30.40/hr.16 x Instinct MI300X 192GB: 15k tokens/s per replica at $38.24/hr.16 x H200 SXM 141GB: 14k tokens/s per replica at $39.68/hr.16 x H200 NVL 141GB: 14k tokens/s per replica at $60.64/hr.
Priced configurations only; 18 more fit without a live price and appear in the table. Marker size grows with the tensor-parallel width. Roofline estimates.
| Rent | |||
|---|---|---|---|
| 4 | ~20k tok/s | Specs → | |
| 8 | 35k tok/s | Specs → | |
TPU 8t announced | 8 | ~11k tok/s | Specs → |
TPU 8i announced | 8 | ~14k tok/s | Specs → |
| 8 | 13k tok/s | Rent → | |
| 8 | 13k tok/s | Rent → | |
GB200 NVL72 GPU 186GB 2 nodes | 16 | 23k tok/s | Rent → |
| 8 | 13k tok/s | Rent → | |
TPU v7 192GB 2 nodes | 16 | ~21k tok/s | Specs → |
| 8 | 13k tok/s | Rent → | |
B200 SXM 180GB 2 nodes | 16 | 22k tok/s | Rent → |
| 8 | 23k tok/s | Specs → | |
B100 SXM 192GB 2 nodes | 16 | 23k tok/s | Specs → |
| 8 | ~9.7k tok/s | Rent → | |
Instinct MI300X 192GB 2 nodes | 16 | ~15k tok/s | Rent → |
| 8 | 15k tok/s | Specs → | |
H200 SXM 141GB 2 nodes | 16 | ~14k tok/s | Rent → |
Data Center GPU Max 1550 128GB 2 nodes | 16 | ~9.3k tok/s | Specs → |
H200 NVL 141GB 2 nodes | 16 | ~14k tok/s | Rent → |
| 16 | ~9.1k tok/s | Specs → | |
| 16 | 11k tok/s | Specs → | |
| 16 | 4.5k tok/s | Specs → | |
Gaudi 3 128GB 2 nodes | 16 | ~11k tok/s | Specs → |
| 16 | ~1.6k tok/s | Specs → | |
| 16 | 4.0k tok/s | Specs → | |
MI250X 128GB 2 nodes | 16 | ~9.3k tok/s | Specs → |
Instinct MI250 128GB 2 nodes | 16 | ~9.3k tok/s | Specs → |
H20 141GB HBM3e 2 nodes | 16 | ~5.1k tok/s | Specs → |
By vendor
At INT4 / FP4 the NVIDIA pick is 16 x H200 SXM 141GB at about 14k tokens/s for $39.68/hr; the AMD pick is 16 x Instinct MI300X 192GB at about 15k tokens/s for $38.24/hr. Per rental dollar AMD delivers 1.15x the tokens of NVIDIA here.
NVIDIA: 344 tokens/s per dollar, 1.2k tokens/s per kW (16 x H200 SXM 141GB).AMD: 394 tokens/s per dollar, 1.3k tokens/s per kW (16 x Instinct MI300X 192GB).Huawei: no live price, no power figure (16 x Ascend 950DT).Intel: no live price, 730 tokens/s per kW (16 x Gaudi 3 128GB).Google: no live price, no power figure (16 x TPU v7 192GB).Other: no live price, 649 tokens/s per kW (16 x Cloud AI 100 Ultra).
Fleet what-if
Compare whole fleets at these settings: a thousand of one part against a hundred of another. Each fleet splits into replicas of its tensor-parallel width; GPUs left over sit idle.
| Fleet | Tokens/s |
|---|---|
1,008 x H200 SXM 141GB | 859k tok/s |
112 x GB200 NVL72 GPU 186GB | 159k tok/s |
1,008 x H200 SXM 141GB: 859k tokens/s, $2,500/hr, 706 kW.112 x GB200 NVL72 GPU 186GB: 159k tokens/s, $1,176/hr, no power figure.
H200 SXM 141GB: 859k tokens/s at 1008 GPUs.GB200 NVL72 GPU 186GB: 1.43M tokens/s at 1008 GPUs.
Size Kimi K3 for a tokens per second target
The calculator starts from your traffic instead of a fleet: set a target rate and read the replica count for every GPU.
Frequently asked questions
How much GPU memory does Kimi K3 need?
The weights take 1,564 GB at the native INT4 / FP4 precision (5,560 GB at BF16, 2,780 GB at FP8, 1,564 GB at INT4). Each concurrent stream adds 27 KB of KV cache per token: 0.1 GB at 4,096 tokens and 0.9 GB at 32,768 tokens.
What is the bare minimum to run Kimi K3?
8 x B300 SXM 262GB at INT4 / FP4 with 4,096 tokens of context and one stream. The cheapest live rental for that footprint is 8 x Instinct MI325X OAM at $30.40 per hour.
How many H100s do you need to run Kimi K3?
more than sixteen H100s at INT4 / FP4 for 4,096 tokens of context and one stream. For 32,768 tokens and 32 concurrent streams the tensor-parallel group is more than sixteen H100s.
How much does Kimi K3 cost per million tokens on an H100?
No live H100 rental price is on file right now, so the cost per token cannot be quoted. The configuration table lists every part that has one.
What precision does Kimi K3 ship in?
The published checkpoint is MXFP4. That 4-bit format is what the lab validated, so the INT4 column is the native one; BF16 figures describe a dequantised copy nobody would deploy.
All figures are roofline estimates from the model's config.json and each GPU's published memory, bandwidth and tensor-core peak, with fixed efficiency factors. Prices are the cheapest live on-demand listing per GPU when the site has one. They are a planning floor; a tuned serving stack can do better. Kimi Linear hybrid: 69 KDA linear-attention layers hold a fixed 434 MB of FP32 state per sequence, and only 24 layers keep an MLA cache (27 KB per token). Ships in MXFP4; 896 routed experts at a reduced width, 16 per token. Active parameters per the Moonshot model card.

