Models
Which model do you want to run?
Pick a model and see the bare minimum to run it, every GPU configuration that fits, AMD against NVIDIA, and what a fleet of a thousand costs, all from the model's own config and live rental prices.
21 models, 73 datacenter GPUs, refreshed with every release.
Popular models
All models
GPU counts are the tensor-parallel width of H100 SXM5 80GB parts needed at native precision: minimum is 4,096 tokens of context and one stream, recommended is 32,768 tokens and 32 streams.
| Model | Parameters | Min H100s |
|---|---|---|
| GLM-5.3 | 753B | 16 |
| GLM-5.2 | 753B | Over 16 |
| DeepSeek-V4-Pro Estimated | 1.60T | 16 |
| DeepSeek-V4-Flash Estimated | 291B | 4 |
| Gemma 4 31B | 31.3B | 1 |
| Gemma 4 26B-A4B | 25.8B | 1 |
| Qwen3.5-122B-A10B | 125B | 4 |
| Qwen3.5-27B | 27.8B | 1 |
| Qwen3.5-397B-A17B | 403B | 16 |
| gpt-oss-120b | 117B | 1 |
| gpt-oss-20b | 20.9B | 1 |
| GLM-4.5 | 358B | 16 |
| Kimi K2 Instruct | 1.03T | 16 |
| Mistral Small 3.2 24B | 24.0B | 1 |
| Qwen3-235B-A22B | 235B | 8 |
| Qwen3-32B | 32.8B | 1 |
| Gemma 3 27B | 27.4B | 1 |
| DeepSeek-R1 | 684B | 16 |
| Phi-4 | 14.7B | 1 |
| Llama 3.3 70B Instruct | 70.6B | 2 |
| Llama 3.1 8B Instruct | 8.0B | 1 |