Models

Which model do you want to run?

Pick a model and see the bare minimum to run it, every GPU configuration that fits, AMD against NVIDIA, and what a fleet of a thousand costs, all from the model's own config and live rental prices.

21 models, 73 datacenter GPUs, refreshed with every release.

Popular models

All models

GPU counts are the tensor-parallel width of H100 SXM5 80GB parts needed at native precision: minimum is 4,096 tokens of context and one stream, recommended is 32,768 tokens and 32 streams.

ModelParametersMin H100s
GLM-5.3 753B 16
GLM-5.2 753B Over 16
DeepSeek-V4-Pro Estimated 1.60T 16
DeepSeek-V4-Flash Estimated 291B 4
Gemma 4 31B 31.3B 1
Gemma 4 26B-A4B 25.8B 1
Qwen3.5-122B-A10B 125B 4
Qwen3.5-27B 27.8B 1
Qwen3.5-397B-A17B 403B 16
gpt-oss-120b 117B 1
gpt-oss-20b 20.9B 1
GLM-4.5 358B 16
Kimi K2 Instruct 1.03T 16
Mistral Small 3.2 24B 24.0B 1
Qwen3-235B-A22B 235B 8
Qwen3-32B 32.8B 1
Gemma 3 27B 27.4B 1
DeepSeek-R1 684B 16
Phi-4 14.7B 1
Llama 3.3 70B Instruct 70.6B 2
Llama 3.1 8B Instruct 8.0B 1