Baseten
Inference is everything.
At a glance
prices checked just now- Cheapest GPU
- $0.63/hr
- NVIDIA T4 16GB
- Fastest GPU
- $9.98/hr
- NVIDIA B200 SXM 180GB
- GPU models
- 6
- 6 configurations
- Certifications
- 5
- SOC 2 Type II, ISO 27001
What's good about Baseten
- Billing Per minute, no idle
- GPU types T4, L4, A10G, A100, H100, B200
- Model APIs Billed per token
- Deployment Cloud-managed or self-hosted
- Named customers Cursor, Notion, Descript, HubSpot
Available GPUs
Cheapest first, per GPU per hour.
| GPU | Configuration | Memory | Price | Region | |
|---|---|---|---|---|---|
| NVIDIA T4 16GB Turing | 1 GPU | 16 GB | $0.63/hr | - | Rent |
| NVIDIA L4 24GB Ada Lovelace | 1 GPU | 24 GB | $0.85/hr | - | Rent |
| NVIDIA A10G 24GB Ampere | 1 GPU | 24 GB | $1.21/hr | - | Rent |
| NVIDIA A100 SXM4 80GB Ampere | 1 GPU | 80 GB | $4.00/hr | - | Rent |
| NVIDIA H100 SXM5 80GB Hopper | 1 GPU | 80 GB | $6.50/hr | - | Rent |
| NVIDIA B200 SXM 180GB Blackwell | 1 GPU | 180 GB | $9.98/hr | - | Rent |
Baseten head to head
Compare any twoAbout Baseten
Baseten runs a managed inference platform for models in production, covering dedicated deployments of custom and fine-tuned models, pre-optimised model APIs billed per token, and training through its Loops SDK. It spans LLMs, image generation, transcription, text-to-speech and embeddings, and offers both cloud-managed and self-hosted deployment with cross-cloud high availability. Compute is billed by the minute against named GPU types including T4, L4, A10G, A100, H100 and B200, and only while a model is actually using it, so there is no idle charge. Named customers include Cursor, Notion, Descript, HubSpot, Writer, Abridge, Clay and OpenEvidence.
Pricing options
Services
- Per minute, no idle Billing
- Charged only while a model is actively using compute.
Compliance and certifications
Independently verified standards governing security, privacy and operations.
Ready to deploy on Baseten?
Visit Baseten





