- Providers
- Baseten
About Baseten
Baseten runs a managed inference platform for models in production, covering dedicated deployments of custom and fine-tuned models, pre-optimised model APIs billed per token, and training through its Loops SDK. It spans LLMs, image generation, transcription, text-to-speech and embeddings, and offers both cloud-managed and self-hosted deployment with cross-cloud high availability. Compute is billed by the minute against named GPU types including T4, L4, A10G, A100, H100 and B200, and only while a model is actually using it, so there is no idle charge. Named customers include Cursor, Notion, Descript, HubSpot, Writer, Abridge, Clay and OpenEvidence.
Why choose Baseten
Charged only while a model is actively using compute.
Services
Pricing options
Available GPUs
No GPU pricing data available yet. Check back soon.