Baseten

Inference is everything.

0 GPUs available

About Baseten

Baseten runs a managed inference platform for models in production, covering dedicated deployments of custom and fine-tuned models, pre-optimised model APIs billed per token, and training through its Loops SDK. It spans LLMs, image generation, transcription, text-to-speech and embeddings, and offers both cloud-managed and self-hosted deployment with cross-cloud high availability. Compute is billed by the minute against named GPU types including T4, L4, A10G, A100, H100 and B200, and only while a model is actually using it, so there is no idle charge. Named customers include Cursor, Notion, Descript, HubSpot, Writer, Abridge, Clay and OpenEvidence.

Why choose Baseten

Billing
Per minute, no idle

Charged only while a model is actively using compute.

T4, L4, A10G, A100, H100, B200
GPU types
Billed per token
Model APIs
Cloud-managed or self-hosted
Deployment
Cursor, Notion, Descript, HubSpot
Named customers

Services

Dedicated model inferencepre-optimised model APIstraining (Loops SDK)self-hosted deployment

Pricing options

Per-minute GPU computeper-token for model APIs

Available GPUs

No GPU pricing data available yet. Check back soon.

Ready to deploy on Baseten?

Inference is everything.

Get started with Baseten