Every model.
One key.
Pick a paid plan and get the entire catalog — every model, every capability tier — behind a single key; Free runs the open-model catalog on that same key. Each plan includes a monthly pool of model usage measured in dollars at public list rates: $45 of usage on a $5 plan, $200 on Pro, $525 on Max.
The open-model catalog, zero commitment. See what one key can do.
- $2 one-time credit to start
- The full open-model catalog, one key
- 30 requests/min · 2 concurrent
- In-flight runs never cut off
A serious month of agent work for the price of a coffee.
- $45 of model usage included every month
- Every model — every capability tier, one key
- 300 requests/min · 10 concurrent
- In-flight runs never cut off
For a daily driver — coding agents that work all day.
- $200 of model usage included every month
- Every model — every capability tier, one key
- 1,000 requests/min · 40 concurrent
- In-flight runs never cut off
5x Pro. For agents that never sleep and teams of one that ship like ten.
- $525 of model usage included every month
- Every model — every capability tier, one key
- 3,000 requests/min · 100 concurrent
- In-flight runs never cut off
No subscription. Load credit and spend it whenever you like — at each model's published rate, zero markup. Credit doesn't expire and doesn't reset, and it keeps working after a monthly pool runs out.
Need more than Max — higher limits, custom usage pools, procurement? Talk to us and we'll size a plan to your fleet.
Included usage is measured in dollars at each model's public list rate — see Plans & limits for how the meter works.
- What counts as included usage?
- Every request is metered in dollars at the model's public list rate — the same per-million-token prices printed on each model's page — and drawn from your plan's monthly pool. A $19 Pro plan carries $200 of usage measured that way: use any model in the catalog and the meter always values it at its public list price.
- Does repeated context burn my included usage?
- No — on supported models, repeated context is cached automatically and counts at just 10% of the model's input list rate. A long agent session that resends the same system prompt and history every turn draws far less from your pool than its raw token count suggests. See Prompt caching in the docs.
- Which models are included?
- Every paid plan covers the entire catalog: every model, every capability tier, no per-model surcharge and no premium gate. Free covers the open-model catalog — the open-weight lineup in full, at the same limits and protocols. Between paid plans, the only thing that changes is how much usage is included and how fast you can push it.
- What happens when I run out?
- In-flight runs finish — nothing is killed mid-stream. As you approach the edge of your plan, the heaviest models pace to your plan while everything else keeps full speed; at your plan's monthly ceiling, new requests pause until the pool refreshes. If you need it back immediately, upgrade takes effect right away; otherwise it refreshes at the start of your next monthly cycle.
- How do rate limits work?
- Each plan carries a requests-per-minute rate and a number of simultaneous in-flight requests. Past either one, the API returns 429 with a retry-after header telling your client exactly how long to wait — every serious SDK handles that automatically. The full table is on the Plans & limits page in the docs.
- Can I try it before paying?
- Yes — the Free plan is a real plan, not a demo: $2 of included usage each month across the open-model catalog, with both API dialects and streaming. Point one agent at the endpoint and see how it runs before you spend a dollar.
Your agent doesn't change.
Everything underneath does.
Start free · no card · Starter from $5/mo