Pricing
Pay only for the tokens you use. Input and output tokens have separate rates, and one private-preview rate card applies to every available model.
| Meter | Rate | Unit |
|---|---|---|
| Input tokens | $0.20 | per 1 million tokens |
| Output tokens | $0.60 | per 1 million tokens |
Worked example
1 million input tokens + 1 million output tokens = $0.80 ($0.20 in + $0.60 out). Try your own numbers in the cost calculator.
FAQ
How does billing work?
You are charged per token — input tokens (your prompt) and output tokens (the model’s response) are metered separately at the rates above. Spend accrues against your organization budget, visible live in the console.
What is a token?
A token is a chunk of text — roughly 4 characters, or about three-quarters of a word in English. Both your prompt and the model’s response are counted in tokens.
Is there a budget cap?
Yes. Set a maximum spend for your organization in the console. When spending reaches the limit, requests pause until you raise it. See rate limits & budgets.
How do these rates compare to OpenAI?
Use the cost calculator to compare against OpenAI’s published rates for your own token volumes. The calculator compares prices; model capabilities differ.
Do you offer dedicated capacity?
Yes — reserved GPUs, higher rate limits, and custom terms are arranged individually during private preview. Email us.