Pricing
Usage-based pricing. You pay per token — no tiers, no seats, no minimums.
| Meter | Rate | Unit |
|---|---|---|
| Input tokens | $0.20 | per 1M tokens |
| Output tokens | $0.60 | per 1M tokens |
Per model
All models currently share one rate.
| Model | Input / 1M | Output / 1M |
|---|---|---|
qwen2.5-7b-instruct | $0.20 | $0.60 |
qwen2.5-3b-instruct | $0.20 | $0.60 |
Worked example
1M input tokens + 1M output tokens = $0.80 ($0.20 in + $0.60 out). Try your own numbers in the cost calculator.
FAQ
How does billing work?
You are charged per token — input tokens (your prompt) and output tokens (the generation) are metered separately at the rates above. Spend accrues against your organization budget, visible live in the console.
What is a token?
A token is a chunk of text — roughly 4 characters, or about three-quarters of a word in English. Both your prompt and the model’s response are counted in tokens.
Is there a budget cap?
Yes. Each organization can set a maximum spend; requests are blocked once the cap is reached until you raise it. See rate limits & budgets.
How do these rates compare to OpenAI?
Use the cost calculator to compare against OpenAI’s published rates for your own token volumes. It is a price-only comparison, not a model-quality claim.