Pricing

Usage-based pricing. You pay per token — no tiers, no seats, no minimums.

MeterRateUnit
Input tokens$0.20per 1M tokens
Output tokens$0.60per 1M tokens

Per model

All models currently share one rate.

ModelInput / 1MOutput / 1M
qwen2.5-7b-instruct$0.20$0.60
qwen2.5-3b-instruct$0.20$0.60

Worked example

1M input tokens + 1M output tokens = $0.80 ($0.20 in + $0.60 out). Try your own numbers in the cost calculator.

FAQ

How does billing work?

You are charged per token — input tokens (your prompt) and output tokens (the generation) are metered separately at the rates above. Spend accrues against your organization budget, visible live in the console.

What is a token?

A token is a chunk of text — roughly 4 characters, or about three-quarters of a word in English. Both your prompt and the model’s response are counted in tokens.

Is there a budget cap?

Yes. Each organization can set a maximum spend; requests are blocked once the cap is reached until you raise it. See rate limits & budgets.

How do these rates compare to OpenAI?

Use the cost calculator to compare against OpenAI’s published rates for your own token volumes. It is a price-only comparison, not a model-quality claim.

Private preview pricing. A single usage-based rate applies today; there are no Startup / Pro / Enterprise tiers.