Docs · Models
Models
JouleCloud serves open-weight models on its own GPU fleet. Below is the catalog available today. The full live list for your org is also visible in the console.
Served models
| Model | Params | Context | Good for | Status |
|---|---|---|---|---|
qwen2.5-3b-instruct | 3B | 32K | High-throughput classification, cheap summarization | available |
qwen2.5-7b-instruct | 7B | 32K | General chat, drafting, extraction, RAG answers | available |
qwen2.5-coder-7b | 7B | 32K | Code generation, refactoring, SQL, structured output | available |
Selecting a model
Set the model field in your request body to the model id, e.g. "model": "qwen2.5-3b-instruct". JouleCloud routes the request across available GPUs that serve the model, favoring locations with cheaper electricity. See Errors for help with model IDs.
Pricing
All served models currently share one usage-based rate: $0.20 per 1M input tokens and $0.60 per 1M output tokens. See Pricing for a worked example and the cost calculator.
Sign in to see the models currently available to your organization.