Docs · Models

Models

JouleCloud serves open-weight models on its own GPU fleet. Below is the catalog available today. The full live list for your org is also visible in the console.

Served models

ModelParamsContextGood forStatus
qwen2.5-3b-instruct3B32KHigh-throughput classification, cheap summarizationavailable
qwen2.5-7b-instruct7B32KGeneral chat, drafting, extraction, RAG answersavailable
qwen2.5-coder-7b7B32KCode generation, refactoring, SQL, structured outputavailable

Selecting a model

Set the model field in your request body to the model id, e.g. "model": "qwen2.5-3b-instruct". JouleCloud routes the request across available GPUs that serve the model, favoring locations with cheaper electricity. See Errors for help with model IDs.

Pricing

All served models currently share one usage-based rate: $0.20 per 1M input tokens and $0.60 per 1M output tokens. See Pricing for a worked example and the cost calculator.

Sign in to see the models currently available to your organization.