Model catalog

Models

JouleCloud hosts these models on its own GPUs. Access every model through the same OpenAI-compatible API at one usage-based rate.

Qwen2.5 Instructavailable
Parameters3B
Context32K
HardwareRTX 3090
High-throughput classification, cheap summarization
$0.20 input · $0.60 output per 1 million tokens · Run it in the playground →
Qwen2.5 Instructavailable
Parameters7B
Context32K
HardwareRTX 3090
General chat, drafting, extraction, RAG answers
$0.20 input · $0.60 output per 1 million tokens · Run it in the playground →
Qwen2.5 Coderavailable
Parameters7B
Context32K
HardwareRTX 3090
Code generation, refactoring, SQL, structured output
$0.20 input · $0.60 output per 1 million tokens · Run it in the playground →
The list above shows the models featured publicly. Sign in and open the console to see the full live catalog for your organization.

served-model automatically routes requests to your organization’s default model. Specify a model ID when you want a particular model.

To choose a model, set the model field to one of the IDs above. See the model docs or cost calculator.