Models

Open-weight models served on JouleCloud’s own GPU fleet, behind one OpenAI-compatible endpoint. Every model shares the same usage-based rate.

qwen2.5-7b-instruct
Qwen2.5 Instructavailable
Parameters7B
Context32K (Qwen2.5 native)
HardwareRTX 3090
General chat, drafting, extraction, RAG answers
$0.20 in · $0.60 out / 1M tokens
qwen2.5-3b-instruct
Qwen2.5 Instructavailable
Parameters3B
Context32K (Qwen2.5 native)
HardwareRTX 3090
High-throughput classification, cheap summarization
$0.20 in · $0.60 out / 1M tokens
Signed-out visitors see the curated served list above. Sign in and open the console to view the full live catalog for your organization — it is queried with your org key.

Selecting a model: set the model field in your request to the id above. See the model docs or cost calculator.