Docs · Understand routing
How JouleCloud routes inference
JouleCloud continuously routes requests toward available GPUs in lower-cost electricity markets.
From request to GPU
- Your client sends a request to the gateway with a served model ID.
- JouleCloud checks model availability, GPU availability, and current electricity prices.
- The selected GPU returns the standard inference response.
Choose an available GPU
Before receiving traffic, every GPU must be online, responsive, and ready for the selected model. As capacity changes, JouleCloud routes across the available fleet.
Price-aware placement
Among available GPUs, JouleCloud sends more requests to lower-cost zones.
Price data updates continuously while the available fleet keeps requests moving.
How price shifts traffic
JouleCloud uses real-time prices first and next-day prices when they provide the freshest signal.
A lower-cost zone must stay cheaper long enough to justify moving traffic. This keeps routing stable and avoids needless model reloads. Your request format never changes.
Continue
- Energy routing — how lower prices align with lower emissions
- Models — choose the logical model ID sent in a request
- Errors — handle unavailable capacity and transient failures