Docs · Understand routing

How JouleCloud routes inference

JouleCloud continuously routes requests toward available GPUs in lower-cost electricity markets.

From request to GPU

  1. Your client sends a request to the gateway with a served model ID.
  2. JouleCloud checks model availability, GPU availability, and current electricity prices.
  3. The selected GPU returns the standard inference response.

Choose an available GPU

Before receiving traffic, every GPU must be online, responsive, and ready for the selected model. As capacity changes, JouleCloud routes across the available fleet.

Price-aware placement

Among available GPUs, JouleCloud sends more requests to lower-cost zones.

Price data updates continuously while the available fleet keeps requests moving.

How price shifts traffic

JouleCloud uses real-time prices first and next-day prices when they provide the freshest signal.

A lower-cost zone must stay cheaper long enough to justify moving traffic. This keeps routing stable and avoids needless model reloads. Your request format never changes.

Continue