Docs · Start
Send your first inference request
Call the OpenAI-compatible Chat Completions endpoint with an organization API key and a served model ID.
1. Create an API key
Each organization gets one gateway key, minted on demand. After you sign up, open the console to reveal, copy, or rotate it. Use it as a bearer token: Authorization: Bearer <key>. See Authentication for details.
2. Choose a model
Set model to one of the served models (see Models). This example uses qwen2.5-3b-instruct.
3. Send the request
Send POST https://api.jouledns.com/v1/chat/completions. The examples below use the same endpoint with cURL and the official OpenAI Python and JavaScript clients.
curl https://api.jouledns.com/v1/chat/completions \
-H "Authorization: Bearer $JC_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"qwen2.5-3b-instruct","messages":[{"role":"user","content":"Hello"}]}'The examples read your key from a
JC_API_KEY environment variable. Never hard-code keys into source you commit.Next
- Routing — how GPU availability and electricity price determine placement
- Energy signals — how price and operational carbon differ
- Chat Completions — request and response shapes
- Streaming — token-by-token responses
- Errors — status codes and the error envelope
- Rate limits — per-org RPM and budget caps
- Playground — try it against the real gateway, in the browser