Docs · API
Streaming supported
Streaming follows the OpenAI Server-Sent Events protocol: set stream: true and read incremental delta chunks.
Gateway SSE streaming is live on
/chat/completions. The repo-local helper libraries remain preview packages and are verified separately across the served catalog.Request
{
"model": "qwen2.5-3b-instruct",
"messages": [{"role": "user", "content": "Count to three."}],
"stream": true
}Response stream
The response is a text/event-stream. Each event is a data: line carrying a JSON chunk; new tokens arrive in choices[0].delta.content. The stream ends with a [DONE] sentinel.
data: {"choices":[{"delta":{"role":"assistant"}}]}
data: {"choices":[{"delta":{"content":"One"}}]}
data: {"choices":[{"delta":{"content":", two"}}]}
data: {"choices":[{"delta":{"content":", three."}}]}
data: {"choices":[{"delta":{},"finish_reason":"stop"}]}
data: [DONE]With the SDKs
Compatible Python and JavaScript clients can handle SSE parsing: pass stream=True (Python) or stream: true (JavaScript), then iterate the returned stream.