Docs · API

Streaming supported

Streaming follows the OpenAI Server-Sent Events protocol: set stream: true and read incremental delta chunks.

Gateway SSE streaming is live on /chat/completions. The repo-local helper libraries remain preview packages and are verified separately across the served catalog.

Request

{
  "model": "qwen2.5-3b-instruct",
  "messages": [{"role": "user", "content": "Count to three."}],
  "stream": true
}

Response stream

The response is a text/event-stream. Each event is a data: line carrying a JSON chunk; new tokens arrive in choices[0].delta.content. The stream ends with a [DONE] sentinel.

data: {"choices":[{"delta":{"role":"assistant"}}]}

data: {"choices":[{"delta":{"content":"One"}}]}

data: {"choices":[{"delta":{"content":", two"}}]}

data: {"choices":[{"delta":{"content":", three."}}]}

data: {"choices":[{"delta":{},"finish_reason":"stop"}]}

data: [DONE]

With the SDKs

Compatible Python and JavaScript clients can handle SSE parsing: pass stream=True (Python) or stream: true (JavaScript), then iterate the returned stream.