API reference
Streaming (SSE)
A chat completion that arrives as it is written: the chunks, what the stream does not carry, how errors show up, and how to read it with curl, Python and JavaScript.
Last updated: 2026-10-03
Add "stream": true to a request to POST /v1/chat/completions. The answer arrives as Server-Sent Events (Content-Type: text/event-stream), in the OpenAI chunk format, while the model writes it.
Streaming is for chat completions only: /v1/responses answers 501 to "stream": true.
curl
curl -N https://api.ai.itsincom.org/v1/chat/completions \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gemma-4-26b",
"stream": true,
"messages": [{"role": "user", "content": "Count from 1 to 10."}]
}'
-N turns off curl's buffering, so you see the chunks as they arrive.
What arrives
Each event is a line starting with data: , followed by an empty line. The last event is data: [DONE].
data: {"id":"req_…","object":"chat.completion.chunk","created":1791028778,"model":"gemma-4-26b","choices":[{"index":0,"delta":{"role":"assistant"},"finish_reason":null}]}
data: {"id":"req_…","object":"chat.completion.chunk","created":1791028778,"model":"gemma-4-26b","choices":[{"index":0,"delta":{"content":"1, 2, 3"},"finish_reason":null}]}
data: {"id":"req_…","object":"chat.completion.chunk","created":1791028778,"model":"gemma-4-26b","choices":[{"index":0,"delta":{},"finish_reason":"stop"}]}
data: [DONE]
- The first chunk carries only the role.
- Then one chunk per piece of text, in
delta.content. With function calling, the pieces of a call arrive indelta.tool_calls. - The last chunk has an empty
deltaand thefinish_reason. - All the chunks of a stream have the same
id.
An answer is complete when the chunk with the finish_reason has arrived, its value is not error, and [DONE] follows it.
What the stream does not carry
- Token counts. There is no
usagein the stream, not even with"stream_options": {"include_usage": true}: the option is accepted and today has no effect. The counts are recorded with the request and appear in your usage in the dashboard. If your code needs them, call withoutstream. - More than one answer. Every chunk has
index: 0. Do not combinestreamwithngreater than 1. - Token probabilities.
logprobsdo not travel in the chunks.
Headers
The headers leave before the first token, and before we know which machine will answer. X-Zona therefore names the zone if the machines that can serve the request are all in one zone; if your key allows several zones with machines for that model, it lists them. The zone that actually served the request is written in its usage record. More in Zones.
Errors
Before the stream starts — an invalid parameter, an unknown model, no machine in your zones, a rate limit — you get an ordinary JSON error with its HTTP status, not a stream. See Errors.
After it has started, the status is already 200:
- if the machine fails before the first piece of text, the request moves to the next machine in your zones, and you notice nothing;
- if no machine can take over, or the failure comes after the text has started, an event with an
errorobject arrives and the stream ends without[DONE]:
data: {"error":{"message":"Upstream backend unavailable. Retry shortly.","type":"bad_gateway","request_id":"req_…"}}
The text you received up to that point is incomplete. The same holds if the connection closes without [DONE], and if the last chunk has finish_reason: "error".
Stopping
To stop a stream, close the connection. There is no way to resume a stream: a new request starts from the beginning, and it is a new request.
Python (OpenAI SDK)
import os
from openai import OpenAI
client = OpenAI(
base_url="https://api.ai.itsincom.org/v1",
api_key=os.environ["API_KEY"],
)
stream = client.chat.completions.create(
model="gemma-4-26b",
messages=[{"role": "user", "content": "Count from 1 to 10."}],
stream=True,
)
for chunk in stream:
text = chunk.choices[0].delta.content if chunk.choices else None
if text:
print(text, end="", flush=True)
print()
JavaScript (Node 18 or later)
Plain fetch, no library. Save it as stream.mjs and run node stream.mjs.
const res = await fetch("https://api.ai.itsincom.org/v1/chat/completions", {
method: "POST",
headers: {
"Authorization": `Bearer ${process.env.API_KEY}`,
"Content-Type": "application/json",
},
body: JSON.stringify({
model: "gemma-4-26b",
stream: true,
messages: [{ role: "user", content: "Count from 1 to 10." }],
}),
// Set a time limit of your own.
signal: AbortSignal.timeout(120_000),
});
// Errors found before the stream starts are ordinary JSON.
if (!res.ok) throw new Error(`${res.status} ${await res.text()}`);
const reader = res.body.getReader();
const decoder = new TextDecoder();
let buffer = "";
let finished = false;
while (!finished) {
const { value, done } = await reader.read();
if (done) break;
buffer += decoder.decode(value, { stream: true });
const lines = buffer.split("\n");
buffer = lines.pop(); // the last line may still be incomplete
for (const line of lines) {
if (!line.startsWith("data: ")) continue;
const data = line.slice(6);
if (data === "[DONE]") { finished = true; break; }
const event = JSON.parse(data);
if (event.error) throw new Error(event.error.message);
const text = event.choices[0].delta.content;
if (text) process.stdout.write(text);
}
}
if (!finished) throw new Error("stream closed before [DONE]: the answer is incomplete");