ITS INCOM AI ITS INCOM AI docs

API reference

Streaming (SSE)

A chat completion that arrives as it is written: the chunks, what the stream does not carry, how errors show up, and how to read it with curl, Python and JavaScript.

Last updated: 2026-10-03

Add "stream": true to a request to POST /v1/chat/completions. The answer arrives as Server-Sent Events (Content-Type: text/event-stream), in the OpenAI chunk format, while the model writes it.

Streaming is for chat completions only: /v1/responses answers 501 to "stream": true.

curl

bash
curl -N https://api.ai.itsincom.org/v1/chat/completions \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gemma-4-26b",
    "stream": true,
    "messages": [{"role": "user", "content": "Count from 1 to 10."}]
  }'

-N turns off curl's buffering, so you see the chunks as they arrive.

What arrives

Each event is a line starting with data: , followed by an empty line. The last event is data: [DONE].

text
data: {"id":"req_…","object":"chat.completion.chunk","created":1791028778,"model":"gemma-4-26b","choices":[{"index":0,"delta":{"role":"assistant"},"finish_reason":null}]}

data: {"id":"req_…","object":"chat.completion.chunk","created":1791028778,"model":"gemma-4-26b","choices":[{"index":0,"delta":{"content":"1, 2, 3"},"finish_reason":null}]}

data: {"id":"req_…","object":"chat.completion.chunk","created":1791028778,"model":"gemma-4-26b","choices":[{"index":0,"delta":{},"finish_reason":"stop"}]}

data: [DONE]
  • The first chunk carries only the role.
  • Then one chunk per piece of text, in delta.content. With function calling, the pieces of a call arrive in delta.tool_calls.
  • The last chunk has an empty delta and the finish_reason.
  • All the chunks of a stream have the same id.

An answer is complete when the chunk with the finish_reason has arrived, its value is not error, and [DONE] follows it.

What the stream does not carry

  • Token counts. There is no usage in the stream, not even with "stream_options": {"include_usage": true}: the option is accepted and today has no effect. The counts are recorded with the request and appear in your usage in the dashboard. If your code needs them, call without stream.
  • More than one answer. Every chunk has index: 0. Do not combine stream with n greater than 1.
  • Token probabilities. logprobs do not travel in the chunks.

Headers

The headers leave before the first token, and before we know which machine will answer. X-Zona therefore names the zone if the machines that can serve the request are all in one zone; if your key allows several zones with machines for that model, it lists them. The zone that actually served the request is written in its usage record. More in Zones.

Errors

Before the stream starts — an invalid parameter, an unknown model, no machine in your zones, a rate limit — you get an ordinary JSON error with its HTTP status, not a stream. See Errors.

After it has started, the status is already 200:

  • if the machine fails before the first piece of text, the request moves to the next machine in your zones, and you notice nothing;
  • if no machine can take over, or the failure comes after the text has started, an event with an error object arrives and the stream ends without [DONE]:
text
data: {"error":{"message":"Upstream backend unavailable. Retry shortly.","type":"bad_gateway","request_id":"req_…"}}

The text you received up to that point is incomplete. The same holds if the connection closes without [DONE], and if the last chunk has finish_reason: "error".

Stopping

To stop a stream, close the connection. There is no way to resume a stream: a new request starts from the beginning, and it is a new request.

Python (OpenAI SDK)

python
import os

from openai import OpenAI

client = OpenAI(
    base_url="https://api.ai.itsincom.org/v1",
    api_key=os.environ["API_KEY"],
)

stream = client.chat.completions.create(
    model="gemma-4-26b",
    messages=[{"role": "user", "content": "Count from 1 to 10."}],
    stream=True,
)

for chunk in stream:
    text = chunk.choices[0].delta.content if chunk.choices else None
    if text:
        print(text, end="", flush=True)
print()

JavaScript (Node 18 or later)

Plain fetch, no library. Save it as stream.mjs and run node stream.mjs.

javascript
const res = await fetch("https://api.ai.itsincom.org/v1/chat/completions", {
  method: "POST",
  headers: {
    "Authorization": `Bearer ${process.env.API_KEY}`,
    "Content-Type": "application/json",
  },
  body: JSON.stringify({
    model: "gemma-4-26b",
    stream: true,
    messages: [{ role: "user", content: "Count from 1 to 10." }],
  }),
  // Set a time limit of your own.
  signal: AbortSignal.timeout(120_000),
});

// Errors found before the stream starts are ordinary JSON.
if (!res.ok) throw new Error(`${res.status} ${await res.text()}`);

const reader = res.body.getReader();
const decoder = new TextDecoder();
let buffer = "";
let finished = false;

while (!finished) {
  const { value, done } = await reader.read();
  if (done) break;
  buffer += decoder.decode(value, { stream: true });

  const lines = buffer.split("\n");
  buffer = lines.pop(); // the last line may still be incomplete

  for (const line of lines) {
    if (!line.startsWith("data: ")) continue;
    const data = line.slice(6);
    if (data === "[DONE]") { finished = true; break; }
    const event = JSON.parse(data);
    if (event.error) throw new Error(event.error.message);
    const text = event.choices[0].delta.content;
    if (text) process.stdout.write(text);
  }
}

if (!finished) throw new Error("stream closed before [DONE]: the answer is incomplete");