ITS INCOM AI ITS INCOM AI docs

Cookbook

Streaming in TypeScript

A chat completion read as it arrives, in TypeScript: a parser that notices errors and incomplete answers, a Next.js route that keeps the key on the server, a React component, and the same with the OpenAI SDK.

Last updated: 2026-10-03

A parser in plain TypeScript, a Next.js route that keeps your key on the server, and a React component that shows the answer as it arrives. The format itself is described in Streaming.

What the code has to handle

See the stream first:

bash
curl -N https://api.ai.itsincom.org/v1/chat/completions \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gemma-4-26b",
    "stream": true,
    "messages": [{"role": "user", "content": "Count from 1 to 5."}]
  }'

-N turns off curl's buffering. What arrives, shortened:

text
data: {"id":"req_…","object":"chat.completion.chunk",…,"choices":[{"index":0,"delta":{"role":"assistant"},"finish_reason":null}]}

data: {"id":"req_…","object":"chat.completion.chunk",…,"choices":[{"index":0,"delta":{"content":"1, 2, 3"},"finish_reason":null}]}

data: {"id":"req_…","object":"chat.completion.chunk",…,"choices":[{"index":0,"delta":{},"finish_reason":"stop"}]}

data: [DONE]

Five facts shape the code below:

  1. Events. Each one is a data: line followed by an empty line. The text is in choices[0].delta.content; the stream ends with data: [DONE].

  2. No event arrives until a machine starts answering. While the request waits for one, the stream is silent: do not put a short timeout on the first event.

  3. Errors come in two forms. What is wrong before the stream starts (a parameter, the key, the model, a rate limit, no machine in your zones) is an ordinary JSON response with an error status. Once the stream has started the status is already 200, and an error arrives as an event, with no [DONE] after it:

    text
    data: {"error":{"message":"Upstream backend unavailable. Retry shortly.","type":"bad_gateway","request_id":"req_…"}}
    
  4. Incomplete answers. A stream that ends without [DONE], or whose last chunk has finish_reason: "error", is an incomplete answer. Show it as such.

  5. No token counts. There is no usage in the stream, not even with "stream_options": {"include_usage": true}: the option is accepted and today has no effect. If your code needs the counts, call without stream and read usage in the response; otherwise they are in your usage records.

Every chunk carries the id of the request: if something goes wrong, quote it to segreteria@itsincom.it.

A parser

Plain TypeScript, for Node 20 or later and for browsers. It calls onText for every piece of text and resolves with the finish_reason; it throws on an error before or inside the stream, and on an incomplete answer.

typescript
// lib/read-chat-stream.ts
export async function readChatStream(
  res: Response,
  onText: (text: string) => void,
): Promise<string> {
  if (!res.ok || !res.body) {
    // An error before the stream: ordinary JSON. The per-address 429 has
    // only a top-level `message`, the other errors an `error` object.
    const body = await res.json().catch(() => null);
    throw new Error(body?.error?.message ?? body?.message ?? `HTTP ${res.status}`);
  }

  const reader = res.body.getReader();
  const decoder = new TextDecoder();
  let buffer = "";
  let finishReason = "";

  for (;;) {
    const { value, done } = await reader.read();
    buffer += done ? decoder.decode() + "\n" : decoder.decode(value, { stream: true });

    const lines = buffer.split("\n");
    buffer = lines.pop() ?? ""; // a line without its end waits for the next read

    for (const line of lines) {
      if (!line.startsWith("data:")) continue;
      const data = line.slice(5).trim();

      if (data === "[DONE]") {
        if (finishReason === "error") throw new Error("Incomplete answer (finish_reason: error).");
        return finishReason;
      }

      const event = JSON.parse(data);
      if (event.error) {
        throw new Error(`${event.error.message} (request ${event.error.request_id})`);
      }
      const choice = event.choices?.[0];
      if (choice?.delta?.content) onText(choice.delta.content);
      if (choice?.finish_reason) finishReason = choice.finish_reason;
    }

    if (done) throw new Error("The stream ended without [DONE]: incomplete answer.");
  }
}

To try it from a terminal, put the function and these lines in one file, main.mts, and run API_KEY=sk-… npx tsx main.mts:

typescript
const res = await fetch("https://api.ai.itsincom.org/v1/chat/completions", {
  method: "POST",
  headers: {
    Authorization: `Bearer ${process.env.API_KEY}`,
    "Content-Type": "application/json",
  },
  body: JSON.stringify({
    model: "gemma-4-26b",
    stream: true,
    messages: [{ role: "user", content: "Count from 1 to 5." }],
  }),
});

const reason = await readChatStream(res, (text) => process.stdout.write(text));
console.log(`\n[finish_reason: ${reason}]`);

Next.js: the key stays on the server

The browser must never see your key. A route handler calls the API with it and passes the stream on unchanged:

typescript
// app/api/chat/route.ts
export async function POST(req: Request) {
  const { messages } = await req.json();

  const upstream = await fetch("https://api.ai.itsincom.org/v1/chat/completions", {
    method: "POST",
    headers: {
      Authorization: `Bearer ${process.env.API_KEY}`,
      "Content-Type": "application/json",
    },
    body: JSON.stringify({ model: "gemma-4-26b", stream: true, messages }),
    signal: req.signal, // abort this call when the browser's request is aborted
  });

  if (!upstream.ok || !upstream.body) {
    // An error before the stream: pass it on as it is, Retry-After included.
    const headers = new Headers({
      "Content-Type": upstream.headers.get("Content-Type") ?? "application/json",
    });
    const retryAfter = upstream.headers.get("Retry-After");
    if (retryAfter) headers.set("Retry-After", retryAfter);
    return new Response(await upstream.text(), { status: upstream.status, headers });
  }

  return new Response(upstream.body, {
    headers: {
      "Content-Type": "text/event-stream",
      "Cache-Control": "no-cache, no-transform",
      "X-Accel-Buffering": "no",
    },
  });
}

Cache-Control: no-cache, no-transform and X-Accel-Buffering: no are the headers the API itself sends, so that proxies pass the stream on as it comes instead of holding it back. Keep them if a proxy such as nginx sits in front of your app.

React: the answer as it arrives

The parser above in lib/read-chat-stream.ts, and a component that uses it:

tsx
"use client";
// app/chat.tsx

import { useEffect, useRef, useState } from "react";
import { readChatStream } from "@/lib/read-chat-stream";

export function Chat() {
  const [answer, setAnswer] = useState("");
  const [error, setError] = useState("");
  const [busy, setBusy] = useState(false);
  const current = useRef<AbortController | null>(null);

  // Leaving the page stops the answer.
  useEffect(() => () => current.current?.abort(), []);

  async function ask(question: string) {
    current.current?.abort(); // a new question stops the previous answer
    const controller = new AbortController();
    current.current = controller;

    setAnswer("");
    setError("");
    setBusy(true);
    try {
      const res = await fetch("/api/chat", {
        method: "POST",
        headers: { "Content-Type": "application/json" },
        body: JSON.stringify({ messages: [{ role: "user", content: question }] }),
        signal: controller.signal,
      });
      await readChatStream(res, (text) => setAnswer((a) => a + text));
    } catch (e) {
      // The text received so far stays on screen, marked as incomplete.
      if (!controller.signal.aborted) setError((e as Error).message);
    } finally {
      if (current.current === controller) setBusy(false);
    }
  }

  return (
    <div>
      <button onClick={() => ask("Count from 1 to 10.")} disabled={busy}>
        Ask
      </button>
      <pre>{answer}</pre>
      {error && <p role="alert">Incomplete answer: {error}</p>}
    </div>
  );
}

Stopping

A stream is stopped by closing the connection; it cannot be resumed, and a new request starts from the beginning. That is why the component aborts its request when a new question comes or the page goes away: the gateway checks, at every piece of text, whether the connection is still open, and stops relaying the answer when it is not.

With the OpenAI SDK

If you use openai-node, the same rules apply. Check that the answer has a finish_reason, because a stream that stops early has none:

typescript
// sdk.mts. npm install openai, then run: API_KEY=sk-… npx tsx sdk.mts
import OpenAI from "openai";

const client = new OpenAI({ baseURL: "https://api.ai.itsincom.org/v1", apiKey: process.env.API_KEY });

const stream = await client.chat.completions.create({
  model: "gemma-4-26b",
  stream: true,
  messages: [{ role: "user", content: "Count from 1 to 5." }],
});

let finishReason: string | null = null;
for await (const chunk of stream) {
  process.stdout.write(chunk.choices[0]?.delta?.content ?? "");
  finishReason = chunk.choices[0]?.finish_reason ?? finishReason;
}
if (!finishReason || finishReason === "error") throw new Error("Incomplete answer.");