API reference
Chat sessions
Conversations kept on the server for a signed-in account: sign in, create, list, rename, delete, and send messages with a streamed answer.
Last updated: 2026-10-03
The developer API is stateless: you send the whole conversation every time, with an API key. This page describes the other API, which keeps the conversations on the server for the account that signed in. It is the API for a chat client — a mobile app, for example — on top of an account of ITS INCOM AI.
- You send only the new message: the history is kept by ITS INCOM AI.
- The answer arrives as a stream of events.
- After the first exchange, the conversation gets a title in the background.
- These are the same conversations as in the web chat of ITS INCOM AI.
Base address and authentication
https://my.ai.itsincom.org/api/v1
Authorization: Bearer <access token>
The token is not an API key: it belongs to a signed-in user, and you get it with the email and the password of the account. An API key (sk-…) gets 401 here.
Signing in
curl -s https://my.ai.itsincom.org/api/v1/auth/login \
-H "Content-Type: application/json" \
-d '{"email": "you@example.com", "password": "…"}'
{
"token": "eyJ…",
"expires_in_seconds": 86400,
"access_token": "eyJ…",
"refresh_token": "eyJ…",
"access_expires_in_seconds": 86400,
"refresh_expires_in_seconds": 2592000,
"user": { "id": "…", "email": "you@example.com", "email_verified": true }
}
(user has a few more fields than shown here.)
- Send
access_tokeninAuthorization: Bearer.tokenandexpires_in_secondsrepeat it, for older clients. - The access token lasts
access_expires_in_seconds— 24 hours by default. The refresh token lastsrefresh_expires_in_seconds— 30 days by default. - A wrong email or password gets
401with codeinvalid_credentials.
To get a new pair before the access token expires, or after:
curl -s https://my.ai.itsincom.org/api/v1/auth/refresh \
-H "Content-Type: application/json" \
-d '{"refresh_token": "'"$REFRESH_TOKEN"'"}'
Every refresh returns a new pair and retires the refresh token you sent. If a retired refresh token is used again, every refresh token of the account is revoked, and the user has to sign in again: we treat it as a stolen token.
POST /auth/logout, with the access token in the header and {"refresh_token": "…"} in the body, revokes the refresh token. The access token keeps working until it expires: delete it on the client.
Everything else on this page needs an account with a verified email. Without it the answer is 403 with code email_unverified.
Endpoints
| Method | Path | What it does |
|---|---|---|
GET |
/chat/sessions |
Your conversations, without their messages |
POST |
/chat/sessions |
Create a conversation |
GET |
/chat/sessions/{id} |
One conversation, with its messages |
PATCH |
/chat/sessions/{id} |
Rename it |
DELETE |
/chat/sessions/{id} |
Delete it, with its messages |
POST |
/chat/sessions/{id}/messages |
Send a message; the answer is streamed |
List
GET /chat/sessions returns a JSON array with up to 200 conversations, the most recent first. messages is always empty here. Conversations archived in the web chat are left out; this API cannot archive.
[
{
"id": "…",
"title": "Notice periods",
"model": "gemma-4-26b",
"created_at": "2026-10-03T08:12:00+00:00",
"updated_at": "2026-10-03T08:14:32+00:00",
"messages": []
}
]
model is the model of the last message sent, or the one given at creation.
Read one
GET /chat/sessions/{id} returns the same object with all the messages, oldest first:
{
"id": "…",
"role": "assistant",
"content": "A notice period is…",
"created_at": "2026-10-03T08:14:32+00:00",
"prompt_tokens": 24,
"completion_tokens": 187,
"latency_ms": 0
}
latency_ms is always 0 here: the time of an answer is in the done event, below. On user messages the token counts are 0.
Create
curl -s https://my.ai.itsincom.org/api/v1/chat/sessions \
-H "Authorization: Bearer $TOKEN" \
-H "Content-Type: application/json" \
-d '{"title": "Contract questions"}'
Both fields are optional: title (up to 255 characters) and model (up to 128). The answer is 201 with the new conversation; use its id for the messages.
- Without a title, the conversation is called
Nuova conversazioneuntil the automatic title replaces it. modelis only a label,qwen2.5:1.5bif you leave it out: every message names its own model.
Rename
PATCH /chat/sessions/{id} with {"title": "…"}, required, up to 255 characters. A title you set yourself — here or at creation — is kept: the automatic title replaces only the default one.
Delete
DELETE /chat/sessions/{id} deletes the conversation and all its messages. There is no undo. The answer is 204, with no body.
Send a message
curl -N https://my.ai.itsincom.org/api/v1/chat/sessions/$CONV_ID/messages \
-H "Authorization: Bearer $TOKEN" \
-H "Content-Type: application/json" \
-d '{
"content": "What is a notice period?",
"model": "gemma-4-26b",
"tier": "medium"
}'
| Field | Required | Notes |
|---|---|---|
content |
yes | The text of the message, up to 8000 characters. Text only. |
model |
yes | A model_id from the catalogue. |
tier |
no | slow, medium, fast or ludicrous. Without it: the tier of the previous message of this conversation, and slow for the first one. It decides which machines may answer: see Tiers. |
backend |
no | A preference for one family of machines, by its internal name: they are tried first, the others after. It is remembered for the next messages of the conversation. You do not need it. |
- The whole history of the conversation goes to the model with every message. Nothing is trimmed: a very long conversation can exceed what the model accepts, and then the answer fails.
- The answer is generated only on machines in the zones offered by ITS INCOM AI, and never moved to another zone. See Zones.
- Your message is saved before the answer starts, the answer when it ends.
The answer is text/event-stream. Every event is a line data: {…} followed by an empty line. For example:
data: {"type":"delta","text":"A notice period "}
data: {"type":"delta","text":"is the time between "}
data: {"type":"done","prompt_tokens":24,"completion_tokens":187,"latency_ms":4823,"outcome":"ok","title_job_queued":true}
type |
Fields | Meaning |
|---|---|---|
delta |
text |
A piece of the answer: append it to what you have. |
error |
code, message |
The answer failed. A done with "outcome": "error" follows. |
done |
prompt_tokens, completion_tokens, latency_ms, outcome, title_job_queued |
The end. latency_ms is the time of the whole answer; outcome is ok or error. |
The code of an error event:
zone_unavailable: no machine in the zones of ITS INCOM AI can serve that model right now. We do not send it elsewhere: retry later, or choose another model. Themessageof this code is in Italian today.model_unavailable: the same, on an installation without zones.upstream_error: anything else. Today this includes amodelthat does not exist: check the name in the catalogue.
Once the stream has started, the HTTP status stays 200 even if an error event arrives. Errors found before it starts are ordinary JSON answers: see Errors. When the answer fails, the conversation keeps an assistant message with the text ⚠️ Errore upstream.
Automatic title
After the first exchange, if the answer is not empty, title_job_queued is true: a background job writes a short title from the first message, on machines in the zones of ITS INCOM AI. Read the conversation again a few seconds later to get it.
A first message shorter than 12 characters becomes the title as it is. If the job cannot generate a title, it uses the beginning of the first message.
Stopping an answer
There is no stop endpoint in this API: to stop an answer, close the connection. The stop button of the web chat works only in the web chat.
A complete client (TypeScript)
Node 18 or later, run as an ES module, with the account's credentials in EMAIL and PASSWORD.
const BASE = 'https://my.ai.itsincom.org/api/v1';
// 1) Sign in
const login = await fetch(`${BASE}/auth/login`, {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({ email: process.env.EMAIL, password: process.env.PASSWORD }),
}).then(r => r.json());
const headers = {
'Authorization': `Bearer ${login.access_token}`,
'Content-Type': 'application/json',
};
// 2) Create a conversation
const conv = await fetch(`${BASE}/chat/sessions`, {
method: 'POST',
headers,
body: JSON.stringify({}),
}).then(r => r.json());
// 3) Send a message and read the stream
const resp = await fetch(`${BASE}/chat/sessions/${conv.id}/messages`, {
method: 'POST',
headers,
body: JSON.stringify({ content: 'What is a notice period?', model: 'gemma-4-26b' }),
});
if (!resp.ok) throw new Error(`HTTP ${resp.status}: ${await resp.text()}`);
const reader = resp.body!.getReader();
const decoder = new TextDecoder();
let buf = '';
let answer = '';
while (true) {
const { done, value } = await reader.read();
if (done) break;
buf += decoder.decode(value, { stream: true });
// Events are separated by an empty line
let end;
while ((end = buf.indexOf('\n\n')) !== -1) {
const frame = buf.slice(0, end);
buf = buf.slice(end + 2);
if (!frame.startsWith('data: ')) continue;
const ev = JSON.parse(frame.slice(6));
if (ev.type === 'delta') answer += ev.text;
if (ev.type === 'error') console.error(ev.code, ev.message);
if (ev.type === 'done' && ev.title_job_queued) {
// The title is written in the background: read it a little later
setTimeout(async () => {
const fresh = await fetch(`${BASE}/chat/sessions/${conv.id}`, { headers }).then(r => r.json());
console.log('title:', fresh.title);
}, 3000);
}
}
}
console.log(answer);
Errors
| Status | When |
|---|---|
401 |
The token is missing, wrong or expired, or the account is disabled. Code unauthorized |
403 |
The email of the account is not verified. Code email_unverified |
404 |
The conversation does not exist, or it is not yours |
422 |
Validation failed: for example content empty or longer than 8000 characters, or model missing. The body lists the fields in errors |
429 |
Too many requests: see below. Retry-After says how long to wait |
Limits
| What | Per minute |
|---|---|
POST /chat/sessions/{id}/messages |
30 |
POST /auth/login |
10 |
POST /auth/refresh |
60 |
They are counted per client address, on one counter shared with the other limited calls of this API — knowledge bases included. People behind the same address share it.
Compared with the developer API
| Developer API | Chat sessions | |
|---|---|---|
| Address | https://api.ai.itsincom.org/v1 |
https://my.ai.itsincom.org/api/v1 |
| Authentication | API key, sk-… |
Token of a signed-in user |
| History | You send it every time | Kept on the server |
| Format | OpenAI | Its own, simpler |
| Zones | Those allowed for the key | Those offered by ITS INCOM AI |
| Stopping an answer | Close the connection | Close the connection |
| Use it for | Your backend, the OpenAI SDKs | A chat client on your own account |