API reference
Chat completions
POST /v1/chat/completions: the OpenAI request and response, with the tier header, the zone header and images.
Last updated: 2026-10-03
POST https://api.ai.itsincom.org/v1/chat/completions
The OpenAI chat completions request and response. Two additions: the X-Siati-Tier header chooses the tier of the request, and the X-Zona header in the response says in which zone it was processed.
Request
curl -i https://api.ai.itsincom.org/v1/chat/completions \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gemma-4-26b",
"messages": [
{"role": "system", "content": "Answer in three sentences."},
{"role": "user", "content": "What is data sovereignty?"}
],
"temperature": 0.7,
"max_completion_tokens": 400
}'
Required: model, and messages with at least one message. The roles are system, user, assistant and tool. Everything else is optional.
Every parameter we accept is in Every parameter: honoured or refused, with what it does. A parameter that is not on that list is refused with 400, and the error names it.
Python
import os
from openai import OpenAI
client = OpenAI(
base_url="https://api.ai.itsincom.org/v1",
api_key=os.environ["API_KEY"],
)
resp = client.chat.completions.create(
model="gemma-4-26b",
messages=[
{"role": "system", "content": "Answer in three sentences."},
{"role": "user", "content": "What is data sovereignty?"},
],
max_completion_tokens=400,
extra_headers={"X-Siati-Tier": "fast"}, # optional
)
print(resp.choices[0].message.content)
JavaScript
import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://api.ai.itsincom.org/v1",
apiKey: process.env.API_KEY,
});
const resp = await client.chat.completions.create(
{
model: "gemma-4-26b",
messages: [{ role: "user", content: "What is data sovereignty?" }],
},
{ headers: { "X-Siati-Tier": "fast" } }, // optional
);
console.log(resp.choices[0].message.content);
Request headers
| Header | |
|---|---|
Authorization: Bearer sk-… |
Required. See Authentication. |
Content-Type: application/json |
Required with curl; the SDKs send it. Without it the body is not read as JSON and you get 422. |
X-Siati-Tier |
slow, medium, fast or ludicrous. Without it, the default tier of the key. An unknown value returns 400. What a tier changes is in Tiers. |
Response
{
"id": "req_288648c0f938503f1236f4d1",
"object": "chat.completion",
"created": 1791028778,
"model": "gemma-4-26b",
"choices": [
{
"index": 0,
"message": { "role": "assistant", "content": "…" },
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 31,
"completion_tokens": 67,
"total_tokens": 98
}
}
Where it differs from OpenAI:
idis the identifier under which we record the request,req_…, not achatcmpl-….modelis themodel_idyou asked for.- There is no
system_fingerprint.
When the model calls a function, message.tool_calls carries the calls and content is null if there is no text: see Guaranteed JSON and function calling. With n greater than 1, choices carries every answer; with logprobs, each choice carries its token probabilities.
Response headers
| Header | |
|---|---|
X-Zona |
The zone where the request was processed. In a stream, if your key allows more than one zone, the list of possible ones: see Streaming. |
X-Request-Id |
An identifier of the HTTP request, on every response. It is not always equal to id: when you write to us, send both. |
X-RateLimit-Limit, X-RateLimit-Remaining |
The limit of your address on this endpoint, not the one of your key: see Rate limits. |
X-Siati-Region |
The region ITS INCOM AI declares, the same on every response. Not every brand sends it; where the request was processed is X-Zona. |
Zones
The request goes only to machines in the zones of your key. If a machine fails before answering, the request moves to the next one, still inside your zones. If none is available there, you get 503 with type zone_unavailable and the zone named in the message: the request is never moved to another zone. More in Zones.
Streaming
Add "stream": true to receive the answer while it is written. The format, what the stream does not carry and how errors arrive are in Streaming.
Images
The models with the vision badge in the catalogue accept images: today gemma-4-26b. The content of a message becomes a list of parts, text and image_url:
curl https://api.ai.itsincom.org/v1/chat/completions \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gemma-4-26b",
"messages": [{
"role": "user",
"content": [
{"type": "text", "text": "What document is this? Extract the date and the amount."},
{"type": "image_url", "image_url": {"url": "data:image/jpeg;base64,…"}}
]
}]
}'
In Python, the data URI of a file:
import base64
with open("invoice.jpg", "rb") as f:
data_uri = "data:image/jpeg;base64," + base64.b64encode(f.read()).decode()
image_url.url takes two forms:
- a base64 data URI, which we recommend: there is nothing to download;
- a public
httporhttpsURL: we download the image ourselves and pass it to the model. It must be reachable from the internet. Private and local addresses are refused with400, and redirects are not followed, so give the final address. If remote URLs are turned off on ITS INCOM AI, the400says so: send a data URI.
| Limit | Value |
|---|---|
| Formats | PNG, JPEG, WebP, GIF |
| Size | 8 MB per image |
| Pixels | 40 megapixels per image |
| Images per request | Depends on the model: gemma-4-26b accepts 2. Above the limit, the 400 states it |
The format is read from the bytes, not from the extension or the declared type: valid base64 that does not contain an image returns 400. A part of another type than text or image_url returns 400 too, and so does an image sent to a model that does not accept images; that error lists the models that do.
Cost: the image becomes tokens, counted in prompt_tokens, so you pay for it at the input price of the model.
Errors
The most common ones. All of them, with what to do, are in Errors.
| HTTP | When |
|---|---|
400 |
A parameter refused or not recognised (named in param), an image we cannot accept, an unknown X-Siati-Tier |
401 |
A problem with the key: see Authentication |
404 |
The model is not in the catalogue (type: model_not_found) |
422 |
model or messages missing, or a value out of range; errors names the field |
429 |
A rate limit: see Rate limits |
502 |
The machines that could serve the request did not answer |
503 |
No machine available in your zones (zone_unavailable) or for that model (model_unavailable), or a model activated on request (model_activation_required) |
Price
Each model has a price for input tokens and one for output tokens, and the tier applies a multiplier. The values are in Pricing.