ITS INCOM AI ITS INCOM AI docs

API reference

Every parameter: honoured or refused

The complete list of what we accept on /v1/chat/completions, what reaches the engine and what we refuse with a 400. No parameter is accepted and ignored.

Last updated: 2026-10-03

On this API a parameter can end up in two ways, and there is no third:

  • it reaches the engine and does what its name says;
  • it is refused with a 400 that names it and says what to use instead.

There is no "accepted and ignored" case, and it is the most useful rule on this page: if the request comes back 200, what you wrote has been applied.

Why we write it so explicitly

Because until 18 August 2026 it was not true. The gateway accepted seven parameters and dropped them without saying so: max_completion_tokens, stop, n, seed, logprobs, top_logprobs, parallel_tool_calls. It also accepted an eighth, top_p, which was even in the table of this documentation. The response came back 200 and looked right.

The worst case was n. The gateway forwarded it, the engine generated the requested answers and billed all of them, and we returned the first one and threw away the others: you paid for tokens you never saw. Worse than ignoring the parameter.

The most expensive case was max_completion_tokens, the new name OpenAI recommends for the token cap. Whoever used it had no cap at all, and found out on the invoice.

It was reported by the team building a Swiss business application on top of this API, with the sentence that decided how the fix is made: "an error is found during development, an ignored cap is found on the invoice".


Honoured

Tested on the engine one by one on 18 August 2026, then through the public gateway.

parameter what it does notes
model which model see the catalogue
messages the conversation roles system, user, assistant, tool
temperature 0–2 lower is more predictable
top_p 0–1 nucleus sampling
max_tokens cap on generated tokens up to 8192
max_completion_tokens the same cap, new name use this one
stop where to stop string or list; we convert the string
n how many answers 1–8; you pay for all of them, and now you see all of them
seed repeatability with everything else equal, the same answer
presence_penalty −2…2
frequency_penalty −2…2
logit_bias per-token bias map token id → −100…100
logprobs token probabilities appear in choices[].logprobs
top_logprobs 0–20 alternatives per token requires logprobs: true
parallel_tool_calls several calls at once, yes or no with false they arrive one at a time
tools / tool_choice function calling see JSON and functions
response_format shape of the answer also with stream: true
stream / stream_options answer in chunks see streaming
user your own reference see below
service_tier auto or default we have a single level; for speed use X-Siati-Tier

The token cap, and its two names

max_tokens and max_completion_tokens are the same cap. If you send one, that one applies. If you send both with the same number, fine. If you send both with different numbers we answer 400: two different caps in the same request mean that whoever wrote it believes in one of them, and guessing which is how a customer ends up paying for our guess.

user, and where it ends up

The user field ends up in the detail of your usage, next to every request. It is for you, to separate your users, departments or customers within a single invoice. It is a string you give us and we give back as it is: we do not link it to anything and do not use it for anything else.

If you are looking for an X-Siati-User header, it is gone: it was documented and never implemented. Use user, which is the OpenAI name and works with the libraries you already have.

logprobs and n are not supported by every model

They work on models served by vLLM — today gemma-4-26b and apertus-70b-instruct. On the others you get a 400 naming the parameter and saying on which models it works, not a 502 and not silence.


Refused, with the reason

parameter why
metadata there is no stored completion to attach labels to. Use user.
store completions are not stored for you to retrieve later. What the request log keeps, and for how long, is in Sovereignty.
functions outdated form: use tools.
function_call outdated form: use tool_choice.
modalities, audio for voice there are dedicated endpoints.
prediction the engine does not do it.
web_search_options our models do not go out to the internet, by design.

Any other parameter

A name we do not recognise gives 400 and is named. That includes typos, which is the most useful part: tempreature used to go through, the temperature stayed at the default and there was no way to notice it from the answer.

The API answers in Italian:

json
{
  "error": {
    "message": "Parametri non riconosciuti: «tempreature». Li rifiutiamo invece di ignorarli, perché un parametro accettato e scartato non si vede dalla risposta. Accettiamo: …",
    "type": "invalid_request_error",
    "param": "tempreature"
  }
}

Transcription: timestamp_granularities

The same rule applies. timestamp_granularities: ["word"] returns per-word timings in words, with start, end and probability of each. Until 18 August the field was accepted and not forwarded, so only segments came back.

Timings exist only inside response_format: "verbose_json": the other two forms have nowhere to put them, and asking for them without verbose_json gives 400 instead of an answer without timings, which could not be told apart from audio in which the words could not be heard.

bash
curl https://api.ai.itsincom.org/v1/audio/transcriptions \
  -H "Authorization: Bearer $API_KEY" \
  -F file=@note.wav \
  -F model=whisper-1 \
  -F response_format=verbose_json \
  -F "timestamp_granularities[]=word"

How we avoid repeating the mistake

The cause was not someone's distraction: request validation returned only the fields it listed, and everything else vanished without an error and without a log line. A defect like that is not found by rereading the code, because there is nothing to see.

Now the list of parameters lives in one place, and an automated test fails if a field is validated without ending up either at the engine or among the refused. We checked it by deliberately adding a forgotten parameter: the test named it.

It does not protect us from not supporting something you need. It protects us from pretending to support it, which is what costs you the most.