API reference
Every parameter: honoured or refused
The complete list of what we accept on /v1/chat/completions, what reaches the engine and what we refuse with a 400. No parameter is accepted and ignored.
Last updated: 2026-10-03
On this API a parameter can end up in two ways, and there is no third:
- it reaches the engine and does what its name says;
- it is refused with a 400 that names it and says what to use instead.
There is no "accepted and ignored" case, and it is the most useful rule on this page: if the request comes back 200, what you wrote has been applied.
Why we write it so explicitly
Because until 18 August 2026 it was not true. The gateway accepted seven
parameters and dropped them without saying so: max_completion_tokens, stop,
n, seed, logprobs, top_logprobs, parallel_tool_calls. It also
accepted an eighth, top_p, which was even in the table of this documentation.
The response came back 200 and looked right.
The worst case was n. The gateway forwarded it, the engine generated the
requested answers and billed all of them, and we returned the first one and
threw away the others: you paid for tokens you never saw. Worse than ignoring
the parameter.
The most expensive case was max_completion_tokens, the new name OpenAI
recommends for the token cap. Whoever used it had no cap at all, and found out
on the invoice.
It was reported by the team building a Swiss business application on top of this API, with the sentence that decided how the fix is made: "an error is found during development, an ignored cap is found on the invoice".
Honoured
Tested on the engine one by one on 18 August 2026, then through the public gateway.
| parameter | what it does | notes |
|---|---|---|
model |
which model | see the catalogue |
messages |
the conversation | roles system, user, assistant, tool |
temperature |
0–2 | lower is more predictable |
top_p |
0–1 | nucleus sampling |
max_tokens |
cap on generated tokens | up to 8192 |
max_completion_tokens |
the same cap, new name | use this one |
stop |
where to stop | string or list; we convert the string |
n |
how many answers | 1–8; you pay for all of them, and now you see all of them |
seed |
repeatability | with everything else equal, the same answer |
presence_penalty |
−2…2 | |
frequency_penalty |
−2…2 | |
logit_bias |
per-token bias | map token id → −100…100 |
logprobs |
token probabilities | appear in choices[].logprobs |
top_logprobs |
0–20 alternatives per token | requires logprobs: true |
parallel_tool_calls |
several calls at once, yes or no | with false they arrive one at a time |
tools / tool_choice |
function calling | see JSON and functions |
response_format |
shape of the answer | also with stream: true |
stream / stream_options |
answer in chunks | see streaming |
user |
your own reference | see below |
service_tier |
auto or default |
we have a single level; for speed use X-Siati-Tier |
The token cap, and its two names
max_tokens and max_completion_tokens are the same cap. If you send one, that
one applies. If you send both with the same number, fine. If you send both with
different numbers we answer 400: two different caps in the same request mean
that whoever wrote it believes in one of them, and guessing which is how a
customer ends up paying for our guess.
user, and where it ends up
The user field ends up in the detail of your usage, next to every request. It
is for you, to separate your users, departments or customers within a single
invoice. It is a string you give us and we give back as it is: we do not link
it to anything and do not use it for anything else.
If you are looking for an X-Siati-User header, it is gone: it was documented
and never implemented. Use user, which is the OpenAI name and works with the
libraries you already have.
logprobs and n are not supported by every model
They work on models served by vLLM — today gemma-4-26b and
apertus-70b-instruct. On the others you get a 400 naming the parameter
and saying on which models it works, not a 502 and not silence.
Refused, with the reason
| parameter | why |
|---|---|
metadata |
there is no stored completion to attach labels to. Use user. |
store |
completions are not stored for you to retrieve later. What the request log keeps, and for how long, is in Sovereignty. |
functions |
outdated form: use tools. |
function_call |
outdated form: use tool_choice. |
modalities, audio |
for voice there are dedicated endpoints. |
prediction |
the engine does not do it. |
web_search_options |
our models do not go out to the internet, by design. |
Any other parameter
A name we do not recognise gives 400 and is named. That includes typos, which
is the most useful part: tempreature used to go through, the temperature
stayed at the default and there was no way to notice it from the answer.
The API answers in Italian:
{
"error": {
"message": "Parametri non riconosciuti: «tempreature». Li rifiutiamo invece di ignorarli, perché un parametro accettato e scartato non si vede dalla risposta. Accettiamo: …",
"type": "invalid_request_error",
"param": "tempreature"
}
}
Transcription: timestamp_granularities
The same rule applies. timestamp_granularities: ["word"] returns per-word
timings in words, with start, end and probability of each. Until 18 August the
field was accepted and not forwarded, so only segments came back.
Timings exist only inside response_format: "verbose_json": the other two forms
have nowhere to put them, and asking for them without verbose_json gives 400
instead of an answer without timings, which could not be told apart from audio
in which the words could not be heard.
curl https://api.ai.itsincom.org/v1/audio/transcriptions \
-H "Authorization: Bearer $API_KEY" \
-F file=@note.wav \
-F model=whisper-1 \
-F response_format=verbose_json \
-F "timestamp_granularities[]=word"
How we avoid repeating the mistake
The cause was not someone's distraction: request validation returned only the fields it listed, and everything else vanished without an error and without a log line. A defect like that is not found by rereading the code, because there is nothing to see.
Now the list of parameters lives in one place, and an automated test fails if a field is validated without ending up either at the engine or among the refused. We checked it by deliberately adding a forgotten parameter: the test named it.
It does not protect us from not supporting something you need. It protects us from pretending to support it, which is what costs you the most.