API reference
What we change in responses
The only two things the gateway changes in what the model produces, and why. Declared, not hidden.
Last updated: 2026-10-03
The gateway changes two things in what the model produces. Nothing else. We write them down here because a service whose point is saying exactly what it does cannot rewrite answers silently.
1. The model's internal delimiters
After a tool call, some models leave in the text the markers they use to separate their internal channels:
<|channel>thought
<channel|>OK. I created the task for the customer Rossi.
Those symbols are protocol, not content: anyone showing that field on a screen
would see them. We remove them, together with the word thought when it stays
alone at the start of a line.
The list of delimiters is closed: <|channel>, <channel|>,
<|start|>, <|end|>, <|message|>, <|im_start|>, <|im_end|> and the
role markers. A generic rule on <|…|> could eat legitimate text — for example
an answer that talks about syntax — and we do not use one.
2. The word "null" instead of the null value
In the arguments of a tool call, on a field that allows null, the model
sometimes writes the four-letter string. A client that checks !== null saves
it, and the database ends up with a record whose description literally says
"null".
We turn into null only these values, and only when the field contains
them entirely, after trimming spaces:
null none nil nessuno n/a undefined (empty string)
The comparison is case-insensitive. A sentence that contains the word is not
touched: "the field was null and must be fixed" stays as it is.
It applies to the arguments of tool calls, where the field has a declared type. We do not touch the free text of the answer.
3. Silence on input
This is not a change to the answer, but it belongs to the same family and it should be said.
On audio without sound, /v1/audio/transcriptions returns empty text
without asking the model. The reason: on digital silence Whisper answers
"Grazie." or "Thank you." with no_speech_prob at 0.0000, that is with full
confidence. On a dictation path an accidental tap on the microphone would turn
into a command.
The check measures the energy of the waveform, not the model's judgement, and triggers below about -50 dBFS — below are digital silence and a closed microphone; above is any speech, even whispered into a distant microphone. It applies only to 16-bit WAV files: compressed formats always go through, because we cannot measure them without decoding and when in doubt, we transcribe. Good audio thrown away is much worse than a hallucination.
What we do not touch
The text of the answer, the numbers, the language, the punctuation, the order of JSON keys. If the model gets a total wrong, that total reaches you as it was written: add up the lines and compare, as you should with any provider.