ITS INCOM AI ITS INCOM AI docs

API reference

What we change in responses

The only two things the gateway changes in what the model produces, and why. Declared, not hidden.

Last updated: 2026-10-03

The gateway changes two things in what the model produces. Nothing else. We write them down here because a service whose point is saying exactly what it does cannot rewrite answers silently.

1. The model's internal delimiters

After a tool call, some models leave in the text the markers they use to separate their internal channels:

text
<|channel>thought
<channel|>OK. I created the task for the customer Rossi.

Those symbols are protocol, not content: anyone showing that field on a screen would see them. We remove them, together with the word thought when it stays alone at the start of a line.

The list of delimiters is closed: <|channel>, <channel|>, <|start|>, <|end|>, <|message|>, <|im_start|>, <|im_end|> and the role markers. A generic rule on <|…|> could eat legitimate text — for example an answer that talks about syntax — and we do not use one.

2. The word "null" instead of the null value

In the arguments of a tool call, on a field that allows null, the model sometimes writes the four-letter string. A client that checks !== null saves it, and the database ends up with a record whose description literally says "null".

We turn into null only these values, and only when the field contains them entirely, after trimming spaces:

text
null    none    nil    nessuno    n/a    undefined    (empty string)

The comparison is case-insensitive. A sentence that contains the word is not touched: "the field was null and must be fixed" stays as it is.

It applies to the arguments of tool calls, where the field has a declared type. We do not touch the free text of the answer.

3. Silence on input

This is not a change to the answer, but it belongs to the same family and it should be said.

On audio without sound, /v1/audio/transcriptions returns empty text without asking the model. The reason: on digital silence Whisper answers "Grazie." or "Thank you." with no_speech_prob at 0.0000, that is with full confidence. On a dictation path an accidental tap on the microphone would turn into a command.

The check measures the energy of the waveform, not the model's judgement, and triggers below about -50 dBFS — below are digital silence and a closed microphone; above is any speech, even whispered into a distant microphone. It applies only to 16-bit WAV files: compressed formats always go through, because we cannot measure them without decoding and when in doubt, we transcribe. Good audio thrown away is much worse than a hallucination.

What we do not touch

The text of the answer, the numbers, the language, the punctuation, the order of JSON keys. If the model gets a total wrong, that total reaches you as it was written: add up the lines and compare, as you should with any provider.