ITS INCOM AI ITS INCOM AI docs

Concepts

Models

How we choose what to serve, what open weights mean here, and which model fits which job.

Last updated: 2026-10-03

ITS INCOM AI serves open-weight models on its own hardware. No closed third-party API in the middle and no opaque costs: you know which model answers, and you can check where it comes from and under which licence.

The catalogue

The up-to-date list lives in the Models catalog, read straight from the database. For every model it shows the zones where it runs, each with its country. A request never leaves the zones allowed for your key: see Zones below.

The table below is the same catalogue, with comments.

Model model_id What it is for Origin
Gemma 4 26B gemma-4-26b General conversation. Accepts images as input. Good at returning structured data from free text: extracting fields from documents 🇺🇸 Google DeepMind
Apertus 70B Instruct apertus-70b-instruct Conversation and answers about your documents. Italian, German, French, English, Romansh 🇨🇭 Swiss AI Initiative — EPFL, ETH Zurich, CSCS
Qwen 2.5 32B qwen2.5:32b The most capable mid-size model: analysis that needs several reasoning steps 🇨🇳 Alibaba
Qwen 2.5 14B qwen2.5:14b Quality and speed balanced. The choice for volume: classification, summaries, drafts 🇨🇳 Alibaba
Qwen 2.5 1.5B qwen2.5:1.5b Simple, repetitive tasks where response time matters 🇨🇳 Alibaba
BGE-M3 bge-m3 Vectors for semantic search, 1024 dimensions, multilingual 🇨🇳 BAAI
BGE-reranker-v2-m3 bge-reranker-v2-m3 Reordering search results, as a second pass 🇨🇳 BAAI

For audio: whisper-1 and whisper-1-hd for transcription, the voices paola, riccardo, amy, ryan, thorsten, siwis for speech. They are documented in Audio API and, as in the OpenAI API, they do not appear in /v1/models.

Zones

A zone is a group of machines in one country, with a declared owner: the platform's own machines, or a partner's. Every model in the catalogue lists its zones with their country, and every API key has the zones it may use — by default all the zones offered by ITS INCOM AI.

  • A request is sent only to machines in the zones allowed for the key, so it never leaves the countries of those zones.
  • If no machine is available there, the answer is 503 with the zone named in the message. The request is never moved to another zone to get an answer.
  • Chat completion responses say where they were processed, in the X-Zona header, and your usage records keep the zone and the site of every request. In a stream the headers leave before a machine is chosen: if your key allows more than one zone, the header lists the possible ones, and the usage record has the exact one.

Zones apply today to text generation: /v1/chat/completions, /v1/completions and /v1/responses. Embeddings, rerank and audio are served by dedicated machines that do not go through zones yet; if it matters for your case, write to segreteria@itsincom.it and we tell you where they run for ITS INCOM AI.

Images as input

Only gemma-4-26b, up to 2 images per request. The vision badge in the catalogue marks the models that accept them; sending an image to a model that does not accept it returns a 400 error listing the models that do — not a generic backend error.

Formats: PNG, JPEG, WebP, GIF. At most 8 MB and 40 megapixels per image. The request format is in Chat completions.

How we choose what to serve

Three criteria:

  1. Licence: usable commercially, without surprises.
  2. Quality: at the top of its size class when it is added.
  3. Sustainability: it must run efficiently on the capacity we have. We do not add a model if it makes the service worse for all the others.

We exclude models that contact the outside world during inference, models whose access is bound in ways we cannot verify, and models whose ecosystem collects user data.

Quantization

Large models may be served quantized: a more compact numerical representation that reduces the memory needed and increases serving capacity, at the cost of a small, measured loss of quality. The precision of each model is stated in the catalogue.

It is invisible to whoever calls the API: the model_id stays the same however the model is loaded, and if we change the quantization your code does not change.

Which one to pick

If you need to Use
Chat, general case gemma-4-26b
Read an image or a scan gemma-4-26b
Extract structured fields from a text gemma-4-26b
Analyse a complex document qwen2.5:32b
Classify or summarise at volume qwen2.5:14b
Answer trivial tasks in milliseconds qwen2.5:1.5b
Search your texts by meaning bge-m3 via /v1/embeddings
Improve the order of search results bge-reranker-v2-m3 via /v1/rerank

Every request names a model_id. Which machine runs it, inside your zones, is up to us.

Tiers and models

The tier does not decide which model you can use: it decides which machines may serve you, the rate limit of your key and a multiplier on the price, which depends on the model. It does not change the order in a machine's queue. Details in Tiers and in the live price list.

If you need a model we do not have

Write to segreteria@itsincom.it. Adding an open-weight model is a mechanical operation, and we do it on request for business customers. A model trained on your own terminology is a project of its own: tell us about it.