Concepts
Models
How we choose what to serve, what open weights mean here, and which model fits which job.
Last updated: 2026-10-03
ITS INCOM AI serves open-weight models on its own hardware. No closed third-party API in the middle and no opaque costs: you know which model answers, and you can check where it comes from and under which licence.
The catalogue
The up-to-date list lives in the Models catalog, read straight from the database. For every model it shows the zones where it runs, each with its country. A request never leaves the zones allowed for your key: see Zones below.
The table below is the same catalogue, with comments.
| Model | model_id |
What it is for | Origin |
|---|---|---|---|
| Gemma 4 26B | gemma-4-26b |
General conversation. Accepts images as input. Good at returning structured data from free text: extracting fields from documents | 🇺🇸 Google DeepMind |
| Apertus 70B Instruct | apertus-70b-instruct |
Conversation and answers about your documents. Italian, German, French, English, Romansh | 🇨🇭 Swiss AI Initiative — EPFL, ETH Zurich, CSCS |
| Qwen 2.5 32B | qwen2.5:32b |
The most capable mid-size model: analysis that needs several reasoning steps | 🇨🇳 Alibaba |
| Qwen 2.5 14B | qwen2.5:14b |
Quality and speed balanced. The choice for volume: classification, summaries, drafts | 🇨🇳 Alibaba |
| Qwen 2.5 1.5B | qwen2.5:1.5b |
Simple, repetitive tasks where response time matters | 🇨🇳 Alibaba |
| BGE-M3 | bge-m3 |
Vectors for semantic search, 1024 dimensions, multilingual | 🇨🇳 BAAI |
| BGE-reranker-v2-m3 | bge-reranker-v2-m3 |
Reordering search results, as a second pass | 🇨🇳 BAAI |
For audio: whisper-1 and whisper-1-hd for transcription, the voices paola, riccardo, amy, ryan, thorsten, siwis for speech. They are documented in Audio API and, as in the OpenAI API, they do not appear in /v1/models.
Zones
A zone is a group of machines in one country, with a declared owner: the platform's own machines, or a partner's. Every model in the catalogue lists its zones with their country, and every API key has the zones it may use — by default all the zones offered by ITS INCOM AI.
- A request is sent only to machines in the zones allowed for the key, so it never leaves the countries of those zones.
- If no machine is available there, the answer is
503with the zone named in the message. The request is never moved to another zone to get an answer. - Chat completion responses say where they were processed, in the
X-Zonaheader, and your usage records keep the zone and the site of every request. In a stream the headers leave before a machine is chosen: if your key allows more than one zone, the header lists the possible ones, and the usage record has the exact one.
Zones apply today to text generation: /v1/chat/completions, /v1/completions and /v1/responses. Embeddings, rerank and audio are served by dedicated machines that do not go through zones yet; if it matters for your case, write to segreteria@itsincom.it and we tell you where they run for ITS INCOM AI.
Images as input
Only gemma-4-26b, up to 2 images per request. The vision badge in the catalogue marks the models that accept them; sending an image to a model that does not accept it returns a 400 error listing the models that do — not a generic backend error.
Formats: PNG, JPEG, WebP, GIF. At most 8 MB and 40 megapixels per image. The request format is in Chat completions.
How we choose what to serve
Three criteria:
- Licence: usable commercially, without surprises.
- Quality: at the top of its size class when it is added.
- Sustainability: it must run efficiently on the capacity we have. We do not add a model if it makes the service worse for all the others.
We exclude models that contact the outside world during inference, models whose access is bound in ways we cannot verify, and models whose ecosystem collects user data.
Quantization
Large models may be served quantized: a more compact numerical representation that reduces the memory needed and increases serving capacity, at the cost of a small, measured loss of quality. The precision of each model is stated in the catalogue.
It is invisible to whoever calls the API: the model_id stays the same however the model is loaded, and if we change the quantization your code does not change.
Which one to pick
| If you need to | Use |
|---|---|
| Chat, general case | gemma-4-26b |
| Read an image or a scan | gemma-4-26b |
| Extract structured fields from a text | gemma-4-26b |
| Analyse a complex document | qwen2.5:32b |
| Classify or summarise at volume | qwen2.5:14b |
| Answer trivial tasks in milliseconds | qwen2.5:1.5b |
| Search your texts by meaning | bge-m3 via /v1/embeddings |
| Improve the order of search results | bge-reranker-v2-m3 via /v1/rerank |
Every request names a model_id. Which machine runs it, inside your zones, is up to us.
Tiers and models
The tier does not decide which model you can use: it decides which machines may serve you, the rate limit of your key and a multiplier on the price, which depends on the model. It does not change the order in a machine's queue. Details in Tiers and in the live price list.
If you need a model we do not have
Write to segreteria@itsincom.it. Adding an open-weight model is a mechanical operation, and we do it on request for business customers. A model trained on your own terminology is a project of its own: tell us about it.