ITS INCOM AI ITS INCOM AI docs

Concepts

Tiers

What a tier changes today: which machines serve you, the rate limit of your key, and the price multiplier.

Last updated: 2026-10-03

A tier is set per request with the X-Siati-Tier header, or per key as its default. The valid values are slow, medium, fast and ludicrous.

Today a tier changes three things:

  1. Which machines may serve the request. Every machine has a weight for each tier, and the request goes only to machines with a weight above zero for your tier, inside your zones.
  2. The rate limit of your key, per minute.
  3. The price multiplier applied to the price of the model.

It does not change the order inside a machine's queue: once a request has reached a machine, it waits like the others. We write it here because the earlier version of this page promised that it did, and it did not.

Limits that apply to everyone

On top of the per-key limit of your tier there is a limit per client address, counted on one counter for all endpoints: chat, completions and responses stop at 60 requests per minute, embeddings and rerank at 120. When you exceed a limit you get 429, and the Retry-After header tells you how long to wait. Details in Rate limits.

How routing works

For a request (tier, model, zones):

  1. Only machines that serve that model, are healthy, are in production, sit in one of your zones, and have a weight above zero for your tier are considered.
  2. They are sorted by weight, then by shortest queue, then by lowest load.
  3. The request goes to the first one. If it fails before the first token, it moves to the next one, always inside your zones.

If no machine can serve your model in your zones, you get 503 naming the zone: we never move you silently to another zone, or to a different model than the one you asked for.

Setting the tier

Per key (default)

You choose it when you create the key, in the dashboard. It cannot be changed afterwards: to move a project to another tier, create a new key and revoke the old one — or send the header below.

Per request

bash
curl https://api.ai.itsincom.org/v1/chat/completions \
  -H "Authorization: Bearer $API_KEY" \
  -H "X-Siati-Tier: fast" \
  -H "Content-Type: application/json" \
  -d '{ "model": "gemma-4-26b", "messages": [{"role": "user", "content": "Hello"}] }'

An unknown value returns 400.

If you need a guarantee, not a preference

Even the highest tier shares machines with other requests. For predictable performance we reserve capacity for your exclusive use, for a monthly fee: your requests queue behind nobody. It is not on the price list because the price is built on the concrete case — write to segreteria@itsincom.it.

Coming from OpenAI

At OpenAI tiers are access levels tied to spending, not a choice per request. The equivalence is rough:

OpenAI Here
Tier 1-2 (default) slow / medium
Tier 3-4 fast
Tier 5 (enterprise) ludicrous
Priority processing Dedicated capacity