ITS INCOM AI ITS INCOM AI docs

API reference

Rate limits

The two limits every /v1 request passes, how each one counts, what its 429 looks like, and how to retry.

Last updated: 2026-10-03

A request to /v1 passes up to two limits, both counted per minute. They are independent: the first one reached answers 429.

Counted per Value Applies to
Per key API key set by the tier of the request every /v1 endpoint, /v1/models included
Per address client IP address fixed per endpoint every /v1 endpoint except /v1/models

Neither of them promises speed. A request that passes both may still wait in a machine's queue, and the tier does not change your place in it: see Tiers.

Per key: set by the tier

Each tier has a number of requests per minute, kept in the database. The values are in the tiers table of Pricing, when that page shows one; the 429 message states the limit as well.

How it counts:

  • One counter per key, for all /v1 endpoints together. Listing the models counts too.
  • The minute is the clock minute. The counter starts again at the beginning of each minute: it is not a sliding window.
  • Each request is compared with the limit of its own tier. If you mix tiers on one key with X-Siati-Tier, they share the counter.

When you exceed it:

http
HTTP/1.1 429 Too Many Requests
Retry-After: 23
Content-Type: application/json

{"error": {"message": "rate limit exceeded (…/min)", "type": "rate_limit_exceeded"}}

Retry-After is the number of seconds until the next minute begins. No response tells you how many requests your key has left.

Per address: fixed per endpoint

Endpoint Requests per minute
POST /v1/chat/completions, /v1/completions, /v1/responses 60
POST /v1/embeddings, /v1/rerank 120
POST /v1/audio/transcriptions 30
POST /v1/audio/speech 120
GET /v1/models not limited

How it counts:

  • Per client address, before the key is checked. Requests with a wrong key count too, and clients behind the same address share the limit.
  • One counter per address for all these endpoints together, which each endpoint compares with its own limit. After 30 chat requests in a minute, transcriptions answer 429 until the minute is over, while embeddings still have 90 to go.
  • The minute starts with the first request counted, not on the clock.

Responses of these endpoints carry two headers about this limit, for the endpoint you called. They describe the limit of your address, not the one of your key:

http
X-RateLimit-Limit: 60
X-RateLimit-Remaining: 59

When you exceed it, the body is not the error object of the rest of the API:

http
HTTP/1.1 429 Too Many Requests
Retry-After: 60
X-RateLimit-Limit: 60
X-RateLimit-Remaining: 0
X-RateLimit-Reset: 1791028838
Content-Type: application/json

{"message": "Too Many Attempts."}

Retry-After is the number of seconds until the window ends; X-RateLimit-Reset is the same moment as a Unix timestamp.

Telling them apart

Per key Per address
Body {"error": {"type": "rate_limit_exceeded", …}} {"message": "Too Many Attempts."}
Retry-After seconds until the next clock minute seconds until the window ends
X-RateLimit-Reset absent present

In both cases, wait Retry-After seconds before you try again.

Retrying

The OpenAI SDKs already retry a failed request a few times on their own; you can change the number of attempts with max_retries in Python, maxRetries in Node.

Without an SDK, read Retry-After and add some randomness, so that many clients do not all come back in the same second:

python
import os
import random
import time

import requests

URL = "https://api.ai.itsincom.org/v1/chat/completions"
HEADERS = {"Authorization": f"Bearer {os.environ['API_KEY']}"}


def post_with_retry(body, attempts=5):
    for attempt in range(attempts):
        r = requests.post(URL, headers=HEADERS, json=body, timeout=120)
        if r.status_code not in (429, 502, 503):
            return r
        if r.status_code == 503 and "model_activation_required" in r.text:
            return r  # retrying will not change it
        wait = float(r.headers.get("Retry-After", 2 ** attempt))
        time.sleep(wait + random.uniform(0, 1))
    return r


r = post_with_retry({
    "model": "gemma-4-26b",
    "messages": [{"role": "user", "content": "Hello"}],
})
print(r.status_code, r.json())

Which errors are worth retrying is in Errors.

The app API

The app API (https://my.ai.itsincom.org/api/v1, see Authentication) has only limits per address, counted in the same way, one counter shared by these endpoints:

Endpoint Requests per minute
POST /auth/login 10
POST /auth/refresh 60
POST /chat/sessions/{id}/messages 30
POST /rag/kb 30
POST /rag/kb/{slug}/docs (upload) 10
POST /rag/kb/{slug}/chat 30
POST /audio/transcribe 30
POST /audio/synthesize 60

If you need more

The per-key limit changes with the tier: compare the values in Pricing. For capacity reserved to you, with nobody ahead of you in the queue, write to segreteria@itsincom.it.