API reference
Rate limits
The two limits every /v1 request passes, how each one counts, what its 429 looks like, and how to retry.
Last updated: 2026-10-03
A request to /v1 passes up to two limits, both counted per minute. They are independent: the first one reached answers 429.
| Counted per | Value | Applies to | |
|---|---|---|---|
| Per key | API key | set by the tier of the request | every /v1 endpoint, /v1/models included |
| Per address | client IP address | fixed per endpoint | every /v1 endpoint except /v1/models |
Neither of them promises speed. A request that passes both may still wait in a machine's queue, and the tier does not change your place in it: see Tiers.
Per key: set by the tier
Each tier has a number of requests per minute, kept in the database. The values are in the tiers table of Pricing, when that page shows one; the 429 message states the limit as well.
How it counts:
- One counter per key, for all
/v1endpoints together. Listing the models counts too. - The minute is the clock minute. The counter starts again at the beginning of each minute: it is not a sliding window.
- Each request is compared with the limit of its own tier. If you mix tiers on one key with
X-Siati-Tier, they share the counter.
When you exceed it:
HTTP/1.1 429 Too Many Requests
Retry-After: 23
Content-Type: application/json
{"error": {"message": "rate limit exceeded (…/min)", "type": "rate_limit_exceeded"}}
Retry-After is the number of seconds until the next minute begins. No response tells you how many requests your key has left.
Per address: fixed per endpoint
| Endpoint | Requests per minute |
|---|---|
POST /v1/chat/completions, /v1/completions, /v1/responses |
60 |
POST /v1/embeddings, /v1/rerank |
120 |
POST /v1/audio/transcriptions |
30 |
POST /v1/audio/speech |
120 |
GET /v1/models |
not limited |
How it counts:
- Per client address, before the key is checked. Requests with a wrong key count too, and clients behind the same address share the limit.
- One counter per address for all these endpoints together, which each endpoint compares with its own limit. After 30 chat requests in a minute, transcriptions answer
429until the minute is over, while embeddings still have 90 to go. - The minute starts with the first request counted, not on the clock.
Responses of these endpoints carry two headers about this limit, for the endpoint you called. They describe the limit of your address, not the one of your key:
X-RateLimit-Limit: 60
X-RateLimit-Remaining: 59
When you exceed it, the body is not the error object of the rest of the API:
HTTP/1.1 429 Too Many Requests
Retry-After: 60
X-RateLimit-Limit: 60
X-RateLimit-Remaining: 0
X-RateLimit-Reset: 1791028838
Content-Type: application/json
{"message": "Too Many Attempts."}
Retry-After is the number of seconds until the window ends; X-RateLimit-Reset is the same moment as a Unix timestamp.
Telling them apart
| Per key | Per address | |
|---|---|---|
| Body | {"error": {"type": "rate_limit_exceeded", …}} |
{"message": "Too Many Attempts."} |
Retry-After |
seconds until the next clock minute | seconds until the window ends |
X-RateLimit-Reset |
absent | present |
In both cases, wait Retry-After seconds before you try again.
Retrying
The OpenAI SDKs already retry a failed request a few times on their own; you can change the number of attempts with max_retries in Python, maxRetries in Node.
Without an SDK, read Retry-After and add some randomness, so that many clients do not all come back in the same second:
import os
import random
import time
import requests
URL = "https://api.ai.itsincom.org/v1/chat/completions"
HEADERS = {"Authorization": f"Bearer {os.environ['API_KEY']}"}
def post_with_retry(body, attempts=5):
for attempt in range(attempts):
r = requests.post(URL, headers=HEADERS, json=body, timeout=120)
if r.status_code not in (429, 502, 503):
return r
if r.status_code == 503 and "model_activation_required" in r.text:
return r # retrying will not change it
wait = float(r.headers.get("Retry-After", 2 ** attempt))
time.sleep(wait + random.uniform(0, 1))
return r
r = post_with_retry({
"model": "gemma-4-26b",
"messages": [{"role": "user", "content": "Hello"}],
})
print(r.status_code, r.json())
Which errors are worth retrying is in Errors.
The app API
The app API (https://my.ai.itsincom.org/api/v1, see Authentication) has only limits per address, counted in the same way, one counter shared by these endpoints:
| Endpoint | Requests per minute |
|---|---|
POST /auth/login |
10 |
POST /auth/refresh |
60 |
POST /chat/sessions/{id}/messages |
30 |
POST /rag/kb |
30 |
POST /rag/kb/{slug}/docs (upload) |
10 |
POST /rag/kb/{slug}/chat |
30 |
POST /audio/transcribe |
30 |
POST /audio/synthesize |
60 |
If you need more
The per-key limit changes with the tier: compare the values in Pricing. For capacity reserved to you, with nobody ahead of you in the queue, write to segreteria@itsincom.it.