Live data
Pricing
How many credits each model uses, read live from the catalogue.
Per model
Every request uses credits from your account, per model, with input and output counted separately: generating tokens costs far more than reading them. Credits per million tokens. Your credits are loaded by your administrator.
| Model | Input (credits / 1M tok) | Output (credits / 1M tok) |
|---|---|---|
apertus-70b-instructApertus 70B |
1.0000 |
3.0000
|
gemma-4-26bGemma 4 26B |
0.3000 |
0.6000
|
How a request is counted
For a request with P prompt tokens and C completion tokens, on a model with input price in and output price out per million tokens, at tier T:
cost = (P × in + C × out) / 1_000_000 × multiplier[T]
Audio has no published price yet: transcription is not billed today, and its price comes back here when the price list, the configuration and the invoice agree.
Tiers
The tier (X-Siati-Tier: slow|medium|fast|ludicrous, or the key's default) decides which machines may serve the request, the rate limit of your key, and a multiplier on the model's price. It does not change the order in a machine's queue. More in Tiers.
| Tier | Rate limit | Price multiplier | Description |
|---|---|---|---|
| slow | 30 req/min | 1.00× |
— |
| medium | 60 req/min | 1.00× |
— |
| fast | 120 req/min | 1.00× |
— |
| ludicrous | 200 req/min | 1.00× |
— |