Changelog
Changelog
Releases and notable platform changes, newest first.
Last updated: 2026-10-03
What changed, when, and why. Newest first.
Between 24 May and 22 September 2026 several changes were not logged here, among them audio, embeddings and rerank, images as input and the Responses API. The pages they affect describe how they work today.
2026-10-03 — Zones, fallback between machines, brands, documentation in English
Zones. A zone is a group of machines in one country, with a declared owner. Every API key has the zones it may use: by default, all the zones its brand offers.
- A text generation request goes only to machines in the zones of its key. If none is available there, the answer is
503and names the zone: the request is never moved to another zone. - Chat completion responses say where they were processed, in the
X-Zonaheader, and the usage records keep the zone and the site of every request. - Only text generation goes through zones for now. Embeddings, rerank and audio do not yet.
- A new machine first enters a test class, where only internal keys reach it, and serves customers after that.
Fallback between machines. If the chosen machine fails before the first token, the request moves to the next one, always inside the zones of the key. Until now the answer in that case was a 502.
Control plane and data planes. The platform is now split in two: a control plane in Switzerland, with the inventory of the machines, the catalogue and the zones; and one data plane per brand, in the brand's country, with accounts, keys, usage records and files. A brand downloads its configuration every minute and, once an hour, sends back only aggregated numbers: calls, errors, tokens and average time per model, machine, endpoint and zone, with no users, keys, texts or IP addresses. siati.ai is the first brand running this way. See Sovereignty.
Brands. The same platform serves several brands, each with its own name, domains, users and payments: by card, or with credits loaded by an administrator. A brand can open sign-up to everyone, or by request: the person asks for access with an address of an allowed domain, confirms it through a signed link, and an administrator approves. The request page does not reveal who already has access.
Documentation. The documentation is moving to English, one page at a time, and fills in the name, the addresses and the data country of each brand. A brand can replace a page with its own. The live pages — catalogue, service status, pricing — are in English: the catalogue shows the zones and the precision of every model, and the pricing page shows the cost formula the code applies, for brands that take payment by card and for those that load credits by hand.
Corrections. Pages that promised more than the service does were corrected:
- A tier does not give priority in a machine's queue. It chooses which machines may serve you, the rate limit of your key and the price multiplier. See Tiers.
- The documentation no longer names models that are not served, such as Mistral Large 2, Qwen 2.5 72B and DeepSeek-R1 Distill, added on 24 May.
- The web chat no longer shows the name and port of a machine in its error messages.
Contents of requests are no longer kept. The request log now keeps only numbers — endpoint, model, tokens, time taken, outcome — for 30 days. The text of requests and answers is kept only for the keys where you turn on "keep messages" in the dashboard, for debugging, encrypted, for 30 days. What the log held before was deleted on 3 October.
Credits. Every path that calls a model now records its usage: the web chat, the playground, knowledge bases (questions and indexing), the chat sessions of the app, and streams the client closes early (with tokens estimated from the text). On brands with prepaid credits each request takes its cost from the balance, and at zero paid calls stop with 402 insufficient_credits.
Fixed.
/v1/completionswith a JSON body answered500. It now answers.- Transcription: a recording that started with a few seconds of silence came back as empty text, because the silence check read only its beginning. It now reads the whole file, and declares silence only after reading all of it. A silent file with
response_format=textanswered500; it now answers with empty text. - After the move to the new machines, earlier on 3 October, uploads over 2 MB were refused. The limit is back to 50 MB per file.
- On brands where access is by request, sign-up through the app API is now refused: access is asked for on the brand's site.
- Every request now has a billing reference generated by us. The
X-Request-Idyou send stays in the logs and in the response, for correlation only.
2026-09-22 — The parameter list, and two false limits removed from the developer page
Reported by someone evaluating the service: the developer page of the site still listed as limits two things fixed on 18 August. They were right, and the page promised more than had been kept: "the list is updated at every release" and "the fix will be followed by its date in the changelog". This entry is that date, a month late.
No longer limits, and removed from the list:
response_format—json_objectandjson_schemawork, also withstream: true. Verified today through the public gateway: a schema withnumeroandtotalereturns{"numero": "2026-114", "totale": 1250.00}, keys and types respected. (Correction, 3 October 2026: on theqwen2.5models,response_formatis applied only without streaming.)top_p— forwarded to the engine. Verified today: attop_p: 0.01with temperature 1.5 the answer is the same three times out of three; at1.0it changes.
What had happened. Until 18 August the validation of the request kept only the fields listed in its rules and silently dropped everything else — top_p included, which was even in the table of this documentation. Fixed at the root: the list of parameters lives in one place, and an automated test fails if a field is validated without reaching either the engine or the list of refused ones. A parameter we do not recognise now gives 400 and is named.
Full list: Every parameter: honoured or refused.
Real limits that remain: max_tokens at 8192; no Assistants API, Realtime, Batch or fine-tuning; no SOC 2 Type II. And schemas constrain the shape of an answer, not its truth.
2026-05-24 — Catalogue expansion and chat improvements
Catalogue — three open-weight models added:
- Mistral Large 2 (123B) — reasoning, code, and many languages, Italian included.
- Qwen 2.5 72B Instruct — the best dense open model of its size when it was added.
- DeepSeek-R1 Distill Llama-70B — a reasoning model that shows its chain of thought in
<think>…</think>blocks.
Llama 3.1 405B removed: in its 4-bit quantised form it lost too much quality on instructed tasks.
Web chat:
- Answers stream token by token, instead of appearing all at once at the end.
- Markdown and math formulas are rendered while the answer arrives.
- The page scrolls with the answer, and stops following it while you scroll up.
- A Stop button interrupts the generation and keeps the partial answer.
- Choosing a model selects the fastest tier available for it.
- Every message shows the time to the first token, the total time and the tokens per second.
- Conversations get a title, generated in the background from the first message.
Chat sessions API: POST /api/v1/chat/sessions/{id}/messages streams the same chunks as the web chat. See Chat sessions.
Outgoing mail moved to a dedicated relay, with authenticated sending and spoofing blocked on the relay.
2026-05-20 — Better retrieval for RAG: reranker, Docling, hybrid search
Three pieces, each for a different weakness of plain vector search:
- Reranker —
bge-reranker-v2-m3. The pipeline takes 30 candidates from the vector store, a cross-encoder scores each of them against the question, and the best 5 are kept. - Docling replaces
pdftotextas the main PDF reader: it keeps tables, the reading order of pages in columns and the captions of figures, and reads scanned pages with OCR.pdftotextstays as the fallback. - Hybrid search — every collection holds dense vectors (BGE-M3) and sparse ones (BM25), combined with Reciprocal Rank Fusion. It finds exact terms — codes, names, acronyms — that dense vectors tend to bury.
Architecture and settings: RAG.
The knowledge bases created before this change were deleted: the old index could not hold both kinds of vectors.
2026-05-19 — RAG goes live, more machines, documentation rebuilt
RAG
- The RAG service is live in the dashboard and in the API.
- Vector store: Qdrant. Embeddings: BGE-M3, multilingual, 1024 dimensions, on a dedicated machine.
- Answers: Apertus 70B by default; any model of the catalogue can be chosen per request.
/dashboard/rag, for those who do not write code: upload PDF, DOCX, Markdown or text files, ask questions, get answers with citations.- API:
POST /api/v1/rag/kb,POST /api/v1/rag/kb/{slug}/docs,POST /api/v1/rag/kb/{slug}/chat.
Machines and routing
- More machines: for embeddings, for Apertus 70B and for the small models.
- Requests are routed by tier and model across several machines: by the weight of each machine for the tier, then by the shortest queue.
- Two inference engines, interchangeable behind the same API.
- A health check every 30 seconds.
- Machines are added and removed from the administration, without a release.
Site and documentation
- The copy of the site rewritten.
- Pricing page: plans billed monthly or yearly, and top-ups in five amounts from 5 to 500 CHF, with an estimate of the tokens they buy.
- This documentation rebuilt from scratch: several pages, search (⌘K), live pages read from the database, copy buttons on code blocks.
Pricing
- Four tiers:
slow,medium,fast,ludicrous. The separate tier for embeddings is gone. - The per-tier surcharge setting was removed: a latent bug in it doubled costs.
- Prices have a single source in the database.
2026-05-17 — Production platform live
- siati.ai is live: site, dashboard (sign-up, billing, playground, API keys), documentation (a single page at first), chat and API, one application on several domains.
- API for the mobile app: sign-in, chat sessions, in-app purchases, billing.
- Payments with Stripe, for subscriptions and top-ups; verification of App Store and Google Play purchases.
2026-05-15 — Backend rebuild started
- The backend is rebuilt from scratch.
- Existing accounts and API keys are carried over from the previous backend: passwords and keys keep working.
- API keys are stored as a salted HMAC-SHA256 hash, never in clear.
- Sign-in with tokens for the mobile app.