Cookbook
Tuning RAG answers
The three things you can change per question, what is fixed, what you can do with the documents, and how to find where a wrong answer comes from.
Last updated: 2026-10-03
There are fewer knobs than you might expect. Per question you choose how many passages the model reads, which model writes the answer, and the tier. The rest — passage size, number of candidates, the instruction to the model — is fixed. This page says what each knob does, what you can do with the documents instead, and how to find where a wrong answer comes from.
What you can change, per question
POST https://my.ai.itsincom.org/api/v1/rag/kb/{slug}/chat
Content-Type: application/json
{
"question": "…",
"top_k": 5,
"model": "gemma-4-26b",
"tier": "medium"
}
| Knob | Default | What it changes |
|---|---|---|
top_k |
5 |
How many of the passages found (up to 30) reach the model: from 1 to 20. Raise it when the answer is spread over several passages; keep it low for a precise fact. More passages also mean more tokens for the model to read. |
model |
apertus-70b-instruct |
The model that writes the answer, from the catalogue; Which one to pick says which fits which job. It does not change which passages are found. |
tier |
medium |
Which machines may write the answer: see Tiers. It does not change the search. |
What is fixed
- Passages: at most 2,048 characters, overlapping by 256, for every knowledge base. There is no setting per knowledge base, in the dashboard or in the API.
- Candidates: the search always keeps up to 30 passages for the rerank.
- The instruction to the model, the limit of 800 tokens on the answer, and the temperature, 0.3.
- The embedding model:
bge-m3, for every knowledge base.
If you need to choose any of these, build your own pipeline with /v1/embeddings and /v1/rerank.
What you can do with the documents
The input is the lever left in your hands.
- Give it text when you have it. TXT and Markdown are read as they are. A DOCX loses its headers, footers and notes, and its tables become plain text. A PDF goes through Docling.
- Run OCR on scanned PDFs before uploading them, for example with
ocrmypdf. Then you do not depend on Docling to read them. - Remove old versions. Two versions of the same contract are both searched, and either can be cited. The same file uploaded twice is indexed twice.
- Ask with the words of the documents. The keyword half of the search matches exact words: names, codes, numbers. Words shorter than 3 characters are ignored by it.
Reading the sources
Every source in the answer looks like this (example values):
{
"score": 0.98,
"dense_score": 0.5,
"text": "…",
"document_filename": "contract.pdf",
"chunk_idx": 12
}
score: the score of the reranker for this passage and your question, from 0 to 1. Sources come sorted by it.dense_score: the score of the first search — the merged ranking of the hybrid search, or the vector similarity if the search fell back to meaning only. It is on another scale: compare it between passages, not withscore.- If the reranker did not answer,
scorerepeatsdense_score, and the order is the one of the first search. chunk_idx: the position of the passage in its document, counting from 0. It is theNof the citation[file name: parte N].
When the answer is wrong
The pipeline can fail in three places. Look at them in this order.
- Reading. Did the text you expect come out of the file at all? A long PDF with very few passages (
chunks_countof the document) is a warning sign. The passages cannot be listed through the API today: ask a question that quotes a sentence of the document word for word, and look for it insources[].text. If it never comes back, delete the document and upload a text version. - Search. The passage exists but is not among the sources: raise
top_k, and use in the question the words the document uses. - Answer. The right passage is in
sourcesand the answer is still wrong: try anothermodel.
An empty sources while your documents are ready is a different case: the search found nothing, or the vector database did not answer. Service status tells you which.
Processing documents again
There is no reindex. To process a document again — after fixing it, or after a change in how documents are read — delete it and upload it again.