ITS INCOM AI ITS INCOM AI docs

Cookbook

Tuning RAG answers

The three things you can change per question, what is fixed, what you can do with the documents, and how to find where a wrong answer comes from.

Last updated: 2026-10-03

There are fewer knobs than you might expect. Per question you choose how many passages the model reads, which model writes the answer, and the tier. The rest — passage size, number of candidates, the instruction to the model — is fixed. This page says what each knob does, what you can do with the documents instead, and how to find where a wrong answer comes from.

What you can change, per question

http
POST https://my.ai.itsincom.org/api/v1/rag/kb/{slug}/chat
Content-Type: application/json

{
  "question": "…",
  "top_k": 5,
  "model": "gemma-4-26b",
  "tier": "medium"
}
Knob Default What it changes
top_k 5 How many of the passages found (up to 30) reach the model: from 1 to 20. Raise it when the answer is spread over several passages; keep it low for a precise fact. More passages also mean more tokens for the model to read.
model apertus-70b-instruct The model that writes the answer, from the catalogue; Which one to pick says which fits which job. It does not change which passages are found.
tier medium Which machines may write the answer: see Tiers. It does not change the search.

What is fixed

  • Passages: at most 2,048 characters, overlapping by 256, for every knowledge base. There is no setting per knowledge base, in the dashboard or in the API.
  • Candidates: the search always keeps up to 30 passages for the rerank.
  • The instruction to the model, the limit of 800 tokens on the answer, and the temperature, 0.3.
  • The embedding model: bge-m3, for every knowledge base.

If you need to choose any of these, build your own pipeline with /v1/embeddings and /v1/rerank.

What you can do with the documents

The input is the lever left in your hands.

  • Give it text when you have it. TXT and Markdown are read as they are. A DOCX loses its headers, footers and notes, and its tables become plain text. A PDF goes through Docling.
  • Run OCR on scanned PDFs before uploading them, for example with ocrmypdf. Then you do not depend on Docling to read them.
  • Remove old versions. Two versions of the same contract are both searched, and either can be cited. The same file uploaded twice is indexed twice.
  • Ask with the words of the documents. The keyword half of the search matches exact words: names, codes, numbers. Words shorter than 3 characters are ignored by it.

Reading the sources

Every source in the answer looks like this (example values):

json
{
  "score": 0.98,
  "dense_score": 0.5,
  "text": "…",
  "document_filename": "contract.pdf",
  "chunk_idx": 12
}
  • score: the score of the reranker for this passage and your question, from 0 to 1. Sources come sorted by it.
  • dense_score: the score of the first search — the merged ranking of the hybrid search, or the vector similarity if the search fell back to meaning only. It is on another scale: compare it between passages, not with score.
  • If the reranker did not answer, score repeats dense_score, and the order is the one of the first search.
  • chunk_idx: the position of the passage in its document, counting from 0. It is the N of the citation [file name: parte N].

When the answer is wrong

The pipeline can fail in three places. Look at them in this order.

  1. Reading. Did the text you expect come out of the file at all? A long PDF with very few passages (chunks_count of the document) is a warning sign. The passages cannot be listed through the API today: ask a question that quotes a sentence of the document word for word, and look for it in sources[].text. If it never comes back, delete the document and upload a text version.
  2. Search. The passage exists but is not among the sources: raise top_k, and use in the question the words the document uses.
  3. Answer. The right passage is in sources and the answer is still wrong: try another model.

An empty sources while your documents are ready is a different case: the search found nothing, or the vector database did not answer. Service status tells you which.

Processing documents again

There is no reindex. To process a document again — after fixing it, or after a change in how documents are read — delete it and upload it again.