ITS INCOM AI ITS INCOM AI docs

Concepts

Knowledge bases (RAG)

Answers from your own documents, with their sources: what happens to a file, how passages are found, what the model is told, and what does not work yet.

Last updated: 2026-10-03

A chat model knows only what was in its training data. With RAG — retrieval-augmented generation — it answers from your documents instead:

  1. Your documents are split into passages, and every passage becomes a vector.
  2. A question becomes a vector too, and the passages closest to it are found.
  3. The best passages go to the model with the question. The model answers from them, and cites them.

Here a set of documents you query together is called a knowledge base.

Where to use it

Not with an API key: the developer API has no /v1/rag/…. To build your own pipeline with a key, the pieces are /v1/embeddings and /v1/rerank.

What happens to a document

Step What happens
Reading PDF: first Docling, which keeps tables and the reading order of columns, and runs OCR on scanned pages. If Docling fails or returns nothing, the text is taken with pdftotext, which does not read scanned pages. DOCX: the text of the main body; headers, footers and notes are not read, and tables become plain text. TXT and Markdown: as they are.
Splitting All spacing, line breaks included, becomes a single space. The text is then cut into passages of at most 2,048 characters, overlapping by 256: at the end of a sentence when possible, otherwise between two words.
Vectors Each passage gets two vectors. A dense one from bge-m3, 1024 dimensions, for meaning. A keyword one: how many times each word appears, leaving out words shorter than 3 characters and the most common Italian, English, German and French words.
Storage The vectors go to a Qdrant collection, one per knowledge base, which weighs each keyword by how rare it is in that knowledge base. The file stays in storage, and the text of the passages is also kept in the database.

Indexing runs in a queue: the upload answers at once, and the document goes from pending to ready — or failed — later.

What happens to a question

  1. The question gets the same two vectors.
  2. Hybrid search. Qdrant takes the 60 passages closest in meaning and the 60 best by keywords, merges the two rankings with Reciprocal Rank Fusion, and keeps the best 30.
  3. Rerank. bge-reranker-v2-m3 reads the question together with each of these passages and scores them. The best top_k stay: 5 unless you ask otherwise, up to 20.
  4. Answer. The passages and the question go to the chat model — apertus-70b-instruct unless you choose another — on machines in the zones offered by ITS INCOM AI.

The keyword half is what finds exact words — a name, a product code, an article number — that the dense vector tends to blur into a general idea. The rerank is what puts the passage that really answers first.

When a piece is missing the pipeline degrades instead of stopping. If the hybrid search fails, it searches by meaning only. If the reranker does not answer, the order of the search is kept. If the vector database does not answer at all, the search finds nothing and the model answers without passages.

What the model is told

A fixed instruction, written in Italian, tells the model to answer only from the passages it is given, to say so when the answer is not there, to cite its sources as [file name: parte N], and to answer in the language of the question. The passages follow, each with its file name, its position in the document (chunk_idx, counting from 0) and its score.

The answer is at most 800 tokens long, at temperature 0.3. You cannot change the instruction: if you need your own, build the pipeline with /v1/embeddings and /v1/rerank.

The instruction makes answers from outside the documents less likely; it does not rule them out. That is why every answer comes with the passages it was given: check them.

Measured

On 22 September 2026 we measured the pipeline on a five-article mandate contract, two passages long. Indexing took 0.20 s; the answers took 0.57 s on average over three questions, each citing two sources, and all three were right. The answers were written by gemma-4-26b.

The document reached the pipeline without the upload, which was out of service that day because the object storage that receives the files had failed. The object storage was replaced on 3 October 2026. Code and method are in Three examples, with measured numbers.

What does not work yet

  • No settings per knowledge base. Passage size and overlap are the same for every knowledge base and cannot be changed, in the dashboard or in the API. Neither can the number of candidates (30), the instruction, or the length of the answer.
  • Renaming a knowledge base is not possible. Deleting one is possible only from the dashboard, and today it leaves the uploaded files in storage: to remove them too, delete the documents first.
  • A failed document is not retried. The passages it indexed before failing stay in the search until you delete it.
  • Scanned PDFs depend on Docling. On the pdftotext fallback they come out empty, and the document fails with Documento vuoto dopo parsing.
  • The status page checks one piece only. Service status shows whether the vector database answers. It does not tell you whether upload, reading or indexing work: if a document stays pending or ends failed, write to segreteria@itsincom.it.
  • Some texts are in Italian: the citation label parte, the fixed answer of a knowledge base with no document ready, and some error messages of failed documents.

Where it runs

The answer is generated only on machines in the zones offered by ITS INCOM AI, and is never moved to another zone. Reading the files, the vectors, the search and the rerank run on dedicated machines that do not go through zones yet. If it matters for your case, write to segreteria@itsincom.it and we tell you where they run for ITS INCOM AI.

When it is the right tool

Good fits:

  • Questions on documents that change slowly: contracts, manuals, regulations, an internal wiki.
  • Finding by meaning and by exact words — codes, names — at the same time.
  • Answers someone has to be able to check: each one comes with its sources.

Poor fits:

  • Live data: give the model a tool to call instead, see OpenAI SDK migration.
  • Changing how the model writes: that is a matter of instructions, or of another model.
  • A few short documents: they may fit in the prompt of an ordinary chat request.