ITS INCOM AI ITS INCOM AI docs

API reference

Embeddings and rerank

POST /v1/embeddings turns texts into vectors with bge-m3; POST /v1/rerank puts candidate texts in order of relevance with bge-reranker-v2-m3.

Last updated: 2026-10-03

Two endpoints for building your own search, with your API key:

  • /v1/embeddings turns texts into vectors, to search by meaning.
  • /v1/rerank takes a query and a list of candidate texts, and puts them in order of relevance.

If you want answers from your documents without building the pipeline yourself, knowledge bases do it for you — with a signed-in user's token, not with an API key.

Embeddings

POST https://api.ai.itsincom.org/v1/embeddings — model bge-m3, multilingual, 1024 dimensions. The request and the response have the OpenAI shape.

bash
curl https://api.ai.itsincom.org/v1/embeddings \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "bge-m3",
    "input": [
      "Either party may end the contract with six months of notice.",
      "Der Vertrag kann mit einer Frist von sechs Monaten gekündigt werden."
    ]
  }'

Parameters

Parameter Type Required Notes
input string or array of strings yes Up to 32 texts per request. An empty text, or an array of token ids, is refused with 400.
model string no Send bge-m3. Today any value is accepted and ignored: the vectors always come from bge-m3, and the response says so in model.
encoding_format string no float (default) or base64 (float32, little-endian).
dimensions integer no Accepted and ignored: vectors always have 1024 dimensions.
user string no Accepted and ignored.

A text longer than 8192 tokens is not refused: it is cut, and the vector describes only its beginning. These are the model's own tokens, not the estimate in usage.

Response

json
{
  "object": "list",
  "data": [
    { "object": "embedding", "index": 0, "embedding": [0.0123, -0.0456, 0.0078] },
    { "object": "embedding", "index": 1, "embedding": [0.0234, -0.0567, 0.0091] }
  ],
  "model": "bge-m3",
  "usage": { "prompt_tokens": 32, "total_tokens": 32 }
}

The vectors are shortened here: each one has 1024 numbers. index is the position of the text in your input.

usage is an estimate — about four characters per token, counted the same way for every language — because the embedding server does not return a count. Embeddings are charged on these input tokens only, at the price in Pricing.

With the OpenAI SDK

python
import os

import numpy as np
from openai import OpenAI

client = OpenAI(base_url="https://api.ai.itsincom.org/v1", api_key=os.environ["API_KEY"])

texts = [
    "Either party may end the contract with six months of notice.",
    "Der Vertrag kann mit einer Frist von sechs Monaten gekündigt werden.",
    "The office is open from Monday to Friday.",
]
resp = client.embeddings.create(model="bge-m3", input=texts)
vectors = np.array([d.embedding for d in resp.data])

# Cosine similarity between every pair of texts
unit = vectors / np.linalg.norm(vectors, axis=1, keepdims=True)
print(np.round(unit @ unit.T, 3))

Rerank

POST https://api.ai.itsincom.org/v1/rerank — model bge-reranker-v2-m3.

A reranker reads the query together with each candidate and gives it a score. It is slower than comparing vectors, and more precise: the usual pattern is to search wide with vectors, then rerank the candidates and keep the best few. Our knowledge bases take up to 30 candidates from the first search and, by default, keep 5.

There is no OpenAI standard for rerank: the request and the response follow the shape used by Cohere and Jina.

bash
curl https://api.ai.itsincom.org/v1/rerank \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "bge-reranker-v2-m3",
    "query": "How much notice is needed to end the contract?",
    "documents": [
      "The fee is paid in two instalments, in January and in July.",
      "Either party may end the contract with six months of notice.",
      "The office is open from Monday to Friday."
    ],
    "top_n": 2,
    "return_documents": true
  }'

Parameters

Parameter Type Required Notes
query string yes Up to 32,000 characters.
documents array of strings yes From 1 to 32 candidates. An empty one is refused with 400.
top_n integer no Return only the best N. Default: all of them.
return_documents boolean no true repeats the text of each result. Default false.
model string no Send bge-reranker-v2-m3. As for embeddings, today any value is accepted and ignored.

A query and a candidate that together exceed 8192 tokens are cut, not refused.

Response

json
{
  "object": "list",
  "model": "bge-reranker-v2-m3",
  "results": [
    { "index": 1, "relevance_score": 0.98, "document": { "text": "Either party may end the contract with six months of notice." } },
    { "index": 0, "relevance_score": 0.01, "document": { "text": "The fee is paid in two instalments, in January and in July." } }
  ],
  "usage": { "total_tokens": 77 }
}

The scores here are an example, not a measurement.

  • results are sorted by relevance_score, from the highest. The score goes from 0 to 1.
  • index is the position of the candidate in your documents.
  • document is there only with return_documents: true.
  • usage.total_tokens is an estimate, about four characters per token: the candidates, plus the query counted once for each candidate, because it is read with each of them. The price is in Pricing.

From Python

The OpenAI SDK has no rerank method: call the endpoint directly.

python
import os

import requests

query = "How much notice is needed to end the contract?"
candidates = [
    "The fee is paid in two instalments, in January and in July.",
    "Either party may end the contract with six months of notice.",
    "The office is open from Monday to Friday.",
]

r = requests.post(
    "https://api.ai.itsincom.org/v1/rerank",
    headers={"Authorization": f"Bearer {os.environ['API_KEY']}"},
    json={"model": "bge-reranker-v2-m3", "query": query, "documents": candidates, "top_n": 2},
    timeout=30,
)
r.raise_for_status()

for hit in r.json()["results"]:
    print(round(hit["relevance_score"], 3), candidates[hit["index"]])

Limits

Two limits apply to both endpoints, and the stricter one wins:

  • The per-minute limit of your key's tier, counted on all your requests with that key. The values are in Pricing. The tier — the key's default, or the X-Siati-Tier header — also sets the price multiplier. It does not change the machines: they are the same for every tier.
  • 120 requests per minute per client address. The counter is shared with the other limited endpoints of the API called from the same address.

Over either limit you get 429 with a Retry-After header. The two limits answer with different bodies: rely on the status and on Retry-After.

There is no larger batch: to embed more than 32 texts, send more requests.

Errors

Status type When
400 invalid_request_error Missing input or query; more than 32 texts or candidates; an empty text; token ids; an unknown tier
401 invalid_request_error The key is missing, wrong, revoked or expired
429 — One of the two limits above
501 not_implemented Rerank is switched off on this installation
502 bad_gateway The embedding or rerank server did not answer. Retry; the body has a request_id to quote if you write to us

Tips

  • Split long texts before embedding them. One vector for a whole long document says little about each part of it, and beyond 8192 tokens the text is cut anyway. Our knowledge bases use passages of at most 2,048 characters, overlapping by 256.
  • Compare vectors with cosine similarity, as in the example above.
  • Do not mix models. A bge-m3 vector can be compared only with other bge-m3 vectors, never with vectors from another provider's model.