API reference
Embeddings and rerank
POST /v1/embeddings turns texts into vectors with bge-m3; POST /v1/rerank puts candidate texts in order of relevance with bge-reranker-v2-m3.
Last updated: 2026-10-03
Two endpoints for building your own search, with your API key:
/v1/embeddingsturns texts into vectors, to search by meaning./v1/reranktakes a query and a list of candidate texts, and puts them in order of relevance.
If you want answers from your documents without building the pipeline yourself, knowledge bases do it for you — with a signed-in user's token, not with an API key.
Embeddings
POST https://api.ai.itsincom.org/v1/embeddings — model bge-m3, multilingual, 1024 dimensions. The request and the response have the OpenAI shape.
curl https://api.ai.itsincom.org/v1/embeddings \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "bge-m3",
"input": [
"Either party may end the contract with six months of notice.",
"Der Vertrag kann mit einer Frist von sechs Monaten gekündigt werden."
]
}'
Parameters
| Parameter | Type | Required | Notes |
|---|---|---|---|
input |
string or array of strings | yes | Up to 32 texts per request. An empty text, or an array of token ids, is refused with 400. |
model |
string | no | Send bge-m3. Today any value is accepted and ignored: the vectors always come from bge-m3, and the response says so in model. |
encoding_format |
string | no | float (default) or base64 (float32, little-endian). |
dimensions |
integer | no | Accepted and ignored: vectors always have 1024 dimensions. |
user |
string | no | Accepted and ignored. |
A text longer than 8192 tokens is not refused: it is cut, and the vector describes only its beginning. These are the model's own tokens, not the estimate in usage.
Response
{
"object": "list",
"data": [
{ "object": "embedding", "index": 0, "embedding": [0.0123, -0.0456, 0.0078] },
{ "object": "embedding", "index": 1, "embedding": [0.0234, -0.0567, 0.0091] }
],
"model": "bge-m3",
"usage": { "prompt_tokens": 32, "total_tokens": 32 }
}
The vectors are shortened here: each one has 1024 numbers. index is the position of the text in your input.
usage is an estimate — about four characters per token, counted the same way for every language — because the embedding server does not return a count. Embeddings are charged on these input tokens only, at the price in Pricing.
With the OpenAI SDK
import os
import numpy as np
from openai import OpenAI
client = OpenAI(base_url="https://api.ai.itsincom.org/v1", api_key=os.environ["API_KEY"])
texts = [
"Either party may end the contract with six months of notice.",
"Der Vertrag kann mit einer Frist von sechs Monaten gekündigt werden.",
"The office is open from Monday to Friday.",
]
resp = client.embeddings.create(model="bge-m3", input=texts)
vectors = np.array([d.embedding for d in resp.data])
# Cosine similarity between every pair of texts
unit = vectors / np.linalg.norm(vectors, axis=1, keepdims=True)
print(np.round(unit @ unit.T, 3))
Rerank
POST https://api.ai.itsincom.org/v1/rerank — model bge-reranker-v2-m3.
A reranker reads the query together with each candidate and gives it a score. It is slower than comparing vectors, and more precise: the usual pattern is to search wide with vectors, then rerank the candidates and keep the best few. Our knowledge bases take up to 30 candidates from the first search and, by default, keep 5.
There is no OpenAI standard for rerank: the request and the response follow the shape used by Cohere and Jina.
curl https://api.ai.itsincom.org/v1/rerank \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "bge-reranker-v2-m3",
"query": "How much notice is needed to end the contract?",
"documents": [
"The fee is paid in two instalments, in January and in July.",
"Either party may end the contract with six months of notice.",
"The office is open from Monday to Friday."
],
"top_n": 2,
"return_documents": true
}'
Parameters
| Parameter | Type | Required | Notes |
|---|---|---|---|
query |
string | yes | Up to 32,000 characters. |
documents |
array of strings | yes | From 1 to 32 candidates. An empty one is refused with 400. |
top_n |
integer | no | Return only the best N. Default: all of them. |
return_documents |
boolean | no | true repeats the text of each result. Default false. |
model |
string | no | Send bge-reranker-v2-m3. As for embeddings, today any value is accepted and ignored. |
A query and a candidate that together exceed 8192 tokens are cut, not refused.
Response
{
"object": "list",
"model": "bge-reranker-v2-m3",
"results": [
{ "index": 1, "relevance_score": 0.98, "document": { "text": "Either party may end the contract with six months of notice." } },
{ "index": 0, "relevance_score": 0.01, "document": { "text": "The fee is paid in two instalments, in January and in July." } }
],
"usage": { "total_tokens": 77 }
}
The scores here are an example, not a measurement.
resultsare sorted byrelevance_score, from the highest. The score goes from 0 to 1.indexis the position of the candidate in yourdocuments.documentis there only withreturn_documents: true.usage.total_tokensis an estimate, about four characters per token: the candidates, plus the query counted once for each candidate, because it is read with each of them. The price is in Pricing.
From Python
The OpenAI SDK has no rerank method: call the endpoint directly.
import os
import requests
query = "How much notice is needed to end the contract?"
candidates = [
"The fee is paid in two instalments, in January and in July.",
"Either party may end the contract with six months of notice.",
"The office is open from Monday to Friday.",
]
r = requests.post(
"https://api.ai.itsincom.org/v1/rerank",
headers={"Authorization": f"Bearer {os.environ['API_KEY']}"},
json={"model": "bge-reranker-v2-m3", "query": query, "documents": candidates, "top_n": 2},
timeout=30,
)
r.raise_for_status()
for hit in r.json()["results"]:
print(round(hit["relevance_score"], 3), candidates[hit["index"]])
Limits
Two limits apply to both endpoints, and the stricter one wins:
- The per-minute limit of your key's tier, counted on all your requests with that key. The values are in Pricing. The tier — the key's default, or the
X-Siati-Tierheader — also sets the price multiplier. It does not change the machines: they are the same for every tier. - 120 requests per minute per client address. The counter is shared with the other limited endpoints of the API called from the same address.
Over either limit you get 429 with a Retry-After header. The two limits answer with different bodies: rely on the status and on Retry-After.
There is no larger batch: to embed more than 32 texts, send more requests.
Errors
| Status | type |
When |
|---|---|---|
400 |
invalid_request_error |
Missing input or query; more than 32 texts or candidates; an empty text; token ids; an unknown tier |
401 |
invalid_request_error |
The key is missing, wrong, revoked or expired |
429 |
— | One of the two limits above |
501 |
not_implemented |
Rerank is switched off on this installation |
502 |
bad_gateway |
The embedding or rerank server did not answer. Retry; the body has a request_id to quote if you write to us |
Tips
- Split long texts before embedding them. One vector for a whole long document says little about each part of it, and beyond 8192 tokens the text is cut anyway. Our knowledge bases use passages of at most 2,048 characters, overlapping by 256.
- Compare vectors with cosine similarity, as in the example above.
- Do not mix models. A
bge-m3vector can be compared only with otherbge-m3vectors, never with vectors from another provider's model.