Cookbook
OpenAI SDK migration
Change the base URL, the key and the model name. What else is different, and what does not exist here.
Last updated: 2026-10-03
If your code uses the OpenAI SDK, three things change: the base URL, the key and the model name. The rest of your code stays as it is, with the differences listed on this page.
| OpenAI | ITS INCOM AI | |
|---|---|---|
| Base URL | https://api.openai.com/v1 |
https://api.ai.itsincom.org/v1 |
| Key | sk-… from OpenAI |
sk-… from the dashboard |
| Model | gpt-4o, … |
an open-weight model from the catalogue |
Chat model names are not translated. A request for gpt-4o answers 404 with model_not_found: it is never sent to another model. Embeddings are the exception, see below.
Python
import os
from openai import OpenAI
# Before:
# client = OpenAI(api_key=os.environ["OPENAI_API_KEY"])
client = OpenAI(
base_url="https://api.ai.itsincom.org/v1",
api_key=os.environ["API_KEY"],
)
resp = client.chat.completions.create(
model="gemma-4-26b",
messages=[{"role": "user", "content": "Hello"}],
)
print(resp.choices[0].message.content)
JavaScript / TypeScript
import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://api.ai.itsincom.org/v1",
apiKey: process.env.API_KEY,
});
const resp = await client.chat.completions.create({
model: "gemma-4-26b",
messages: [{ role: "user", content: "Hello" }],
});
console.log(resp.choices[0].message.content);
Which SDK calls work
| SDK call | Endpoint | Notes |
|---|---|---|
chat.completions.create |
/v1/chat/completions |
streaming, tools, response_format, images on the models that accept them |
completions.create |
/v1/completions |
one prompt per request |
responses.create |
/v1/responses |
no streaming, no stored conversations: see Responses API |
embeddings.create |
/v1/embeddings |
bge-m3 only: see Embeddings below |
audio.transcriptions.create |
/v1/audio/transcriptions |
whisper-1 or whisper-1-hd |
audio.speech.create |
/v1/audio/speech |
our voices, wav or mp3: see Audio below |
models.list |
/v1/models |
the catalogue, audio models included |
/v1/rerank exists too, but the OpenAI SDK has no method for it: call it over HTTP, see Rerank.
Everything else does not exist here: Assistants and threads, Batch, Files, fine-tuning, image generation, moderation, Realtime, vector stores. Those paths answer 404 with code: unknown_endpoint. For answers over your own documents there is RAG, which signs in with your account rather than the API key; or build your own on /v1/embeddings and /v1/rerank.
Models
| If your code says | Use | Notes |
|---|---|---|
a chat model: gpt-4o, gpt-4o-mini, gpt-4.1, … |
gemma-4-26b, or another chat model |
which one for which job: Models |
| a chat model, with images in the messages | a model marked vision in the catalogue |
Images as input |
text-embedding-3-small, text-embedding-3-large, text-embedding-ada-002 |
bge-m3 |
1024 dimensions |
whisper-1, gpt-4o-transcribe |
whisper-1 or whisper-1-hd |
other names are refused |
tts-1, tts-1-hd |
the same | the name is not checked: one speech engine serves every request |
dall-e-3, gpt-image-1 |
— | no image generation |
Chat completions: what is different
- Unknown parameters are refused. A parameter we do not recognise gets a
400that names it, typos included. These are refused on purpose:store,metadata,functionsandfunction_call(usetoolsandtool_choice),modalities,audio,prediction,web_search_options. The full list is in Every parameter. - Some parameters work only on some models.
n,logprobs,top_logprobs,logit_biasandparallel_tool_callsare refused with a400on the models that cannot apply them. Some libraries sendn: 1orlogprobs: falseby default: if that400appears, drop them from the request. - The token cap is
max_tokensormax_completion_tokens, up to 8192. If you send both with different values, the answer is400. - Streaming sends the same chunks and ends with
data: [DONE]. There is no final chunk withusage, even withstream_options: {"include_usage": true}: the tokens of each request are in Usage in the dashboard. - Errors. An unknown model gives
404 model_not_found. No machine available in your zones gives503withtype: zone_unavailableand the zone in the message: the request is not sent to another zone. The limits per key and per client address give429withRetry-After(Rate limits). The SDKs retry429and5xxtwice by default (max_retries). Some messages are in Italian: in code, check the status andtype, not the text. All the codes: Errors.
Tools (function calling)
The same shape as OpenAI: tools, tool_choice, and tool_calls in the answer.
tools = [{
"type": "function",
"function": {
"name": "get_weather",
"description": "Current weather for a city",
"parameters": {
"type": "object",
"properties": {"city": {"type": "string"}},
"required": ["city"],
},
},
}]
resp = client.chat.completions.create(
model="gemma-4-26b",
messages=[{"role": "user", "content": "What is the weather in Lugano?"}],
tools=tools,
tool_choice="auto",
)
print(resp.choices[0].message.tool_calls)
Send the result back as a message with role: "tool" and the tool_call_id, as with OpenAI. Schemas, guaranteed JSON and their limits: JSON and functions.
Tier and zone from the SDK
The tier is a request header. The zone where a chat completion was processed comes back in the X-Zona header.
client = OpenAI(
base_url="https://api.ai.itsincom.org/v1",
api_key=os.environ["API_KEY"],
default_headers={"X-Siati-Tier": "fast"}, # every request
)
raw = client.chat.completions.with_raw_response.create(
model="gemma-4-26b",
messages=[{"role": "user", "content": "Hello"}],
extra_headers={"X-Siati-Tier": "slow"}, # this request only
)
print(raw.headers.get("X-Zona"))
resp = raw.parse()
const { data, response } = await client.chat.completions
.create(
{ model: "gemma-4-26b", messages: [{ role: "user", content: "Hello" }] },
{ headers: { "X-Siati-Tier": "fast" } },
)
.withResponse();
console.log(response.headers.get("x-zona"));
What a tier changes: Tiers. What a zone is: Zones.
Embeddings
emb = client.embeddings.create(
model="bge-m3",
input=["first text", "second text"],
)
print(len(emb.data[0].embedding)) # 1024
- Re-embed your whole corpus. Vectors from OpenAI models and from
bge-m3are not comparable, and they have different lengths: do not mix them in one index. - 1024 dimensions, always.
dimensionsis accepted and has no effect. inputis a string or a list of up to 32 strings. Lists of token ids are refused.encoding_formatisfloatorbase64. The Python SDK asks forbase64by default and decodes it for you.modelis not checked today: every request is answered bybge-m3, and the response says so inmodel. Writebge-m3anyway.
Audio
Transcription takes whisper-1 or whisper-1-hd, and response_format json, verbose_json or text: srt and vtt are not available.
Speech uses our voices, listed in Audio. OpenAI voice names (alloy, echo, nova, …) are accepted and mapped to one of our voices, female or male, in the language given by language: a field of ours, Italian by default. For English text, send it:
speech = client.audio.speech.create(
model="tts-1",
voice="alloy",
input="Hello from the other side.",
response_format="mp3",
extra_body={"language": "en"},
)
speech.write_to_file("hello.mp3")
The default format is wav, where OpenAI's is mp3: ask for mp3 if you need it. Only wav and mp3 exist. speed is accepted and has no effect today.
LangChain
import os
from langchain_openai import ChatOpenAI, OpenAIEmbeddings
llm = ChatOpenAI(
base_url="https://api.ai.itsincom.org/v1",
api_key=os.environ["API_KEY"],
model="gemma-4-26b",
)
llm.invoke("Hello")
emb = OpenAIEmbeddings(
base_url="https://api.ai.itsincom.org/v1",
api_key=os.environ["API_KEY"],
model="bge-m3",
check_embedding_ctx_length=False, # send text, not token ids
chunk_size=32, # at most 32 texts per request
)
Without those two settings OpenAIEmbeddings sends token ids and up to 1000 texts per request, and both are refused.
LlamaIndex
import os
from llama_index.llms.openai_like import OpenAILike
llm = OpenAILike(
api_base="https://api.ai.itsincom.org/v1",
api_key=os.environ["API_KEY"],
model="gemma-4-26b",
is_chat_model=True,
)
OpenAILike is the LlamaIndex class for OpenAI-compatible endpoints. If you use tools through LlamaIndex, add is_function_calling_model=True.
Moving one call at a time
The SDK binds a client to one base URL, so an OpenAI client and a ITS INCOM AI client can live side by side in the same code. Move one call, compare the answers, then move the next.