ITS INCOM AI ITS INCOM AI docs

API reference

Audio: transcription and speech

POST /v1/audio/transcriptions and POST /v1/audio/speech, in the shape of OpenAI's: what they accept, what they return, and what does not work yet.

Last updated: 2026-10-03

Two endpoints, in the shape of OpenAI's:

Endpoint Does Engine
POST /v1/audio/transcriptions audio → text Whisper large-v3-turbo (whisper-1) or Whisper large-v3 (whisper-1-hd)
POST /v1/audio/speech text → audio Piper, six voices in four languages

The OpenAI SDKs work with them: change base_url and the key. The examples below use:

bash
export API_KEY="sk-…"

Transcription

POST https://api.ai.itsincom.org/v1/audio/transcriptions, as a multipart form.

bash
curl https://api.ai.itsincom.org/v1/audio/transcriptions \
  -H "Authorization: Bearer $API_KEY" \
  -F file=@message.m4a \
  -F model=whisper-1 \
  -F language=en
json
{ "text": "Hello, I wanted to ask whether we can meet tomorrow." }

Python

python
import os

from openai import OpenAI

client = OpenAI(base_url="https://api.ai.itsincom.org/v1", api_key=os.environ["API_KEY"])

with open("message.m4a", "rb") as f:
    result = client.audio.transcriptions.create(
        model="whisper-1",
        file=f,
        language="en",  # optional: detected from the audio if omitted
    )

print(result.text)

TypeScript (Node 20 or later, on your server: the key must not reach a browser)

typescript
// transcribe.mts. Run: API_KEY=sk-… npx tsx transcribe.mts
import { openAsBlob } from "node:fs";

const form = new FormData();
form.append("file", await openAsBlob("message.m4a"), "message.m4a");
form.append("model", "whisper-1");

const res = await fetch("https://api.ai.itsincom.org/v1/audio/transcriptions", {
  method: "POST",
  headers: { Authorization: `Bearer ${process.env.API_KEY}` },
  body: form,
});
const body = await res.json();
if (!res.ok) throw new Error(body.error?.message ?? body.message);
console.log(body.text);

Fields

Field
file required The audio, up to 25 MB. See accepted formats.
model optional whisper-1 (default) runs Whisper large-v3-turbo; whisper-1-hd runs Whisper large-v3. Any other value gives 400.
language optional A two-letter code: it, en, de, fr… Without it, Whisper detects the language.
prompt optional Up to 1,024 characters of context: names and terms you expect, spelled the way you want them back.
response_format optional json (default), text or verbose_json.
temperature optional From 0 to 1.
timestamp_granularities[] optional word, segment or both. Only with verbose_json.

prompt, language and response_format sent empty count as absent. SDKs often send an empty variable instead of leaving the field out, and that used to make the whole transcription fail.

What comes back

With json, the default, the object shown above: {"text": "…"}.

With text, the bare text, as Content-Type: text/plain; charset=utf-8.

With verbose_json, also the language, the duration in seconds and the segments with their timings. Segments are passed on as the engine returns them; the example shows only some of their fields:

json
{
  "task": "transcribe",
  "language": "en",
  "duration": 4.32,
  "text": "Hello, how are you today? I wanted to ask whether we can meet tomorrow.",
  "segments": [
    { "id": 0, "start": 0.0, "end": 2.1, "text": " Hello, how are you today?" },
    { "id": 1, "start": 2.1, "end": 4.32, "text": " I wanted to ask whether we can meet tomorrow." }
  ]
}

Per-word timings

timestamp_granularities[]=word adds words at the top level of the response, with the start, end and probability of every word. It is what you need to highlight the exact point in the audio, cut a recording, or show where recognition was uncertain.

bash
curl https://api.ai.itsincom.org/v1/audio/transcriptions \
  -H "Authorization: Bearer $API_KEY" \
  -F file=@note.wav \
  -F model=whisper-1 \
  -F response_format=verbose_json \
  -F "timestamp_granularities[]=word"
json
{
  "words": [
    { "start": 0.0, "end": 0.42, "word": " Hello,", "probability": 0.97 },
    { "start": 0.42, "end": 0.71, "word": " how", "probability": 0.99 }
  ]
}

Timings exist only inside verbose_json: the other two forms have nowhere to put them. Asking for them with another format gives 400, instead of an answer without timings that you could not tell apart from audio in which no words were heard. Until 18 August 2026 the field was accepted and not forwarded.

A measured example, with 55 timed words, is in Three examples, with measured numbers.

Which model

whisper-1 is the default and the faster of the two. If it makes mistakes on your audio — noise, phone calls, strong accents — run the same files through whisper-1-hd and compare.

The model is loaded on demand: the first request after a long pause can take several seconds longer.

Accepted formats

The check looks at the content of the file, not at its name or at the Content-Type you declare. A renamed file does not pass; a file with the wrong Content-Type passes if what is inside is a format we accept.

These are the accepted types, as detected from the content:

text
audio/wav  audio/wave  audio/x-wav  audio/vnd.wave
audio/mpeg  audio/mp3  audio/x-mpeg
audio/mp4  audio/m4a  audio/x-m4a  video/mp4
audio/ogg  video/ogg  application/ogg  audio/opus
audio/webm  video/webm
audio/flac  audio/x-flac
audio/3gpp  video/3gpp  audio/amr

Why some start with video/. WebM, MP4 and Ogg are containers, and the type detected from the content does not say whether there is video inside: a voice-only recording is detected as video/webm or video/mp4. Browsers record in these containers (Chrome in WebM, Safari in MP4), so refusing video/webm would mean refusing dictation from a browser. Until 18 August 2026 it did.

When a file is refused, you get 400 and the message names the type that was detected and lists the formats we accept. It is in Italian today:

json
{"error":{"message":"Formato audio non supportato per «dictation.avi»: il contenuto del file è stato riconosciuto come video/x-msvideo. Accettiamo WAV, MP3, M4A, MP4, Ogg, WebM, FLAC e 3GPP. …","type":"invalid_request_error"}}

Audio without sound

A file without sound returns empty text, with status 200, without asking the model. It is not an error: the file is valid, there is just nothing to transcribe. On digital silence Whisper answers "Thank you." with full confidence, and on a dictation path an accidental tap on the microphone would otherwise become a command. The threshold and the reasoning are in What we change in responses.

How it works today:

  • Only 16-bit PCM WAV is measured. Compressed formats always go to the model.
  • Only the beginning of the file is measured: the first 32,000 samples, which is two seconds of 16 kHz mono audio and less than a second at 44.1 or 48 kHz. A WAV that starts with that much near-silence (below about -50 dBFS) is treated as silent, even if someone speaks afterwards. If your recordings can start with a long pause, trim it, or send a compressed format.
  • With verbose_json, the empty answer has duration set to 0 and language set to the one you sent, or it if you sent none.
  • With response_format=text, a silent WAV returns 500 today instead of an empty body. It is a defect on our side; with json and verbose_json the answer is correct.
  • Only this endpoint does it. The app endpoint /api/v1/audio/transcribe sends every file to the model.

Speech

POST https://api.ai.itsincom.org/v1/audio/speech, with a JSON body. The answer is the audio file.

bash
curl https://api.ai.itsincom.org/v1/audio/speech \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "voice": "amy",
    "input": "Hello! This is a test of speech synthesis.",
    "response_format": "mp3"
  }' \
  --output hello.mp3

Python, with the client from the transcription example:

python
res = client.audio.speech.create(
    model="tts-1",          # the SDK requires it; the value is ignored
    voice="amy",
    input="Hello! This is a test of speech synthesis.",
    response_format="mp3",  # the default here is wav, not mp3 as at OpenAI
)
res.write_to_file("hello.mp3")

TypeScript (Node 20 or later)

typescript
// speak.mts. Run: API_KEY=sk-… npx tsx speak.mts
import { writeFile } from "node:fs/promises";

const res = await fetch("https://api.ai.itsincom.org/v1/audio/speech", {
  method: "POST",
  headers: {
    Authorization: `Bearer ${process.env.API_KEY}`,
    "Content-Type": "application/json",
  },
  body: JSON.stringify({ voice: "amy", input: "Hello!", response_format: "mp3" }),
});
if (!res.ok) {
  const body = await res.json();
  throw new Error(body.error?.message ?? body.message);
}
await writeFile("hello.mp3", new Uint8Array(await res.arrayBuffer()));

Fields

Field
input required The text, up to 5,000 characters.
voice optional One of our voices or an OpenAI voice name. Without it, the default voice of language.
language optional Ours, not OpenAI's. A two-letter code that picks the default voice and the language OpenAI voice names are served in. Default it.
response_format optional wav (default) or mp3. Other formats give 400.
model optional Accepted and ignored, whatever its value: there is one engine.
speed optional Accepted between 0.5 and 2.0, and ignored: the audio always comes at normal speed.

Here too, a field that is not in the table is ignored, not refused.

Voices

Voice Language
paola Italian female; default for it
riccardo Italian male; a lower-quality voice model (x_low)
amy English (US) female; default for en
ryan English (US) male
thorsten German (Germany) male; default for de
siwis French (France) female; default for fr

Voice names are case-sensitive. An unknown voice gives 400, and the message lists the valid ones.

OpenAI voice names work too, in upper or lower case. They carry only the character of the voice; the language comes from language:

  • alloy, nova, shimmer, coral, sage → the female voice of the language;
  • echo, onyx, fable, ash, ballad, verse → the male voice.

German has only a male voice and French only a female one: there you get that voice whatever the name asks for. So alloy with "language": "en" is amy, and alloy with no language is paola.

Other languages. There is no voice for Romansh, or for any language not in the table: with such a language, the text is read by an Italian voice.

What comes back

The audio, as Content-Type: audio/wav or audio/mpeg. The type always matches the content: if the MP3 encoder were missing you would get 502, never a WAV labelled as MP3. MP3 is several times smaller than WAV, which is the reason to ask for it on a mobile network.

There is no streaming: the whole audio is generated before the answer starts. For spoken replies, split the text into sentences and request them one at a time, so the first can play while the next is generated.


Limits

  • Size: 25 MB per audio file; 5,000 characters per speech request.
  • Per key: the per-minute limit of the request's tier, on one counter with all your other /v1 calls.
  • Per client address: 30 requests per minute on transcriptions, 120 on speech. The counter belongs to the address, not to the key, and it is shared with the other /v1 endpoints: once an address has made 30 requests to any of them in the same minute, its transcriptions get 429 until the minute is over. People behind one public address, such as an office or a school, share it.

How both limits count, and how to tell their 429 apart, is in Rate limits.

Errors

Status type When
400 invalid_request_error Something in the request is refused: a missing or invalid field, a file over 25 MB or in a format we do not accept, word timings without verbose_json, an unknown voice, a text over 5,000 characters. The message says what.
401 invalid_request_error, with code invalid_api_key The key is missing, malformed, revoked or expired, or the account is disabled.
429 see below Over one of the limits.
502 bad_gateway The transcription or speech engine did not answer, or answered with an error.

On 429, Retry-After says how many seconds to wait. The two limits answer with different bodies: the per-key limit with the usual error object ("type": "rate_limit_exceeded"), the per-address limit with only {"message": "Too Many Attempts."}.

On 502, retry a few times with growing pauses. When the engine refuses a file, that also comes back as 502 today: if the same file keeps failing, the file is the likely cause.

Messages are partly in Italian today: branch on the status and on type, not on the text. To report a problem, write to segreteria@itsincom.it with the request_id of the error if there is one, otherwise the X-Request-Id header of the response. The other statuses of the API are in Errors.

Price

Transcription is not billed today. This page used to publish a price per minute; on 1 October 2026 we found that the configured price, the public price list and the invoice did not agree. No audio price is published here until the three match: see Pricing.

Audio calls that reach an engine appear in your usage records, in the dashboard.


A voice conversation

Transcribe what the user said, answer with a chat model, read the answer aloud:

python
import os

from openai import OpenAI

client = OpenAI(base_url="https://api.ai.itsincom.org/v1", api_key=os.environ["API_KEY"])

with open("question.wav", "rb") as f:
    heard = client.audio.transcriptions.create(model="whisper-1", file=f, language="en")

if not heard.text:
    raise SystemExit("Nothing was said.")  # a silent WAV gives empty text

answer = client.chat.completions.create(
    model="gemma-4-26b",
    messages=[{"role": "user", "content": heard.text}],
    max_tokens=300,  # a short answer: speech takes up to 5,000 characters
).choices[0].message.content

client.audio.speech.create(
    model="tts-1",
    voice="amy",
    input=answer,
    response_format="mp3",
).write_to_file("answer.mp3")

Names and terms. If the audio contains names or technical terms, put them in prompt, spelled the way you want them back:

python
with open("meeting.m4a", "rb") as f:
    client.audio.transcriptions.create(
        model="whisper-1-hd",
        file=f,
        prompt="Kowalczyk, Qdrant, rerank, embeddings",
    )

In apps that sign their users in

Apps that sign their users in use the app API at https://my.ai.itsincom.org/api/v1, with a user token instead of a key: see Authentication. The same engines are there, with other names:

App API Key API
POST /api/v1/audio/transcribe, file in audio POST /v1/audio/transcriptions, file in file
POST /api/v1/audio/synthesize, text in text, format in format POST /v1/audio/speech, input, response_format
GET /api/v1/audio/voices none
bash
curl https://my.ai.itsincom.org/api/v1/audio/transcribe \
  -H "Authorization: Bearer $USER_TOKEN" \
  -F audio=@note.m4a \
  -F with_segments=true
json
{ "text": "…", "language": "en", "duration_seconds": 3.4, "model": "whisper-1", "segments": [] }
bash
curl https://my.ai.itsincom.org/api/v1/audio/voices \
  -H "Authorization: Bearer $USER_TOKEN"
json
{
  "voices": [
    { "name": "paola", "language": "it", "gender": "female" },
    { "name": "riccardo", "language": "it", "gender": "male" }
  ],
  "defaults_by_language": { "it": "paola", "en": "amy", "de": "thorsten", "fr": "siwis", "rm": "paola" }
}

How they differ from the key API:

  • transcribe accepts audio, model, language, prompt and with_segments. There is no temperature, no per-word timing and no silence check, and an empty field is refused instead of counting as absent.
  • synthesize accepts text, voice, language and format (wav or mp3). Without language, the default voice follows the language of the user's account. The answer adds X-Voice, the voice used, and X-Duration-Seconds, an estimate of the length.
  • Errors have the app shape, {"error": {"code": "…", "message": "…"}}: 422 validation_error with the fields, 400 invalid_voice, 502 upstream_unavailable.
  • Limits are per client address, like those of /v1: 30 transcriptions and 60 syntheses per minute, on one counter shared with the other app endpoints that have a limit (the list). voices has none.