API reference
Audio: transcription and speech
POST /v1/audio/transcriptions and POST /v1/audio/speech, in the shape of OpenAI's: what they accept, what they return, and what does not work yet.
Last updated: 2026-10-03
Two endpoints, in the shape of OpenAI's:
| Endpoint | Does | Engine |
|---|---|---|
POST /v1/audio/transcriptions |
audio → text | Whisper large-v3-turbo (whisper-1) or Whisper large-v3 (whisper-1-hd) |
POST /v1/audio/speech |
text → audio | Piper, six voices in four languages |
The OpenAI SDKs work with them: change base_url and the key. The examples
below use:
export API_KEY="sk-…"
Transcription
POST https://api.ai.itsincom.org/v1/audio/transcriptions, as a multipart form.
curl https://api.ai.itsincom.org/v1/audio/transcriptions \
-H "Authorization: Bearer $API_KEY" \
-F file=@message.m4a \
-F model=whisper-1 \
-F language=en
{ "text": "Hello, I wanted to ask whether we can meet tomorrow." }
Python
import os
from openai import OpenAI
client = OpenAI(base_url="https://api.ai.itsincom.org/v1", api_key=os.environ["API_KEY"])
with open("message.m4a", "rb") as f:
result = client.audio.transcriptions.create(
model="whisper-1",
file=f,
language="en", # optional: detected from the audio if omitted
)
print(result.text)
TypeScript (Node 20 or later, on your server: the key must not reach a browser)
// transcribe.mts. Run: API_KEY=sk-… npx tsx transcribe.mts
import { openAsBlob } from "node:fs";
const form = new FormData();
form.append("file", await openAsBlob("message.m4a"), "message.m4a");
form.append("model", "whisper-1");
const res = await fetch("https://api.ai.itsincom.org/v1/audio/transcriptions", {
method: "POST",
headers: { Authorization: `Bearer ${process.env.API_KEY}` },
body: form,
});
const body = await res.json();
if (!res.ok) throw new Error(body.error?.message ?? body.message);
console.log(body.text);
Fields
| Field | ||
|---|---|---|
file |
required | The audio, up to 25 MB. See accepted formats. |
model |
optional | whisper-1 (default) runs Whisper large-v3-turbo; whisper-1-hd runs Whisper large-v3. Any other value gives 400. |
language |
optional | A two-letter code: it, en, de, fr… Without it, Whisper detects the language. |
prompt |
optional | Up to 1,024 characters of context: names and terms you expect, spelled the way you want them back. |
response_format |
optional | json (default), text or verbose_json. |
temperature |
optional | From 0 to 1. |
timestamp_granularities[] |
optional | word, segment or both. Only with verbose_json. |
prompt, language and response_format sent empty count as absent. SDKs
often send an empty variable instead of leaving the field out, and that used to
make the whole transcription fail.
What comes back
With json, the default, the object shown above: {"text": "…"}.
With text, the bare text, as Content-Type: text/plain; charset=utf-8.
With verbose_json, also the language, the duration in seconds and the
segments with their timings. Segments are passed on as the engine returns them;
the example shows only some of their fields:
{
"task": "transcribe",
"language": "en",
"duration": 4.32,
"text": "Hello, how are you today? I wanted to ask whether we can meet tomorrow.",
"segments": [
{ "id": 0, "start": 0.0, "end": 2.1, "text": " Hello, how are you today?" },
{ "id": 1, "start": 2.1, "end": 4.32, "text": " I wanted to ask whether we can meet tomorrow." }
]
}
Per-word timings
timestamp_granularities[]=word adds words at the top level of the response,
with the start, end and probability of every word. It is what you need to
highlight the exact point in the audio, cut a recording, or show where
recognition was uncertain.
curl https://api.ai.itsincom.org/v1/audio/transcriptions \
-H "Authorization: Bearer $API_KEY" \
-F file=@note.wav \
-F model=whisper-1 \
-F response_format=verbose_json \
-F "timestamp_granularities[]=word"
{
"words": [
{ "start": 0.0, "end": 0.42, "word": " Hello,", "probability": 0.97 },
{ "start": 0.42, "end": 0.71, "word": " how", "probability": 0.99 }
]
}
Timings exist only inside verbose_json: the other two forms have nowhere to put
them. Asking for them with another format gives 400, instead of an answer
without timings that you could not tell apart from audio in which no words were
heard. Until 18 August 2026 the field was accepted and not forwarded.
A measured example, with 55 timed words, is in Three examples, with measured numbers.
Which model
whisper-1 is the default and the faster of the two. If it makes mistakes on
your audio — noise, phone calls, strong accents — run the same files through
whisper-1-hd and compare.
The model is loaded on demand: the first request after a long pause can take several seconds longer.
Accepted formats
The check looks at the content of the file, not at its name or at the
Content-Type you declare. A renamed file does not pass; a file with the wrong
Content-Type passes if what is inside is a format we accept.
These are the accepted types, as detected from the content:
audio/wav audio/wave audio/x-wav audio/vnd.wave
audio/mpeg audio/mp3 audio/x-mpeg
audio/mp4 audio/m4a audio/x-m4a video/mp4
audio/ogg video/ogg application/ogg audio/opus
audio/webm video/webm
audio/flac audio/x-flac
audio/3gpp video/3gpp audio/amr
Why some start with video/. WebM, MP4 and Ogg are containers, and the type
detected from the content does not say whether there is video inside: a
voice-only recording is detected as video/webm or video/mp4. Browsers record
in these containers (Chrome in WebM, Safari in MP4), so refusing video/webm
would mean refusing dictation from a browser. Until 18 August 2026 it did.
When a file is refused, you get 400 and the message names the type that
was detected and lists the formats we accept. It is in Italian today:
{"error":{"message":"Formato audio non supportato per «dictation.avi»: il contenuto del file è stato riconosciuto come video/x-msvideo. Accettiamo WAV, MP3, M4A, MP4, Ogg, WebM, FLAC e 3GPP. …","type":"invalid_request_error"}}
Audio without sound
A file without sound returns empty text, with status 200, without asking
the model. It is not an error: the file is valid, there is just nothing to
transcribe. On digital silence Whisper answers "Thank you." with full
confidence, and on a dictation path an accidental tap on the microphone would
otherwise become a command. The threshold and the reasoning are in
What we change in responses.
How it works today:
- Only 16-bit PCM WAV is measured. Compressed formats always go to the model.
- Only the beginning of the file is measured: the first 32,000 samples, which is two seconds of 16 kHz mono audio and less than a second at 44.1 or 48 kHz. A WAV that starts with that much near-silence (below about -50 dBFS) is treated as silent, even if someone speaks afterwards. If your recordings can start with a long pause, trim it, or send a compressed format.
- With
verbose_json, the empty answer hasdurationset to0andlanguageset to the one you sent, oritif you sent none. - With
response_format=text, a silent WAV returns500today instead of an empty body. It is a defect on our side; withjsonandverbose_jsonthe answer is correct. - Only this endpoint does it. The app endpoint
/api/v1/audio/transcribesends every file to the model.
Speech
POST https://api.ai.itsincom.org/v1/audio/speech, with a JSON body. The answer is the audio
file.
curl https://api.ai.itsincom.org/v1/audio/speech \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"voice": "amy",
"input": "Hello! This is a test of speech synthesis.",
"response_format": "mp3"
}' \
--output hello.mp3
Python, with the client from the transcription example:
res = client.audio.speech.create(
model="tts-1", # the SDK requires it; the value is ignored
voice="amy",
input="Hello! This is a test of speech synthesis.",
response_format="mp3", # the default here is wav, not mp3 as at OpenAI
)
res.write_to_file("hello.mp3")
TypeScript (Node 20 or later)
// speak.mts. Run: API_KEY=sk-… npx tsx speak.mts
import { writeFile } from "node:fs/promises";
const res = await fetch("https://api.ai.itsincom.org/v1/audio/speech", {
method: "POST",
headers: {
Authorization: `Bearer ${process.env.API_KEY}`,
"Content-Type": "application/json",
},
body: JSON.stringify({ voice: "amy", input: "Hello!", response_format: "mp3" }),
});
if (!res.ok) {
const body = await res.json();
throw new Error(body.error?.message ?? body.message);
}
await writeFile("hello.mp3", new Uint8Array(await res.arrayBuffer()));
Fields
| Field | ||
|---|---|---|
input |
required | The text, up to 5,000 characters. |
voice |
optional | One of our voices or an OpenAI voice name. Without it, the default voice of language. |
language |
optional | Ours, not OpenAI's. A two-letter code that picks the default voice and the language OpenAI voice names are served in. Default it. |
response_format |
optional | wav (default) or mp3. Other formats give 400. |
model |
optional | Accepted and ignored, whatever its value: there is one engine. |
speed |
optional | Accepted between 0.5 and 2.0, and ignored: the audio always comes at normal speed. |
Here too, a field that is not in the table is ignored, not refused.
Voices
| Voice | Language | |
|---|---|---|
paola |
Italian | female; default for it |
riccardo |
Italian | male; a lower-quality voice model (x_low) |
amy |
English (US) | female; default for en |
ryan |
English (US) | male |
thorsten |
German (Germany) | male; default for de |
siwis |
French (France) | female; default for fr |
Voice names are case-sensitive. An unknown voice gives 400, and the message
lists the valid ones.
OpenAI voice names work too, in upper or lower case. They carry only the
character of the voice; the language comes from language:
alloy,nova,shimmer,coral,sage→ the female voice of the language;echo,onyx,fable,ash,ballad,verse→ the male voice.
German has only a male voice and French only a female one: there you get that
voice whatever the name asks for. So alloy with "language": "en" is amy,
and alloy with no language is paola.
Other languages. There is no voice for Romansh, or for any language not in
the table: with such a language, the text is read by an Italian voice.
What comes back
The audio, as Content-Type: audio/wav or audio/mpeg. The type always matches
the content: if the MP3 encoder were missing you would get 502, never a WAV
labelled as MP3. MP3 is several times smaller than WAV, which is the reason to
ask for it on a mobile network.
There is no streaming: the whole audio is generated before the answer starts. For spoken replies, split the text into sentences and request them one at a time, so the first can play while the next is generated.
Limits
- Size: 25 MB per audio file; 5,000 characters per speech request.
- Per key: the per-minute limit of the request's tier, on one counter with all your other
/v1calls. - Per client address:
30requests per minute on transcriptions,120on speech. The counter belongs to the address, not to the key, and it is shared with the other/v1endpoints: once an address has made 30 requests to any of them in the same minute, its transcriptions get429until the minute is over. People behind one public address, such as an office or a school, share it.
How both limits count, and how to tell their 429 apart, is in
Rate limits.
Errors
| Status | type |
When |
|---|---|---|
400 |
invalid_request_error |
Something in the request is refused: a missing or invalid field, a file over 25 MB or in a format we do not accept, word timings without verbose_json, an unknown voice, a text over 5,000 characters. The message says what. |
401 |
invalid_request_error, with code invalid_api_key |
The key is missing, malformed, revoked or expired, or the account is disabled. |
429 |
see below | Over one of the limits. |
502 |
bad_gateway |
The transcription or speech engine did not answer, or answered with an error. |
On 429, Retry-After says how many seconds to wait. The two limits answer
with different bodies: the per-key limit with the usual error object
("type": "rate_limit_exceeded"), the per-address limit with only
{"message": "Too Many Attempts."}.
On 502, retry a few times with growing pauses. When the engine refuses a file,
that also comes back as 502 today: if the same file keeps failing, the file is
the likely cause.
Messages are partly in Italian today: branch on the status and on type, not on
the text. To report a problem, write to segreteria@itsincom.it with the request_id
of the error if there is one, otherwise the X-Request-Id header of the
response. The other statuses of the API are in Errors.
Price
Transcription is not billed today. This page used to publish a price per minute; on 1 October 2026 we found that the configured price, the public price list and the invoice did not agree. No audio price is published here until the three match: see Pricing.
Audio calls that reach an engine appear in your usage records, in the dashboard.
A voice conversation
Transcribe what the user said, answer with a chat model, read the answer aloud:
import os
from openai import OpenAI
client = OpenAI(base_url="https://api.ai.itsincom.org/v1", api_key=os.environ["API_KEY"])
with open("question.wav", "rb") as f:
heard = client.audio.transcriptions.create(model="whisper-1", file=f, language="en")
if not heard.text:
raise SystemExit("Nothing was said.") # a silent WAV gives empty text
answer = client.chat.completions.create(
model="gemma-4-26b",
messages=[{"role": "user", "content": heard.text}],
max_tokens=300, # a short answer: speech takes up to 5,000 characters
).choices[0].message.content
client.audio.speech.create(
model="tts-1",
voice="amy",
input=answer,
response_format="mp3",
).write_to_file("answer.mp3")
Names and terms. If the audio contains names or technical terms, put them in
prompt, spelled the way you want them back:
with open("meeting.m4a", "rb") as f:
client.audio.transcriptions.create(
model="whisper-1-hd",
file=f,
prompt="Kowalczyk, Qdrant, rerank, embeddings",
)
In apps that sign their users in
Apps that sign their users in use the app API at https://my.ai.itsincom.org/api/v1,
with a user token instead of a key: see
Authentication. The same engines are
there, with other names:
| App API | Key API |
|---|---|
POST /api/v1/audio/transcribe, file in audio |
POST /v1/audio/transcriptions, file in file |
POST /api/v1/audio/synthesize, text in text, format in format |
POST /v1/audio/speech, input, response_format |
GET /api/v1/audio/voices |
none |
curl https://my.ai.itsincom.org/api/v1/audio/transcribe \
-H "Authorization: Bearer $USER_TOKEN" \
-F audio=@note.m4a \
-F with_segments=true
{ "text": "…", "language": "en", "duration_seconds": 3.4, "model": "whisper-1", "segments": [] }
curl https://my.ai.itsincom.org/api/v1/audio/voices \
-H "Authorization: Bearer $USER_TOKEN"
{
"voices": [
{ "name": "paola", "language": "it", "gender": "female" },
{ "name": "riccardo", "language": "it", "gender": "male" }
],
"defaults_by_language": { "it": "paola", "en": "amy", "de": "thorsten", "fr": "siwis", "rm": "paola" }
}
How they differ from the key API:
- transcribe accepts
audio,model,language,promptandwith_segments. There is notemperature, no per-word timing and no silence check, and an empty field is refused instead of counting as absent. - synthesize accepts
text,voice,languageandformat(wavormp3). Withoutlanguage, the default voice follows the language of the user's account. The answer addsX-Voice, the voice used, andX-Duration-Seconds, an estimate of the length. - Errors have the app shape,
{"error": {"code": "…", "message": "…"}}:422validation_errorwith the fields,400invalid_voice,502upstream_unavailable. - Limits are per client address, like those of
/v1: 30 transcriptions and 60 syntheses per minute, on one counter shared with the other app endpoints that have a limit (the list).voiceshas none.