Skip to content
kenari.

Transcriptions

Turn an audio file into text. The endpoint is compatible with OpenAI Audio Transcriptions, so the official OpenAI SDKs work with the base URL https://kenari.id/v1. Only speech-to-text models are served here. Sending any other model returns 400.

POST /v1/audio/transcriptions

Find the speech-to-text models with:

Terminal window
curl -s https://kenari.id/v1/models \
| jq '.data[] | select(.endpoints | index("audio_transcription")) | {id, pricing_lines}'

The request is multipart/form-data.

FieldTypeRequiredDescription
modelstringyesSpeech-to-text model id, for example whisper-large-v3-turbo.
filefileyesThe audio to transcribe.
languagestringnoLanguage code of the audio, for example id or en. Detected automatically when left out.
promptstringnoText that guides spelling of names and terms or the style of the transcript.
temperaturenumbernoSampling temperature. Passed to the provider unchanged.
response_formatstringnojson (default), verbose_json or text.

kenari does not check the audio format. Which formats work depends on the model. The whole request, file included, can be up to 32 MiB, and a larger body is rejected with 413. A provider has 120 seconds to answer, so split very long recordings into several files.

timestamp_granularities is not supported. Other fields are ignored.

With json, or when response_format is left out, the response holds only the text:

{ "text": "Hello, this is an audio transcription." }

With verbose_json, the response adds task, language, duration in seconds, and segments with timing for each segment. Fields the model does not report are left out.

{
"text": "Hello, this is an audio transcription.",
"task": "transcribe",
"language": "en",
"duration": 8.5,
"segments": [
{
"id": 0,
"seek": 0,
"start": 0.0,
"end": 3.2,
"text": "Hello, this is an audio transcription.",
"tokens": [50364, 1234],
"temperature": 0.0,
"avg_logprob": -0.28,
"compression_ratio": 1.23,
"no_speech_prob": 0.008
}
]
}

kenari rebuilds the response from the fields shown above, so any other field the model returns is not passed through. A words array appears only if the model returns one.

With text, the response body is the plain transcript, not JSON.

srt and vtt are not supported. They return the default json response.

The examples use recording.mp3. Replace it with the path to your own recording.

Terminal window
curl https://kenari.id/v1/audio/transcriptions \
-H "Authorization: Bearer $KENARI_API_KEY" \
-F "file=@recording.mp3" \
-F "model=whisper-large-v3-turbo" \
-F "language=en"
import os
from openai import OpenAI
client = OpenAI(
base_url="https://kenari.id/v1",
api_key=os.environ["KENARI_API_KEY"],
)
with open("recording.mp3", "rb") as audio:
transcript = client.audio.transcriptions.create(
model="whisper-large-v3-turbo",
file=audio,
language="en",
)
print(transcript.text)
import fs from "node:fs";
import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://kenari.id/v1",
apiKey: process.env.KENARI_API_KEY,
});
const transcript = await client.audio.transcriptions.create({
model: "whisper-large-v3-turbo",
file: fs.createReadStream("recording.mp3"),
language: "en",
});
console.log(transcript.text);

Transcription is billed per second of audio, rounded up. The reservation is based on the duration kenari estimates from the file size, so a request can fail with 402 even when the audio is short. See How billing works for how the reservation and the charge work, and the model’s price in Models and pricing.

StatusCodeWhen
400bad_requestThe model is not a speech-to-text model, file or model is missing, the multipart body is malformed or cut off, or the provider rejected the audio.
413noneThe request body is above 32 MiB. The response is plain text.

See Errors for the shared codes, such as insufficient_balance and all_providers_failed.