Skip to content
kenari.

Text to speech

Turn text into spoken audio. The endpoint is compatible with OpenAI Audio Speech, so the official OpenAI SDKs work with the base URL https://kenari.id/v1. Only text-to-speech models are served here. Sending any other model returns 400.

POST /v1/audio/speech

The body is JSON.

FieldTypeRequiredDescription
modelstringyesText-to-speech model id, for example gemini-3-1-flash-tts.
inputstringyesThe text to speak. See Text length limit.
voicestringnoVoice name. Names differ per model. Leave it out for the model’s default voice. See Choosing a voice.
response_formatstringnomp3, wav, pcm, opus, aac or flac. Not every model produces all of them. Leave it out for the model’s default. See Choosing a format.
speednumbernoSpeaking speed. Applied only by models that support it.
languagestringnoLanguage code, for example id or en. Detected automatically when left out.

The response body is the raw audio, not JSON. The Content-Type follows the format you asked for.

FormatContent-Type
mp3audio/mpeg
wavaudio/wav
pcmaudio/pcm
opusaudio/ogg
aacaudio/aac
flacaudio/flac

Every model has its own voice names, and they are not interchangeable. kenari passes voice to the model as you send it. The voices a model offers are listed in the voices field of that model in GET /v1/models:

Terminal window
curl -s https://kenari.id/v1/models \
| jq '.data[] | select(.endpoints | index("audio_speech")) | {id, voices}'
  • The first entry in voices is the default, used when you leave voice out.
  • Names are matched without regard to case.
  • A model without a voices field has no recorded list. kenari does not check voice for it and passes the value to the model.
  • A voice outside the list returns 400 with the allowed names, and nothing is charged.

There is no separate field for speaking style. Put the direction in front of the text, for example Read this like a newscaster: ....

The formats a model can produce are in the formats field of that model in GET /v1/models:

Terminal window
curl -s https://kenari.id/v1/models \
| jq '.data[] | select(.endpoints | index("audio_speech")) | {id, formats}'
  • The first entry in formats is the default, used when you leave response_format out.
  • Names are matched without regard to case.
  • A format outside the list returns 400 with the allowed names, and nothing is charged. kenari never swaps in a different format, because the Content-Type would then describe the wrong file.
  • A model without a formats field accepts mp3, wav and pcm. Any other value gives you mp3, and the Content-Type is audio/mpeg. So opus, aac and flac work only on models that list them in formats.

Synthesis is not streamed, and each provider enforces its own time limit, so text that is too long cannot finish. When a model records a limit, it appears as max_input_chars in GET /v1/models, and longer text returns 400 before anything is charged. Without that field kenari does not cap the length, but a very long request can still fail on the provider’s time limit. Split long text into several requests.

Terminal window
curl https://kenari.id/v1/audio/speech \
-H "Authorization: Bearer $KENARI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gemini-3-1-flash-tts",
"input": "Hello, welcome to kenari.",
"voice": "Kore",
"response_format": "mp3"
}' \
--output speech.mp3
import os
from openai import OpenAI
client = OpenAI(
base_url="https://kenari.id/v1",
api_key=os.environ["KENARI_API_KEY"],
)
response = client.audio.speech.create(
model="gemini-3-1-flash-tts",
input="Hello, welcome to kenari.",
voice="Kore",
response_format="mp3",
)
response.write_to_file("speech.mp3")
import fs from "node:fs";
import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://kenari.id/v1",
apiKey: process.env.KENARI_API_KEY,
});
const response = await client.audio.speech.create({
model: "gemini-3-1-flash-tts",
input: "Hello, welcome to kenari.",
voice: "Kore",
response_format: "mp3",
});
fs.writeFileSync("speech.mp3", Buffer.from(await response.arrayBuffer()));

Text to speech is billed per 1,000 characters of input, rounded up. The amount is held when the request is accepted and released if no audio is delivered. See How billing works and the model’s price in Models and pricing.

StatusCodeWhen
400bad_requestThe model is not a text-to-speech model, input is empty, voice or response_format is not available for the model, or input is longer than max_input_chars.
402insufficient_balanceThe balance does not cover the request.
503all_providers_failed, upstream_errorNo provider could produce the audio. Nothing is charged. Retry.

See Errors for every other code.