Skip to content
kenari.

Music

Make a complete song from lyrics, or an instrumental track from a text description. The response is a JSON object that carries the audio as base64. Only music models are served here. Sending any other model returns 400.

POST /v1/music/generations

Music models are marked with music in the endpoints field of GET /v1/models:

Terminal window
curl -s https://kenari.id/v1/models \
| jq '.data[] | select(.endpoints | index("music")) | {id, max_lyrics_chars, max_prompt_chars, max_duration_secs, pricing_lines}'

If the call returns nothing, no music model is available at the moment.

The body is JSON.

FieldTypeRequiredDescription
modelstringyesMusic model id.
lyricsstringconditionalThe words to sing. Required for a sung track, which is when instrumental is false or left out.
promptstringconditionalStyle and mood of the track. Required when instrumental is true. Optional on a sung track, where it describes the style.
instrumentalbooleannotrue makes a track without vocals. Default false. When true, lyrics is ignored.
response_formatstringnoOnly mp3 is available, and it is the default. wav and pcm return 400.
streambooleannoReserved. true returns 400.
lyrics_optimizerbooleannoReserved. true returns 400.

The song is in data[0].b64_json, and its format is in data[0].format:

{
"data": [
{ "b64_json": "<base64 of the mp3 bytes>", "format": "mp3" }
]
}

data always has exactly one item and format is always mp3. The audio is not a download link, so decode the base64 and write it to a file.

Because a generation is slow, the body can arrive in two shapes.

  • Leading spaces. While the song is generated, the gateway writes one space every 20 seconds to keep the connection open, then sends the JSON. Parse the whole body as JSON, which ignores leading whitespace. Do not treat it as raw audio.
  • Late errors. A failure in the first moments keeps its real status: 400, 402, 429 or 503. A failure after the gateway has already sent 200 comes back as HTTP 200 with a standard error object, for example {"error": {"code": "...", "message": "...", "param": null, "type": "..."}}. Always check for an error field before decoding the audio.

lyrics is the text that gets sung, not a description of the song. Sending “a relaxed pop song about Jakarta” as lyrics makes the model sing that sentence. To describe a track instead, set instrumental to true and describe it in prompt.

Section markers help the model structure a song:

[Verse]
City lights come on one by one
I am walking home with a new story
[Chorus]
Tonight belongs to us
Until the morning comes

Lengths are counted in Unicode characters, not bytes, so accented text is not cut short. The limits differ per model. When a model records them, GET /v1/models shows max_lyrics_chars and max_prompt_chars on its entry. Without those fields, the defaults apply: 3,500 characters for lyrics and 2,000 for prompt. Text over the limit returns 400 before anything is charged.

The length of a song is decided by the model. There is no duration field, so you cannot ask for a shorter or longer song. Where the model records it, max_duration_secs in GET /v1/models gives the longest song in seconds. It is informational: no request is rejected because of it, and real songs are usually shorter.

Most songs take two to three minutes, and longer songs take longer. The gateway keeps the connection open and waits up to 15 minutes before giving up. Set your client timeout to at least 5 minutes, and a generous value such as 15 minutes is safest. A client that stops at 3 minutes will sometimes cancel a song that was about to finish.

The first line picks a music model from the list. Set MODEL yourself if you already know the id.

Terminal window
MODEL=$(curl -s https://kenari.id/v1/models \
| jq -r '[.data[] | select(.endpoints | index("music"))][0].id')
curl https://kenari.id/v1/music/generations \
-H "Authorization: Bearer $KENARI_API_KEY" \
-H "Content-Type: application/json" \
-d "{
\"model\": \"$MODEL\",
\"lyrics\": \"[Verse]\nCity lights come on one by one\n\n[Chorus]\nTonight belongs to us\"
}" \
--max-time 900 \
-o response.json
# Stop on an error object, then decode only a non-empty audio string
jq -e 'has("error") | not' response.json > /dev/null || { jq .error response.json; exit 1; }
jq -er '.data[0].b64_json | select(type == "string" and length > 0)' response.json | base64 -d > song.mp3

An instrumental track:

Terminal window
curl https://kenari.id/v1/music/generations \
-H "Authorization: Bearer $KENARI_API_KEY" \
-H "Content-Type: application/json" \
-d "{
\"model\": \"$MODEL\",
\"instrumental\": true,
\"prompt\": \"relaxed lo-fi, soft piano, night rain\"
}" \
--max-time 900 \
-o response.json
jq -e 'has("error") | not' response.json > /dev/null || { jq .error response.json; exit 1; }
jq -er '.data[0].b64_json | select(type == "string" and length > 0)' response.json | base64 -d > instrumental.mp3

There is no OpenAI SDK method for music, so use requests.

import base64
import os
import requests
headers = {"Authorization": f"Bearer {os.environ['KENARI_API_KEY']}"}
models = requests.get("https://kenari.id/v1/models", timeout=30).json()["data"]
model = next(m["id"] for m in models if "music" in m["endpoints"])
response = requests.post(
"https://kenari.id/v1/music/generations",
headers=headers,
json={
"model": model,
"lyrics": "[Verse]\nCity lights come on one by one\n\n[Chorus]\nTonight belongs to us",
},
timeout=900,
)
body = response.json() # leading spaces are ignored
if "error" in body:
raise RuntimeError(body["error"]["message"])
with open("song.mp3", "wb") as f:
f.write(base64.b64decode(body["data"][0]["b64_json"]))
import fs from "node:fs";
const headers = {
Authorization: `Bearer ${process.env.KENARI_API_KEY}`,
"Content-Type": "application/json",
};
const { data: models } = await (await fetch("https://kenari.id/v1/models")).json();
const model = models.find((m) => m.endpoints.includes("music")).id;
const response = await fetch("https://kenari.id/v1/music/generations", {
method: "POST",
headers,
body: JSON.stringify({
model,
lyrics: "[Verse]\nCity lights come on one by one\n\n[Chorus]\nTonight belongs to us",
}),
signal: AbortSignal.timeout(900_000),
});
const body = JSON.parse(await response.text()); // leading spaces are ignored
if (body.error) throw new Error(body.error.message);
fs.writeFileSync("song.mp3", Buffer.from(body.data[0].b64_json, "base64"));

Music is billed a flat price per song, whatever its length. A failed generation is not charged. Each model’s price per song is in Models and pricing. See How billing works for the rest.

StatusCodeWhen
400bad_requestThe model is not a music model, the field required for your track type is missing, lyrics or prompt is over the limit, or response_format is not mp3.
402insufficient_balanceThe balance does not cover the price of one song.
429rate_limit_exceededThe account’s music concurrency limit is reached. See Rate limits. Wait for the Retry-After header, then retry. Nothing is charged.
503all_providers_failed, upstream_errorNo provider could produce the song. Nothing is charged. Retry.

Errors after the gateway has sent 200 arrive inside the response body. See Response. For every other code, see Errors.