Skip to content
kenari.

Rate limits

kenari limits how fast you can call the API in a few places. This page lists each limit, the error it returns, and how to retry. For the error formats and every other code, see Errors.

StatusCodeRetry-AfterWhat to do
429rate_limit_exceededNo, except for musicWait a few seconds and retry. Space your requests out.
429free_quota_rpmYesWait for the Retry-After seconds, then retry.
429free_quota_dailyYes, until 00:00 UTCDo not retry today. Use a paid model, top up, or wait for the reset.
429plan_limit_reachedNoDo not retry in a loop. Turn on PAYG or the free fallback, or wait for the quota window to roll.
429upstream_errorNoA provider is rate limiting. Wait and retry.
503all_providers_failedYes, when providers are paused or busyWait for the Retry-After seconds, then retry.

A Retry-After header holds a number of seconds. When a response has no Retry-After, use exponential backoff.

The codes in this table are the OpenAI-format error codes, returned by Chat completions, Responses and the other endpoints. Messages errors have no code field. They carry error.type, and several of the limits above arrive there as rate_limit_error. On Messages, decide by the HTTP status and the Retry-After header instead. See Errors.

Each account has a per-minute request limit for paid models. It is shared by Chat completions, Messages and Responses. Over the limit, the request fails with 429 and rate_limit_exceeded, and no Retry-After header. The limit is set on your account, so it can differ from one account to another. Requests to :free models count against the free-model limits below instead, and the other endpoints do not count against this limit.

The three Usage endpoints share a limit of 60 requests per minute for each account. Over the limit, the request fails with 429 and rate_limit_exceeded, and no Retry-After header. It is counted apart from the account limit above.

Each login key can create 20 per-tool keys per minute. Over the limit, the request fails with 429 and rate_limit_exceeded, and no Retry-After header. Listing and revoking keys are not limited.

Every account has a per-minute limit and a daily quota for :free models. The values depend on your account tier. Free models lists the tiers, and GET /api/public/pricing returns the current numbers.

  • Per minute. Over the limit, the request fails with 429 and free_quota_rpm, and Retry-After says how many seconds remain until the limit window resets.
  • Per day. The daily quota counts successful requests and resets at 00:00 UTC. Over the quota, the request fails with 429 and free_quota_daily, and Retry-After is the number of seconds until the reset. Topping up or subscribing raises the quota.
  • Bursts. Free requests that arrive close together are spaced out and served one after another, not rejected. A sustained flood is rejected with 429 and rate_limit_exceeded.
  • Token budget. A free model can carry a total token budget. When it is used up, that model returns 429 and rate_limit_exceeded until the budget window rolls.

A successful free-model response carries X-RateLimit-Limit, X-RateLimit-Remaining and X-RateLimit-Reset when the model has a per-minute limit. X-RateLimit-Reset is a Unix timestamp in seconds.

A subscription covers its models with quota windows. When a window is used up and neither PAYG nor the free fallback is on, requests to covered models fail with 429 and plan_limit_reached, and no Retry-After header. The message of this error is in Indonesian, so branch on code (on Messages, on the status and error.type). Turn on PAYG or the free fallback in your subscription settings, or wait for the window to roll. See Subscriptions.

Each account can run two music generations at the same time. A third request fails with 429 and rate_limit_exceeded, and Retry-After is set to 60 seconds. Retry once one of your generations finishes. See Music.

When a model’s provider rate limits kenari, you get 429 and upstream_error, with no Retry-After header. When a provider times out or fails, you get 503 and upstream_error. When every provider for a model is temporarily paused or at capacity, you get 503 and all_providers_failed, with a Retry-After header. A route or a BYOK fallback can move the request to another provider before you see any of these. See Routing and routes and BYOK.

A shared API key is paced as a whole, so the people using a share cannot use up the owner’s account limit. Each client IP address is also limited to 30 requests per minute on one shared key, on /v1 and on the MCP server alike. Over that limit, the response is 429 with the plain-text body too many requests, try again shortly, not JSON. See Authentication and API keys.

On Messages, apply the same rules by HTTP status, because there is no code. Retry only errors that can clear on their own: 429 with rate_limit_exceeded, free_quota_rpm or upstream_error, and 503. Do not retry free_quota_daily, plan_limit_reached or any 402, because they stay the same until you change something.

  1. If the response has a Retry-After header, wait that many seconds.
  2. Otherwise wait with exponential backoff and jitter, for example 1, 2, 4 and 8 seconds, each plus a random fraction of a second.
  3. Stop after a few attempts and surface the error.

The official OpenAI and Anthropic SDKs already retry 429 and 5xx responses a couple of times by default and honor Retry-After. Raise their retry count with max_retries if that is enough. This example shows the logic when you want control over it.

import os
import random
import time
from openai import APIStatusError, OpenAI
client = OpenAI(
base_url="https://kenari.id/v1",
api_key=os.environ["KENARI_API_KEY"],
max_retries=0,
)
RETRYABLE = {"rate_limit_exceeded", "free_quota_rpm", "upstream_error", "all_providers_failed"}
def chat(messages, model="step-3-7-flash:free", attempts=5):
for attempt in range(attempts):
try:
return client.chat.completions.create(model=model, messages=messages)
except APIStatusError as error:
retryable = error.status_code in (429, 503) and error.code in RETRYABLE
if not retryable or attempt == attempts - 1:
raise
retry_after = error.response.headers.get("retry-after")
delay = float(retry_after) if retry_after else min(30, 2**attempt)
time.sleep(delay + random.random())
reply = chat([{"role": "user", "content": "Hello!"}])
print(reply.choices[0].message.content)

Do not retry a request that failed in the middle of a stream by resuming it. Send it again as a new request.