Rate limits
kenari limits how fast you can call the API in a few places. This page lists each limit, the error it returns, and how to retry. For the error formats and every other code, see Errors.
What each limit returns
Section titled “What each limit returns”| Status | Code | Retry-After | What to do |
|---|---|---|---|
| 429 | rate_limit_exceeded | No, except for music | Wait a few seconds and retry. Space your requests out. |
| 429 | free_quota_rpm | Yes | Wait for the Retry-After seconds, then retry. |
| 429 | free_quota_daily | Yes, until 00:00 UTC | Do not retry today. Use a paid model, top up, or wait for the reset. |
| 429 | plan_limit_reached | No | Do not retry in a loop. Turn on PAYG or the free fallback, or wait for the quota window to roll. |
| 429 | upstream_error | No | A provider is rate limiting. Wait and retry. |
| 503 | all_providers_failed | Yes, when providers are paused or busy | Wait for the Retry-After seconds, then retry. |
A Retry-After header holds a number of seconds. When a response has no Retry-After, use exponential backoff.
The codes in this table are the OpenAI-format error codes, returned by Chat completions, Responses and the other endpoints. Messages errors have no code field. They carry error.type, and several of the limits above arrive there as rate_limit_error. On Messages, decide by the HTTP status and the Retry-After header instead. See Errors.
Account requests per minute
Section titled “Account requests per minute”Each account has a per-minute request limit for paid models. It is shared by Chat completions, Messages and Responses. Over the limit, the request fails with 429 and rate_limit_exceeded, and no Retry-After header. The limit is set on your account, so it can differ from one account to another. Requests to :free models count against the free-model limits below instead, and the other endpoints do not count against this limit.
Usage reads
Section titled “Usage reads”The three Usage endpoints share a limit of 60 requests per minute for each account. Over the limit, the request fails with 429 and rate_limit_exceeded, and no Retry-After header. It is counted apart from the account limit above.
Per-tool key creation
Section titled “Per-tool key creation”Each login key can create 20 per-tool keys per minute. Over the limit, the request fails with 429 and rate_limit_exceeded, and no Retry-After header. Listing and revoking keys are not limited.
Free-model limits
Section titled “Free-model limits”Every account has a per-minute limit and a daily quota for :free models. The values depend on your account tier. Free models lists the tiers, and GET /api/public/pricing returns the current numbers.
- Per minute. Over the limit, the request fails with
429andfree_quota_rpm, andRetry-Aftersays how many seconds remain until the limit window resets. - Per day. The daily quota counts successful requests and resets at 00:00 UTC. Over the quota, the request fails with
429andfree_quota_daily, andRetry-Afteris the number of seconds until the reset. Topping up or subscribing raises the quota. - Bursts. Free requests that arrive close together are spaced out and served one after another, not rejected. A sustained flood is rejected with
429andrate_limit_exceeded. - Token budget. A free model can carry a total token budget. When it is used up, that model returns
429andrate_limit_exceededuntil the budget window rolls.
A successful free-model response carries X-RateLimit-Limit, X-RateLimit-Remaining and X-RateLimit-Reset when the model has a per-minute limit. X-RateLimit-Reset is a Unix timestamp in seconds.
Subscription quota windows
Section titled “Subscription quota windows”A subscription covers its models with quota windows. When a window is used up and neither PAYG nor the free fallback is on, requests to covered models fail with 429 and plan_limit_reached, and no Retry-After header. The message of this error is in Indonesian, so branch on code (on Messages, on the status and error.type). Turn on PAYG or the free fallback in your subscription settings, or wait for the window to roll. See Subscriptions.
Music concurrency
Section titled “Music concurrency”Each account can run two music generations at the same time. A third request fails with 429 and rate_limit_exceeded, and Retry-After is set to 60 seconds. Retry once one of your generations finishes. See Music.
Provider limits and outages
Section titled “Provider limits and outages”When a model’s provider rate limits kenari, you get 429 and upstream_error, with no Retry-After header. When a provider times out or fails, you get 503 and upstream_error. When every provider for a model is temporarily paused or at capacity, you get 503 and all_providers_failed, with a Retry-After header. A route or a BYOK fallback can move the request to another provider before you see any of these. See Routing and routes and BYOK.
Shared keys
Section titled “Shared keys”A shared API key is paced as a whole, so the people using a share cannot use up the owner’s account limit. Each client IP address is also limited to 30 requests per minute on one shared key, on /v1 and on the MCP server alike. Over that limit, the response is 429 with the plain-text body too many requests, try again shortly, not JSON. See Authentication and API keys.
Retry safely
Section titled “Retry safely”On Messages, apply the same rules by HTTP status, because there is no code. Retry only errors that can clear on their own: 429 with rate_limit_exceeded, free_quota_rpm or upstream_error, and 503. Do not retry free_quota_daily, plan_limit_reached or any 402, because they stay the same until you change something.
- If the response has a
Retry-Afterheader, wait that many seconds. - Otherwise wait with exponential backoff and jitter, for example 1, 2, 4 and 8 seconds, each plus a random fraction of a second.
- Stop after a few attempts and surface the error.
The official OpenAI and Anthropic SDKs already retry 429 and 5xx responses a couple of times by default and honor Retry-After. Raise their retry count with max_retries if that is enough. This example shows the logic when you want control over it.
import osimport randomimport time
from openai import APIStatusError, OpenAI
client = OpenAI( base_url="https://kenari.id/v1", api_key=os.environ["KENARI_API_KEY"], max_retries=0,)
RETRYABLE = {"rate_limit_exceeded", "free_quota_rpm", "upstream_error", "all_providers_failed"}
def chat(messages, model="step-3-7-flash:free", attempts=5): for attempt in range(attempts): try: return client.chat.completions.create(model=model, messages=messages) except APIStatusError as error: retryable = error.status_code in (429, 503) and error.code in RETRYABLE if not retryable or attempt == attempts - 1: raise retry_after = error.response.headers.get("retry-after") delay = float(retry_after) if retry_after else min(30, 2**attempt) time.sleep(delay + random.random())
reply = chat([{"role": "user", "content": "Hello!"}])print(reply.choices[0].message.content)Do not retry a request that failed in the middle of a stream by resuming it. Send it again as a new request.