What is an AI gateway

An AI gateway is one door standing in front of many models. Your app talks to a single endpoint with a single key, and the job of picking a path belongs to the gateway.
The more useful question is usually not what it is but when you need one. That answer starts to show up the moment your app calls models from more than one provider.
The tiring part is not the model
There are plenty of good models. What wears you down is the admin around them, which nobody warns you about.
It usually starts with OpenAI, because you were already using that SDK. Then some task turns out nicer on Anthropic, so you sign up again. After that comes Google, DeepSeek, and whoever launches next month. Every provider arrives with its own entourage: its own dashboard, its own key, and its own limits.
Seven accounts means seven places a balance can quietly hit zero, seven keys to rotate the day one of them leaks, and seven invoices falling due on their own schedules. The failure modes are all different too, and you usually find out when an important request comes back an error at your busiest hour.
None of that is the worst part, though. One provider goes down and your whole feature goes quiet with it, not because you ran out of ideas, but because somebody else owns the only path your app has.
Work should not stop because one provider is having a bad night.
A gateway, without the jargon
Your app never needs to know who is healthy this morning. You send a request to one address, and the gateway reads which model you asked for, picks a backend that can answer, and hands the reply back in a shape you already recognize. When the first path is blocked by a limit, a timeout, or an error, it moves to the next one without you swapping a key and without you writing try/except for every brand you have ever used.
There are two shapes, and both of them deliberately copy something that already exists:
- OpenAI-compatible. The base URL replaces
https://api.openai.com/v1, so the SDK you already use keeps working. The address and the key are the only things that change. - Anthropic-style. The Messages endpoint, with Anthropic-shaped requests. This is what Claude Code and other clients that speak that protocol expect.
A good gateway will not make you learn a new format. It stands in front of the ones you already know.
Straight to the provider vs through a gateway
| Straight to the provider | Through a gateway | |
|---|---|---|
| Accounts | One account per brand | One account at the gateway |
| Keys | One key per brand, sometimes per project | One key for many models |
| Switch models | Change the base URL, sometimes the SDK | Change the model field |
| If one provider dies | Work stops, or you write your own fallback | The gateway switches paths |
| Billing | One invoice per brand | One billing system, one currency |
| Analytics | Split across dashboards | One request history |
Going straight to a provider makes sense while you are on one brand. A gateway starts earning its keep the moment a second brand shows up, or the moment those seven tabs start to feel like unpaid work.
One thing worth being clear about: a gateway is not a new model and it is no smarter than Claude or GPT. All it handles is routing, keys, and fallback.
What a gateway handles, and what it leaves alone
What it handles:
- One authentication. Every request of yours uses one gateway key.
- A model catalog. One list of ids pulled from a public endpoint rather than hardcoded into your repo.
- Routing. Pick a healthy path, then move when that path fails.
- Metering. Input tokens, output tokens, and cost, recorded per request.
What it leaves alone:
- Your prompt is still your prompt.
- Upstream content policy applies exactly as before.
- The model still belongs to its provider, and the gateway does not modify it.
So if what you need really is one specific brand, with no fallback and no shared usage record, call that provider directly. A gateway earns its place when you would rather not have your work depend on a single provider.
One key, two API shapes
On paper, “OpenAI-compatible” sounds like enough of an explanation. In practice there is a small trap that catches almost everyone at least once.
OpenAI clients, whether through the SDK or a curl to /chat/completions, use a base URL that already includes /v1, something like https://example.id/v1, and the SDK appends /chat/completions behind it. Anthropic clients are less consistent. Claude Code appends /v1 on its own, so the base URL you configure has to be without /v1, while parts of the Anthropic SDK docs use a base URL that already ends in /v1. Pair them wrong and what you get back is 405, not a friendly message explaining where you went astray.
Two shapes to remember:
- OpenAI-compatible: base URL ends in
/v1. - Claude Code (Anthropic protocol): base URL without
/v1.
One key covers both, and the only thing that changes is the address you paste.
One request, two shapes, one key
The same request can go to two different paths without touching the key at all.
OpenAI-compatible:
curl https://kenari.id/v1/chat/completions \
-H "Authorization: Bearer kn-..." \
-H "Content-Type: application/json" \
-d '{"model":"step-3-7-flash:free","messages":[{"role":"user","content":"Hello!"}]}'
Anthropic-style, through the Messages endpoint, as in the docs:
curl https://kenari.id/v1/messages \
-H "Authorization: Bearer kn-..." \
-H "Content-Type: application/json" \
-d '{"model":"step-3-7-flash:free","max_tokens":512,"messages":[{"role":"user","content":"Halo!"}]}'
The path, the messages shape, and the reply shape are what differ. The host, the kn- key, and the model id are identical, because the gateway is the one translating to the backend. You should not have to carry seven keys just because the world settled on two API dialects.
If you build your own route in the dashboard, put that route name in the model field instead of a catalog id. Our docs use opus-hemat as the example. Nothing changes on the client side.
kenari: the local instance of this idea
kenari is a gateway of exactly this kind, and the bill lands in Rupiah.
You get one kn- key, an OpenAI-compatible endpoint at https://kenari.id/v1, and for Claude Code you point at the root https://kenari.id with no /v1, because that client adds /v1 itself. The Messages shape lives at POST /v1/messages.
Four things set us apart from a gateway in general:
- Pay per token. Your balance goes down by the tokens you actually spend, not by the month.
- Automatic routing. Every metered model already has a route. kenari picks a healthy backend, then moves when that path is blocked, hits a limit, or suddenly gets slow.
- Last-resort fallback. When the external paths run out, the kenari pool can be the final step, and there is a built-in
kenari-freeroute pointing at free models. - BYOK. Already have your own OpenAI or Anthropic key? Register it. Requests served by your key do not debit your kenari balance, though the analytics still record them.
It is one instance of the idea: one endpoint, one key, one bill.
One thing to note before you start. The model list changes, so please do not hardcode it. GET https://kenari.id/v1/models is public and needs no key, and every entry there carries its id, prices in micro-Rupiah per 1 million tokens, and a flag for whether that variant is :free.
:free models and your balance
A single model id can have two faces. step-3-7-flash debits your balance and step-3-7-flash:free does not, at the cost of a per-minute limit and a daily allowance per account. We keep the exact numbers on the plans page, because the operator can change them at any time and we would rather this article not become a source of stale numbers.
The free path is best-effort, so please do not make it your only production path. It exists so you can test before topping up.
Send a paid request on an empty balance and the server answers HTTP 402 insufficient_balance. That is not a bug, the balance really is short.
If you have never sent a request to kenari at all, the steps are in your first request from zero.
When a gateway is overkill
If you call one model, from one provider, on an account that already works, a gateway is just one extra hop. The cost is small, but sometimes you genuinely do not need it.
A gateway starts to pay off once you are changing models more than once a month, or once some provider has already left you waiting. The same goes for when you want a single history for the whole team instead of seven CSVs to reconcile by hand.
The good news is that you do not have to move everything at once. Change the base URL and the key in one script first, and if the answers look the same, move the rest across.