Prompt caching
Prompt caching lets a provider reuse the start of a prompt it has already processed. When many requests begin with the same text, such as a long system prompt, a set of tools or a large document, the provider reads the repeated part from the cache at a lower price. This usually shortens the time until the response begins. The provider does the caching. kenari passes the cache counts through in usage and bills them on their own price lines.
How a request gets cached
Section titled “How a request gets cached”A cache matches the beginning of a prompt, so the order of your prompt decides what can be reused. Put content that stays the same first: system prompt, tool definitions, reference documents. Put what changes last: the user’s question. A change anywhere in the cached part breaks the match from that point on.
What you send depends on the API format you call. Some providers cache a repeated prefix on their own. Others cache only the parts marked with cache_control.
| Format you call | What to send |
|---|---|
| Chat completions | Nothing. If the provider needs marks, kenari marks the system prompt and the last message for you. |
| Responses | Nothing. kenari marks the system prompt and the last message for you when the provider needs marks. |
| Messages | Your own cache_control marks. They pass through untouched, and kenari adds none. A request that also lists a kenari: server tool is rebuilt by kenari, so your marks are replaced by the two automatic ones. |
You cannot see which style a provider uses, so read usage after the request to learn whether anything was cached. On Messages, cache_control is ignored when the provider caches on its own. Providers also set a minimum prompt length for caching, and a shorter prompt is never cached.
This Messages request marks a long system prompt for caching. handbook.txt stands for any text file long enough to reach the provider’s minimum:
import os
import anthropic
client = anthropic.Anthropic( base_url="https://kenari.id", api_key=os.environ["KENARI_API_KEY"],)
with open("handbook.txt") as file: handbook = file.read()
message = client.messages.create( model="claude-sonnet-5-5", max_tokens=512, system=[ {"type": "text", "text": handbook, "cache_control": {"type": "ephemeral"}}, ], messages=[{"role": "user", "content": "How many days of annual leave do I get?"}],)print(message.usage)Send the same request again shortly after. If the model caches prompts, message.usage.cache_read_input_tokens is above 0 on the second response. If it stays 0, nothing was cached.
Read the cache in usage
Section titled “Read the cache in usage”Every format reports cache reads in its own field.
| Format | Cache reads | Cache writes |
|---|---|---|
| Chat completions | usage.prompt_tokens_details.cached_tokens | Not reported |
| Messages | usage.cache_read_input_tokens | usage.cache_creation_input_tokens |
| Responses | usage.input_tokens_details.cached_tokens | Not reported |
The formats count differently:
- On Chat completions and Responses,
prompt_tokensorinput_tokensis the whole prompt, andcached_tokensis the part of it that came from the cache. - On Messages,
input_tokensexcludes the cached tokens. The whole prompt isinput_tokens + cache_read_input_tokens + cache_creation_input_tokens.
The Messages cache fields appear only when the provider reports them. On a non-streaming Chat completions response, cached_tokens is 0 when nothing was cached. In a stream, prompt_tokens_details is present only when something was cached.
How cache is priced
Section titled “How cache is priced”Each model has Cache read and Cache write prices next to Input and Output, under pricing in GET /v1/models. kenari bills cached tokens at those prices. How a missing price falls back to Input, and the full formula, are in How billing works.
Free cache on subscription plans
Section titled “Free cache on subscription plans”Some plans do not count cached tokens against your quota for the models on their free cache list. Catalog prices do not change, but eligible cache tokens are left out of quota usage and out of the balance charge that follows a spent window. See Subscriptions.
Keep a conversation together
Section titled “Keep a conversation together”To make follow-up requests reuse the same cached prefix, send a prompt_cache_key string on Chat completions or Responses, and use the same value for every request of one conversation. A valid, nonempty key you send wins over the one kenari would derive from the conversation. kenari trims surrounding whitespace, and ignores an empty key or one longer than 512 bytes, so the derived key applies instead.
Two habits keep caching working:
- Do not change the system prompt or the first user message between turns.
- Start the conversation with text. A conversation whose first user message has no text, such as an image alone, gets no automatic key, so send a
prompt_cache_keyyourself.
Limits and pitfalls
Section titled “Limits and pitfalls”- A cache entry expires after a time limit that the provider sets. The first request after that is a miss.
- Not every model caches prompts. Check
usageinstead of assuming. - The first request of a prompt is a miss by definition. Hits start with the second request.
- A hit is never guaranteed. Read
usageto see what happened, and do not build logic on a request being cached.