Skip to content
kenari.

Prompt caching

Prompt caching lets a provider reuse the start of a prompt it has already processed. When many requests begin with the same text, such as a long system prompt, a set of tools or a large document, the provider reads the repeated part from the cache at a lower price. This usually shortens the time until the response begins. The provider does the caching. kenari passes the cache counts through in usage and bills them on their own price lines.

A cache matches the beginning of a prompt, so the order of your prompt decides what can be reused. Put content that stays the same first: system prompt, tool definitions, reference documents. Put what changes last: the user’s question. A change anywhere in the cached part breaks the match from that point on.

What you send depends on the API format you call. Some providers cache a repeated prefix on their own. Others cache only the parts marked with cache_control.

Format you callWhat to send
Chat completionsNothing. If the provider needs marks, kenari marks the system prompt and the last message for you.
ResponsesNothing. kenari marks the system prompt and the last message for you when the provider needs marks.
MessagesYour own cache_control marks. They pass through untouched, and kenari adds none. A request that also lists a kenari: server tool is rebuilt by kenari, so your marks are replaced by the two automatic ones.

You cannot see which style a provider uses, so read usage after the request to learn whether anything was cached. On Messages, cache_control is ignored when the provider caches on its own. Providers also set a minimum prompt length for caching, and a shorter prompt is never cached.

This Messages request marks a long system prompt for caching. handbook.txt stands for any text file long enough to reach the provider’s minimum:

import os
import anthropic
client = anthropic.Anthropic(
base_url="https://kenari.id",
api_key=os.environ["KENARI_API_KEY"],
)
with open("handbook.txt") as file:
handbook = file.read()
message = client.messages.create(
model="claude-sonnet-5-5",
max_tokens=512,
system=[
{"type": "text", "text": handbook, "cache_control": {"type": "ephemeral"}},
],
messages=[{"role": "user", "content": "How many days of annual leave do I get?"}],
)
print(message.usage)

Send the same request again shortly after. If the model caches prompts, message.usage.cache_read_input_tokens is above 0 on the second response. If it stays 0, nothing was cached.

Every format reports cache reads in its own field.

FormatCache readsCache writes
Chat completionsusage.prompt_tokens_details.cached_tokensNot reported
Messagesusage.cache_read_input_tokensusage.cache_creation_input_tokens
Responsesusage.input_tokens_details.cached_tokensNot reported

The formats count differently:

  • On Chat completions and Responses, prompt_tokens or input_tokens is the whole prompt, and cached_tokens is the part of it that came from the cache.
  • On Messages, input_tokens excludes the cached tokens. The whole prompt is input_tokens + cache_read_input_tokens + cache_creation_input_tokens.

The Messages cache fields appear only when the provider reports them. On a non-streaming Chat completions response, cached_tokens is 0 when nothing was cached. In a stream, prompt_tokens_details is present only when something was cached.

Each model has Cache read and Cache write prices next to Input and Output, under pricing in GET /v1/models. kenari bills cached tokens at those prices. How a missing price falls back to Input, and the full formula, are in How billing works.

Some plans do not count cached tokens against your quota for the models on their free cache list. Catalog prices do not change, but eligible cache tokens are left out of quota usage and out of the balance charge that follows a spent window. See Subscriptions.

To make follow-up requests reuse the same cached prefix, send a prompt_cache_key string on Chat completions or Responses, and use the same value for every request of one conversation. A valid, nonempty key you send wins over the one kenari would derive from the conversation. kenari trims surrounding whitespace, and ignores an empty key or one longer than 512 bytes, so the derived key applies instead.

Two habits keep caching working:

  • Do not change the system prompt or the first user message between turns.
  • Start the conversation with text. A conversation whose first user message has no text, such as an image alone, gets no automatic key, so send a prompt_cache_key yourself.
  • A cache entry expires after a time limit that the provider sets. The first request after that is a miss.
  • Not every model caches prompts. Check usage instead of assuming.
  • The first request of a prompt is a miss by definition. Hits start with the second request.
  • A hit is never guaranteed. Read usage to see what happened, and do not build logic on a request being cached.