Skip to content
kenari.

Read documents

Attach a PDF to an ordinary chat request, on any model. kenari reads the document before the request is routed and gives the model what it contains: the text of every page that has a text layer, and an image of every page that does not, such as a scan or a page that is only a figure. Page order is kept. No plugin is needed.

This costs nothing extra. What the model receives is billed as its normal input tokens, under whatever pays for the request, so a subscription plan that covers the model covers the document too.

For documents that need real OCR, such as handwriting or poor scans, name the ocr engine. That reading is billed per page from your balance and is never covered by a plan. POST /v1/ocr does the same reading with no model in the loop.

Add a file content part to a user message. file_data is a data URL in the form data:<type>;base64,<contents>. The data URL must name the media type. A data URL without one, such as data:;base64,..., is refused. A plain http URL is not fetched, so send the contents yourself. A request can carry several documents.

The shell builds the request body, because curl cannot read a file into JSON by itself. This needs jq. Replace invoice.pdf with your own file.

Terminal window
{ printf 'data:application/pdf;base64,'; base64 < invoice.pdf | tr -d '\n'; } > invoice.dataurl
jq -n --rawfile data invoice.dataurl '{
model: "step-3-7-flash:free",
messages: [{
role: "user",
content: [
{type: "text", text: "What is the total on this invoice?"},
{type: "file", file: {filename: "invoice.pdf", file_data: $data}}
]
}]
}' | curl https://kenari.id/v1/chat/completions \
-H "Authorization: Bearer $KENARI_API_KEY" \
-H "Content-Type: application/json" \
-d @-

The model in the example is free, so the example runs with no balance. On POST /v1/messages, a document block with a base64 PDF source is read the same way. Responses requests do not take documents. See Images and files as input.

  • application/pdf: pages with a text layer go as text, pages without one go as images.
  • text/plain, text/markdown, text/csv, text/html and application/json: sent as text.
  • Images such as image/png, image/jpeg and image/webp: sent as an image.

Any other type is refused with 400 before anything is charged.

  • About 15 MB per document and 100 pages per PDF.
  • Up to 20 pages without a text layer per request. Each one is sent as an image, roughly 1,000 to 1,500 input tokens. A longer scan is refused with 400: split it, or use the ocr engine.
  • A document that makes the request too long for the model’s context is refused with 400 and no charge. The message gives the estimated token count and the model’s context size, so you can pick a larger model or send fewer pages.
  • A damaged or password-protected PDF is refused with 400.

The file-parser plugin is optional. Name an engine only when you want something other than the default.

"plugins": [{"id": "file-parser", "pdf": {"engine": "ocr"}}]
engineWhat happensCost
none or autoText pages as text, other pages as images. Several documents per requestModel input tokens only
pdf-text (alias cloudflare-ai)Text layer only. A PDF with pages that have no text is refusedModel input tokens only
nativeThe file goes to the model untouched, for a model that reads files itselfModel input tokens only
ocr (alias mistral-ocr)Paid OCR, then the text goes to the model. One document per requestPer page from your balance, plus model tokens

With native, a model that does not read that kind of file itself refuses the request with 400 and no charge, and images must be sent as image_url parts. Any other engine name is refused with 400.

The sections from here to Keys with restrictions describe the ocr engine.

With the ocr engine, the response carries the extracted text next to the model’s answer, in message.annotations. The request is the same as above plus "plugins": [{"id": "file-parser", "pdf": {"engine": "ocr"}}].

{
"choices": [{
"message": {
"role": "assistant",
"content": "The total is Rp 417,000.",
"annotations": [{
"type": "file",
"file": {
"hash": "sha256:a3b6919a...",
"name": "invoice.pdf",
"content": [{"type": "text", "text": "# INVOICE\n..."}],
"confidence": 0.9892,
"low_confidence": false,
"reuse_id": "ocr_8f286cec-075b-4083-a462-fc89170feb4f"
}
}]
}
}]
}

content holds the extracted text, hash identifies it, and reuse_id is present when the reading was stored for reuse. Annotations are added to the non-streaming response only. With stream: true the document is still read and billed, but the text and the reuse_id are not returned, so use a non-streaming request when you need them.

confidence is the mean per-page score from the reading engine. It is null when the engine gives no score. low_confidence is true when a score is below kenari’s threshold, and a missing score does not always set it. Treat it as a warning, not a guarantee. Correctly and incorrectly read documents score in overlapping ranges, and the engine tends to fill an unreadable spot with a plausible value instead of leaving it blank. Have a person check any figure you read from a scan, whatever low_confidence says.

Reuse an OCR reading without another page charge

Section titled “Reuse an OCR reading without another page charge”

Send the reuse_id back in the plugin on the next turn, with no file part:

{
"model": "step-3-7-flash:free",
"messages": [{"role": "user", "content": "Who is the seller?"}],
"plugins": [{"id": "file-parser", "pdf": {"engine": "ocr", "reuse_id": "ocr_8f286cec-075b-4083-a462-fc89170feb4f"}}]
}

Substitute the reuse_id that your own response returned.

That turn does not read the document again and adds no page charge. A request carries a file or a reuse_id, never both. A reuse_id works only for the account that received it, and it works on both endpoints: one from a chat turn can be used on /v1/ocr, and the other way round. If the chat model fails after the document was read, the page charge stays. Sending the same document again on the same account is not billed twice, as long as the first reading was stored, which means its response carried a reuse_id.

The default reading, pdf-text and native have no charge of their own. The document’s text and page images are the model’s input tokens, billed like any other input, and a plan that covers the model covers them.

The ocr engine bills per page read, from your balance, and an image counts as one page. Plans never cover it, even when a plan covers the model that answers. If your balance cannot cover the document, kenari refuses the request with 402 and insufficient_balance before it reads anything, and nothing is charged. One OCR request produces two rows in usage, one for the reading and one for the model’s answer. See How billing works for balance and price dimensions, and Errors for the error.

This section applies to the ocr engine and /v1/ocr. A key restricted to certain models can use OCR only if you ticked Document reading under Paid capabilities when you created it. The default reading only uses model tokens, so a key that may use a model may also attach documents to it. A shared API key can reuse only readings that the same key made, so a holder of a shared key cannot read back documents from your other keys. See Authentication and API keys for key restrictions and sharing.

POST /v1/ocr returns the extracted text and its metadata, with no model generating a reply, so you pay the per-page fee and no token charge. See OCR for the request and response.