Skip to content
kenari.

Read documents

Attach a PDF or document image to an ordinary chat request. kenari reads the document to text before the request is routed, then gives that text to the model you picked. Any model works, including models that cannot see images at all.

The document itself never reaches the model. What the model receives is the extracted text.

Document reading has two entry points. Use the file-parser plugin on POST /v1/chat/completions when you want a model to answer something about the document. Use the POST /v1/ocr endpoint when you want the text and nothing else, with no model charge on top of the per-page fee.

Add two things to POST /v1/chat/completions: a content part of type file, and the file-parser plugin.

Terminal window
curl https://kenari.id/v1/chat/completions \
-H "Authorization: Bearer $KENARI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "step-3-7-flash",
"messages": [{
"role": "user",
"content": [
{"type": "text", "text": "What is the total on this invoice?"},
{"type": "file", "file": {
"filename": "invoice.pdf",
"file_data": "data:application/pdf;base64,JVBERi0xLjcK..."
}}
]
}],
"plugins": [{"id": "file-parser", "pdf": {"engine": "ocr"}}]
}'

file_data is a data URL: data:<type>;base64,<contents>. A plain http URL is not fetched, so upload the contents yourself.

One request carries one document. Send a second document as a separate request.

application/pdf, image/png, image/jpeg, image/webp, image/gif, image/tiff. Anything else is refused with status 400 before any charge.

A file content part without the file-parser plugin is refused with 400. That is deliberate: without the plugin the document would travel to a model that cannot read it, and you would still pay for an invented answer.

engine currently accepts only ocr.

Alongside the model’s answer, the response carries the extracted text in annotations.

{
"choices": [{
"message": {
"role": "assistant",
"content": "The total is Rp 417,000.",
"annotations": [{
"type": "file",
"file": {
"hash": "sha256:a3b6919a...",
"name": "invoice.pdf",
"content": [{"type": "text", "text": "# INVOICE\n..."}],
"confidence": 0.9892,
"low_confidence": false,
"reuse_id": "ocr_8f286cec-075b-4083-a462-fc89170feb4f"
}
}]
}
}]
}

A streamed response (stream: true) does not carry this annotation. The parse still runs and is still billed, but the text and the reuse_id are only available on a non-streamed response.

low_confidence is a warning, not a guarantee

Section titled “low_confidence is a warning, not a guarantee”

confidence is the mean per-page score from the reading engine. low_confidence is true when that score falls below kenari’s floor, or when the engine reported no score at all.

Treat it as a warning only. In our measurements, correctly read documents and incorrectly read ones had overlapping score ranges: one receipt whose contents were entirely wrong still scored above four documents that were read perfectly. The engine tends to invent a plausible value rather than leave it blank, so figures read out of a scanned document need a human check whatever low_confidence says.

Keep the reuse_id from the first answer and send it back on the next turn, with no file content part:

{
"model": "step-3-7-flash",
"messages": [{"role": "user", "content": "Who is the seller?"}],
"plugins": [{"id": "file-parser", "pdf": {"engine": "ocr", "reuse_id": "ocr_8f286cec-..."}}]
}

That turn does not re-read the document and adds no page charge. A reuse_id is only valid for the account that received it.

Re-sending the exact same document is not billed twice either, so you can retry a failed request without paying again.

POST /v1/ocr returns the extracted text only, with no model in the loop. The per-page price is the same as the plugin path, and there is no token charge.

Terminal window
curl https://kenari.id/v1/ocr \
-H "Authorization: Bearer $KENARI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"file": {
"filename": "invoice.pdf",
"file_data": "data:application/pdf;base64,JVBERi0xLjcK..."
}
}'

The request carries file or reuse_id, never both. file holds filename and file_data. file_data is a data URL (data:<type>;base64,<contents>) and is base64 only: a plain http URL is not fetched, so upload the contents yourself. The accepted types match the plugin: application/pdf, image/png, image/jpeg, image/webp, image/gif, image/tiff. engine is optional, and only ocr is accepted today.

The response:

{
"id": "req_8f286cec-075b-4083",
"pages": 3,
"cost_micro_idr": 900000,
"name": "invoice.pdf",
"hash": "sha256:a3b6919a...",
"content": [{"type": "text", "text": "# INVOICE\n..."}],
"confidence": 0.9892,
"low_confidence": false,
"reuse_id": "ocr_8f286cec-075b-4083-a462-fc89170feb4f"
}

content is the extracted text. pages is what was actually read and is what the charge is based on. cost_micro_idr is the fee in micro-Rupiah (Rupiah times 1,000,000), and is zero when the request only replayed a reuse_id. confidence and low_confidence mean the same as in the plugin response.

Two limits apply, and either can refuse first. The page limit is set by kenari and is 100 pages today. The document size limit is about 15 MB. A document past either is refused, charged nothing, and not read.

A reuse_id works in both directions. One issued by a plugin turn in a chat request can be replayed on this endpoint, and one issued here can be used in a chat turn. Both serve the same stored reading with no extra page charge. A reuse_id is only valid for the account that received it.

A shared key can only reuse a reading it made itself

Section titled “A shared key can only reuse a reading it made itself”

A key you shared with someone else can only reuse a reading that key itself made. Any other key you own that is not a share can still reuse anything on your account. If a shared key tries to redeem a reuse_id your own key created, the request is refused with status 400, the document is not read, and the holder has to upload the document themselves and pay.

That is deliberate. A share is a limited grant, and reuse_ids ride in exported chat logs. If a shared key could redeem a reuse_id from another key, its holder would get the text of documents they never uploaded, for free.

Billed per page read, from the same Rupiah balance as token usage. The per-page price is in the dashboard.

Pages are counted from what the engine actually read, not from the file size. An image counts as one page.

The reading charge is separate from the model’s token charge: one request produces two rows in Usage, one for the reading and one for the model’s answer.

There is a page limit per document, set by kenari and 100 pages today. There is also a document size limit of about 15 MB. The two limits stand on their own: whichever is exceeded refuses the document with no charge.

An API key restricted to certain models cannot use document reading until you grant it. When you create the key or the share, tick Pembacaan dokumen in the capability list. This covers keys you share with other people too.

The reason is that document reading is its own charge, separate from the model’s. A key you limited to a few models is not automatically allowed to run up a per-page cost, because that was not part of what you granted when you made the key.

A key with no model restriction is unaffected and can use document reading straight away. The capability list only appears for a key that restricts models.

An older key that lists ocr among its models keeps working with no change.

A refused key gets a 400 on both the plugin path and the endpoint, is charged nothing, and its document is never read. The same rule covers reuse_id: a key that is not allowed cannot use an earlier reading either, even though reuse is free.