Skip to content
kenari.

Images and files as input

Some models read images directly. You attach the image to a message, and the model answers about it in the same request. Documents work on any model: kenari reads a PDF, text file or image file first and gives the model its text and page images. This page covers images on all three chat formats and documents on Chat completions and Messages. Read documents has the limits and the engines.

GET /v1/models lists what each model accepts under modalities.input. The call is public, so you can filter it without a key:

Terminal window
curl -s https://kenari.id/v1/models \
| jq -r '.data[] | select((.modalities.input // []) | index("image")) | .id'

Replace "image" with "pdf" to find models that read PDFs themselves.

Value in modalities.inputThe model reads
imageImages.
pdfPDF files, by the model itself.
fileAny document. Every chat model lists it, because kenari reads attached documents for any model.

The list can also contain audio and video, which this page does not cover. The examples below use step-3-7-flash:free, which reads images and costs nothing to call.

Add an image_url part next to your text. The url is either a data URL or a web address.

import base64
import os
from openai import OpenAI
client = OpenAI(
base_url="https://kenari.id/v1",
api_key=os.environ["KENARI_API_KEY"],
)
with open("receipt.jpg", "rb") as file:
image = base64.b64encode(file.read()).decode()
response = client.chat.completions.create(
model="step-3-7-flash:free",
messages=[
{
"role": "user",
"content": [
{"type": "text", "text": "What is the total on this receipt?"},
{"type": "image_url", "image_url": {"url": f"data:image/jpeg;base64,{image}"}},
],
}
],
)
print(response.choices[0].message.content)

A web address works the same way: {"type": "image_url", "image_url": {"url": "https://example.com/receipt.jpg"}}.

kenari does not download the image. With a web address, the model’s provider fetches it, and not every provider does. When a provider cannot, the request fails with 400 and the message remote image URL not supported by upstream. A data URL works on every model that reads images, so use one when you need a request to work everywhere.

On /v1/messages an image is a block with a source of type base64 or url. Point the Anthropic SDK at https://kenari.id, without /v1.

import base64
import os
import anthropic
client = anthropic.Anthropic(
base_url="https://kenari.id",
api_key=os.environ["KENARI_API_KEY"],
)
with open("receipt.jpg", "rb") as file:
image = base64.b64encode(file.read()).decode()
message = client.messages.create(
model="step-3-7-flash:free",
max_tokens=1024,
messages=[
{
"role": "user",
"content": [
{"type": "image", "source": {"type": "base64", "media_type": "image/jpeg", "data": image}},
{"type": "text", "text": "What is the total on this receipt?"},
],
}
],
)
print(next(block.text for block in message.content if block.type == "text"))

A url source is {"type": "url", "url": "https://example.com/receipt.jpg"}, with the same caveat about web addresses as above. Any other source type, such as an uploaded file id, is ignored without an error, so the model answers without the image.

On /v1/responses an image is an input_image part whose image_url is a string, either a data URL or a web address:

{
"model": "step-3-7-flash:free",
"input": [
{
"role": "user",
"content": [
{"type": "input_text", "text": "What is the total on this receipt?"},
{"type": "input_image", "image_url": "data:image/jpeg;base64,/9j/4AAQ..."}
]
}
]
}

On Chat completions a document is a file part. file_data is a data URL. This works on any model: kenari reads the document before routing, so the model does not need to read files itself.

import base64
import os
from openai import OpenAI
client = OpenAI(
base_url="https://kenari.id/v1",
api_key=os.environ["KENARI_API_KEY"],
)
with open("invoice.pdf", "rb") as file:
document = base64.b64encode(file.read()).decode()
response = client.chat.completions.create(
model="step-3-7-flash:free",
messages=[
{
"role": "user",
"content": [
{"type": "text", "text": "What is the invoice number?"},
{
"type": "file",
"file": {
"filename": "invoice.pdf",
"file_data": f"data:application/pdf;base64,{document}",
},
},
],
}
],
)
print(response.choices[0].message.content)

kenari sends the model the text of every page that has a text layer and an image of every page that does not. You can attach several documents to one request. Accepted types are PDF, text/plain, text/markdown, text/csv, text/html, application/json and images. The limits are about 15 MB and 100 pages per document, and 20 pages without a text layer per request. A document that makes the request too long for the model’s context is refused with 400 and no charge, and the message gives the token counts. See Read documents for the limits and the engines.

Send images as image_url parts when you can. An image in a file part is sent to the model as an image, but with the native engine it is refused with 400. A file part without a file object is refused too.

Two other formats behave differently:

  • /v1/responses does not accept input_file and answers 400.
  • On /v1/messages, a document block with a base64 source is read the same way as a file part, on any model. A document block with a url source reaches only models served through an Anthropic-style API, and for other models it is dropped without an error. Send the PDF as base64.

With the native engine, the file goes to the model untouched. This needs a model that reads files itself. file in modalities.input does not tell you that, because every chat model lists it, and pdf marks a model that reads PDFs itself. kenari checks the model’s own capability:

  • A model that reads PDFs takes PDF files.
  • A model that reads any document takes any document.
  • Any other model: the request is refused with 400 before anything is charged, because the document would otherwise reach a model that cannot read it.
"plugins": [{"id": "file-parser", "pdf": {"engine": "native"}}]

A refusal reads like this:

{
"error": {
"message": "model 'step-3-7-flash:free' does not read application/pdf file attachments. Send the file to a model that reads files, or add the file-parser plugin to read it with OCR (billed per page from your balance)",
"type": "bad_request",
"param": null,
"code": "bad_request"
}
}

The paid ocr engine reads a scan or handwriting to text with OCR. It is billed per page from your balance and never covered by a plan. See Read documents.

Images and files that the model reads are billed as Input tokens of that model, at the model’s normal rate. The token count depends on the model, so read the usage in the response: usage.prompt_tokens on Chat completions, usage.input_tokens on Messages and Responses. On Messages, add the cache fields to get the whole prompt, as described in Prompt caching. A subscription plan that covers the model covers these tokens too, and that includes documents that kenari reads for you. Only the ocr engine is billed per page instead. See How billing works.

  • The whole request body is limited to 32 MiB. A larger request is refused with 413. Base64 adds about a third to the size of a file.
  • kenari does not resize images or count them. Limits on image size and image count come from the model and its provider.
  • kenari does not check an image against modalities.input. A model that does not list image can fail with an error or answer without looking at the image, so pick a model from the list.
  • A conversation resends its images on every request, and each request is billed for them again. Drop images from the history once the model has used them.