Skip to content
kenari.

Structured output

Structured output means the model answers with JSON your code can parse, instead of free text. On Chat completions you ask for it with response_format. kenari forwards that field to some models and drops it for others, and it does not check the result. Treat response_format as a request to the model, and validate every reply in your own code.

{"type": "json_object"} asks for a single valid JSON object. Also tell the model in the prompt to answer with JSON, because some providers require it.

Terminal window
curl https://kenari.id/v1/chat/completions \
-H "Authorization: Bearer $KENARI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "step-3-7-flash:free",
"response_format": {"type": "json_object"},
"messages": [
{"role": "system", "content": "Reply with a JSON object and nothing else."},
{"role": "user", "content": "Extract vendor and total: Invoice from Toko Maju, total 417000 rupiah."}
]
}'

{"type": "json_schema"} asks for JSON that matches a JSON Schema. Give the schema a name and put the schema itself under schema.

{
"model": "step-3-7-flash:free",
"response_format": {
"type": "json_schema",
"json_schema": {
"name": "invoice",
"schema": {
"type": "object",
"properties": {
"vendor": {"type": "string"},
"total": {"type": "integer"}
},
"required": ["vendor", "total"]
}
}
},
"messages": [
{"role": "user", "content": "Extract the invoice fields: Invoice from Toko Maju, total 417000 rupiah."}
]
}

kenari does not interpret response_format. What happens depends on the API format the provider uses to serve the model. This is not the endpoint you call, and GET /v1/models does not show it, so plan to validate and retry on every model:

The model is served throughWhat happens to response_format
An OpenAI-style APIIt is forwarded as you sent it. Whether the model enforces it is up to the model.
An Anthropic-style API or a Responses-style APIIt is not sent. The model answers from your prompt alone.
Any API, and the request also lists a kenari: server toolIt is not sent.

kenari never returns an error because a model ignores the field, and GET /v1/models does not say which models enforce it. If a provider rejects the field, you get a 400. A reply can therefore be valid JSON, invalid JSON, or JSON wrapped in a code fence.

Request JSON, validate the reply against your schema, and send the validation error back when it fails. This works on every model.

import os
from openai import OpenAI
from pydantic import BaseModel, ValidationError
class Invoice(BaseModel):
vendor: str
total: int
client = OpenAI(
base_url="https://kenari.id/v1",
api_key=os.environ["KENARI_API_KEY"],
)
def extract_invoice(text: str, attempts: int = 3) -> Invoice:
messages = [
{"role": "system", "content": "Reply with one JSON object and nothing else."},
{"role": "user", "content": f"Extract vendor and total from: {text}"},
]
for _ in range(attempts):
response = client.chat.completions.create(
model="step-3-7-flash:free",
messages=messages,
response_format={
"type": "json_schema",
"json_schema": {"name": "invoice", "schema": Invoice.model_json_schema()},
},
)
reply = response.choices[0].message.content or ""
try:
return Invoice.model_validate_json(reply)
except ValidationError as error:
messages += [
{"role": "assistant", "content": reply},
{"role": "user", "content": f"That reply was invalid: {error}. Reply with corrected JSON only."},
]
raise RuntimeError("no valid JSON after retries")
print(extract_invoice("Invoice from Toko Maju, total 417000 rupiah."))

Sending response_format as well as validating adds no separate fee and helps on the models that honor it. Each retry is a new request, billed as usual on paid models.

Forcing a tool call requests arguments shaped by a schema, and works on any model with tool_call set to true. Define one tool whose input_schema is your schema, force it with tool_choice, and read the arguments from the call. On Messages:

import os
import anthropic
client = anthropic.Anthropic(
base_url="https://kenari.id",
api_key=os.environ["KENARI_API_KEY"],
)
response = client.messages.create(
model="step-3-7-flash:free",
max_tokens=1024,
tools=[
{
"name": "record_invoice",
"description": "Record the fields of an invoice.",
"input_schema": {
"type": "object",
"properties": {
"vendor": {"type": "string"},
"total": {"type": "integer"},
},
"required": ["vendor", "total"],
},
}
],
tool_choice={"type": "tool", "name": "record_invoice"},
messages=[{"role": "user", "content": "Invoice from Toko Maju, total 417000 rupiah."}],
)
invoice = next(block.input for block in response.content if block.type == "tool_use")
print(invoice)

block.input is already parsed into an object. You do not run the tool or send a result back. The same pattern works on Chat completions with a function tool and tool_choice: {"type": "function", "function": {"name": "record_invoice"}}, where the JSON is in tool_calls[0].function.arguments. Validate it all the same, and handle a reply that has no tool call. See Function calling.

  • /v1/responses does not support structured output. text.format with json_object or json_schema is rejected with 400. Use Chat completions or Messages.
  • A reply cut off by max_tokens is invalid JSON. On Chat completions, check that finish_reason is not length. On Messages, check that stop_reason is not max_tokens. Reasoning models spend part of max_tokens on thinking, so leave room. See Reasoning.
  • Do not rely on schema keywords such as minimum or pattern being enforced. Models differ, so check them in your own code.