Structured output
Structured output means the model answers with JSON your code can parse, instead of free text. On Chat completions you ask for it with response_format. kenari forwards that field to some models and drops it for others, and it does not check the result. Treat response_format as a request to the model, and validate every reply in your own code.
Ask for JSON
Section titled “Ask for JSON”{"type": "json_object"} asks for a single valid JSON object. Also tell the model in the prompt to answer with JSON, because some providers require it.
curl https://kenari.id/v1/chat/completions \ -H "Authorization: Bearer $KENARI_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "step-3-7-flash:free", "response_format": {"type": "json_object"}, "messages": [ {"role": "system", "content": "Reply with a JSON object and nothing else."}, {"role": "user", "content": "Extract vendor and total: Invoice from Toko Maju, total 417000 rupiah."} ] }'Ask for a schema
Section titled “Ask for a schema”{"type": "json_schema"} asks for JSON that matches a JSON Schema. Give the schema a name and put the schema itself under schema.
{ "model": "step-3-7-flash:free", "response_format": { "type": "json_schema", "json_schema": { "name": "invoice", "schema": { "type": "object", "properties": { "vendor": {"type": "string"}, "total": {"type": "integer"} }, "required": ["vendor", "total"] } } }, "messages": [ {"role": "user", "content": "Extract the invoice fields: Invoice from Toko Maju, total 417000 rupiah."} ]}What kenari does with response_format
Section titled “What kenari does with response_format”kenari does not interpret response_format. What happens depends on the API format the provider uses to serve the model. This is not the endpoint you call, and GET /v1/models does not show it, so plan to validate and retry on every model:
| The model is served through | What happens to response_format |
|---|---|
| An OpenAI-style API | It is forwarded as you sent it. Whether the model enforces it is up to the model. |
| An Anthropic-style API or a Responses-style API | It is not sent. The model answers from your prompt alone. |
Any API, and the request also lists a kenari: server tool | It is not sent. |
kenari never returns an error because a model ignores the field, and GET /v1/models does not say which models enforce it. If a provider rejects the field, you get a 400. A reply can therefore be valid JSON, invalid JSON, or JSON wrapped in a code fence.
The reliable pattern: validate and retry
Section titled “The reliable pattern: validate and retry”Request JSON, validate the reply against your schema, and send the validation error back when it fails. This works on every model.
import os
from openai import OpenAIfrom pydantic import BaseModel, ValidationError
class Invoice(BaseModel): vendor: str total: int
client = OpenAI( base_url="https://kenari.id/v1", api_key=os.environ["KENARI_API_KEY"],)
def extract_invoice(text: str, attempts: int = 3) -> Invoice: messages = [ {"role": "system", "content": "Reply with one JSON object and nothing else."}, {"role": "user", "content": f"Extract vendor and total from: {text}"}, ] for _ in range(attempts): response = client.chat.completions.create( model="step-3-7-flash:free", messages=messages, response_format={ "type": "json_schema", "json_schema": {"name": "invoice", "schema": Invoice.model_json_schema()}, }, ) reply = response.choices[0].message.content or "" try: return Invoice.model_validate_json(reply) except ValidationError as error: messages += [ {"role": "assistant", "content": reply}, {"role": "user", "content": f"That reply was invalid: {error}. Reply with corrected JSON only."}, ] raise RuntimeError("no valid JSON after retries")
print(extract_invoice("Invoice from Toko Maju, total 417000 rupiah."))Sending response_format as well as validating adds no separate fee and helps on the models that honor it. Each retry is a new request, billed as usual on paid models.
Structured output with a tool
Section titled “Structured output with a tool”Forcing a tool call requests arguments shaped by a schema, and works on any model with tool_call set to true. Define one tool whose input_schema is your schema, force it with tool_choice, and read the arguments from the call. On Messages:
import os
import anthropic
client = anthropic.Anthropic( base_url="https://kenari.id", api_key=os.environ["KENARI_API_KEY"],)
response = client.messages.create( model="step-3-7-flash:free", max_tokens=1024, tools=[ { "name": "record_invoice", "description": "Record the fields of an invoice.", "input_schema": { "type": "object", "properties": { "vendor": {"type": "string"}, "total": {"type": "integer"}, }, "required": ["vendor", "total"], }, } ], tool_choice={"type": "tool", "name": "record_invoice"}, messages=[{"role": "user", "content": "Invoice from Toko Maju, total 417000 rupiah."}],)
invoice = next(block.input for block in response.content if block.type == "tool_use")print(invoice)block.input is already parsed into an object. You do not run the tool or send a result back. The same pattern works on Chat completions with a function tool and tool_choice: {"type": "function", "function": {"name": "record_invoice"}}, where the JSON is in tool_calls[0].function.arguments. Validate it all the same, and handle a reply that has no tool call. See Function calling.
Limits and pitfalls
Section titled “Limits and pitfalls”/v1/responsesdoes not support structured output.text.formatwithjson_objectorjson_schemais rejected with400. Use Chat completions or Messages.- A reply cut off by
max_tokensis invalid JSON. On Chat completions, check thatfinish_reasonis notlength. On Messages, check thatstop_reasonis notmax_tokens. Reasoning models spend part ofmax_tokenson thinking, so leave room. See Reasoning. - Do not rely on schema keywords such as
minimumorpatternbeing enforced. Models differ, so check them in your own code.