Skip to main content
Every field of the Chat Completions protocol is forwarded to the model as sent, and whether the model honours it is up to the model. Check the model in GET /v1/models before relying on a feature.

Image input

Models with image in input_modalities read images passed in messages. Send content as an array of parts:

Structured output (JSON)

response_format is passed to the model unchanged. Models that support it return JSON matching the schema:
If a model does not support response_format, the API may return 400 unsupported_parameter, or the model may ignore the field. In that case, describe the format in the prompt.

Tool calling: the full loop

1

Describe the tools

Pass them in tools. If the model decides to call a function, the response has finish_reason: "tool_calls".
2

Run the function yourself

Take the name and arguments from message.tool_calls.
3

Return the result

Append the model’s message and a role: "tool" message with the same tool_call_id to the history, then send the request again.

Prompt caching

Some models cache a repeated prompt. The number of tokens served from cache is in usage.prompt_tokens_details.cached_tokens. If the catalog lists cached_input_per_mtok for the model, these tokens are billed at that price; otherwise at the regular input price. To hit the cache more often, keep the unchanging part (system prompt, instructions, documents) at the start of the messages and the changing part at the end.