GET /v1/models before relying on a feature.
Image input
Models withimage in input_modalities read images passed in messages. Send content as an array of parts:
Structured output (JSON)
response_format is passed to the model unchanged. Models that support it return JSON matching the schema:
Tool calling: the full loop
1
Describe the tools
Pass them in
tools. If the model decides to call a function, the response has finish_reason: "tool_calls".2
Run the function yourself
Take the name and arguments from
message.tool_calls.3
Return the result
Append the model’s message and a
role: "tool" message with the same tool_call_id to the history, then send the request again.Prompt caching
Some models cache a repeated prompt. The number of tokens served from cache is inusage.prompt_tokens_details.cached_tokens. If the catalog lists cached_input_per_mtok for the model, these tokens are billed at that price; otherwise at the regular input price.
To hit the cache more often, keep the unchanging part (system prompt, instructions, documents) at the start of the messages and the changing part at the end.
