> ## Documentation Index
> Fetch the complete documentation index at: https://docs.ascn.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Model features

> Image input, JSON output, tool calling, prompt caching.

Every field of the Chat Completions protocol is forwarded to the model as sent, and whether the model honours it is up to the model. Check the model in `GET /v1/models` before relying on a feature.

## Image input

Models with `image` in `input_modalities` read images passed in `messages`. Send `content` as an array of parts:

```python theme={null}
response = client.chat.completions.create(
    model="openai/gpt-4o-mini",
    messages=[{
        "role": "user",
        "content": [
            {"type": "text", "text": "What is in this picture?"},
            {"type": "image_url", "image_url": {"url": "https://example.com/photo.jpg"}},
        ],
    }],
)
```

## Structured output (JSON)

`response_format` is passed to the model unchanged. Models that support it return JSON matching the schema:

```python theme={null}
response = client.chat.completions.create(
    model="openai/gpt-4o-mini",
    messages=[{"role": "user", "content": "Name the capital of France and its population."}],
    response_format={
        "type": "json_schema",
        "json_schema": {
            "name": "city",
            "schema": {
                "type": "object",
                "properties": {
                    "city": {"type": "string"},
                    "population": {"type": "integer"},
                },
                "required": ["city", "population"],
            },
        },
    },
)
```

<Tip>
  If a model does not support `response_format`, the API may return `400 unsupported_parameter`, or the model may ignore the field. In that case, describe the format in the prompt.
</Tip>

## Tool calling: the full loop

<Steps>
  <Step title="Describe the tools">
    Pass them in `tools`. If the model decides to call a function, the response has `finish_reason: "tool_calls"`.
  </Step>

  <Step title="Run the function yourself">
    Take the name and arguments from `message.tool_calls`.
  </Step>

  <Step title="Return the result">
    Append the model's message and a `role: "tool"` message with the same `tool_call_id` to the history, then send the request again.
  </Step>
</Steps>

```python theme={null}
import json

tools = [{
    "type": "function",
    "function": {
        "name": "get_weather",
        "description": "Current weather for a city.",
        "parameters": {"type": "object", "properties": {"city": {"type": "string"}}, "required": ["city"]},
    },
}]
messages = [{"role": "user", "content": "What is the weather in Berlin?"}]

first = client.chat.completions.create(model="openai/gpt-4o-mini", messages=messages, tools=tools)
call = first.choices[0].message.tool_calls[0]
args = json.loads(call.function.arguments)

result = {"city": args["city"], "temp_c": 18}  # your function

messages.append(first.choices[0].message)
messages.append({"role": "tool", "tool_call_id": call.id, "content": json.dumps(result)})

final = client.chat.completions.create(model="openai/gpt-4o-mini", messages=messages, tools=tools)
print(final.choices[0].message.content)
```

## Prompt caching

Some models cache a repeated prompt. The number of tokens served from cache is in `usage.prompt_tokens_details.cached_tokens`. If the catalog lists `cached_input_per_mtok` for the model, these tokens are billed at that price; otherwise at the regular input price.

To hit the cache more often, keep the unchanging part (system prompt, instructions, documents) at the start of the messages and the changing part at the end.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.