Model catalog
GET /v1/models is the live catalog. Every entry carries the model’s name, context window, accepted inputs and your price. GET /v1/models/{model} returns a single entry.
A model name has the form vendor/model, for example openai/gpt-4o-mini. Case and separators do not matter: anthropic/claude-haiku-4.5 and anthropic/claude-haiku-4-5 are the same model.
string
required
The value to put in
model.string
Display name.
string
required
The model’s vendor.
string
text — served by /v1/chat/completions and /v1/responses; image, video, music, audio — by the matching /v1/{kind}/jobs.integer
Context window in tokens (prompt plus answer).
integer
Largest answer the model can produce, when known.
string[]
text, image, audio, video, file.object
required
Your price in USD. Text models:
input_per_mtok, output_per_mtok and cached_input_per_mtok (per million tokens). Media models: unit (image, video, second, request, 1k_characters) and per_unit for the cheapest configuration.Cost
Requests are charged in US dollars, counted in micro-dollars: 1 µ0.000001. Text:- A price per million tokens is also the price per token in µ$.
- Token counts come from the model’s own
usagereport, returned to you unchanged. - Any request that consumed tokens costs at least 1 µ$; the total is rounded up.
- Hidden reasoning tokens are billed as output tokens.
cost. estimated_cost is the forecast at start. Duration, resolution, sound and tier raise the price. A failed job costs nothing.
Balance holds
Before a request runs, its largest possible cost is held against your balance, so parallel requests cannot together spend more than you have.- Text — the prompt plus
max_tokensof answer (4096 if unset). When the request ends, the hold is replaced by the actual cost. - Media — the job’s cost. The first time a configuration is used (a new model, a first 4K clip), the hold carries a margin. A failed job releases its hold at once.
402 insufficient_balance before any work begins. meta gives required_micro_usd, balance_micro_usd and held_micro_usd.
Reasoning models
o-series, DeepSeek R1 and thinking variants, Gemini Pro and Claude thinking spend part ofmax_tokens on hidden reasoning. Give them max_tokens of 2000 or more, or the answer may be cut off with finish_reason: "length".
