One current model from every family approved for public discovery. The catalog refreshes automatically as providers release models and update metadata or pricing.
Claude Opus 5 is the successor to Opus 4.8 for complex agentic coding and enterprise work — a step change on deep reasoning, long-horizon agentic execution, and test-time compute scaling at Opus 4.8 pricing. 1M context window, 128K max output. Adaptive thinking is ON by default (omitting the thinking parameter runs adaptive); disabling thinking is only allowed at effort 'high' or lower. Raw thinking tokens are never returned. Full effort ladder (low/medium/high/xhigh/max), 512-token prompt-cache minimum, elevated cybersecurity safeguards (responses can return stop_reason 'refusal'). Separate rate-limit bucket from the Opus 4.x pool.
Context1M
Familyclaude
Input$5/M
Output$25/M
AgentChatCodeSummarizationTool UseReasoningVisionJSON modePrompt cache
Next-generation Sonnet-tier Claude model with the best combination of speed and intelligence. 1M context window, 128K max output, always-on adaptive thinking, and new tokenizer (~30% more tokens for the same text). Extended thinking (manual thinking budgets) removed; setting temperature/top_p/top_k to non-default values returns 400. Pricing is $2/$10 per MTok input/output — announced at launch as introductory pricing and made the standard price on August 10, 2026.
Context1M
Familyclaude
Input$2/M
Output$10/M
AgentChatCodeSummarizationTool UseReasoningVisionJSON modePrompt cache
Fast and cost-effective Claude model. Near-Sonnet 4.5 performance at roughly one-third the cost. Best for high-volume, latency-sensitive tasks. Supports extended thinking; no adaptive thinking and no effort parameter.
Context200K
Familyclaude
Input$1/M
Output$5/M
AgentChatCodeSummarizationTool UseReasoningVisionJSON modePrompt cache
OpenAI's next-generation flagship model. Previewed June 26, 2026; generally available July 9, 2026 across ChatGPT, Codex, and the API. Launches with OpenAI's most robust safety stack to date. Introduces predictable prompt caching with explicit cache breakpoints and 30-min minimum cache life.
Context1M
Familygpt
Input$5/M
Output$30/M
AgentChatCodeSummarizationTool UseVisionReasoningJSON modePrompt cache
OpenAI's smartest and most precise model — uses extra compute to think harder and provide consistently better answers on complex tasks. Available for Responses API requests, including through the Batch API.
Context1M
Familygpt
Input$30/M
Output$180/M
AgentChatCodeSummarizationTool UseVisionReasoningJSON mode
Balanced GPT-5.6 model for everyday work, competitive with GPT-5.5 while being ~2x cheaper. Previewed June 26, 2026; generally available July 9, 2026.
Context1M
Familygpt
Input$2/M
Output$12/M
AgentChatCodeSummarizationTool UseVisionReasoningJSON modePrompt cache
—
Context1M
Familygemini
Input$2.08/M
Output$10.39/M
AgentChatCodeSummarizationMultimodalReasoningVisionTool UseJSON modePrompt cache
xAI's frontier Grok model for coding, agentic tasks, and knowledge work. 500k context, text and image input, function calling, structured outputs, and configurable reasoning effort (low/medium/high/xhigh, default high).
Context500K
Familygrok
Input$2/M
Output$6/M
AgentChatCodeReasoningVisionTool UseJSON modePrompt cache
Mistral's flagship 675B MoE model (41B active). Open-weight (Apache 2.0) general-purpose multimodal model with 256K context and vision input. Released by Mistral on 2025-12-02; available on Vertex AI Model Garden as self-deploy (open-weight) since 2025-12-09; not currently offered as a fully-managed Vertex partner-model API.
Context256K
Familymistral
Input$0.50/M
Output$1.50/M
AgentChatCodeSummarizationTool UseVisionJSON modePrompt cache