Product

DiscoverShipLaunchModelsPricing

Company

AboutFAQHelp

Legal

Privacy policyTerms of serviceSecurity

Workflows

Claude Code orchestrationCodex workflow automationSecurity for AI coding agentsAI code review agentsMulti-agent software deliveryChangelog-to-launch automationAI social post generationAI roadmap generation

Connect

DiscordYouTubeLinkedInGitHub
© 2026 AI Expediteai@aiexpedite.com
AI Expedite LogoAI Expedite
PricingFAQAbout
DiscordYouTubeGitHub

Models

One platform. Every frontier model.

Pick the right brain for the job. Route the easy stuff to a fast, cheap model and reserve the expensive reasoners for the work that actually needs them.

9Public models
5Providers
1MTokens of context

How to read the cost meter

$ = cheap. $$$$$ = expensive.

The dollar-sign meter on each card ranks every listed model against the others by output-token cost. Five $ marks the priciest on this page; one $ marks the cheapest. Hover the card to see the actual per-million-token rates.

$$$$$Cheapest$$$$$Low cost$$$$$Mid$$$$$Premium$$$$$Most expensive

Featured catalog

The latest public models

One current model from every family approved for public discovery. The catalog refreshes automatically as providers release models and update metadata or pricing.

Claude Opus 5

Anthropic
$$$$$Premium

Claude Opus 5 is the successor to Opus 4.8 for complex agentic coding and enterprise work — a step change on deep reasoning, long-horizon agentic execution, and test-time compute scaling at Opus 4.8 pricing. 1M context window, 128K max output. Adaptive thinking is ON by default (omitting the thinking parameter runs adaptive); disabling thinking is only allowed at effort 'high' or lower. Raw thinking tokens are never returned. Full effort ladder (low/medium/high/xhigh/max), 512-token prompt-cache minimum, elevated cybersecurity safeguards (responses can return stop_reason 'refusal'). Separate rate-limit bucket from the Opus 4.x pool.

Context1M
Familyclaude
Input$5/M
Output$25/M
AgentChatCodeSummarizationTool UseReasoningVisionJSON modePrompt cache

Claude Sonnet 5

Anthropic
$$$$$Mid

Next-generation Sonnet-tier Claude model with the best combination of speed and intelligence. 1M context window, 128K max output, always-on adaptive thinking, and new tokenizer (~30% more tokens for the same text). Extended thinking (manual thinking budgets) removed; setting temperature/top_p/top_k to non-default values returns 400. Pricing is $2/$10 per MTok input/output — announced at launch as introductory pricing and made the standard price on August 10, 2026.

Context1M
Familyclaude
Input$2/M
Output$10/M
AgentChatCodeSummarizationTool UseReasoningVisionJSON modePrompt cache

Claude Haiku 4.5

Anthropic
$$$$$Low cost

Fast and cost-effective Claude model. Near-Sonnet 4.5 performance at roughly one-third the cost. Best for high-volume, latency-sensitive tasks. Supports extended thinking; no adaptive thinking and no effort parameter.

Context200K
Familyclaude
Input$1/M
Output$5/M
AgentChatCodeSummarizationTool UseReasoningVisionJSON modePrompt cache

GPT-5.6 Sol

OpenAI
$$$$$Most expensive

OpenAI's next-generation flagship model. Previewed June 26, 2026; generally available July 9, 2026 across ChatGPT, Codex, and the API. Launches with OpenAI's most robust safety stack to date. Introduces predictable prompt caching with explicit cache breakpoints and 30-min minimum cache life.

Context1M
Familygpt
Input$5/M
Output$30/M
AgentChatCodeSummarizationTool UseVisionReasoningJSON modePrompt cache

GPT-5.5 Pro

OpenAI
$$$$$Most expensive

OpenAI's smartest and most precise model — uses extra compute to think harder and provide consistently better answers on complex tasks. Available for Responses API requests, including through the Batch API.

Context1M
Familygpt
Input$30/M
Output$180/M
AgentChatCodeSummarizationTool UseVisionReasoningJSON mode

GPT-5.6 Terra

OpenAI
$$$$$Premium

Balanced GPT-5.6 model for everyday work, competitive with GPT-5.5 while being ~2x cheaper. Previewed June 26, 2026; generally available July 9, 2026.

Context1M
Familygpt
Input$2/M
Output$12/M
AgentChatCodeSummarizationTool UseVisionReasoningJSON modePrompt cache

Gemini 3.6 Flash

Google
$$$$$Mid

—

Context1M
Familygemini
Input$2.08/M
Output$10.39/M
AgentChatCodeSummarizationMultimodalReasoningVisionTool UseJSON modePrompt cache

Grok 4.6

xAI
$$$$$Low cost

xAI's frontier Grok model for coding, agentic tasks, and knowledge work. 500k context, text and image input, function calling, structured outputs, and configurable reasoning effort (low/medium/high/xhigh, default high).

Context500K
Familygrok
Input$2/M
Output$6/M
AgentChatCodeReasoningVisionTool UseJSON modePrompt cache

Mistral Large 3

Mistral
$$$$$Cheapest

Mistral's flagship 675B MoE model (41B active). Open-weight (Apache 2.0) general-purpose multimodal model with 256K context and vision input. Released by Mistral on 2025-12-02; available on Vertex AI Model Garden as self-deploy (open-weight) since 2025-12-09; not currently offered as a fully-managed Vertex partner-model API.

Context256K
Familymistral
Input$0.50/M
Output$1.50/M
AgentChatCodeSummarizationTool UseVisionJSON modePrompt cache

How we think about cost

The right model for the right task

You don't need the most expensive model for every step. AI Expedite lets you route by stage — drafting on cheap models, reasoning on premium ones — and shows the true cost of every execution.

$$$$$High volume

Use lower-cost models for drafts, classifications, fan-out, and boilerplate. They stretch each credit furthest when the task is straightforward.

$$$$$General workloads

Balanced models are a strong default for production workflows: capable, responsive, and economical enough for repeated use.

$$$$$Hard reasoning

Reserve premium reasoners for difficult debugging, architectural decisions, and work where deeper analysis justifies the higher token rate.

Mix and match across providers

No vendor lock-in. Switch a step from Claude to Gemini to GPT-5 without rewriting your agent.

Start free

Pricing is sourced from the platform catalog in USD per 1M tokens and refreshes at least hourly.