LLM API catalog
Every frontier LLM API, priced side by side.
Compare GPT, Claude, Gemini, DeepSeek, Grok and more on live per-token prices, context windows and input types — then call any of them through one OpenAI-compatible API key.
- LLMs available
- 31
- Model makers
- 7
- Lowest input / 1M tokens
- $0.08
- Largest context window
- 1.05M
Live catalog · prices sync from the platform every 5 minutes
01/Price spectrum
What one million input tokens costs.
Each dot is a model at its cheapest available channel. The faint trail leads back to its base price — the gap is what multi-channel routing saves you.
- Cheapest input
- $0.08GPT 5.6 Luna
- Biggest channel saving
- −90%Claude Fable 5
- Median saving vs base
- −60%
02/Catalog
All LLM APIs
Sorted by release, newest first. Prices are USD per 1M tokens at the cheapest channel; a struck-through figure is the base price.
31 models
| Model | Provider | Max output | Inputs | → | |||
|---|---|---|---|---|---|---|---|
Claude Sonnet 5.5 claude-sonnet-5-5· Anthropic· 1M | Anthropic | 1M | 128K |
| $0.60−70% | $3.00 | |
Claude Opus 5.5 claude-opus-5-5· Anthropic· 1M | Anthropic | 1M | 128K |
| $1.20−70% | $6.00 | |
DeepSeek Flash deepseek-flash· DeepSeek· 1M | DeepSeek | 1M | 384K |
| $0.15−50% | $0.60 | |
GPT 6 Astra gpt-6-astra· OpenAI· 1M | OpenAI | 1M | 128K |
| $4.00−60% | $20.00 | |
Gemini 3.8 Flash gemini-3.8-flash· Google· 1M | 1M | 66K |
| $0.60−60% | $3.00 | ||
Grok 4.6 grok-4.6· xAI· 500K | xAI | 500K | 500K |
| $1.40−30% | $4.20 | |
DeepSeek V4 Pro deepseek-v4-pro· DeepSeek· 1M | DeepSeek | 1M | 384K |
| $0.945−30% | $2.84 | |
DeepSeek V4 Flash deepseek-v4-flash· DeepSeek· 1M | DeepSeek | 1M | 384K |
| $0.315−30% | $0.945 | |
Gemini 3.6 Flash gemini-3.6-flash· Google· 1M | 1M | 66K |
| $0.60−60% | $3.00 | ||
Claude Opus 5 claude-opus-5· Anthropic· 1M | Anthropic | 1M | 128K |
| $1.50−70% | $7.50 | |
Kimi K3 kimi-k3· Kimi· 1M | Kimi | 1M | 262K |
| $2.10−30% | $10.50 | |
GLM 5.2 glm-5.2· Z.AI· 1M | Z.AI | 1M | 128K |
| $0.98−30% | $3.08 | |
Grok 4.5 grok-4.5· xAI· 500K | xAI | 500K | 500K |
| $0.20−90% | $0.60 | |
GPT 5.6 Luna gpt-5.6-luna· OpenAI· 1M | OpenAI | 1M | 128K |
| $0.08−60% | $0.48 | |
GPT 5.6 Terra gpt-5.6-terra· OpenAI· 1M | OpenAI | 1M | 128K |
| $0.80−60% | $4.80 | |
GPT 5.5 gpt-5.5· OpenAI· 1M | OpenAI | 1M | 128K |
| $1.00−80% | $6.00 | |
GPT 5.6 Sol gpt-5.6-sol· OpenAI· 1M | OpenAI | 1M | 128K |
| $2.00−60% | $12.00 | |
Claude Sonnet 5 claude-sonnet-5· Anthropic· 1M | Anthropic | 1M | 64K |
| $0.60−70% | $3.00 | |
Claude Opus 4.8 claude-opus-4-8· Anthropic· 1M | Anthropic | 1M | 128K |
| $1.50−70% | $7.50 | |
Claude Fable 5 claude-fable-5· Anthropic· 1M | Anthropic | 1M | 128K |
| $1.00−90% | $5.00 | |
Gemini 3.5 Flash gemini-3.5-flash· Google· 1.05M | 1.05M | 66K |
| $0.60−60% | $3.60 | ||
Claude Opus 4.7 claude-opus-4-7· Anthropic· 1M | Anthropic | 1M | 128K |
| $1.50−70% | $7.50 | |
GPT 5.4 nano gpt-5.4-nano· OpenAI· 400K | OpenAI | 400K | 128K |
| $0.16−20% | $1.00 | |
GPT 5.4 mini gpt-5.4-mini· OpenAI· 400K | OpenAI | 400K | 128K |
| $0.15−80% | $0.90 | |
Gemini 3.1 Flash-Lite Preview gemini-3.1-flash-lite-preview· Google· 1.05M | 1.05M | 66K |
| $0.10−60% | $0.60 | ||
Gemini 3.1 Pro Preview gemini-3.1-pro-preview· Google· 1.05M | 1.05M | 66K |
| $0.80−60% | $4.80 | ||
Claude Opus 4.6 claude-opus-4-6· Anthropic· 1M | Anthropic | 1M | 128K |
| $1.50−70% | $7.50 | |
GPT 5.4 Pro gpt-5.4-pro· OpenAI· 1.05M | OpenAI | 1.05M | 128K |
| $18.00−40% | $108.00 | |
GPT 5.4 gpt-5.4· OpenAI· 1.05M | OpenAI | 1.05M | 128K |
| $0.50−80% | $3.00 | |
Claude Sonnet 4.6 claude-sonnet-4-6· Anthropic· 1M | Anthropic | 1M | 64K |
| $0.90−70% | $4.50 | |
Claude Haiku 4.5 claude-haiku-4-5· Anthropic· 200K | Anthropic | 200K | 64K |
| $0.30−70% | $1.50 |
03/Field guide
How to choose an LLM API
Four questions that settle most model decisions faster than a benchmark chart.
Match the tier to the task
Flagship models earn their price on multi-step reasoning, agentic coding and long-form writing. Flash, mini, lite and nano tiers are built for classification, extraction, routing and high-volume chat — often at a tenth of the cost with acceptable quality.
Read input and output separately
Output tokens usually cost three to eight times more than input. Summarisation and RAG are input-heavy; drafting and code generation are output-heavy. Estimate your ratio before comparing headline prices.
Treat context as a budget
A one-million-token window is a ceiling, not a target. Several models switch to a higher long-context rate past a threshold, and prompt caching makes repeated context far cheaper than resending it.
Keep switching cheap
Every model here answers on the OpenAI-compatible Chat Completions endpoint, and Anthropic Messages and native Gemini formats are supported too. Changing models is a one-string change, so you can A/B test without a rewrite.
04/FAQ
LLM API questions, answered
Can I call these LLMs with the OpenAI SDK?
Yes. Point the OpenAI SDK’s base_url at the API address shown in your console and pass any model ID from the table above. Anthropic Messages and the native Gemini API formats are also supported, so existing Claude and Gemini code needs little to no change.
How is LLM usage billed?
Pay as you go, per token, with no monthly fee or subscription. Input, output and cached tokens are priced separately. Many models are served by several channels at different rates, and the table shows the lowest one so you can pick by budget.
What is a channel, and why are some prices discounted?
A channel is an upstream route that serves the same model. Requests are spread across healthy channels by weight or round-robin; when one fails with a retryable error the request moves to the next. Discounted channels let you trade a little routing choice for a much lower price.
Which LLM has the largest context window?
Sort the table by “Largest context” to see the current leaders. Most frontier models now accept around one million tokens; check the max output column too, because output length is capped separately.
How do I start calling an LLM?
Sign up, create an API key in the console and copy a cURL, Python or Node example from any model page. New accounts get free trial credits, no credit card required.
In a few minutes, your first AI video.
Sign up, grab a key, and call image, video and language models from one endpoint. Register now for free trial credits, no credit card required.