Data
AI model API pricing 2026: price data for 19 models from 8 vendors
Checked CSV CC BY 4.0
AI model API pricing in 2026 spans a 40x range. On 2026-10-04, a typical request of 1,000 input tokens and 500 output tokens cost $0.00035 on GPT-6 Luna and $0.014 on Claude Opus 5.5; across the 19 models in this dataset the median is $2.94 per 1,000 requests. Every price below is the standard list price per million tokens, checked on the same day against the vendor's own endpoint where one exists, with its source linked.
The dataset covers the 19 models from 8 vendors that AIHub offers: OpenAI, Anthropic, Google, Meta, Mistral, Perplexity, DeepSeek and xAI. It is published as a table, a CSV and a chart, all free to reuse with attribution. The methodology section says exactly how each number was taken and, just as important, what was not measured.
Key findings
- Claude Opus 5.5 costs 40 times as much as GPT-6 Luna for the same request: $14.00 against $0.35 per 1,000 requests.
- Output is the expensive half. The median model charges 5x as much for an output token as for an input token, and 11 of 19 charge 5 times or more. Even when the prompt is twice as long as the answer, output is about 71% of the cost.
- Frontier prices have converged. GPT-6 Sol and Claude Sonnet 5 are both 2 dollars in and 10 dollars out per million tokens; Gemini 3.1 Pro is 2 and 12. The spread at the top is in the premium tier (Claude Opus 5.5) rather than between vendors.
- Open-weight models have no single price. DeepSeek V4.1 Flash is served by 28 hosts with input prices from $0.003 to $0.45 per million tokens, a 150x spread for the same weights.
- 4 of 19 headline prices changed in 11 days. Between 2026-09-23 and 2026-10-04, only one of those was a vendor repricing its own model; the rest were the aggregator headline following a different cheapest host.
- 5 models charge more for long prompts. Past 200,000 or 272,000 input tokens their rate per token rises, so the table price only holds below that threshold.
API prices for 19 models, per million tokens
Input and output are US dollars per million tokens at the standard tier. Cached input is the price of re-reading a prompt prefix the provider has cached, where the endpoint offers it. "1,000 requests" is the cost of 1,000 requests of 1,000 input and 500 output tokens. Models marked with an asterisk have no vendor endpoint, so the price is the median across hosts.
| Model | Input | Output | Cached input | Context | 1,000 requests | Source |
|---|---|---|---|---|---|---|
| GPT-6 Sol | $2.00 | $10.00 | $0.20 | 1.05M | $7.00 | OpenAI |
| GPT-6 Luna | $0.10 | $0.50 | $0.01 | 1.05M | $0.35 | OpenAI |
| GPT-5.4 Mini | $0.75 | $4.50 | $0.075 | 400K | $3.00 | OpenAI |
| Claude Opus 5.5 | $4.00 | $20.00 | $0.20 | 1M | $14.00 | Anthropic |
| Claude Sonnet 5 | $2.00 | $10.00 | $0.20 | 1M | $7.00 | Anthropic |
| Claude Haiku 4.5 | $1.00 | $5.00 | $0.10 | 200K | $3.50 | Anthropic |
| Gemini 3.1 Pro | $2.00 | $12.00 | $0.20 | 1.05M | $8.00 | Google AI Studio |
| Gemini 3.8 Flash | $0.75 | $3.75 | $0.075 | 1.05M | $2.63 | Google AI Studio |
| Gemini 3.5 Flash Lite | $0.30 | $2.50 | $0.03 | 1.05M | $1.55 | Google AI Studio |
| Llama 4 Maverick * | $0.31 | $0.925 | n/a | 1.05M | $0.77 | 4 hosts |
| Llama 3.3 70B * | $0.45 | $0.72 | n/a | 131K | $0.81 | 10 hosts |
| Mistral Medium 3.5 | $1.50 | $7.50 | n/a | 262K | $5.25 | Mistral |
| Mistral Small | $0.15 | $0.60 | $0.015 | 262K | $0.45 | Mistral |
| Perplexity Sonar Pro | $3.00 | $15.00 | n/a | 200K | $10.50 | Perplexity |
| Perplexity Sonar | $1.00 | $1.00 | n/a | 127K | $1.50 | Perplexity |
| DeepSeek V4 Pro * | $1.358 | $3.1675 | n/a | 1.05M | $2.94 | 16 hosts |
| DeepSeek V4.1 Flash | $0.30 | $1.20 | $0.006 | 1.05M | $0.90 | DeepSeek |
| Grok 4.7 | $2.00 | $6.00 | $0.50 | 500K | $5.00 | xAI |
| Grok 4.3 | $1.25 | $2.50 | $0.20 | 1M | $2.50 | xAI |
Why input and output are priced so differently
Reading a prompt is cheap for the provider because the whole input is processed in one parallel pass. Writing an answer is not: every output token needs its own pass through the model. That is why the output column is the one to watch. For a chat-style request with a 1,000-token prompt and a 500-token answer, the answer is the smaller half of the text and still about 71% of the cost on the median model. Reasoning models make this sharper, because the hidden reasoning they produce before answering is billed as output too.
The ratio is not uniform. Perplexity Sonar charges the same for input and output, while Gemini 3.5 Flash Lite charges more than 8 times as much for output. If your workload is mostly reading (classification, extraction, summarising long documents into short answers), a model with a low input price matters more than the headline output figure. If it is mostly writing (drafting, code generation), the output price is nearly the whole bill.
The same model, different prices
Closed models are sold at one list price by their vendor and resold at that price or a little above by cloud platforms; regional and data-residency endpoints typically add about 10 percent. Open-weight models are different, because anyone with GPUs can host them. These are the models in the dataset with three or more hosts on OpenRouter, ranked by how far apart the cheapest and most expensive host are on input price.
| Model | Hosts | Input range | Output range |
|---|---|---|---|
| DeepSeek V4.1 Flash | 28 | $0.003 to $0.45 | $0.18 to $2.40 |
| Llama 3.3 70B | 10 | $0.10 to $1.04 | $0.32 to $2.253 |
| DeepSeek V4 Pro | 16 | $0.2067 to $1.91 | $0.4176 to $4.20 |
| Llama 4 Maverick | 4 | $0.1875 to $0.35 | $0.6525 to $1.15 |
| GPT-6 Sol | 3 | $2.00 to $2.20 | $10.00 to $11.00 |
| Claude Opus 5.5 | 5 | $4.00 to $4.40 | $20.00 to $22.00 |
| Claude Sonnet 5 | 5 | $2.00 to $2.20 | $10.00 to $11.00 |
| Claude Haiku 4.5 | 4 | $1.00 to $1.10 | $5.00 to $5.50 |
| GPT-6 Luna | 3 | $0.10 to $0.11 | $0.50 to $0.55 |
The cheapest host is not always the one to use. The lowest prices often come with a shorter context window, a quantized model, or a host that is frequently unavailable, and some cheap listings pair a near-zero input price with an output price twice the median. Compare both columns for your own input-to-output mix before picking a host on its headline number.
How fast prices move
AIHub copies OpenRouter's headline price for every model into its own catalog. Comparing the copy from 2026-09-23 with the headline on 2026-10-04 shows how much a price quoted in an article can drift in under two weeks.
| Model | 2026-09-23 | 2026-10-04 | What changed |
|---|---|---|---|
| Llama 3.3 70B | $0.10 / $0.32 | $0.22 / $0.50 | cheapest host changed |
| DeepSeek V4 Pro | $0.9396 / $1.8792 | $0.2088 / $0.4176 | cheapest host changed |
| DeepSeek V4.1 Flash | $0.14 / $0.42 | $0.003 / $2.40 | cheapest host changed |
| Grok 4.7 | $1.60 / $4.80 | $2.00 / $6.00 | vendor price change |
Prices are input / output in dollars per million tokens. The practical lesson for anyone budgeting an AI feature: write the date next to every price you quote, and re-check before a launch or a contract renewal. The other 15 models held their price over the same period.
Long prompts and cached prompts
Two adjustments change the effective price more than most comparisons admit. The first is a long-context surcharge: GPT-6 Sol, GPT-6 Luna, Gemini 3.1 Pro, Grok 4.7 and Grok 4.3 switch to a higher rate once a single request passes 200K or 272K input tokens: input doubles and output rises by half (OpenAI, Google) or doubles (xAI). If you send whole codebases or long contracts in one request, the table price understates your cost.
The second runs the other way. 13 of the 19 models offer cached input, where a prompt prefix you send repeatedly (a long system prompt, a reference document) is billed at a fraction of the normal input price on later requests, most often a tenth. A chatbot that sends the same 3,000-token instructions with every message can cut its input bill by most of that amount, which is why the cached column can matter more than the input column for production apps.
What a month of usage costs
Per-million prices are hard to feel. Here is the same arithmetic as a monthly bill: 3,000 requests a month, about 100 a day, each with 1,000 tokens in and 500 out, at the standard price and with no caching.
| Model | Per request | 3,000 requests a month |
|---|---|---|
| Claude Opus 5.5 | $0.014 | $42.00 |
| GPT-6 Sol | $0.007 | $21.00 |
| Gemini 3.1 Pro | $0.008 | $24.00 |
| Grok 4.7 | $0.005 | $15.00 |
| GPT-6 Luna | $0.00035 | $1.05 |
At this volume the frontier models land between $15.00 and $42.00 a month, the same range as a 20 dollar consumer chat subscription, while the smallest models cost about a dollar. That trade is the comparison made in Claude Pro vs ChatGPT Plus and in the ChatGPT Plus alternative guide. The catch is that real conversations resend earlier turns, so a ten-turn chat costs far more than ten single requests; how to compare AI models side by side shows the effect on a real test.
Methodology
Source. Every price was read from OpenRouter's public endpoints API (/api/v1/models/{id}/endpoints) on 2026-10-04, one call per model, by a script that writes this page's dataset. Each
model's source link points to its OpenRouter page, where the same endpoint list is
visible.
Which price counts. For each model we took the vendor's own endpoint (16 of 19 models): OpenAI for GPT, Anthropic for Claude, Google AI Studio for Gemini, and the vendor itself for Mistral, Perplexity, DeepSeek and xAI. Discounted and premium service tiers (flex, fast, priority and batch) were excluded, and where a vendor lists several standard endpoints the lowest was used, which removes regional surcharges. The 3 models with no vendor endpoint (both Llama models and DeepSeek V4 Pro) use the median of their hosts' standard prices.
Spot check. Anthropic's own pricing page, read the same day, lists Claude Opus 5.5 at 4 and 20 dollars, Claude Sonnet 5 at 2 and 10, and Claude Haiku 4.5 at 1 and 5 per million input and output tokens, with cached reads at a tenth of input for Sonnet and Haiku. All match the dataset. The other vendors' pages were not checked by hand.
The request. "1,000 requests" assumes 1,000 input and 500 output tokens per request, with no caching, no tools and no reasoning tokens. It is a convenient yardstick, not a typical workload; recompute with your own mix from the CSV.
What was not measured.
- Quality, speed and reliability. A cheaper model that needs two attempts is not cheaper.
- Tokenizer differences. Vendors count tokens differently, so the same text can be a different number of tokens on each model and per-token prices are not perfectly comparable.
- Reasoning tokens. Models that think before answering bill that hidden text as output, which can multiply the real cost of a request.
- Discounts and surcharges: batch and flex tiers (often half price), committed-use and enterprise contracts, regional and data-residency premiums, long-context rates beyond the thresholds noted above, and taxes.
- Tool and media fees: web search, image and audio input, and code execution are priced separately and not included.
- Consumer subscriptions such as ChatGPT Plus or Claude Pro, free tiers, and prices in currencies other than US dollars.
- Which host an aggregator actually routes a request to. A router may send traffic to a cheaper or more expensive host than the reference price used here.
Disclosure. AIHub sells access to these 19 models on a subscription, which is why it tracks their prices. The prices on this page are the vendors' and hosts' list prices, not AIHub's, and nothing on the page is a paid placement.
Reuse and citation
The table, the CSV and the chart are published under Creative Commons Attribution 4.0. Use them in articles, slides or your own tools, with a link back. Suggested citation:
AIHub, "AI model API pricing 2026", prices checked 2026-10-04, https://aihub.group/data/ai-model-api-pricing/ (CC BY 4.0).
This page is updated when prices change materially, and each update keeps its own dated CSV. For how the frontier models compare on actual work rather than price, see ChatGPT vs Claude vs Gemini.