Back to AIHub

Data

AI model API pricing 2026: price data for 19 models from 8 vendors

Checked CSV CC BY 4.0

AI model API pricing in 2026 spans a 40x range. On 2026-10-04, a typical request of 1,000 input tokens and 500 output tokens cost $0.00035 on GPT-6 Luna and $0.014 on Claude Opus 5.5; across the 19 models in this dataset the median is $2.94 per 1,000 requests. Every price below is the standard list price per million tokens, checked on the same day against the vendor's own endpoint where one exists, with its source linked.

The dataset covers the 19 models from 8 vendors that AIHub offers: OpenAI, Anthropic, Google, Meta, Mistral, Perplexity, DeepSeek and xAI. It is published as a table, a CSV and a chart, all free to reuse with attribution. The methodology section says exactly how each number was taken and, just as important, what was not measured.

Key findings

  • Claude Opus 5.5 costs 40 times as much as GPT-6 Luna for the same request: $14.00 against $0.35 per 1,000 requests.
  • Output is the expensive half. The median model charges 5x as much for an output token as for an input token, and 11 of 19 charge 5 times or more. Even when the prompt is twice as long as the answer, output is about 71% of the cost.
  • Frontier prices have converged. GPT-6 Sol and Claude Sonnet 5 are both 2 dollars in and 10 dollars out per million tokens; Gemini 3.1 Pro is 2 and 12. The spread at the top is in the premium tier (Claude Opus 5.5) rather than between vendors.
  • Open-weight models have no single price. DeepSeek V4.1 Flash is served by 28 hosts with input prices from $0.003 to $0.45 per million tokens, a 150x spread for the same weights.
  • 4 of 19 headline prices changed in 11 days. Between 2026-09-23 and 2026-10-04, only one of those was a vendor repricing its own model; the rest were the aggregator headline following a different cheapest host.
  • 5 models charge more for long prompts. Past 200,000 or 272,000 input tokens their rate per token rises, so the table price only holds below that threshold.
Bar chart of the cost of 1,000 API requests (1,000 input and 500 output tokens each) for 19 AI models, checked 2026-10-04. Claude Opus 5.5 $14.00, Perplexity Sonar Pro $10.50, Gemini 3.1 Pro $8.00, GPT-6 Sol $7.00, Claude Sonnet 5 $7.00, Mistral Medium 3.5 $5.25, Grok 4.7 $5.00, Claude Haiku 4.5 $3.50, GPT-5.4 Mini $3.00, DeepSeek V4 Pro $2.94, Gemini 3.8 Flash $2.63, Grok 4.3 $2.50, Gemini 3.5 Flash Lite $1.55, Perplexity Sonar $1.50, DeepSeek V4.1 Flash $0.90, Llama 3.3 70B $0.81, Llama 4 Maverick $0.77, Mistral Small $0.45, GPT-6 Luna $0.35.
Cost of 1,000 requests of 1,000 input and 500 output tokens, standard list prices, 2026-10-04. Download the chart (PNG) or the data (CSV). Both are CC BY 4.0: reuse them anywhere with a link to this page.

API prices for 19 models, per million tokens

Input and output are US dollars per million tokens at the standard tier. Cached input is the price of re-reading a prompt prefix the provider has cached, where the endpoint offers it. "1,000 requests" is the cost of 1,000 requests of 1,000 input and 500 output tokens. Models marked with an asterisk have no vendor endpoint, so the price is the median across hosts.

Model Input Output Cached input Context 1,000 requests Source
GPT-6 Sol $2.00 $10.00 $0.20 1.05M $7.00 OpenAI
GPT-6 Luna $0.10 $0.50 $0.01 1.05M $0.35 OpenAI
GPT-5.4 Mini $0.75 $4.50 $0.075 400K $3.00 OpenAI
Claude Opus 5.5 $4.00 $20.00 $0.20 1M $14.00 Anthropic
Claude Sonnet 5 $2.00 $10.00 $0.20 1M $7.00 Anthropic
Claude Haiku 4.5 $1.00 $5.00 $0.10 200K $3.50 Anthropic
Gemini 3.1 Pro $2.00 $12.00 $0.20 1.05M $8.00 Google AI Studio
Gemini 3.8 Flash $0.75 $3.75 $0.075 1.05M $2.63 Google AI Studio
Gemini 3.5 Flash Lite $0.30 $2.50 $0.03 1.05M $1.55 Google AI Studio
Llama 4 Maverick * $0.31 $0.925 n/a 1.05M $0.77 4 hosts
Llama 3.3 70B * $0.45 $0.72 n/a 131K $0.81 10 hosts
Mistral Medium 3.5 $1.50 $7.50 n/a 262K $5.25 Mistral
Mistral Small $0.15 $0.60 $0.015 262K $0.45 Mistral
Perplexity Sonar Pro $3.00 $15.00 n/a 200K $10.50 Perplexity
Perplexity Sonar $1.00 $1.00 n/a 127K $1.50 Perplexity
DeepSeek V4 Pro * $1.358 $3.1675 n/a 1.05M $2.94 16 hosts
DeepSeek V4.1 Flash $0.30 $1.20 $0.006 1.05M $0.90 DeepSeek
Grok 4.7 $2.00 $6.00 $0.50 500K $5.00 xAI
Grok 4.3 $1.25 $2.50 $0.20 1M $2.50 xAI

Why input and output are priced so differently

Reading a prompt is cheap for the provider because the whole input is processed in one parallel pass. Writing an answer is not: every output token needs its own pass through the model. That is why the output column is the one to watch. For a chat-style request with a 1,000-token prompt and a 500-token answer, the answer is the smaller half of the text and still about 71% of the cost on the median model. Reasoning models make this sharper, because the hidden reasoning they produce before answering is billed as output too.

The ratio is not uniform. Perplexity Sonar charges the same for input and output, while Gemini 3.5 Flash Lite charges more than 8 times as much for output. If your workload is mostly reading (classification, extraction, summarising long documents into short answers), a model with a low input price matters more than the headline output figure. If it is mostly writing (drafting, code generation), the output price is nearly the whole bill.

The same model, different prices

Closed models are sold at one list price by their vendor and resold at that price or a little above by cloud platforms; regional and data-residency endpoints typically add about 10 percent. Open-weight models are different, because anyone with GPUs can host them. These are the models in the dataset with three or more hosts on OpenRouter, ranked by how far apart the cheapest and most expensive host are on input price.

Model Hosts Input range Output range
DeepSeek V4.1 Flash 28 $0.003 to $0.45 $0.18 to $2.40
Llama 3.3 70B 10 $0.10 to $1.04 $0.32 to $2.253
DeepSeek V4 Pro 16 $0.2067 to $1.91 $0.4176 to $4.20
Llama 4 Maverick 4 $0.1875 to $0.35 $0.6525 to $1.15
GPT-6 Sol 3 $2.00 to $2.20 $10.00 to $11.00
Claude Opus 5.5 5 $4.00 to $4.40 $20.00 to $22.00
Claude Sonnet 5 5 $2.00 to $2.20 $10.00 to $11.00
Claude Haiku 4.5 4 $1.00 to $1.10 $5.00 to $5.50
GPT-6 Luna 3 $0.10 to $0.11 $0.50 to $0.55

The cheapest host is not always the one to use. The lowest prices often come with a shorter context window, a quantized model, or a host that is frequently unavailable, and some cheap listings pair a near-zero input price with an output price twice the median. Compare both columns for your own input-to-output mix before picking a host on its headline number.

How fast prices move

AIHub copies OpenRouter's headline price for every model into its own catalog. Comparing the copy from 2026-09-23 with the headline on 2026-10-04 shows how much a price quoted in an article can drift in under two weeks.

Model 2026-09-23 2026-10-04 What changed
Llama 3.3 70B $0.10 / $0.32 $0.22 / $0.50 cheapest host changed
DeepSeek V4 Pro $0.9396 / $1.8792 $0.2088 / $0.4176 cheapest host changed
DeepSeek V4.1 Flash $0.14 / $0.42 $0.003 / $2.40 cheapest host changed
Grok 4.7 $1.60 / $4.80 $2.00 / $6.00 vendor price change

Prices are input / output in dollars per million tokens. The practical lesson for anyone budgeting an AI feature: write the date next to every price you quote, and re-check before a launch or a contract renewal. The other 15 models held their price over the same period.

Long prompts and cached prompts

Two adjustments change the effective price more than most comparisons admit. The first is a long-context surcharge: GPT-6 Sol, GPT-6 Luna, Gemini 3.1 Pro, Grok 4.7 and Grok 4.3 switch to a higher rate once a single request passes 200K or 272K input tokens: input doubles and output rises by half (OpenAI, Google) or doubles (xAI). If you send whole codebases or long contracts in one request, the table price understates your cost.

The second runs the other way. 13 of the 19 models offer cached input, where a prompt prefix you send repeatedly (a long system prompt, a reference document) is billed at a fraction of the normal input price on later requests, most often a tenth. A chatbot that sends the same 3,000-token instructions with every message can cut its input bill by most of that amount, which is why the cached column can matter more than the input column for production apps.

What a month of usage costs

Per-million prices are hard to feel. Here is the same arithmetic as a monthly bill: 3,000 requests a month, about 100 a day, each with 1,000 tokens in and 500 out, at the standard price and with no caching.

Model Per request 3,000 requests a month
Claude Opus 5.5 $0.014 $42.00
GPT-6 Sol $0.007 $21.00
Gemini 3.1 Pro $0.008 $24.00
Grok 4.7 $0.005 $15.00
GPT-6 Luna $0.00035 $1.05

At this volume the frontier models land between $15.00 and $42.00 a month, the same range as a 20 dollar consumer chat subscription, while the smallest models cost about a dollar. That trade is the comparison made in Claude Pro vs ChatGPT Plus and in the ChatGPT Plus alternative guide. The catch is that real conversations resend earlier turns, so a ten-turn chat costs far more than ten single requests; how to compare AI models side by side shows the effect on a real test.

Methodology

Source. Every price was read from OpenRouter's public endpoints API (/api/v1/models/{id}/endpoints) on 2026-10-04, one call per model, by a script that writes this page's dataset. Each model's source link points to its OpenRouter page, where the same endpoint list is visible.

Which price counts. For each model we took the vendor's own endpoint (16 of 19 models): OpenAI for GPT, Anthropic for Claude, Google AI Studio for Gemini, and the vendor itself for Mistral, Perplexity, DeepSeek and xAI. Discounted and premium service tiers (flex, fast, priority and batch) were excluded, and where a vendor lists several standard endpoints the lowest was used, which removes regional surcharges. The 3 models with no vendor endpoint (both Llama models and DeepSeek V4 Pro) use the median of their hosts' standard prices.

Spot check. Anthropic's own pricing page, read the same day, lists Claude Opus 5.5 at 4 and 20 dollars, Claude Sonnet 5 at 2 and 10, and Claude Haiku 4.5 at 1 and 5 per million input and output tokens, with cached reads at a tenth of input for Sonnet and Haiku. All match the dataset. The other vendors' pages were not checked by hand.

The request. "1,000 requests" assumes 1,000 input and 500 output tokens per request, with no caching, no tools and no reasoning tokens. It is a convenient yardstick, not a typical workload; recompute with your own mix from the CSV.

What was not measured.

  • Quality, speed and reliability. A cheaper model that needs two attempts is not cheaper.
  • Tokenizer differences. Vendors count tokens differently, so the same text can be a different number of tokens on each model and per-token prices are not perfectly comparable.
  • Reasoning tokens. Models that think before answering bill that hidden text as output, which can multiply the real cost of a request.
  • Discounts and surcharges: batch and flex tiers (often half price), committed-use and enterprise contracts, regional and data-residency premiums, long-context rates beyond the thresholds noted above, and taxes.
  • Tool and media fees: web search, image and audio input, and code execution are priced separately and not included.
  • Consumer subscriptions such as ChatGPT Plus or Claude Pro, free tiers, and prices in currencies other than US dollars.
  • Which host an aggregator actually routes a request to. A router may send traffic to a cheaper or more expensive host than the reference price used here.

Disclosure. AIHub sells access to these 19 models on a subscription, which is why it tracks their prices. The prices on this page are the vendors' and hosts' list prices, not AIHub's, and nothing on the page is a paid placement.

Reuse and citation

The table, the CSV and the chart are published under Creative Commons Attribution 4.0. Use them in articles, slides or your own tools, with a link back. Suggested citation:

AIHub, "AI model API pricing 2026", prices checked 2026-10-04, https://aihub.group/data/ai-model-api-pricing/ (CC BY 4.0).

This page is updated when prices change materially, and each update keeps its own dated CSV. For how the frontier models compare on actual work rather than price, see ChatGPT vs Claude vs Gemini.

Frequently asked questions

Which AI model API is the cheapest in 2026?

In this dataset, checked 2026-10-04, the cheapest model is GPT-6 Luna at $0.10 per million input tokens and $0.50 per million output tokens, which puts 1,000 typical requests at $0.35. The most expensive is Claude Opus 5.5 at $14.00 for the same 1,000 requests.

How much does 1 million tokens cost?

It depends on the model and on whether the tokens are input or output. Across 19 models the input price runs from $0.10 to $4.00 per million tokens and the output price from $0.50 to $20.00. Frontier models from OpenAI, Anthropic and Google sit between 2 and 4 dollars in and 10 and 20 dollars out.

Why do output tokens cost more than input tokens?

Generating a token takes a full pass through the model for every token produced, while input tokens are processed in parallel. Vendors price that in: the median model here charges 5x as much per output token as per input token, and 11 of 19 charge 5 times or more. For a request with twice as much input as output, output is still about 71% of the bill.

Why does the same model have different prices on different sites?

Open-weight models such as Llama and DeepSeek are served by many independent hosts, and each sets its own price. DeepSeek V4.1 Flash is listed by 28 hosts on OpenRouter, with input prices from $0.003 to $0.45 per million tokens. Aggregators often show the cheapest host as the model's headline price, so that number moves whenever a host joins, leaves or changes its rate.

Featured on Nick Launches