
Topics Covered:
Helpful tools for this guide
Table of Contents
- Quick Answer: What Does 1 Million Tokens Cost?
- Why Isn't There One AI Token Price?
- AI Token Cost Comparison — Verified September 26, 2026
- How Much Do OpenAI Tokens Cost?
- How Much Do Claude Tokens Cost?
- How Much Do Gemini Tokens Cost?
- Which Model Would We Actually Use?
- 7 Worked Cost Examples
- How Caching Cuts Your Bill
- How to Calculate Your Blended Rate
- Frequently Asked Questions
- How much does 1 million tokens cost right now?
- What happened to Claude Opus 5?
- Is Claude Sonnet 5 still $2 per million input tokens?
- Will Gemini 3.6 Flash get more expensive?
- What's the fastest way to cut my AI token bill this month?
- Final Takeaway
Last verified: September 26, 2026. What changed since this was first published (Aug 25, 2026):
- Claude Opus 4.8 → replaced twice: Opus 5 (Jul 24), then Opus 5.5 (Sep 22) at $4/$20, 20% cheaper than Opus 5
- Claude Sonnet 5's planned Sept 1 price increase to $3/$15 was cancelled — still $2/$10
- Gemini 3.6 Flash dropped from $1.50/$7.50 to an introductory $0.75/$3.75, rising again Jan 1, 2027
AI pricing looks simple until you try to calculate a real bill. Input and output tokens are billed at different rates, and those rates change more often than most guides admit — this page was checked against Anthropic, OpenAI, and Google's own pricing pages on September 26, 2026, not copied from another article. Right now, current-generation models range from $0.20 to $10 per million input tokens and $1.20 to $50 per million output tokens. The model, the direction (input vs. output), and the date you're reading this all decide what you actually pay.
Quick Answer: What Does 1 Million Tokens Cost?
There is no single price. Every provider splits input and output pricing, and every model tier has its own rate. GPT-5.6 Luna costs $0.20 per million input tokens and $1.20 for output. Claude Fable 5.1, Anthropic's most capable model, costs $10 for input and $50 for output — 50x more per input token than Luna. The same million tokens can cost a few cents or tens of dollars depending on which model sent the bill.
Quick rule: never estimate AI cost from token count alone. You need the model, the input volume, and the expected output volume.
Why Isn't There One AI Token Price?
Treat tokens like electricity: knowing your usage tells you nothing until you know the rate per unit, and every model has its own rate. Providers split pricing further still — a cached prompt, a batch job, and a live request can cost different amounts for the exact same tokens. That's why a search for "AI token cost" surfaces several different correct-sounding numbers instead of one.
AI Token Cost Comparison — Verified September 26, 2026
| Model | Input / MTok | Output / MTok | Note |
|---|---|---|---|
| GPT-5.6 Luna | $0.20 | $1.20 | Promotional rate through Nov 21, 2026 |
| GPT-5.6 Terra | $2.00 | $12.00 | Promotional rate through Nov 21, 2026 |
| GPT-5.6 Sol | $4.00 | $20.00 | Promotional; standard rate is $5 / $30 |
| Claude Haiku 4.5 | $1.00 | $5.00 | Unchanged since launch |
| Claude Sonnet 5 | $2.00 | $10.00 | Planned Sept 1 increase to $3/$15 was cancelled |
| Claude Opus 5.5 | $4.00 | $20.00 | Replaced Opus 5 ($5/$25) on Sept 22, 2026 |
| Claude Fable 5.1 | $10.00 | $50.00 | Frontier tier |
| Gemini 3.5 Flash-Lite | $0.30 | $2.50 | Cheapest current Gemini model |
| Gemini 3.6 Flash | $0.75 | $3.75 | Introductory; rises to $1.50/$7.50 on Jan 1, 2027 |
| Gemini 3.1 Pro Preview | $2.00 / $4.00 | $12.00 / $18.00 | Higher rate applies above 200K input tokens |
Sources: Anthropic pricing, OpenAI pricing, and Google Gemini pricing. Confirm current rates before budgeting a production workload — three of these ten prices changed in the last month alone.
How Much Do OpenAI Tokens Cost?
OpenAI's GPT-5.6 family ships in three tiers instead of one flat rate. Sol is the flagship at a promotional $4 input / $20 output per MTok (standard rate $5/$30 resumes after November 21). Terra sits in the middle at $2/$12, pitched as GPT-5.5-level performance at roughly half the cost. Luna is the volume tier at $0.20/$1.20 — cheap enough that even careless prompting rarely becomes a real bill.
How Much Do Claude Tokens Cost?
Anthropic quietly replaced its mid-flagship model on September 22 — Claude Opus 5.5 now costs $4 input / $20 output per MTok, 20% cheaper than Opus 5's $5/$25 on both input and output. If a guide you're reading still quotes Opus 5 or Opus 4.8, it's already out of date.
Claude Sonnet 5 stays at its introductory $2/$10 rate — Anthropic had scheduled a jump to $3/$15 for September 1, then cancelled it. Claude Haiku 4.5 remains the cheapest Claude model at $1/$5, and Claude Fable 5.1 sits at the top at $10/$50 for the hardest reasoning tasks.
How Much Do Gemini Tokens Cost?
Gemini 3.6 Flash currently costs $0.75 input / $3.75 output per MTok — but that's introductory. Google's own pricing page confirms it doubles to $1.50/$7.50 on January 1, 2027, so anything you build on it now gets meaningfully pricier in a few months unless you re-check then. Gemini 3.5 Flash-Lite is the budget option at $0.30/$2.50, and Gemini 3.1 Pro Preview charges $2/$12 for prompts at or under 200K tokens, jumping to $4/$18 above that — one of the few models here where prompt length itself changes your rate.
Which Model Would We Actually Use?
Prices alone don't tell you what to pick. For most production apps as of this update, Claude Sonnet 5 at $2/$10 is the best cost-to-capability tradeoff we'd reach for first — it's priced like a mid-tier model but benchmarks closer to what used to require a flagship. Reserve Claude Opus 5.5 or GPT-5.6 Sol for tasks where Sonnet's output needs a manual rework pass; if you're not seeing that, you're likely overpaying. For high-volume, low-stakes work — classification, tagging, routing — GPT-5.6 Luna's $0.20 input rate is hard to beat on pure economics, though Gemini 3.5 Flash-Lite is worth benchmarking against it for your specific task before committing.
7 Worked Cost Examples
| Workload | Model | Tokens (in/out) | Estimated cost |
|---|---|---|---|
| Customer support reply bot, 10K conversations/mo | GPT-5.6 Luna | 5M in / 2M out | $1.00 + $2.40 = $3.40 |
| Daily blog post generation | Claude Sonnet 5 | 3K in / 1.5K out | $0.006 + $0.015 = $0.021/post |
| Codebase refactor session | Claude Opus 5.5 | 150K in / 8K out | $0.60 + $0.16 = $0.76 |
| Research agent, long documents | Gemini 3.1 Pro (>200K) | 250K in / 10K out | $1.00 + $0.18 = $1.18 |
| High-volume data tagging, 1M rows | Gemini 3.5 Flash-Lite | 50M in / 5M out | $15.00 + $12.50 = $27.50 |
| Monthly chatbot at scale | GPT-5.6 Terra | 20M in / 5M out | $40.00 + $60.00 = $100.00 |
| Frontier reasoning task, one-off | Claude Fable 5.1 | 20K in / 5K out | $0.20 + $0.25 = $0.45 |
These are calculated at list-price rates above — not pulled from an actual invoice. Your real bill will differ once caching, retries, and system-prompt overhead are factored in.
How Caching Cuts Your Bill
Every provider here discounts repeated input. Claude Opus 5.5 cache reads run about $0.20 per MTok against a $4 base input rate — roughly a 95% discount on anything reused, like a long system prompt or reference document. OpenAI's GPT-5.6 tiers price cached input at 80–90% off standard, and Gemini's context caching works the same way plus an hourly storage fee for keeping the cache warm. If your app resends the same instructions or documents on every call, caching is the biggest lever you're not pulling.
How to Calculate Your Blended Rate
Input cost = input tokens ÷ 1,000,000 × input rateOutput cost = output tokens ÷ 1,000,000 × output rate
Example: an app sending 5 million input tokens and 1 million output tokens through Claude Sonnet 5 in a month pays 5 × $2 = $10 for input and 1 × $10 = $10 for output — $20 total, even though output is only a sixth of the volume. That pattern holds across nearly every model here: output tokens typically cost 5x input tokens, so long generated answers get expensive fast even with modest prompts.
Frequently Asked Questions
How much does 1 million tokens cost right now?
Between $0.20 and $10 for input, and $1.20 to $50 for output, depending on the model. See the comparison table above for exact current rates.
What happened to Claude Opus 5?
Anthropic replaced it with Claude Opus 5.5 on September 22, 2026, at a 20% lower rate ($4/$20 vs. $5/$25) — the same pattern Anthropic has followed with every Opus refresh.
Is Claude Sonnet 5 still $2 per million input tokens?
Yes. Anthropic had planned to raise it to $3/$15 on September 1, 2026, but cancelled that increase.
Will Gemini 3.6 Flash get more expensive?
Yes — its $0.75/$3.75 rate is introductory and doubles to $1.50/$7.50 on January 1, 2027, per Google's own pricing page.
What's the fastest way to cut my AI token bill this month?
Turn on prompt caching for anything you send repeatedly (system prompts, reference docs), switch high-volume low-stakes tasks to the cheapest tier your provider offers, and check whether your workload's output length — not its prompt length — is what's actually driving the bill.
Final Takeaway
Model pricing in this space now shifts every few weeks — three of the ten models in this guide changed price in the last month. Check the model's current rate directly before budgeting a project, and pull your real prompt and response sizes into the AI Token Counter & Cost Calculator instead of estimating from a price card alone.
Related Articles
Continue with closely related CountFlows guides.
AI
Why ChatGPT, Claude, and Gemini Stop Mid-Sentence
Discover the three unrelated reasons AI models stop mid-sentence and the exact steps to fix output caps, context window overflows, and connection issues.
AI
Is an Em Dash a Sign of AI? Why AI Uses Em Dashes (2026)
Learn why em dashes are associated with AI writing, why AI tools use them, whether they can identify AI-generated text, and how to remove them when needed.
AI
Why AI Chatbots Can't Count Syllables (And How to Fix Them)
AI chatbots can explain the 5-7-5 haiku rule perfectly, yet still produce lines with the wrong syllable count. Learn why tokens and sounds do not match, why AI-generated lyrics often fail to fit melodies, and how to fix the problem with an external syllable counter.

