
This post compares per-token API prices across seven providers as of August 2026, and shows why the headline rate is rarely what you end up paying.
Providers quote USD per million tokens for input (prompt) and output (reply). Headline rates are a starting point; caching, long-context surcharges, modality coverage, and output multiples move real bills further.
This page is a snapshot from early August 2026. Token prices change frequently; confirm rates at each provider’s pricing page before budgeting.
Seven providers: Anthropic (Claude), OpenAI (GPT), Google (Gemini), DeepSeek, Moonshot (Kimi), Alibaba (Qwen), xAI (Grok). A combined table sorted by input price follows the per-provider sections.
Conventions:
- Context window — maximum tokens (prompt + reply) the model holds at once. A dash (—) means no published headline window.
- 262K vs 256K, 1.05M vs 1M — decimal vs binary unit conventions for the same capacity.
1 Data provenance
Figures are compiled from secondary sources (pricing trackers and explainer posts; full list under References), except the Anthropic table, which was checked against Anthropic’s published documentation.
- Objective: compare rough price bands, long-context surcharges, modality coverage, and output multiples.
- Not suitable for: procurement or volume projections — use first-party pricing pages.
- Downstream impact: a stale rate misstates budget and tier selection.
Reliable for band-level conclusions that survive ~10% figure error. Not a billing source of truth.
2 Anthropic
Anthropic prices on a tiered ladder (small → frontier) with no long-context surcharge on models 4.6+ (flat rate regardless of prompt length). Haiku 4.5 predates that guarantee but has a 200K window with no long-context tier.
| Model | Input $/M | Output $/M | Context | Notes |
|---|---|---|---|---|
| Haiku 4.5 | 1.00 | 5.00 | 200K | Cheapest tier; the one current model not at 1M |
| Sonnet 5 | 2.00 | 10.00 | 1M | Intro rate through 31 Aug 2026; then 3.00 / 15.00 |
| Opus 5 | 5.00 | 25.00 | 1M | |
| Fable 5 | 10.00 | 50.00 | 1M | Frontier tier |
| Mythos 5 | 10.00 | 50.00 | 1M | Same capabilities and price as Fable 5, but not generally available — Project Glasswing participants only |
Modalities: text, image, PDF/documents.
Discounts:
- Prompt caching — ~90% off cached input reads; cache writes cost 1.25× base input (5-minute TTL) or 2× (1-hour TTL). Pays from second request at 5-minute TTL, third at 1-hour TTL.
- Batch API — 50% off input and output for deferred jobs.
3 OpenAI
Widest price range among Western majors. Long-context surcharge above 272K tokens. Audio sold as a separate model, not folded into text tiers.
| Model | Input $/M | Output $/M | Context | Notes |
|---|---|---|---|---|
| GPT-5.6 Luna | 0.20 | 1.20 | 1.05M | Budget tier |
| GPT-5.6 Terra | 2.00 | 12.00 | 1.05M | Mid tier |
| GPT-5.6 Sol | 5.00 | 30.00 | 1.05M | Flagship |
| GPT-Realtime-2.1 | 32.00 | 64.00 | — | Audio in / audio out, per M audio tokens |
Modalities: text and image on GPT-5.6 line; audio via GPT-Realtime-2.1 only.
Long-context surcharge: above 272K tokens, input 2× and output 1.5×.
4 Google Gemini
Audio and video are first-class inputs on every tier. Also sells image and video generation per unit. Long-context surcharge at a lower threshold than OpenAI.
| Model | Input $/M | Output $/M | Context | Notes |
|---|---|---|---|---|
| Gemini 3.5 Flash-Lite | 0.30 | 2.50 | — | Cheapest tier |
| Gemini 3.6 Flash | 1.50 | 7.50 | — | Mid tier |
| Gemini 3.1 Pro (≤200K) | 2.00 | 12.00 | 2M | Rate below the 200K threshold |
| Gemini 3.1 Pro (>200K) | 4.00 | 18.00 | 2M | Long-context rate |
Modalities: text, image, audio, video. Image generation from $0.02/image; video generation from $0.15/second.
Discounts: Batch 50% off. Caching charges rent on storage ($1–4.50 per M tokens per hour), not a write multiple — break-even depends on hold time and reuse.
5 DeepSeek
Lowest output price in this comparison; steepest cache-hit rate. Text-only; no batch tier.
| Model | Input $/M | Output $/M | Context | Notes |
|---|---|---|---|---|
| V4-Flash | 0.14 | 0.28 | 1M | Cache-hit input $0.0028 |
| V4-Pro | 0.435 | 0.87 | 1M | Cache-hit input $0.0036 |
Modalities: text.
Caveats: no batch tier; 2× peak-hour pricing announced without start date. V4-Pro cache-hit rate (99.2% off base) is unusually deeper than V4-Flash (98%) — verify at source before relying on it.
Against Western flagships, V4-Flash input is 36× cheaper than Opus 5; against budget tiers the gap falls to ~1.4× vs GPT-5.6 Luna.
6 Moonshot Kimi
Mid-range pricing; native video input (only non-Google provider here with video). Weakest batch tier on this page (40% discount vs 50% elsewhere); flagship K3 excluded from batch.
| Model | Input $/M | Output $/M | Context | Notes |
|---|---|---|---|---|
| K2.5 | 0.60 | 3.00 | 262K | Batch eligible |
| K2.6 / K2.7 Code | 0.95 | 4.00 | 256K | Batch eligible |
| K3 | 3.00 | 15.00 | 1M | Flagship; cache-hit input $0.30; not batch eligible |
Modalities: text, image, video.
Discounts: batch bills at 60% of standard rate (40% discount); K2.5/K2.6/K2.7 only.
7 Alibaba Qwen
Headline numbers are international (Singapore) rates; Mainland China endpoint is 60–70% cheaper with no free quota.
| Model | Input $/M | Output $/M | Context | Notes |
|---|---|---|---|---|
| Qwen3.5 Flash | 0.10 | 0.40 | — | Lowest input price in this comparison |
| Qwen3.5 Plus | 0.40 | 2.40 | 1M | |
| Qwen3.8-Max | 2.00 | 6.00 | 1M | Flagship; cached input $0.25 |
Modalities: text on Flash (some image variants); text and image on Plus and Max.
8 xAI Grok
Differentiator is live X/Twitter data as a first-class input. Output multiple (output/input ratio) is 2–3×; most Western text tiers sit at 5–6×; Gemini 3.5 Flash-Lite reaches 8.3×.
| Model | Input $/M | Output $/M | Context | Notes |
|---|---|---|---|---|
| Grok 4.1 Fast | 0.20 | 0.50 | 2M | Joint-largest context here, tied with Gemini 3.1 Pro |
| Grok 4.3 | 1.25 | 2.50 | 1M | |
| Grok 4.5 | 2.00 | 6.00 | 500K | Flagship; cached input $0.50; no batch discount |
Modalities: text, image, live X/Twitter data.
9 Combined table
Sorted by input price. Five models share $2.00 input across five providers.
| Provider | Model | Input $/M | Output $/M | Context | Modalities | Notes |
|---|---|---|---|---|---|---|
| Alibaba | Qwen3.5 Flash | 0.10 | 0.40 | — | text | some image variants |
| DeepSeek | V4-Flash | 0.14 | 0.28 | 1M | text | |
| OpenAI | GPT-5.6 Luna | 0.20 | 1.20 | 1.05M | text, image | |
| xAI | Grok 4.1 Fast | 0.20 | 0.50 | 2M | text, image, live X | |
| Gemini 3.5 Flash-Lite | 0.30 | 2.50 | — | text, image, audio, video | ||
| Alibaba | Qwen3.5 Plus | 0.40 | 2.40 | 1M | text, image | |
| DeepSeek | V4-Pro | 0.435 | 0.87 | 1M | text | |
| Moonshot | K2.5 | 0.60 | 3.00 | 262K | text, image, video | |
| Moonshot | K2.6 / K2.7 Code | 0.95 | 4.00 | 256K | text, image, video | same window as K2.5, binary units |
| Anthropic | Haiku 4.5 | 1.00 | 5.00 | 200K | text, image, PDF | |
| xAI | Grok 4.3 | 1.25 | 2.50 | 1M | text, image, live X | |
| Gemini 3.6 Flash | 1.50 | 7.50 | — | text, image, audio, video | ||
| Anthropic | Sonnet 5 | 2.00 | 10.00 | 1M | text, image, PDF | intro rate to 31 Aug 2026; then 3.00 / 15.00 |
| OpenAI | GPT-5.6 Terra | 2.00 | 12.00 | 1.05M | text, image | |
| Gemini 3.1 Pro | 2.00 | 12.00 | 2M | text, image, audio, video | rate at ≤200K | |
| Alibaba | Qwen3.8-Max | 2.00 | 6.00 | 1M | text, image | |
| xAI | Grok 4.5 | 2.00 | 6.00 | 500K | text, image, live X | |
| Moonshot | K3 | 3.00 | 15.00 | 1M | text, image, video | |
| Gemini 3.1 Pro | 4.00 | 18.00 | 2M | text, image, audio, video | rate above 200K | |
| Anthropic | Opus 5 | 5.00 | 25.00 | 1M | text, image, PDF | |
| OpenAI | GPT-5.6 Sol | 5.00 | 30.00 | 1.05M | text, image | |
| Anthropic | Fable 5 | 10.00 | 50.00 | 1M | text, image, PDF | |
| Anthropic | Mythos 5 | 10.00 | 50.00 | 1M | text, image, PDF | ‡ not generally available |
| OpenAI | GPT-Realtime-2.1 | 32.00 | 64.00 | — | audio | † per M audio tokens |
† Audio and text tokens are not interchangeable; this row is not like-for-like with text-token rows.
‡ Project Glasswing participants only.
10 Structural billing factors
Four factors move bills beyond headline input rates:
- Caching — 75–99% off cached input; dominates cost for stable system prompts. Cache writes cost more than base input (1.25–2× on Anthropic). Low-reuse workloads may pay more with caching enabled.
- Long-context surcharges — OpenAI 2× input / 1.5× output past 272K; Google past 200K. Anthropic charges one rate throughout on 1M-context models.
- Modality coverage — hard boundary. Gemini: audio/video on every tier; Kimi: image/video; DeepSeek and most Qwen: text-only.
- Output multiples — 2× at DeepSeek and Grok 4.3; 8.3× at Gemini 3.5 Flash-Lite. On output-heavy workloads, the ratio sets the bill.
Chinese labs lead on input vs Western flagships (one to two orders of magnitude) but gap vs Western budget tiers falls under 2×; no output-multiple advantage.
11 Constraints
Figures reflect rates as reported by sources below as of 8 August 2026. Intro rates expire (Sonnet 5: 31 August 2026); announced changes (DeepSeek peak-hour multiplier) may land without warning; regional endpoints diverge.
Reference list is almost entirely secondary. Anthropic’s documentation is the only first-party page checked directly. Confirm every rate at the provider before committing budget.
Headline. Rates. Mislead. Caching. Surcharges. Modalities. Multiples. Decide. Bills. Verify. Before. Budgeting.
12 References
Anthropic
- Anthropic platform pricing documentation
- CloudZero: Claude pricing breakdown
- MetaCTO: Anthropic API pricing — full breakdown of costs and integration
- BenchLM: Anthropic API pricing tracker
OpenAI
- Finout: OpenAI pricing in 2026
- DevTK: OpenAI API pricing guide 2026
- BenchLM: OpenAI API pricing tracker
- AI Pricing Guru: OpenAI pricing
Google (Gemini)
- Fello AI: Gemini pricing explained
- CloudZero: Gemini pricing breakdown
- BenchLM: Google API pricing tracker
- CostGoat: Gemini API pricing
DeepSeek
- Morph: DeepSeek API overview and rates
- Fello AI: DeepSeek pricing explained
- BenchLM: DeepSeek API pricing tracker
- CostGoat: DeepSeek API pricing
Moonshot (Kimi)
- Puter: Kimi API pricing tutorial
- BenchLM: Moonshot API pricing tracker
- SecondTalent: every Kimi AI model explained and compared
- CostGoat: Kimi API pricing
Alibaba (Qwen)
- Puter: Qwen API pricing tutorial
- eesel: Qwen pricing guide
- BenchLM: Alibaba API pricing tracker
- YottaLabs: Qwen 3.8 API access, token plans and pricing 2026
xAI (Grok)