LLM API Pricing Compared: Claude, GPT, Gemini, and the Chinese Challengers (August 2026)

A snapshot of input/output token costs and supported modalities across the major AI API providers

Per-token rates across seven providers as of August 2026 — and the four structural differences (caching, long-context surcharges, modality coverage, output multiples) that move real bills further than the headline number does.
AI
LLM
Pricing
Author

Ravi Kalia

Published

August 8, 2026

LLM API Pricing Compared

This post compares per-token API prices across seven providers as of August 2026, and shows why the headline rate is rarely what you end up paying.

Providers quote USD per million tokens for input (prompt) and output (reply). Headline rates are a starting point; caching, long-context surcharges, modality coverage, and output multiples move real bills further.

This page is a snapshot from early August 2026. Token prices change frequently; confirm rates at each provider’s pricing page before budgeting.

Seven providers: Anthropic (Claude), OpenAI (GPT), Google (Gemini), DeepSeek, Moonshot (Kimi), Alibaba (Qwen), xAI (Grok). A combined table sorted by input price follows the per-provider sections.

Conventions:

1 Data provenance

Figures are compiled from secondary sources (pricing trackers and explainer posts; full list under References), except the Anthropic table, which was checked against Anthropic’s published documentation.

  • Objective: compare rough price bands, long-context surcharges, modality coverage, and output multiples.
  • Not suitable for: procurement or volume projections — use first-party pricing pages.
  • Downstream impact: a stale rate misstates budget and tier selection.

Reliable for band-level conclusions that survive ~10% figure error. Not a billing source of truth.

2 Anthropic

Anthropic prices on a tiered ladder (small → frontier) with no long-context surcharge on models 4.6+ (flat rate regardless of prompt length). Haiku 4.5 predates that guarantee but has a 200K window with no long-context tier.

Model Input $/M Output $/M Context Notes
Haiku 4.5 1.00 5.00 200K Cheapest tier; the one current model not at 1M
Sonnet 5 2.00 10.00 1M Intro rate through 31 Aug 2026; then 3.00 / 15.00
Opus 5 5.00 25.00 1M
Fable 5 10.00 50.00 1M Frontier tier
Mythos 5 10.00 50.00 1M Same capabilities and price as Fable 5, but not generally available — Project Glasswing participants only

Modalities: text, image, PDF/documents.

Discounts:

  • Prompt caching — ~90% off cached input reads; cache writes cost 1.25× base input (5-minute TTL) or 2× (1-hour TTL). Pays from second request at 5-minute TTL, third at 1-hour TTL.
  • Batch API — 50% off input and output for deferred jobs.

3 OpenAI

Widest price range among Western majors. Long-context surcharge above 272K tokens. Audio sold as a separate model, not folded into text tiers.

Model Input $/M Output $/M Context Notes
GPT-5.6 Luna 0.20 1.20 1.05M Budget tier
GPT-5.6 Terra 2.00 12.00 1.05M Mid tier
GPT-5.6 Sol 5.00 30.00 1.05M Flagship
GPT-Realtime-2.1 32.00 64.00 Audio in / audio out, per M audio tokens

Modalities: text and image on GPT-5.6 line; audio via GPT-Realtime-2.1 only.

Long-context surcharge: above 272K tokens, input 2× and output 1.5×.

4 Google Gemini

Audio and video are first-class inputs on every tier. Also sells image and video generation per unit. Long-context surcharge at a lower threshold than OpenAI.

Model Input $/M Output $/M Context Notes
Gemini 3.5 Flash-Lite 0.30 2.50 Cheapest tier
Gemini 3.6 Flash 1.50 7.50 Mid tier
Gemini 3.1 Pro (≤200K) 2.00 12.00 2M Rate below the 200K threshold
Gemini 3.1 Pro (>200K) 4.00 18.00 2M Long-context rate

Modalities: text, image, audio, video. Image generation from $0.02/image; video generation from $0.15/second.

Discounts: Batch 50% off. Caching charges rent on storage ($1–4.50 per M tokens per hour), not a write multiple — break-even depends on hold time and reuse.

5 DeepSeek

Lowest output price in this comparison; steepest cache-hit rate. Text-only; no batch tier.

Model Input $/M Output $/M Context Notes
V4-Flash 0.14 0.28 1M Cache-hit input $0.0028
V4-Pro 0.435 0.87 1M Cache-hit input $0.0036

Modalities: text.

Caveats: no batch tier; 2× peak-hour pricing announced without start date. V4-Pro cache-hit rate (99.2% off base) is unusually deeper than V4-Flash (98%) — verify at source before relying on it.

Against Western flagships, V4-Flash input is 36× cheaper than Opus 5; against budget tiers the gap falls to ~1.4× vs GPT-5.6 Luna.

6 Moonshot Kimi

Mid-range pricing; native video input (only non-Google provider here with video). Weakest batch tier on this page (40% discount vs 50% elsewhere); flagship K3 excluded from batch.

Model Input $/M Output $/M Context Notes
K2.5 0.60 3.00 262K Batch eligible
K2.6 / K2.7 Code 0.95 4.00 256K Batch eligible
K3 3.00 15.00 1M Flagship; cache-hit input $0.30; not batch eligible

Modalities: text, image, video.

Discounts: batch bills at 60% of standard rate (40% discount); K2.5/K2.6/K2.7 only.

7 Alibaba Qwen

Headline numbers are international (Singapore) rates; Mainland China endpoint is 60–70% cheaper with no free quota.

Model Input $/M Output $/M Context Notes
Qwen3.5 Flash 0.10 0.40 Lowest input price in this comparison
Qwen3.5 Plus 0.40 2.40 1M
Qwen3.8-Max 2.00 6.00 1M Flagship; cached input $0.25

Modalities: text on Flash (some image variants); text and image on Plus and Max.

8 xAI Grok

Differentiator is live X/Twitter data as a first-class input. Output multiple (output/input ratio) is 2–3×; most Western text tiers sit at 5–6×; Gemini 3.5 Flash-Lite reaches 8.3×.

Model Input $/M Output $/M Context Notes
Grok 4.1 Fast 0.20 0.50 2M Joint-largest context here, tied with Gemini 3.1 Pro
Grok 4.3 1.25 2.50 1M
Grok 4.5 2.00 6.00 500K Flagship; cached input $0.50; no batch discount

Modalities: text, image, live X/Twitter data.

9 Combined table

Sorted by input price. Five models share $2.00 input across five providers.

Provider Model Input $/M Output $/M Context Modalities Notes
Alibaba Qwen3.5 Flash 0.10 0.40 text some image variants
DeepSeek V4-Flash 0.14 0.28 1M text
OpenAI GPT-5.6 Luna 0.20 1.20 1.05M text, image
xAI Grok 4.1 Fast 0.20 0.50 2M text, image, live X
Google Gemini 3.5 Flash-Lite 0.30 2.50 text, image, audio, video
Alibaba Qwen3.5 Plus 0.40 2.40 1M text, image
DeepSeek V4-Pro 0.435 0.87 1M text
Moonshot K2.5 0.60 3.00 262K text, image, video
Moonshot K2.6 / K2.7 Code 0.95 4.00 256K text, image, video same window as K2.5, binary units
Anthropic Haiku 4.5 1.00 5.00 200K text, image, PDF
xAI Grok 4.3 1.25 2.50 1M text, image, live X
Google Gemini 3.6 Flash 1.50 7.50 text, image, audio, video
Anthropic Sonnet 5 2.00 10.00 1M text, image, PDF intro rate to 31 Aug 2026; then 3.00 / 15.00
OpenAI GPT-5.6 Terra 2.00 12.00 1.05M text, image
Google Gemini 3.1 Pro 2.00 12.00 2M text, image, audio, video rate at ≤200K
Alibaba Qwen3.8-Max 2.00 6.00 1M text, image
xAI Grok 4.5 2.00 6.00 500K text, image, live X
Moonshot K3 3.00 15.00 1M text, image, video
Google Gemini 3.1 Pro 4.00 18.00 2M text, image, audio, video rate above 200K
Anthropic Opus 5 5.00 25.00 1M text, image, PDF
OpenAI GPT-5.6 Sol 5.00 30.00 1.05M text, image
Anthropic Fable 5 10.00 50.00 1M text, image, PDF
Anthropic Mythos 5 10.00 50.00 1M text, image, PDF ‡ not generally available
OpenAI GPT-Realtime-2.1 32.00 64.00 audio † per M audio tokens

† Audio and text tokens are not interchangeable; this row is not like-for-like with text-token rows.

‡ Project Glasswing participants only.

10 Structural billing factors

Four factors move bills beyond headline input rates:

  • Caching — 75–99% off cached input; dominates cost for stable system prompts. Cache writes cost more than base input (1.25–2× on Anthropic). Low-reuse workloads may pay more with caching enabled.
  • Long-context surcharges — OpenAI 2× input / 1.5× output past 272K; Google past 200K. Anthropic charges one rate throughout on 1M-context models.
  • Modality coverage — hard boundary. Gemini: audio/video on every tier; Kimi: image/video; DeepSeek and most Qwen: text-only.
  • Output multiples — 2× at DeepSeek and Grok 4.3; 8.3× at Gemini 3.5 Flash-Lite. On output-heavy workloads, the ratio sets the bill.

Chinese labs lead on input vs Western flagships (one to two orders of magnitude) but gap vs Western budget tiers falls under 2×; no output-multiple advantage.

11 Constraints

Figures reflect rates as reported by sources below as of 8 August 2026. Intro rates expire (Sonnet 5: 31 August 2026); announced changes (DeepSeek peak-hour multiplier) may land without warning; regional endpoints diverge.

Reference list is almost entirely secondary. Anthropic’s documentation is the only first-party page checked directly. Confirm every rate at the provider before committing budget.

Headline. Rates. Mislead. Caching. Surcharges. Modalities. Multiples. Decide. Bills. Verify. Before. Budgeting.

12 References

Anthropic

OpenAI

Google (Gemini)

DeepSeek

Moonshot (Kimi)

Alibaba (Qwen)

xAI (Grok)