Claude vs Gemini vs ChatGPT vs DeepSeek is usually answered with a vibe rather than a number. This piece tries to fix that: every price below comes from the company’s own pricing page, every benchmark is dated and attributed to whoever published it, and every table states its assumptions so you can redo the arithmetic yourself. One disclosure before anything else — this article is itself written by Claude, an Anthropic product, comparing Anthropic against three of its direct competitors. Every figure here is sourced rather than recalled, and where a competitor beats Claude on a cited number, that is reported exactly as plainly as the reverse.
How this comparison was put together
Pricing figures for Claude vs Gemini vs ChatGPT vs DeepSeek were pulled directly from each company’s live documentation on 20 September 2026: Anthropic’s official model-pricing table, OpenAI’s API pricing reference, Google’s Gemini Developer API pricing page, and DeepSeek’s own API docs. Benchmark numbers are taken from each company’s launch announcement and labelled as self-reported, because AI labs do not run identical evaluation harnesses and a number a company chooses to publish on its own launch day is a different kind of claim from a number an outside party reproduces. Wherever an independent evaluation existed — in this research, exactly one did — it is called out as such.
The four models at a glance
| Company | Flagship model | Released | Context window |
|---|---|---|---|
| Anthropic | Claude Opus 5 | 24 July 2026 | 1,000,000 tokens |
| OpenAI | GPT-6 Astra | 3 September 2026 | 1,050,000 tokens (128K max output) |
| Gemini 3.1 Pro | 19 February 2026 (preview) | 1,048,576 tokens (65,536 max output) | |
| DeepSeek | DeepSeek V4 Pro | GA 12–13 August 2026 | 1,000,000 tokens (384K max output) |
Anthropic also ships Claude Fable 5.1 and Claude Sonnet 5 as parallel, differently-priced lines rather than one single “best” model, and the same is true across every company here — OpenAI sells GPT-6 Astra alongside the cheaper GPT-5.6 Sol/Terra/Luna tier, Google sells Gemini 3.1 Pro alongside the Flash line, and DeepSeek sells V4 Pro alongside the cheaper V4 Flash. The flagship comparison above is a starting point, not the whole picture — the pricing tables below cover the full line-up from each company.
API pricing, compared with the actual arithmetic
| Model | Input | Output | Notes |
|---|---|---|---|
| Claude Opus 5 | $5 / MTok | $25 / MTok | Flat rate to 1M context; no tiered pricing |
| GPT-6 Astra | $10 / MTok | $50 / MTok | Doubles to $20/$75 (input/output) past 272K tokens |
| Gemini 3.1 Pro | $2 / MTok (≤200K) | $12 / MTok (≤200K) | Rises to $4/$18 past 200K tokens |
| DeepSeek V4 Pro | $0.66–$1.32 / MTok | $1.98–$3.96 / MTok | Off-peak vs peak (01:00–04:00, 06:00–10:00 UTC weekdays) |
Those per-token rates are hard to feel until they are run against an actual request, so here is the computed cost of one specific, realistic job: a 100,000-input-token request (roughly a 150-page document) that returns a 2,000-token answer, at standard rates with no caching or discounts applied.
| Model | Input cost | Output cost | Total |
|---|---|---|---|
| Claude Opus 5 | 100,000 × $5/1M = $0.50 | 2,000 × $25/1M = $0.05 | $0.55 |
| GPT-6 Astra | 100,000 × $10/1M = $1.00 | 2,000 × $50/1M = $0.10 | $1.10 |
| Gemini 3.1 Pro | 100,000 × $2/1M = $0.20 | 2,000 × $12/1M = $0.024 | $0.224 |
| DeepSeek V4 Pro (peak) | 100,000 × $1.32/1M = $0.132 | 2,000 × $3.96/1M = $0.0079 | $0.140 |
| DeepSeek V4 Pro (off-peak) | 100,000 × $0.66/1M = $0.066 | 2,000 × $1.98/1M = $0.0040 | $0.070 |
Assumptions: cache-miss input pricing throughout (no prompt caching applied to any model), standard/global inference routing, no batch-API discount. Batch and cache discounts would lower every figure here, roughly proportionally.
On this specific, stated workload, GPT-6 Astra costs almost exactly eight times more than DeepSeek V4 Pro at off-peak rates for what is, on paper, a comparable job. That gap narrows a lot once you apply prompt caching — Claude’s cache reads run at roughly a tenth of base input price, and Gemini and DeepSeek both offer similar mechanisms — but the sticker-price spread between the cheapest and most expensive flagship model in this comparison is real and worth knowing before you pick a default.
Consumer subscription pricing (US)
| Company | Free | Mid tier | Top consumer tier |
|---|---|---|---|
| Anthropic (Claude) | Yes — Sonnet & Haiku | Pro: $20/mo ($17/mo annual) | Max: from $100/mo (5x or 20x Pro usage) |
| OpenAI (ChatGPT) | Yes | Go: $8/mo · Plus: $20/mo | Pro: $200/mo — new sign-ups paused since 10 Sept 2026 |
| Google (Gemini) | Yes | Google AI Pro: $19.99/mo | Google AI Ultra: $99.99/mo (5x) or $199.99/mo (20x) |
| DeepSeek | Yes — entire product | No paid consumer tier exists | No paid consumer tier exists |
DeepSeek is the outlier worth sitting with: there is no Plus, no Pro, no premium consumer plan at all. The chat app at chat.deepseek.com and its mobile apps are free, including reasoning mode, and DeepSeek monetises purely through its API. Compare that against OpenAI, which paused new sign-ups and upgrades to its $200 ChatGPT Pro tier on 10 September 2026 — a genuinely unusual move for a subscription product, and one worth watching rather than glossing over.
India pricing: the market this site's audience actually pays in
All three paid competitors have localised pricing into rupees during 2026, and the differences are not cosmetic currency conversions — they reflect real product and payments decisions specific to the Indian market.
| Company | India pricing | Notable detail |
|---|---|---|
| Claude (Anthropic) | Pro: ₹2,399/mo (₹2,000/mo effective, billed annually) · Max 5x: ₹11,999/mo · Max 20x: ₹23,999/mo | No UPI support yet — card or app-store billing only, as of the July 2026 India localisation |
| ChatGPT (OpenAI) | Go: ₹399/mo · Plus: ₹1,999/mo · Pro: ₹19,900/mo | Go has run as a free 12-month promotion in India since 4 November 2025; OpenAI has not publicly committed to an end date |
| Gemini (Google) | AI Plus (India-only tier): ₹399/mo · AI Pro: ₹1,950/mo (first month free) · AI Ultra: ₹24,500/mo | AI Plus is a India-specific tier Google created in December 2025 specifically to undercut ChatGPT Go on price |
| DeepSeek | Free everywhere, India included | No India-specific tier needed — there is nothing to localise |
Benchmark claims, as each company reports them
Direct benchmark comparison across Claude vs Gemini vs ChatGPT vs DeepSeek is genuinely hard, because the four companies do not converge on which benchmarks to publish. Here is what each one actually put in its own launch materials — not what third-party trackers or leaked screenshots claimed, which in at least one case (Gemini 3.1 Pro, below) turned out to disagree with the company’s own blog post.
| Company | Headline self-reported number | Source |
|---|---|---|
| Anthropic | Claude Opus 5: 96.0% on SWE-bench Verified | anthropic.com/news, 24 Jul 2026 |
| OpenAI | GPT-6 Astra: 96.0% on GPQA Diamond; claims to “saturate” ARC-AGI-3 at 99.9% | openai.com/index/gpt-6-astra, 3 Sep 2026 |
| Gemini 3.1 Pro: 77.1% on ARC-AGI-2, “more than double” the prior Gemini 3 Pro | blog.google, 19 Feb 2026 | |
| DeepSeek | (No independent self-reported comparison table published; see NIST evaluation below) | — |
The one independent evaluation in this comparison
Every number above is self-reported. Exactly one genuinely independent, cross-model evaluation turned up in this research: NIST’s Center for AI Standards and Innovation (CAISI) published an assessment of DeepSeek V4 Pro on 1 May 2026, benchmarking it directly against GPT-5.5 and Claude Opus 4.6 using its own Item Response Theory methodology across five domains.
| Benchmark | DeepSeek V4 Pro | GPT-5.5 | Claude Opus 4.6 |
|---|---|---|---|
| GPQA-Diamond | 90% | 96% | 91% |
| OTIS-AIME-2025 | 97% | 100% | 92% |
| PUMaC 2024 | 96% | 96% | 95% |
| SWE-Bench Verified | 74% | 81% | 79% |
| FrontierScience | 74% | 79% | 72% |
| IRT-estimated Elo | 800 ± 28 | 1260 ± 28 | 999 ± 27 |
These are 2026-era models (GPT-5.5, Opus 4.6) predating the September 2026 flagships covered elsewhere in this piece — NIST had not evaluated GPT-6 Astra, Gemini 3.1 Pro, or Claude Opus 5 at time of publication. Treat this table as the most trustworthy relative comparison available, on slightly older models.
CAISI’s own conclusion is worth quoting rather than paraphrasing: DeepSeek V4 is “the most capable PRC AI model evaluated by CAISI to date,” but its capabilities “lag behind the frontier by about 8 months” on this composite. That is a meaningfully different picture from either DeepSeek’s own price-focused marketing or the assumption that a large Elo or benchmark gap automatically follows from DeepSeek’s far lower price — the gap on raw capability, by this specific independent measure, is real but narrower than the 15x price difference might suggest.
Human-preference rankings: what Arena shows, and a data problem worth disclosing
Arena (formerly LMArena, rebranded after a $150 million Series A in January 2026) runs the most widely cited human-preference leaderboard in the industry — over 10 million anonymous head-to-head votes, cited by OpenAI, Anthropic and Google alike in their own launch posts. It should be a clean, independent tiebreaker. In practice, fetching its live text leaderboard three times over roughly twenty minutes during this research returned three different sets of top-ranked models and different exact Elo scores each time — almost certainly because it is a heavy, dynamically-rendered table that a page-reading tool struggles to parse consistently, not because the actual rankings were shifting that fast.
A live leaderboard is not a citation. It is a pointer to where the citation currently lives — and the honest version of “Model X ranks #1 on Arena” always carries an access date.
Where each one actually wins
Pulling the verified numbers above together, the four models sort into genuinely different use cases rather than a single ranked list — this is analysis and synthesis, not a fifth primary source, and is labelled as such.
| Model | Strongest verified case |
|---|---|
| Claude Opus 5 | Coding/agentic work (96.0% self-reported SWE-bench Verified) and the largest gap-free context window at a flat per-token rate |
| GPT-6 Astra | Highest self-reported abstract-reasoning and science scores (ARC-AGI-3, GPQA Diamond, FrontierMath) — at the highest price in this comparison |
| Gemini 3.1 Pro | Best balance of price and context-window ceiling among the closed-source frontier models, with the cheapest sub-200K-token rate of the three flagships |
| DeepSeek V4 Pro | Cost per token, by a wide margin — with NIST’s independent evaluation putting the capability gap at roughly eight months, not the years the price gap might imply |
How to actually choose
1. Work out your real token volume before comparing sticker prices. The arithmetic above shows the pricing gap is only meaningful at scale; casual use makes any of the four effectively free against a subscription.
2. Match the task to the verified strength, not the marketing headline. Coding and long-document work favour Claude’s context window and SWE-bench number; abstract reasoning and science favour GPT-6 Astra’s self-reported scores; tight budgets and open-weight flexibility favour DeepSeek, backed by an independent evaluation rather than just a low price.
3. Check India-specific pricing directly before subscribing, since three of the four companies changed their India pricing or launched India-specific tiers within the past year, and at least one promotional offer (ChatGPT Go) has an unconfirmed end date.
4. Do not trust a single Arena screenshot. Treat any “ranked #1 on Arena” claim, including implicitly favourable ones toward Claude, as something to verify at the live URL with today’s date, not as a settled fact.
5. Re-run this comparison before you commit long-term. Every model named here launched or was priced within the twelve months before this piece was written; at the pace this industry moves, treat any comparison — this one included — as accurate as of its stated date, not as evergreen truth.
Final verdict
There is no single winner in Claude vs Gemini vs ChatGPT vs DeepSeek, and every company’s own launch materials are, unsurprisingly, built to suggest otherwise. What the verified numbers actually show is four different bets: OpenAI is spending the most on raw reasoning and charging the most for it; Anthropic is charging a middle price for the largest practical context window and the strongest self-reported coding score; Google is undercutting both on price at the flagship tier while publishing the thinnest set of comparable benchmarks; and DeepSeek is giving away the consumer product entirely and pricing its API low enough that, per the one independent evaluation in this research, the capability-per-dollar case is genuinely strong rather than merely cheap.
Our read, stated plainly and including the conflict of interest disclosed at the top: if cost is not a constraint and the work is coding-heavy, Claude Opus 5’s numbers hold up. If raw reasoning on novel problems is the job, GPT-6 Astra’s self-reported scores are the most aggressive in this comparison, at a price to match. If the workload is high-volume and cost-sensitive, DeepSeek V4 Pro is the one competitor here with independent, government-run evidence behind its price-to-capability claim, not just a marketing page. Gemini 3.1 Pro sits as the value option among the three closed-source frontier labs — competitive on the one benchmark Google chose to publish, and notably quieter than its rivals about the rest.
Frequently asked
- Which is cheaper: Claude, Gemini, ChatGPT or DeepSeek?
- On API pricing, DeepSeek V4 Pro is dramatically cheaper — roughly 4 to 8 times cheaper than Claude Opus 5 and GPT-6 Astra on a comparable 100,000-input/2,000-output-token request, per official pricing pages checked on 20 September 2026. On consumer subscriptions, DeepSeek has no paid tier at all; among the paid options, Gemini’s Google AI Pro ($19.99/mo) and Claude Pro ($20/mo, $17/mo annual) sit close together, with ChatGPT Plus also at $20/mo.
- Which AI model has the biggest context window?
- Claude Opus 5, Gemini 3.1 Pro and DeepSeek V4 Pro all offer approximately 1 million token context windows. GPT-6 Astra offers slightly more at 1.05 million tokens, but doubles its per-token price for any request over 272,000 tokens — Claude’s 1M window carries no such tiered pricing.
- Is DeepSeek actually as good as Claude, ChatGPT or Gemini?
- By the only independent evaluation found in this research — NIST’s CAISI assessment, published 1 May 2026 — DeepSeek V4 Pro trailed the US frontier (GPT-5.5 and Claude Opus 4.6, both slightly older models) by roughly eight months on a five-benchmark composite. That is a real gap, but a much smaller one than the price difference alone would suggest.
- Does ChatGPT Go actually cost nothing in India?
- As of this guide’s research, ChatGPT Go has run as a free promotion in India since 4 November 2025 for a 12-month period per account, per OpenAI’s original announcement and TechCrunch’s reporting. OpenAI has not published a public end date, and its own Help Center page on the promotion could not be directly accessed during this research, so readers should confirm current terms at chatgpt.com/pricing before assuming free access continues indefinitely.
- Which model actually ranks highest on independent benchmarks?
- It depends which independent source you trust. Arena’s human-preference leaderboard has consistently shown Claude’s Fable and Opus lines near the top across repeated checks, though this guide flags that the exact live scores were inconsistent across repeated fetches of the same page. On NIST’s CAISI evaluation — the only government-run, non-vendor benchmark found in this research — GPT-5.5 outscored both Claude Opus 4.6 and DeepSeek V4 Pro on 4 of 5 tracked benchmarks, though that evaluation predates the September 2026 flagship models covered elsewhere in this guide.
- Why does this guide keep saying "self-reported"?
- Because AI companies do not run a shared, standardised benchmark suite, and a number a company publishes about its own model on its own launch day is a different kind of evidence than a number an independent party reproduces. This guide labels every benchmark by who reported it and when, and flags the one instance (NIST’s DeepSeek evaluation) where an independent, non-vendor source exists.
