Across 69 models on BenchLM's SWE-bench Verified board, retrieved 2026-09-07, Claude Opus 5 leads at 96%, with Claude Fable 5 one point behind at 95%.
DeepSeek-V4-Flash and GPT-5.6 Luna price the same coding task 45 to 90 times cheaper than the top two Claude tiers.
For most agent workflows, the plan's usage cap, not the per-token rate, decides how far a session runs before it stops. Match your pick to the use case below.
LLM Coding models compared
Model | SWE-bench Verified | Input $/M | Output $/M | Context | Tier |
|---|---|---|---|---|---|
Claude Opus 5 | 96%† | $5.00 | $25.00 | 1M | Flagship |
Claude Fable 5 | 95%† | $10.00 | $50.00 | 1M | Deep-reasoning flagship |
Claude Sonnet 5 | 85.2%† | $2.00 | $10.00 | 1M | Mid |
Claude Haiku 4.5 | E | $1.00 | $5.00 | 1M | Budget |
GPT-5.6 Sol | n/a‡ | $2.00 | $10.00 | not stated | Flagship |
GPT-5.6 Terra | n/a‡ | $1.00 | $6.00 | not stated | Mid |
GPT-5.6 Luna | n/a‡ | $0.10 | $0.60 | not stated | Budget |
Gemini 2.5 Pro | 63.8%† | $1.25 to $2.50 | $10.00 to $15.00 | not stated | Previous-gen flagship |
Gemini 3.1 Pro Preview | E | $2.00 to $4.00 | $12.00 to $18.00 | not stated | Flagship (preview) |
DeepSeek-V4-Pro | E | $0.435 | $0.87 | not stated | Open-weight, hosted |
DeepSeek-V4-Flash | E | $0.14 | $0.28 | 1M | Open-weight, budget |
Qwen3-Coder-Next | E | not published | not published | not stated | Open-weight, on-device |
Table note: † SWE-bench Verified score reported by BenchLM.ai, a third-party benchmark aggregator, not independently verified by LLM Waves Research and not self-reported by the vendor. ‡ OpenAI reports GPT-5.6 Sol, Terra and Luna against SWE-bench Pro, a harder, different benchmark than SWE-bench Verified; the two scores are not comparable and no SWE-bench Verified figure for the GPT-5.6 family was found this session. E: insufficient data, fewer sources than the reporting threshold, not ranked. Google's tiered input and output prices split at 200,000 tokens of prompt length, with the higher rate applying above that line. Price direction: lower is better.
Claude Opus 5 leads the ranked field by 1 point over Claude Fable 5 and by more than 10 points over the next tier, Claude Sonnet 5 at 85.2%.
Gemini 2.5 Pro, the only Google model with an independently sourced SWE-bench Verified figure this session, trails the Claude flagships by more than 30 points at 63.8%.
DeepSeek-V4-Flash prices the same task at $0.14 per million input tokens against $10.00 for Claude Fable 5, a 71-fold gap on the input side alone.
Basis and method
This comparison uses Measurement Basis 2. Every figure below was collected from a published vendor pricing page, a vendor model card or announcement, an arXiv technical report, or a named third-party benchmark aggregator, never produced by testing run by LLM Waves Research. Pricing traces to each vendor's own current documentation: Anthropic for Claude, OpenAI for GPT-5.6, Google for Gemini, DeepSeek for its own models.
Coding benchmark scores come from two sources that do not agree on which test to run: OpenAI self-reports SWE-bench Pro in its own GPT-5.6 announcement, while the SWE-bench Verified figures for Claude and Gemini models come from BenchLM.ai, a third-party leaderboard that evaluates models independently of the vendors.
Every source is listed by name, with its URL and retrieval date, in the raw data file and on the methodology page.
Decision axis: models are ranked first on SWE-bench Verified where an independently sourced score exists, then on price per coding task at a fixed 50,000-input, 5,000-output token profile, since a leaderboard rank with no price attached answers half the question a developer is asking.
Ranked bar chart of SWE-bench Verified scores across seven of the models in this comparison.

Percent scored on SWE-bench Verified, seven models, source BenchLM.ai, retrieved 2026-09-07
Claude Opus 5 scores 96% on SWE-bench Verified per BenchLM.ai, one point ahead of Claude Fable 5 at 95% and more than 30 points ahead of Gemini 2.5 Pro at 63.8%.
The gap between Claude Opus 5 and Claude Fable 5 sits inside a plausible margin of run-to-run variance on a 500-instance benchmark.
Treat the two as a tied top group rather than a single winner, and read the 10-plus point gap down to Claude Sonnet 5 as the more reliable signal in this data.
Cost per task, not cost per token
Every vendor pricing page quotes a rate per million tokens. None of the 13 buying-guide roundups reviewed for this comparison convert that rate into what a single coding task actually costs, and readers are searching for exactly that conversion.
The formula used here: a task is set at 50,000 input tokens, a realistic size for a model reading a small repository's relevant files plus a prompt, and 5,000 output tokens, a realistic size for a patch or a short function with explanation.
Cost per task equals the input token count divided by one million, times the input rate, plus the output token count divided by one million, times the output rate.
The formula and the two token counts are fixed across every model in this table, so the comparison isolates price, not task size.
Under this formula, Claude Fable 5 costs $0.75 per task. Claude Opus 5 costs $0.375, half of Fable, despite scoring within a point of it on SWE-bench Verified, the trade Anthropic itself points to in its own release material.
Claude Sonnet 5 and GPT-5.6 Sol both land at $0.15. GPT-5.6 Luna costs $0.008; DeepSeek-V4-Flash costs $0.0084, with cache hits on repeated context pushing that lower still.
At 500 tasks a month, a plausible volume for one active developer running an agent through a workday, the same workload costs $187.50 on Claude Opus 5, $75 on Claude Sonnet 5 or GPT-5.6 Sol, and under $5 on GPT-5.6 Luna or DeepSeek-V4-Flash.
The 45-fold to 90-fold spread is the number the per-token rate alone never shows.

Claude Fable 5 costs $0.75 per computed coding task against $0.008 for GPT-5.6 Luna and $0.0084 for DeepSeek-V4-Flash, a 90-fold spread on the same fixed task profile.
LLM Waves Research computed this figure by applying the published per-token rates above to a fixed task profile. It does not appear in this form on any of the pricing pages, benchmark boards or roundups checked for this comparison.
The subscription cap decides more than the price does
None of the 13 roundups reviewed for this comparison mention that a developer working through a Claude Pro or Max plan, or a ChatGPT Plus or Pro plan, hits a usage ceiling with nothing to do with the metered API price above.
Anthropic's own usage documentation confirms the mechanism without publishing the number: a Claude Pro or Max session runs against a five-hour rolling window, and a separate weekly cap resets on its own schedule for Claude Opus models specifically, distinct from the cap on every other Claude tier.
The support article confirms both limits exist and explains how to check consumption in account settings, but states neither the session nor the weekly limit in hours or messages for either plan.
OpenAI publishes exact numbers, but only for its most expensive tier. A $200-per-month ChatGPT Pro plan carries 200 messages a week for the chat model OpenAI's help center names "GPT-6 Pro," plus a separate 170-messages-a-day allowance for GPT-5.6 Sol Pro, with a combined daily ceiling of 200 across both; a $100-per-month Pro plan shares one pool of 50 messages a week across the same two models. The $20 Plus plan gets Medium and High reasoning effort but not the Extra High or Pro-tier models, and states no specific message count at all.
Not recommended for anyone routing high-volume agent workloads through a fixed-price consumer plan without checking that plan's cap first.
A workload that fits inside a five-hour session on the metered API can stall mid-task on a subscription plan that meters differently, and at least one of the two providers here does not publish the number needed to plan around it.
This site's full model index lists metered API pricing separately from any plan cap for every model tracked here.
SWE-bench Verified and SWE-bench Pro are not the same test
A comparison that lines up the SWE-bench Verified score for Claude against the headline benchmark number for GPT-5.6 is comparing two different tests.
OpenAI's own GPT-5.6 announcement reports SWE-bench Pro, not SWE-bench Verified, for Sol, Terra and Luna: 64.6%, 63.4% and 62.7% respectively.
SWE-bench Pro uses a harder, separately curated instance set than SWE-bench Verified, and a score on one does not convert to a score on the other.
No independently sourced SWE-bench Verified figure for any GPT-5.6 variant turned up in this session's collection, which means the 96% figure for Claude Opus 5 and the 64.6% figure for GPT-5.6 Sol cannot be placed on the same axis without a conversion neither vendor nor any aggregator checked here has published.
Readers comparing "GPT-5.6 versus Claude on coding benchmarks" are, at present, comparing a number that exists for one family against a number that does not exist for the other on the same test.

SWE-bench Pro scores, self-reported by OpenAI, versus SWE-bench Verified scores, reported by BenchLM.ai, by model family, 2026-09-07
The GPT-5.6 family reports SWE-bench Pro scores between 62.7% and 64.6%. The Claude and Gemini figures in this comparison are SWE-bench Verified scores from a different, non-equivalent test, and no SWE-bench Verified figure for GPT-5.6 was found this session.
The open-weight tier: what actually ships a price
This site's own open-source coding roundup covers the wider open-weight field; the two picks below are the ones with a confirmed architecture or hosted price as of this collection.
Qwen3-Coder-Next activates 3 billion of its 80 billion total parameters per inference pass, according to the Qwen team's own arXiv technical report, small enough for a single high-end consumer GPU rather than a data-center cluster, the entire appeal of the open-weight tier for a developer who wants no per-token bill at all.
Two secondary sources disagree on its SWE-bench score, one reporting 70% and another 74%, and neither traces back to the technical report or to an aggregator with a published methodology, so this comparison carries no ranked score for it. No hosted API price was confirmed this session either.
DeepSeek-V4-Pro and DeepSeek-V4-Flash are hosted, not self-hosted, and their appeal runs the other way: a price low enough that the API bill approaches zero without local inference at all.
DeepSeek's own pricing page lists cache-hit input at $0.0028 per million tokens against $0.435 for a cache miss, an 84% discount for repeated context such as a stable system prompt or a file the model has already read in the same session.
Neither DeepSeek-V4 model carries an independently sourced SWE-bench Verified score here; DeepSeek V3, a previous generation, scored 42% on BenchLM's board, a figure that describes a retired model and is not reused as a V4 claim.
Hosted pricing across the flagship and budget tiers in this comparison does not scale in proportion between input and output tokens, which the chart below makes visible in a way the pricing pages themselves do not.

Input and output price per million tokens, US dollars, eleven models, retrieved 2026-09-07
Claude Fable 5 charges $50.00 per million output tokens against $0.60 for GPT-5.6 Luna, an 83-fold gap on output alone, wider than the corresponding input-price gap.
What happens when a coding agent fails mid-task
None of the roundups checked for this comparison describe what a developer actually sees when an agent session breaks.
Three failure modes recur across the vendor documentation gathered here, with no standard error message shared across providers: a context-window overflow, where the harness drops earlier turns rather than failing cleanly and a partial diff results; a malformed diff, where the returned patch does not apply against the current file state and forces a manual fix; and a rate-limit lockout mid-session, where a five-hour or weekly cap lands partway through a task with no partial-credit resumption built into the plan mechanics above.
Each provider's documentation covers only its own mechanism, so a developer working across more than one model reconciles three separate failure vocabularies by hand.
Model-to-agent compatibility is not the same question as model quality
A SWE-bench score describes performance against a fixed evaluation harness, not performance inside Cursor, GitHub Copilot, Cline, or a custom agent loop with its own tool-calling format.
A high score does not guarantee a model's tool-call syntax matches what a specific coding agent expects, and a mismatch here surfaces as a silently failed tool call rather than a benchmark-visible error.
Only one of the 13 roundups checked for this comparison addresses the distinction, and only from a production-pipeline angle. Treat the figures above as a starting filter, not a guarantee of behavior inside any single agent.
What we did not measure
This comparison did not run any model against any benchmark, coding task, or production workload. It did not test latency, tokens-per-second, or time-to-first-token, all of which affect an agentic coding session independently of accuracy or price.
It did not evaluate non-English coding prompts, on-device inference speed for Qwen3-Coder-Next, or the coding accuracy of Claude Haiku 4.5, for which no independently sourced score was found.
It excludes GPT-6 Astra, priced at $5.00 input and $25.00 output per million tokens, since no coding benchmark figure for it was found this session, and Claude Mythos 5, priced identically to Fable 5, since vendor material does not document how its intended use differs from that of Fable 5 for a coding reader.
Every figure above traces to a named source with a URL and retrieval date in the raw data file and on the methodology page; nothing here was estimated or interpolated where a source was missing.
Use case verdicts
Four models in this comparison carry both an independently sourced SWE-bench Verified score and a computed cost per task: Claude Opus 5, Claude Fable 5, Claude Sonnet 5 and Gemini 2.5 Pro.
Plotting price against score for just those four is what turns the ranked table into a value pick, not only an accuracy pick.

Cost per computed task, US dollars, against SWE-bench Verified percent, four models with both figures available, 2026-09-07
Claude Opus 5 scores 96% at $0.375 per task, beating Claude Fable 5 on both price and score; Gemini 2.5 Pro is the cheapest of the four at $0.1125 per task but trails on SWE-bench Verified by more than 30 points.
For developers running a personal coding agent daily
Best value goes to GPT-5.6 Terra at $0.08 per task: it scores 63.4% on SWE-bench Pro, 1.2 points under the 64.6% Sol scores, close enough to skip paying nearly double for a $0.15 task on Sol for day-to-day changes; step up to Sol or Claude Opus 5 once a task is large or unfamiliar enough that the accuracy gap is worth paying for.
For startups metering API spend directly
Best value goes to DeepSeek-V4-Flash at $0.0084 per task, cache hits pushing repeated context lower still; situational, since no independently sourced SWE-bench Verified figure exists for the V4 family, so the accuracy trade against Claude or GPT-5.6 is unquantified here.
For enterprise teams standardizing on one vendor for compliance reasons
Top pick goes to Claude Opus 5 at 96% SWE-bench Verified, the only model in this comparison with both a top-ranked score and a full commercial support and pricing structure from a single vendor.
For high-volume batch workloads, such as a nightly pass across a large codebase
Best value goes to GPT-5.6 Luna at $0.008 per task, cheap enough that volume rarely binds; not recommended once task complexity approaches what Sol or Opus 5 was benchmarked against, since Luna scores 62.7% on SWE-bench Pro, trailing Sol by almost two points on a benchmark already harder than SWE-bench Verified.
For self-hosting on a single high-end GPU
Top pick by default goes to Qwen3-Coder-Next on architecture alone: 80 billion parameters total, 3 billion active per pass. Insufficient data on its exact accuracy, since no independently sourced SWE-bench score was found, so this pick is provisional pending a confirmed figure.
For low-latency, low-context single-file edits
Insufficient data across the board, since this comparison excluded time-to-first-token and tokens-per-second figures for every model; readers optimizing for speed rather than accuracy or price need a separate, speed-focused comparison.
For non-English codebases and comments
Insufficient data across the board; this comparison collected no non-English-specific coding benchmark for any of the twelve models checked.
What's the best open-weight LLM for coding?
Qwen3-Coder-Next, on architecture and cost alone: it activates 3 billion of 80 billion parameters per inference pass, small enough to run on a single high-end consumer GPU, and carries no per-token API bill when self-hosted.
This comparison found no independently sourced SWE-bench score for it, so the pick describes what it is built to do, not a confirmed accuracy ranking against the hosted flagships above.
Which model is best for agentic, multi-step coding tasks?
Claude Opus 5, on the only independently sourced SWE-bench Verified score in this comparison that clears 95%, at 96%.
Claude Fable 5 scores within a point of it at half again the price, and the gap between the two sits inside plausible run-to-run variance on a 500-instance benchmark, so treat them as a tied top group rather than a single winner for multi-step work.
Is Claude better than GPT for coding?
On the only figures collected this session, the 96% SWE-bench Verified score for Claude Opus 5 and the 64.6% SWE-bench Pro score for GPT-5.6 Sol cannot be compared directly, since the two scores come from different, non-equivalent benchmarks.
No SWE-bench Verified figure for GPT-5.6 Sol and no SWE-bench Pro figure for Claude Opus 5 were found in this session's sources, so "better" depends on which test a reader trusts, not on a number this comparison can put on one axis.
How much do the best coding LLMs actually cost per month?
At a fixed 500-task monthly workload, computed at 50,000 input and 5,000 output tokens per task, the field in this comparison ranges from under $5 a month on GPT-5.6 Luna or DeepSeek-V4-Flash to $187.50 a month on Claude Opus 5, before any subscription-plan usage cap is factored in. The per-token rate alone understates this range, since it never states a task size.
What happens when I hit my plan's rate limit mid-task?
Rate-limit behavior differs by provider, and neither publishes a full recovery path. Anthropic's own documentation confirms a five-hour session cap plus a separate weekly cap for Claude Opus models, without stating the numeric limit for either.
OpenAI publishes exact weekly and daily message counts for its $100 and $200 ChatGPT Pro tiers but states no specific limit for the $20 Plus tier in its own help center article.
In both cases, a session that hits its cap stops rather than downgrading gracefully to a cheaper model automatically.
How is coding ability actually benchmarked?
SWE-bench Verified and SWE-bench Pro both score a model on whether it can resolve real, closed GitHub issues by producing a patch that passes the project's existing test suite, but SWE-bench Pro draws from a harder, separately curated instance set.
A score on one benchmark does not convert to the other, which is why this comparison lists them in separate columns rather than one ranked list, per the gap section above.
The numbers above describe September 7, 2026, and nothing after it
Every figure in this comparison is a snapshot. Pricing pages change without changelogs, model rosters add and retire tiers inside a single quarter, and a subscription cap Anthropic or OpenAI has not published today may appear in an updated support article tomorrow.
