Claude Opus 5 resolves 96.0% of SWE-bench Verified tasks against DeepSeek-V4-Pro's 80.6%, a 15.4-point gap, measured September 8, 2026.
DeepSeek-V4-Pro's off-peak output rate of $1.98 per million tokens runs 4 to 13 times below Claude's three published tiers, depending on which DeepSeek rate and which Claude model get compared.
Neither vendor leads on both accuracy and price at once.
Model | Output price† | Input price† | SWE-bench Verified | SWE-bench Pro | Context window |
|---|---|---|---|---|---|
Claude Opus 5 | $25.00/M | $5.00/M | 96.0% | 79.2% | 1,000,000 |
Claude Sonnet 5 | $10.00/M | $2.00/M | 85.2%‡ | 63.2%‡ | ◇ |
DeepSeek-V4-Pro (off-peak) | $1.98/M | $0.66/M | 80.6% | 55.4% | 1,000,000 |
DeepSeek-V4-Flash (off-peak) | $0.66/M | $0.22/M | 79.0%‡ | n/a | 1,000,000 |
Table note: † DeepSeek's off-peak rate applies 16:00 to 08:00 UTC; its peak rate, shown in Chart 1, runs roughly double. ‡ third-party leaderboard reading (BenchLM.ai and llm-stats.com, named here per Class B sourcing rules, never linked), used only where Anthropic or DeepSeek has not published a matching figure directly. ◇ not independently confirmed this session. n/a: not published by any source checked.
Four numbers carry the table. Claude Opus 5 leads every accuracy column. DeepSeek-V4-Pro leads every price column, at either of its two rates.
Claude Sonnet 5 sits between the two flagships on cost per output token, at $10.00 against Opus 5's $25.00 and V4-Pro's $1.98. Context windows tie at 1,000,000 tokens everywhere a figure was confirmed directly against a vendor source.
Basis and method
We aggregated published figures from both vendors and two named third-party trackers, and ran no benchmark, latency test, or prompt suite. Pricing comes from DeepSeek's API documentation and its change-log entry for the August 16, 2026 rate adjustment, cross-checked against two independent pricing trackers reporting identical dollar figures. SWE-bench scores for DeepSeek-V4-Pro come from its own Hugging Face model card; Claude Opus 5's come from Anthropic's July 24, 2026 launch materials and system card.
Where neither vendor published a figure, BenchLM.ai and llm-stats.com are named and dated. No arithmetic beyond simple ratio division, Claude price divided by DeepSeek price at matched market tier, was performed. Full source URLs and retrieval dates sit in the raw data file linked at the close, and on the methodology page.
What this comparison ranks on
Three axes decide the verdicts below: cost per output token at DeepSeek's peak and off-peak rates against Claude's three flat rates, accuracy on SWE-bench Verified and SWE-bench Pro, and where each vendor stores and processes account data.

DeepSeek's off-peak and peak rates against Claude's three flat prices, in USD per million output tokens, collected September 8, 2026. Claude Opus 5's $25.00 sits 6.3 times above DeepSeek-V4-Pro's peak rate and 12.6 times above its off-peak rate. (Dark variant: 01-output-price-by-tier-dark.png. Dataset: 01-output-price-by-tier.csv.)
DeepSeek prices by the hour; Claude prices by the model
DeepSeek runs two clocks, not one. Since 16:00 UTC on August 16, 2026, DeepSeek's pricing documentation has split every rate into a peak window and an off-peak window, off-peak running exactly half of peak.
DeepSeek-V4-Pro costs $1.32 per million input tokens and $3.96 output tokens at peak; $0.66 and $1.98 off-peak. DeepSeek-V4-Flash costs $0.44 and $1.32 at peak; $0.22 and $0.66 off-peak. A cached prompt costs far less again, a roughly 30-fold discount on a cache hit versus a miss, shown in Chart 2. Peak hours are 01:00 - 04:00 and 06:00 - 10:00 UTC, Monday through Friday (all other hours are off-peak).
Claude carries none of this structure. Anthropic prices Opus 5, Sonnet 5, and Haiku 4.5 the same around the clock, at $5.00, $2.00, and $1.00 input, respectively, with output running five times the input rate on every tier.
DeepSeek has shipped several model generations since 2023; the full release history sits on its creator page.
That structural difference is why flat "DeepSeek costs 90 times less" claims, still circulating for this exact search, do not hold up against DeepSeek's current rate card.
Match Opus 5 against V4-Pro's off-peak rate and the real gap is 12.6 times. Match it against V4-Pro's peak rate and the gap narrows to 6.3 times.
Match Haiku 4.5, Claude's cheapest tier, against V4-Flash's peak rate and the gap shrinks to 3.8 times. None resembles 90.

SWE-bench Verified and SWE-bench Pro measure different things
SWE-bench Verified and SWE-bench Pro share a name and a task format. They do not share a difficulty floor. OpenAI built SWE-bench Verified by having 93 professional developers screen the original SWE-bench set by hand, removing ambiguous issues and unreliable environments, cutting 68.3% of the pool and leaving 500 tasks weighted toward simpler fixes.
Scale AI built SWE-bench Pro as a separate benchmark: 1,865 tasks across 41 repositories, many drawn from copyleft or private codebases chosen to resist training-data contamination, averaging 107.4 lines changed across 4.1 files.
That construction gap, not a different scoring method, is why every model here drops 17 to 25 points between the two. DeepSeek-V4-Pro, Max reasoning mode, falls from 80.6% to 55.4%. Claude Sonnet 5 falls from 85.2% to 63.2%.
Claude Opus 5 falls from 96.0% to 79.2%, the smallest drop, and still leads both outright. Weigh SWE-bench Pro more heavily than Verified for a real, multi-file refactor. Verified alone flatters every model on this table. Anthropic has released four Claude generations since 2023; its full model timeline is tracked separately.

Where each vendor actually stores a prompt
DeepSeek names its jurisdiction directly. Its privacy policy states twice, in the main storage section and again in its EEA supplement, that user data is collected, processed, and stored in the People's Republic of China, under operating entity Hangzhou DeepSeek Artificial Intelligence Co., Ltd.
Three governments have already acted on that fact. Italy's data protection authority ordered DeepSeek's chatbot blocked "with immediate effect" in February 2025, citing unsatisfactory answers on data storage. South Korea suspended new app downloads that same month over data-protection law gaps. Australia barred DeepSeek from government devices in February 2025 on national-security advice; personal-device use by staff was not restricted.
Anthropic describes a different default. Its data residency documentation states that inference runs in "any available geography" unless a customer opts into a US-only setting, and training on customer inputs is off by default across the API, Claude for Work, and Claude Gov. Anthropic offers no EU-specific residency directly; EU customers reach regional deployment only through AWS Bedrock, GCP Vertex, or Microsoft Foundry.
One more fact belongs here, since it touches trust rather than accuracy. Anthropic's account, published February 23, 2026, states that DeepSeek generated over 150,000 exchanges with Claude through fraudulent accounts, part of a larger pattern it also attributes to Moonshot AI and MiniMax, calling it a terms-of-service violation.
That claim comes from Anthropic itself, not an independent finding. No independent forensic confirmation of the exchange count turned up, and no DeepSeek response was found. Read it as a disclosed allegation, not an adjudicated fact.
Self-hosting DeepSeek's open weights does not obviously beat the API
DeepSeek ships its weights under an MIT license, downloadable from Hugging Face with no restriction on commercial use. Claude carries no open-weight tier at any price. That asymmetry only matters once the hardware bill lands.
DeepSeek's 671B-parameter open-weight generation needs roughly 625GB of VRAM to run at usable speed, per Inferbase's published cost analysis, which prices an 8-GPU H100 node at $23.92 an hour on RunPod, working out to $3.65 per million output tokens amortized. Lambda, CoreWeave, and AWS price the same node class higher, up to $7.32 on AWS.
DeepSeek-V4-Pro is a larger model again, 1.6 trillion total parameters against the 671B generation Inferbase measured, and needs more hardware, not less.
Set those rental rates against DeepSeek's current price and self-hosting weakens further: off-peak API access already costs $1.98 per million output tokens, below every rental option measured. Only sustained volume in the tens of billions of tokens monthly, or a compliance rule that excludes China-based processing outright, moves the math toward renting the hardware.

What this comparison did not measure
LLM Waves Research ran no latency test, no throughput test, and no output-quality benchmark against any Claude or DeepSeek model for this comparison.
Every accuracy and price figure above is aggregated from vendor documentation, vendor model cards, and named third-party trackers, retrieved September 8, 2026, not measured independently by this team.
Which model fits which reader
For developers
Situational. V4-Pro's off-peak rate suits high-volume iteration; Opus 5 suits the harder ticket where the 15.4-point Verified gap changes the outcome. Build during DeepSeek's peak window and the price gap narrows to 6.3 times, not 12.6. Anthropic's cache minimums and thinking-token costs run deeper than this piece covers; the full Claude API pricing breakdown has the tool-use and extended-thinking mechanics.
For startups
Best value: DeepSeek-V4-Flash off-peak. Disqualifier: not for a team contractually barred from China-processed data.
For enterprise
Situational. Claude reaches EU residency only through AWS Bedrock, GCP Vertex, or Microsoft Foundry. DeepSeek offers none directly.
For high volume
Best value: DeepSeek-V4-Flash. Plan around its 2,500-concurrent-connection cap, not a requests-per-minute ceiling; DeepSeek publishes no RPM figure.
For low latency
Insufficient data. Neither vendor published a comparable latency figure, and this article ran no timing test.
For on-device
Not recommended, either vendor. V4-Pro's footprint needs multi-GPU clusters, not a laptop. Claude has no open-weight tier at all.
For self-hosting
Situational, DeepSeek only, since Claude has no open-weight tier. The compute math above shows it rarely beats the API outright.
For non-English
Insufficient data. No vendor-published non-English accuracy figure turned up for either model.
Is DeepSeek or Claude better for coding?
Claude Opus 5 scores higher on both SWE-bench Verified, 96.0% against 80.6%, and SWE-bench Pro, 79.2% against 55.4%. DeepSeek-V4-Pro's off-peak rate costs 12.6 times less per output token. A team billing per fix should weigh that price gap against how often the extra 15.4 points changes whether a fix ships clean.
Which vendor is more private, and where is data actually processed?
DeepSeek's privacy policy names China as the storage and processing location. Anthropic defaults to a global inference footprint with a US-only opt-in and offers no EU-specific residency directly, only through AWS, GCP, or Microsoft regional deployments.
Neither structure is private by default in the strict sense; both require an active choice to restrict where data lands.
Can I self-host DeepSeek, and is that actually cheaper than the API?
DeepSeek's weights are open under an MIT license. Cloud GPU rental for the hardware class its 671B-parameter generation needs runs $3.65 to $7.32 per million output tokens, against DeepSeek's off-peak API rate of $1.98.
Self-hosting only wins at very high sustained volume or where a compliance rule forces it.
Does DeepSeek's API have a free tier?
DeepSeek's documentation does not state a guaranteed free API quota in the pages checked for this article. Third-party trackers report a roughly 5-million-token signup credit, but a second tracker disputes that it is universal.
DeepSeek's free chat product, separate from the metered API, is the only free access confirmed directly from DeepSeek.
Recheck both vendors' pricing pages before building on this comparison
DeepSeek changed its entire rate card once already in 2026, and Anthropic announced, then canceled, a Sonnet 5 price increase set for September 1, 2026. A price captured on September 8, 2026, is a snapshot, not a standing fact.
