DeepSeek and ChatGPT split by workload, not by a single winner.
DeepSeek-V4-Pro prices output tokens at $1.98 per million off peak, roughly 10 to 25 times below ChatGPT's GPT-6 Astra and GPT-5.6 Sol tiers, yet above OpenAI's own GPT-5.6 Luna at $1.20.
Figures collected through September 15, 2026.
Price, context and access by model
Rates are per million tokens, listed input then output (lower is better), for the tiers compared throughout this piece; the full DeepSeek-V4-Pro model entry carries additional specification detail beyond price.
DeepSeek's peak window runs 01:00 to 04:00 and 06:00 to 10:00 UTC, seven hours a day; all other hours are off-peak.
Model | Input | Output | Context window | Max output | Consumer access |
|---|---|---|---|---|---|
DeepSeek-V4-Pro, off peak | $0.66 | $1.98 | u | u | DeepSeek app, free |
DeepSeek-V4-Pro, peak | $1.32 | $3.96 | u | u | DeepSeek app, free |
DeepSeek-V4-Flash, off peak | $0.22 | $0.66 | u | u | DeepSeek app, free |
GPT-6 Astra, up to 272,000 input tokens | $10.00 | $50.00 | 1,050,000 | 128,000 | ChatGPT Pro |
GPT-5.6 Sol, promotional through Nov 21, 2026 | $4.00 | $20.00 | 1,050,000 | 128,000 | ChatGPT Plus |
GPT-5.6 Terra, up to 272,000 input tokens | $2.00 | $12.00 | 1,050,000 | 128,000 | not sold standalone |
GPT-5.6 Luna, up to 272,000 input tokens | $0.20 | $1.20 | 1,050,000 | 128,000 | ChatGPT Go |
Table note: u marks a figure not independently confirmed this session; it will be filled once a source is verified. † marks a vendor-reported figure carried at face value. GPT-5.6 Sol's standard, non-promotional rate is reported at $5.00 input / $30.00 output† on OpenAI's marketing page, which conflicts with the developer documentation's promotional figure above; the developer documentation is treated as authoritative here because it is the page developers bill against.
DeepSeek-V4-Pro's off-peak output price sits 25 times below GPT-6 Astra's $50.00 and 10 times below GPT-5.6 Sol's promotional $20.00.
Against GPT-5.6 Luna, the comparison reverses: OpenAI's own cheapest current tier charges $1.20 per million output tokens, 39 percent less than DeepSeek-V4-Pro's off-peak rate.
The 9x and 18x figures visible on other DeepSeek vs ChatGPT pages compare DeepSeek-V4-Pro against whichever ChatGPT tier that page happened to price, not the full four-tier OpenAI ladder above.
Basis and method
Published figures, not new tests, carry this comparison. DeepSeek's rates come from its own API pricing documentation, collected between September 11 and September 15, 2026 (UTC); OpenAI's rates are covered separately below, with full source and retrieval-date detail on the methodology page. The self-hosting and cache-blend figures below are original arithmetic performed on those published rates, with the formula shown inline at each calculation. No figure here came from testing DeepSeek or ChatGPT directly.
Decision axes: this comparison ranks DeepSeek and ChatGPT on price per token at declared tiers, jurisdiction of the underlying weights, and the reliability and benchmark evidence a reader can actually verify.

Off-peak, peak and the tier that actually competes with DeepSeek
DeepSeek's headline price only holds for seven-eighths of the day. During its 01:00 to 04:00 and 06:00 to 10:00 UTC peak windows, DeepSeek-V4-Pro doubles to $1.32 input and $3.96 output.
A reader running batch jobs on a fixed daily schedule can land entirely inside or entirely outside that window by accident.
GPT-5.6 Terra, priced at $2.00 input and $12.00 output on OpenAI's developer pricing page, is the OpenAI tier closest in ambition to DeepSeek-V4-Pro. It still costs six times DeepSeek-V4-Pro's off-peak output rate and three times its peak rate.
GPT-5.6 Luna is the tier that actually undercuts DeepSeek on output price, and it does so by running a smaller, cheaper model than the one most DeepSeek vs ChatGPT comparisons put on the other side of the table.
The peak/off-peak split and its origin get fuller treatment in the site's DeepSeek API pricing coverage.
If a workload runs entirely inside DeepSeek's peak window, DeepSeek-V4-Pro is the wrong pick for a cost-sensitive reader; GPT-5.6 Luna or a shifted schedule both undercut it.

What cache hits actually save
DeepSeek lists roughly a 30 times discount on cached input tokens, dropping the off-peak $0.66 rate to about $0.022 per million. That number only applies to the share of a request that hits cache.
The headline rate was never the real one. Blended input cost = (1 minus cache-hit share) times $0.66, plus cache-hit share times $0.022.
At a 50 percent hit rate, the effective input price is $0.34 per million, near half the headline.
At 80 percent, a realistic figure for a workload reusing the same system prompt, it drops to $0.15.
At 95 percent it reaches $0.05. None of these numbers appear on DeepSeek's pricing page; they follow from applying its own published discount to its own published base rate.

Self-hosting the open weights, priced against renting the API
DeepSeek-V4-Pro ships as 671B-class open weights under an MIT license, which means a reader can run it without touching DeepSeek's endpoint at all.
An 8-way H100 80GB node, a common configuration for a model of that size, rents for roughly $2.00 per GPU-hour on current cloud markets, or about $11,680 per month at continuous use.
At an assumed sustained throughput of 4,000 output tokens per second, that node can produce about 10.5 billion output tokens across a full month if kept busy the entire time, putting the floor cost near $1.11 per million tokens.
Effective cost per million tokens = monthly node cost, divided by utilization share times monthly token capacity, which reduces to $1.11 divided by utilization share once the node's own fixed cost and capacity are locked in.
Utilization decides the winner. Break-even utilization against any API rate equals $1.11 divided by that rate.
Against the off-peak output rate of $1.98, that is 1.11 divided by 1.98, or about 56 percent.
Against the peak rate of $3.96, exactly double the off-peak rate, it is 1.11 divided by 3.96, or about 28 percent, half the off-peak threshold for the same doubling reason.
Below those thresholds, DeepSeek's own API is cheaper than running the weights yourself. For a reader whose objection to DeepSeek is jurisdiction rather than price, the same MIT-licensed weights run on non-Chinese infrastructure without a self-hosting project at all.
A third-party inference provider, Fireworks AI, hosts DeepSeek-V4-class models on US infrastructure at rates close to DeepSeek's own, and DeepInfra offers a comparable option; either resolves the "my data goes to China" objection that stops readers cold on other pages in this comparison without ever pricing the fix.

Which benchmark score is real
Three different SWE-bench Verified scores for DeepSeek-V4-Pro circulate across pages covering this comparison: 80.6 percent, 73.1 percent, and approximately 91.2 percent.
DeepSeek's own model card states 80.6 percent SWE-bench Verified and 55.4 percent SWE-bench Pro; that is the figure this article uses, because it is the only one of the three tied to a named, checkable source rather than a secondary aggregator.
GPT-6 Astra's own launch material does not report a SWE-bench Verified score at all.
It headlines OSWorld 2.0, ARC-AGI-3, GPQA Diamond, Terminal-Bench 4.0, DeepSWE v1.1, and FrontierCode 1.1 instead, which means DeepSeek and ChatGPT share no benchmark that both vendors disclose on their own terms.
A reader comparing "DeepSeek's 80.6 against ChatGPT's coding score" is comparing a number to a number that does not exist yet.
Whether it is even up
DeepSeek does not publish a public uptime figure, and no independent, dated availability report exists to fill that gap. What does exist is volume: "server busy" is the single most repeated phrase in reader complaints about DeepSeek across the community threads this research surveyed, with one user reporting the message fifteen times in one session.
That is a reported signal from users, not a measured statistic, and it is treated that way here. Neither vendor publishes an audited figure. OpenAI publishes a public status page for ChatGPT and the API, which at minimum gives a reader somewhere to check before committing a production workload; DeepSeek does not offer an equivalent page.
More on how OpenAI structures its releases sits on the site's OpenAI creator page. The two consumer apps also differ on caps: DeepSeek's app carries no published message limit, while ChatGPT's free tier is capped at 10 messages per 5 hours, a fact that shapes which one a non-paying reader can actually rely on day to day.

What this article did not measure
This article did not run a latency test, a throughput test, or an output-quality benchmark against any model discussed. The self-hosting and cache-blend figures are arithmetic on vendor-published rates, not measurements of either product in operation.
Uptime is reported from user complaints, not from an independently logged availability record.
Which one for which reader
For a developer building at scale
DeepSeek-V4-Pro off peak, with its cache discount applied to any repeated system prompt. Top pick on price against every ChatGPT tier except Luna.
Wrong pick if the workload runs inside DeepSeek's daily peak window without a way to shift it.
For a startup watching runway
GPT-5.6 Luna is the best-value pick: it beats DeepSeek-V4-Pro's own off-peak output rate while running on the same 1,050,000-token context window as OpenAI's top-priced tiers.
Wrong pick if the task needs the heavier reasoning Luna was not built to deliver.
For a regulated enterprise buyer
ChatGPT, or DeepSeek's MIT-licensed weights run through a non-Chinese host such as Fireworks AI or DeepInfra. Situational: DeepSeek's own China-based endpoint is not recommended where data residency is a hard requirement, and the same weights elsewhere remove that specific objection without changing the price story.
For self-hosting
DeepSeek-V4-Pro's open weights, above roughly 56 percent sustained node utilization. Not recommended below that threshold; renting the API is cheaper until utilization clears the break-even point calculated above.
Track utilization before committing to hardware.
For non-English use
Insufficient data. Neither vendor's own documentation in this research supplies a dated, sourced non-English benchmark for the tiers compared here, and no substitute figure was inferred to fill the gap.
Is DeepSeek actually cheaper than ChatGPT once cache hits and peak hours are counted in
Only against ChatGPT's two most expensive tiers. DeepSeek-V4-Pro off peak, with a high cache-hit rate applied, beats GPT-6 Astra and GPT-5.6 Sol by a wide margin, and it still beats GPT-5.6 Terra.
It does not beat GPT-5.6 Luna, OpenAI's own budget tier, on raw output price.
Can you self-host DeepSeek, and at what volume does it beat paying the API
Above roughly 56 percent sustained utilization of an 8-way H100 node against the off-peak API rate, or above roughly 28 percent against the peak rate, using the assumptions and formula stated in the self-hosting section above.
Below those thresholds, DeepSeek's own API costs less than running the weights.
Is DeepSeek safe to use for business data
DeepSeek's privacy policy names China as its storage and processing location. Italy, South Korea, and Australia took government action against the service in February 2025, and Taiwan, India, the Czech Republic, and more than 17 US states have since restricted or banned it in some official capacity.
Readers for whom that is disqualifying can run the same MIT-licensed weights through a non-Chinese host instead of dismissing the model outright.
Which is better for coding, DeepSeek or ChatGPT
Unresolved, because the two vendors do not publish a shared benchmark. DeepSeek's own model card reports 80.6 percent on SWE-bench Verified; GPT-6 Astra's launch material does not mention that benchmark at all.
A reader who needs a real answer should run both models against a sample of their own tasks rather than trust either vendor's chosen scoreboard.
Can DeepSeek replace ChatGPT, or do you end up running both
For a single cost-sensitive workload with a schedule that avoids DeepSeek's peak window, DeepSeek alone is workable. For anything that also needs a checkable uptime record, non-English benchmark coverage, or a subscription-priced consumer app rather than metered API billing, most readers in this research end up running both rather than replacing one with the other.
DeepSeek's peak-hour doubling, OpenAI's own Luna undercut, and the missing shared benchmark are all facts that hold today and can all move before either vendor's next release. Recheck the numbers above against the sources in the raw data file before building a production decision on them.
