DeepInfra sells DeepSeek-V4-Flash-0731 at $0.08 input and $0.18 output per million tokens, 72% below DeepSeek's own off-peak direct rate, the lowest standing flat price among six routes checked from August 25 to September 2, 2026.

DeepSeek's own off-peak rate wins instead on DeepSeek-V4-Pro-0813, the larger model. A reader who only needs Flash-tier pricing can stop here.

Six routes to DeepSeek-V4-Flash-0731 and DeepSeek-V4-Pro-0813, priced the same day

Route

Flash-0731 input ($/1M)

Flash-0731 output ($/1M)

Pro-0813 input ($/1M)

Pro-0813 output ($/1M)

Verdict

DeepInfra

$0.08

$0.18

$1.30

$2.60

Best value, Flash

DeepSeek direct (off-peak)

$0.22

$0.66

$0.66

$1.98

Best value, Pro

OpenRouter (discounted route)

$0.05

$0.10

$0.58

$1.74

Situational†

Together

$0.14

$0.28

$1.32

$3.96

Close second

Novita

$0.14

$0.28

$1.32

$3.96

Close second‡

Fireworks

n/r

n/r

n/a

n/a

Insufficient data

Table note: n/r means no standard inference rate was found on the page checked. n/a means the row was excluded from ranking rather than scored last. † the discounted OpenRouter route is not the same figure OpenRouter shows on its own undated model pages, detailed below. ‡ Novita's pricing page also lists DeepSeek-V4-Flash-0731 a second time at $0.44 input and $1.32 output, identical to its listing for a separate vision-preview model, an inconsistency unresolved as of the retrieval date. Direction of goodness: lower is better on every price column.

Three numbers carry the finding. DeepInfra prices DeepSeek-V4-Flash-0731 at $0.08 per million input tokens and $0.18 per million output tokens, the cheapest standing rate this comparison found.

DeepSeek's own off-peak direct rate for DeepSeek-V4-Pro-0813 is $0.66 input and $1.98 output, cheaper than every reseller's flat rate for the same model.

Fireworks carries no comparable standard inference price in this table, marked as insufficient data rather than ranked last.

Basis and method

LLM Waves Research collected every figure in this article from the named provider's own pricing or documentation page between 2026-08-25 and 2026-09-02 UTC. This is Basis 2: aggregation, not measurement.

No figure in this article was produced by running a prompt against any endpoint; every number is copied from a published rate card and checked against the arithmetic shown in the tables.

The blended price column weights one input token to four output tokens, a common ratio for multi-turn chat and agent workloads, computed as 0.2 times the input rate plus 0.8 times the output rate.

LLM Waves Research received no pre-release access, no free credits, and no rate-limit exemption from any provider named here, and no provider saw this article before publication.

Figures were collected by an automated retrieval script (v2.3.1) and spot-checked by hand against the primary page for each of the six routes; prose was written and edited by the LLM Waves Research team. Full source URLs and retrieval timestamps live in the methodology page and the linked raw data file.


We rank six routes below on blended price per million tokens for DeepSeek-V4-Flash-0731 and DeepSeek-V4-Pro-0813, then check peak-hour surcharges, cache discounts, the July migration, and the compliance question enterprise buyers ask before any of that pricing matters.

What does the DeepSeek API actually cost per model, per million tokens?

DeepSeek-V4-Flash-0731 (API ID deepseek-v4-flash-0731) is the small, fast model; DeepSeek-V4-Pro-0813 (API ID deepseek-v4-pro-0813) is the larger reasoning-capable model. Both list a context window of 1,000,000 tokens on DeepSeek's own pricing documentation, and both bill by the token rather than by the request.

We price six routes to those two models only, here. A wider comparison across providers and models beyond DeepSeek runs on the broader LLM API provider comparison.

DeepSeek's rolling aliases deepseek-v4-flash and deepseek-v4-pro point to the current snapshot without a dated suffix.

Three of the five resellers in this comparison price the dated snapshot and the rolling alias differently, which is one reason the same model name carries different numbers on different pages.

DigitalOcean's route through OpenRouter and OpenRouter's own headline listing for deepseek-v4-pro both show $0.87 input and $1.74 output, yet OpenRouter's own listing page for the dated snapshot deepseek-v4-pro-0813 shows $0.58 input and $1.74 output on the same domain, retrieved the same day.

Two pages, one company, two prices. A reader would reasonably assume both describe the same product. Nobody publishes a rate card that names the snapshot and the alias side by side and says which one a new integration actually calls by default.

DeepSeek's peak surcharge doubles the off-peak price

DeepSeek's own documentation states that off-peak rates run at half the peak rate, not a discount bolted onto a base price but the base price itself moving on a clock.

Peak hours run 01:00 to 04:00 UTC and 06:00 to 10:00 UTC, Monday through Friday, which lands during the Chinese business day and, for most United States developers working standard hours, during the American night and early morning.

Off-peak output for DeepSeek-V4-Pro-0813 costs $1.98 per million tokens; peak output for the same model costs $3.96.

That is not a rounding difference. A team running batch jobs on a schedule it controls can choose the off-peak window deliberately; a team serving live user traffic around the clock pays the blended average of both windows whether it tracks the schedule or not.

DeepSeek's own pricing page lists DeepSeek-V4-Pro-0813 output at $3.96 per million tokens at peak and $1.98 off-peak, and DeepSeek-V4-Flash-0731 output at $1.32 peak against $0.66 off-peak, an exact halving in both cases.

Cached input costs 97% less than a cache miss

Cache math is where DeepSeek's own direct rate stops looking expensive. Off-peak cache-hit input for DeepSeek-V4-Flash-0731 costs $0.007 per million tokens against $0.22 for a cache miss, a reduction of 96.8%.

DeepSeek-V4-Pro-0813 shows the same pattern at a higher price point, $0.022 per million cache-hit tokens against $0.66 on a miss, a 96.7% cut. Every reseller in this comparison also lists a cached rate.

DeepInfra's own pricing page lists its cheapest cached input at $0.016 for Flash and $0.10 for Pro, both higher than DeepSeek's own $0.007 and $0.022. Nobody beats that floor.

A workload that resends the same system prompt and tool schema on every call, the common shape of an agent loop, earns almost all of that discount. A workload of one-off, unrepeated prompts earns none of it, whichever provider processes the request.


A cache hit on DeepSeek-V4-Pro-0813 costs $0.022 per million input tokens against $0.66 for a cache miss at off-peak rates, the widest reseller-beating gap measured in this comparison.

What happened to deepseek-chat and deepseek-reasoner, and what breaks if an integration still calls them

DeepSeek's own change log states that the legacy model names deepseek-chat and deepseek-reasoner would be discontinued three months out, announced 2026-04-24 for a 2026-07-24 cutover.

During the transition, deepseek-chat pointed to DeepSeek-V4-Flash in non-thinking mode and deepseek-reasoner pointed to the same model in thinking mode.

Reasoning stopped being a separate model name and became a request parameter on a single model family.

An integration written against deepseek-reasoner before July 24 needs two changes, not one: swap the model string to deepseek-v4-flash or deepseek-v4-pro, then set the thinking parameter explicitly, because the old alias no longer carries that setting implicitly.

DeepSeek's documentation does not describe what error response a call to the retired names now returns; that detail was not confirmed this session and is not repeated here as fact.

Is the DeepSeek API safe or compliant enough for enterprise use?

DeepSeek's own privacy policy states that user data is collected, processed, and stored in the People's Republic of China, and that data may be used "to train and improve our technology, such as our machine learning models and algorithms."

The same policy also states a right to opt out of that training use, exercised by writing to privacy@deepseek.com, a detail some third-party privacy write-ups omit when they describe the policy as having no opt-out.

Compare that against OpenAI's own data-controls documentation, which states that data sent to its API "is not used to train or improve OpenAI models" unless a customer opts in, a policy in effect since March 1, 2023.

DeepSeek defaults to training-in with an opt-out; OpenAI defaults to training-out with an opt-in. Read the policy, not the summary. That distinction, not an adjective, is the actual compliance question.

One caveat matters for this comparison specifically: the data-location and training terms in DeepSeek's privacy policy describe the API operated directly by DeepSeek.

A reseller hosting the same open-weight model on its own infrastructure is a separate data-processing relationship governed by that reseller's own terms, not DeepSeek's privacy policy, and this article did not check each of the five resellers' own data-processing agreements individually.

An enterprise buyer with a data-residency requirement should read the specific reseller's DPA before assuming the China-routing question is settled either way by what DeepSeek's own policy says.

DeepSeek-V4-Pro-0813 undercuts every flagship rival on output

Set against three flagship-tier rivals at list price, DeepSeek-V4-Pro-0813 through DeepInfra costs $2.34 blended per million tokens. Gemini 3.1 Pro Preview costs $10.00 blended at its sub-200,000-token input tier.

Claude Opus 5 costs $21.00 blended. GPT-5.6 Sol costs $25.00 blended, the most expensive of the four.

Every one of those three flagship competitors lists a published rate card of its own, so this is a reported comparison, not a characterization: the gap runs from 4.3 times to 10.7 times DeepSeek's blended cost, depending on which flagship model a team already pays for.

DeepSeek-V4-Pro-0813 through DeepInfra lists at $2.34 blended per million tokens against $10.00 for Gemini 3.1 Pro Preview, $21.00 for Claude Opus 5, and $25.00 for GPT-5.6 Sol, all figures from each vendor's own pricing page.

DeepSeek-V4-Flash-0731 costs a fraction of every budget rival

The same pattern holds harder at the budget tier. DeepSeek-V4-Flash-0731 through DeepInfra costs $0.16 blended. GPT-5.6 Luna, priced on OpenAI's own rate card, costs $1.00 blended, 84% more. Gemini 3.5 Flash-Lite costs $2.06 blended.

Claude Haiku 4.5 costs $4.20 blended, 26 times DeepSeek-V4-Flash-0731's rate. A team already budgeting for a rival's cheapest tier is likely still overpaying against DeepSeek's cheapest tier, whichever route to DeepSeek it uses; even DeepSeek's own most expensive listed route in this comparison, its off-peak direct rate at $0.57 blended, still beats every non-DeepSeek model checked here.

DeepSeek-V4-Flash-0731 through DeepInfra lists at $0.16 blended per million tokens against $1.00 for GPT-5.6 Luna, $2.06 for Gemini 3.5 Flash-Lite, and $4.20 for Claude Haiku 4.5.

What this comparison did not measure

We priced publicly listed API rates for six named routes to DeepSeek-V4-Flash-0731 and DeepSeek-V4-Pro-0813; we did not measure output quality, latency, uptime, or benchmark accuracy for any provider named here.

We did not independently verify each reseller's default reasoning setting, quantization, or context-window ceiling, so part of the price spread between routes may reflect a configuration difference rather than a pure margin difference.

We did not check the current free-tier credit amount or expiry against any account console.

We did not confirm what error response DeepSeek's API now returns for the retired deepseek-chat and deepseek-reasoner names. We did not audit any of the five resellers' own data-processing agreements.

Which route fits which reader

For developers prototyping against both models

DeepInfra, best value, since it lists both DeepSeek-V4-Flash-0731 and DeepSeek-V4-Pro-0813 at a standing flat rate with no peak-hour clock to track under one account.

For startups watching a per-token budget

DeepInfra's rate for DeepSeek-V4-Flash-0731 is $0.08 input. Output costs $0.18. Top pick on cost among every standing rate checked in this comparison.

For high-volume, cache-heavy agent loops

DeepSeek's own direct off-peak rate, top pick for this use case. Its cache-hit floor of $0.007 to $0.022 undercuts every reseller's cached rate found this session, even though its uncached rate loses to DeepInfra on Flash-tier pricing alone.

For low latency

Insufficient data. Every latency figure surfaced in this research traced either to a provider's own blog about its own service or to an unverified comparative claim about a competitor, and this comparison does not repeat a competitor claim it has not independently sourced.

For enterprise buyers with a data-residency requirement

Situational, conditioned on which route is in question. DeepSeek's direct API is governed by a privacy policy that places data processing in China with an opt-out for training use.

A reseller route is governed by that reseller's own DPA, not checked individually here, so the constraint is: read the specific contract before assuming either answer.

For self-hosting instead of any hosted API

Situational, conditioned on token volume clearing whatever a team's own inference hardware costs. Five independent infrastructure providers hosting DeepSeek-V4-Flash-0731 and DeepSeek-V4-Pro-0813 on their own hardware indicates the weights are licensed for redistribution.

This comparison did not price hardware or self-hosting operating cost, so treat this as a threshold question rather than a number.

Which DeepSeek API is the most affordable?

DeepSeek-V4-Flash-0731 through DeepInfra, at $0.08 input and $0.18 output per million tokens. That is the cheapest standing, non-promotional rate found across the six routes checked in this comparison.

Which is the cheapest API?

Of every model priced in this comparison, DeepSeek-V4-Flash-0731 through DeepInfra is cheapest overall. Blended cost runs $0.16 per million tokens against $1.00 for the cheapest non-DeepSeek model checked, GPT-5.6 Luna. Not close.

Which DeepSeek API is free?

None of the six routes priced in this comparison is free at production volume. DeepSeek has offered free API credit to new accounts in the past; this comparison did not verify the current grant amount or expiry against DeepSeek's own account console, so treat any specific figure quoted elsewhere as unconfirmed until checked there directly.

Does DeepSeek still charge peak and off-peak rates?

Yes: DeepSeek's own pricing documentation lists peak hours running 01:00 to 04:00 UTC and 06:00 to 10:00 UTC on weekdays, with peak pricing running exactly double the off-peak price on every model this comparison checked.

Does DeepSeek have rate limits?

DeepSeek's own pricing documentation states that concurrency limits vary by model in a range of 500 to 2,500 requests, untested under load in this comparison.

Is self-hosting DeepSeek cheaper than the API?

Insufficient data to answer with a number. Self-hosting only beats a $0.08 to $0.18 per-million-token API rate past a volume threshold this comparison did not calculate, since that threshold depends on hardware cost, utilization, and a team's own engineering time, none of which this comparison priced.

How does DeepSeek billing actually work?

Every model bills by the token, not by the request: input tokens split into a cache-hit rate and a cache-miss rate, output tokens bill at a separate rate, and both rates double during the documented peak window and return to the listed rate outside it.

DeepSeek's consumer chat product at its own website is a separate, non-API product from the billing described in this article; nothing in this comparison prices that consumer product.

Recheck this before the next invoice

Every number above is a snapshot of a rate card that changes without a changelog entry on most of the pages that quote it. Five of six routes in this comparison do not publish a peak or off-peak distinction at all, meaning their flat rate could already reflect one, the other, or neither by the time this loads for a reader.

Run the current rate through the account you actually bill against before committing a production workload to any figure in this article.