Fireworks AI and DeepInfra both serve DeepSeek-V4-Pro-0813 over an OpenAI-compatible API. On serverless pricing retrieved 2026-08-25, DeepInfra charges 34% less per token, and the gap triples to 267% on the cheaper Flash tier. If price alone decides the pick, DeepInfra wins and no further reading is required.

Serverless pricing on shared models (per 1M tokens, standard tier, retrieved 2026-08-25)

Model

Fireworks input

Fireworks output

DeepInfra input

DeepInfra output

Cheaper

DeepSeek-V4-Pro-0813

$1.74

$3.48

$1.30

$2.60

DeepInfra

DeepSeek-V4-Flash-0731

$0.22

$0.66

$0.08

$0.18

DeepInfra

Llama 3.3 70B

—

—

$0.10

$0.32

not ranked

Table note: the dash marks a price not published on the provider's pricing page on 2026-08-25. Llama 3.3 70B is excluded from the cheaper column rather than ranked last, since Fireworks' pricing documentation lists no serverless rate for it on that date.

DeepSeek-V4-Pro-0813 costs $1.74 per 1M input tokens and $3.48 per 1M output tokens on Fireworks' serverless pricing documentation, against $1.30 and $2.60 on DeepInfra's pricing page, a 34% gap on both figures. DeepSeek-V4-Flash-0731 shows a wider split: $0.66 per 1M output tokens on Fireworks against $0.18 on DeepInfra, a 267% difference.

Fireworks' current serverless pricing page does not list Llama 3.3 70B; DeepInfra prices that model at $0.10 input and $0.32 output per 1M tokens.

The percentage gap is not constant across the two DeepSeek tiers, and the chart below makes that visible at a glance.

The Flash-tier gap is nearly eight times wider in percentage terms than the Pro-tier gap, even though the dollar amounts at stake on Flash are small.

Basis and method

This comparison uses aggregated figures collected from each provider's own pricing and documentation pages, from Artificial Analysis, and from LLM Waves' own inference tracking pages.

No figure below was produced by a benchmark run for this article. Every number carries its source, its URL, and the date it was retrieved.

Differences between the two providers reflect what each documents publicly, not an independent verification of either provider's claims.

What each source measures. Provider pricing pages measure the rate each provider publishes for itself, not an independently timed request. Artificial Analysis and the LLM Waves inference pages measure a live, rolling snapshot of output speed and latency across each provider's catalog; neither page carries its own fixed measurement date, so treat the speed figures below as directional rather than pinned to 2026-08-25 specifically.

Arithmetic performed. Percentage gaps between the two providers on an identical model version are computed as (higher price divided by lower price, minus one). GPU hourly rate comparisons are computed the same way.

Statement. No figure in this article was produced by testing. Every number is either a provider's own published rate or a third party's published measurement, cited at its source.

Provenance. LLM Waves Research holds no pre-release access, free credits, or rate-limit exemptions from Fireworks AI or DeepInfra. Neither provider saw this article before publication.

Methodology page: How LLM Waves measures inference providers.

This comparison ranks Fireworks AI against DeepInfra on five decision axes: per-token price on shared models, dedicated GPU price, caching and batch discount depth, rate limit design, and documented compliance and support structure.

What per-token price doesn't show

Prompt cache discounts

Prompt caching bills a discounted rate when a request reuses an already-processed prompt prefix. Fireworks prices cached input for DeepSeek-V4-Pro-0813 at $0.145 per 1M tokens against $1.74 for a fresh prompt, a 91.7% discount built into the standard rate. DeepInfra prices the same cached tokens at $0.10 against $1.30 standard, a 92.3% discount, functionally identical to Fireworks on this model.

The difference sits one layer down. DeepInfra also sells a Prompt Cache Retention add-on that guarantees a 5-minute or 1-hour cache window, for a 1.25x or 2x write premium on top of the standard input rate.

That add-on is documented as live on only two models, Nemotron-3-Ultra-550B-A55B and Kimi-K2.7-Code, as of 2026-08-25. Fireworks documents no equivalent paid guarantee; its discount applies automatically wherever a cache hit occurs, with no window to purchase.

Batch discounts

Batch processing shows a clearer split. Fireworks' batch inference documentation states a 50% discount off standard per-token pricing, stacking an additional 50% off cached tokens within the same job. DeepInfra's batch API announcement states a flat 20% discount.

Fireworks batch jobs run within a selectable 12 to 72 hour expiration window; DeepInfra returns batch results within a fixed 24 hours. A workload heavy on repeated, cacheable prompts saves more on Fireworks; a workload that needs a guaranteed same-day turnaround gets a firmer number from DeepInfra.

Dedicated GPU pricing runs 2.6 to 3.2 times higher on Fireworks

Both providers rent dedicated GPUs by the hour for workloads that outgrow shared serverless capacity. On rates retrieved 2026-08-25, Fireworks' pricing page lists an H100 80GB at $7.00 per hour against $2.20 on DeepInfra, a 3.18x gap. Fireworks has announced a rate increase to $8.00 per hour effective 2026-09-01; DeepInfra's rate carried no announced change as of the 2026-08-25 retrieval.

GPU

Fireworks (2026-08-25)

DeepInfra (2026-08-25)

Ratio

H100 80GB

$7.00/hr

$2.20/hr

3.18x

H200 141GB

$7.00/hr

$2.69/hr

2.60x

B200 180GB

$10.00/hr

$3.69/hr

2.71x

Table note: Fireworks rates rise across this table on 2026-09-01, to $8.00, $8.00, and $13.00 per hour respectively. DeepInfra published no equivalent change as of the 2026-08-25 retrieval.

The ratio holds within a narrow band, 2.6x to 3.2x, across three GPU classes rather than spiking on one card. That consistency across hardware classes is a normalization this article computed directly from each provider's published rate card; none of the comparison pages we reviewed while building this piece combined dedicated GPU pricing into a single ratio.

The chart below plots that same rate card across all three GPU classes side by side.

No single GPU class drives the gap; a team budgeting for dedicated capacity on either provider can apply roughly the same multiplier regardless of which card it rents.

How many models each provider actually hosts

Fireworks' own model catalog page states it serves "hundreds of models" without a fixed count.

DeepInfra's documentation states "100+ open source models" across text, embeddings, vision, and speech categories. Neither figure resolves to a specific number, and third-party trackers do not agree with each other either.

The LLM Waves Fireworks inference page lists 8 tracked models, and the LLM Waves DeepInfra inference page lists 69, both retrieved 2026-08-25. The independent tracker llm-stats.com lists 11 active Fireworks models from 8 organizations; pricepertoken.com lists 71 chat models from 17 authors on DeepInfra.

Every count in this paragraph measures a different thing: a snapshot of actively indexed or benchmarked models, not the full deployable catalog either provider states it hosts.

Two providers can list the same model under different names and versions, which is one reason a shared model count is hard to pin down. LLM Waves plans a dedicated explainer on that naming mismatch; no live article covered it on 2026-08-25.

Placing LLM Waves' own count next to each provider's official catalog language shows how far apart "tracked" and "advertised" can sit.

DeepInfra's tracked count runs roughly six to eight times higher than Fireworks' across both sources, a gap too consistent to be a fluke of one tracker's methodology.

Output speed: vendor claims against independent measurement

Fireworks documents its own inference engine, FireAttention, now on its fourth published version. Fireworks' FireAttention V4 announcement reports over 250 tokens per second on 8 NVIDIA B200 GPUs running DeepSeek V3-0324 at FP4 precision, a vendor claim describing a specific hardware and quantization configuration, not a general throughput guarantee.

Artificial Analysis, an independent tracker, lists Nemotron 3.5 Lightning as the fastest model on both platforms in its current snapshot: 554 tokens per second on Fireworks against approximately 450 tokens per second on DeepInfra. LLM Waves' own catalog-wide median tells a different story: 116 tokens per second across Fireworks' 8 tracked models against 73 tokens per second across DeepInfra's 69, a comparison of a small curated catalog against a much larger one rather than a like-for-like speed test. Treat both figures as directional; neither source stamps a fixed measurement date.

How the two providers throttle usage

Rate limit design

Fireworks documents a default serverless ceiling of 21.6 million total prompt tokens per minute, adaptive by account spend tier, with a Priority tier that reduces 503 errors under load rather than raising the numeric ceiling.

DeepInfra documents a default limit of 200 concurrent requests per model, which converts to roughly 12,000 requests per minute for one-second requests.

The two designs suit different traffic shapes. A token-ceiling model favors a smaller number of large, high-volume requests with a predictable aggregate budget. A concurrency-slot model favors many small, fast requests running in parallel until the per-model slot count is exhausted.

Neither documentation states a general HTTP status code behavior for the other provider's failure mode, so a team switching providers should test its own request pattern against each limit rather than assume equivalence.

Output token ceilings

DeepInfra's chat documentation caps maximum output at 16,384 tokens for most models, a hard ceiling stated on its own page.

Fireworks' documentation does not publish an equivalent global output ceiling; this article found no comparable figure and marks Fireworks insufficient data on this point rather than assuming no limit exists.

Compliance, data retention, and what neither provider publishes

Fireworks' data security documentation lists SOC 2 Type II and HIPAA compliance, plus ISO 27001, 27701, and 42001 certification, alongside a zero data retention option and encryption in transit and at rest.

DeepInfra's terms page documents a zero data retention commitment of its own, stating that customer data is not used to train or fine-tune models except when a customer requests fine-tuning for its own use; this review found no equivalent published certification list on DeepInfra's own pages, which is a gap in what DeepInfra documents, not a claim that no such certification exists.

Neither provider publishes a numeric uptime SLA, such as a stated 99.9% availability figure, on its own documentation. Fireworks instead publishes response-time support tiers: one hour for production-breaking issues, four business hours for high severity, eight business hours for medium, and two business days for low severity.

DeepInfra's terms describe a 24-hour hardware replacement commitment for dedicated deployments, with no equivalent response-time tier structure found. A team that needs a contractual uptime number will not get one from either provider's public documentation.

Fine-tuning and dedicated deployment

Fireworks prices LoRA supervised fine-tuning at $0.50 per 1M training tokens for models up to 16 billion parameters, scaling to $10 to $40 per 1M tokens above 300 billion.

Fireworks documents hosting up to 100 fine-tuned adapters per deployment at no added serving cost, billed at the base model's per-token rate. DeepInfra documents LoRA support within its private deployment product, billed by GPU hour rather than by training token; this review found no equivalent per-token fine-tuning rate table on DeepInfra's documentation.

What we did not measure

This article did not run inference requests against either provider. It did not independently time latency, verify uptime, or reproduce either provider's own speed claims. Every speed and latency figure above is either a provider's own published claim or a third party's live snapshot, clearly labeled as one or the other. Pricing for models outside the three compared here, and for regions or currencies other than the published USD rate, was not collected.

Which provider fits which reader

For developers evaluating both, DeepInfra is the top pick on raw catalog breadth, listing 69 models on LLM Waves' own tracking page against 8 on Fireworks, alongside the cheaper per-token rate measured above. If a project needs Fireworks-specific tooling such as its Grammar Mode structured output or FireAttention-tuned proprietary models, Fireworks is the better starting point instead.

For startups, DeepInfra is best value: its DeepStart program grants 1 billion free tokens to companies that have raised between $250,000 and $10 million and were founded within the last two years. A startup outside that window gets only a $1 evaluation credit from Fireworks, not a production budget.

For high-volume batch workloads, Fireworks is the top pick, on a 50% batch discount that stacks with its cache discount, against DeepInfra's flat 20%. A workload that cannot tolerate up to a 72-hour completion window should check DeepInfra's fixed 24-hour turnaround instead.

For regulated or enterprise buyers, Fireworks is the top pick on documented compliance, listing SOC 2 Type II, HIPAA, and three ISO certifications where DeepInfra's own pages list none. A buyer for whom price outweighs paperwork should still price DeepInfra directly, since its lower per-token and per-GPU rates stand regardless of the certification gap.

For dedicated, self-hosted deployment, DeepInfra is the top pick, with an H100 80GB priced at $2.20 per hour against $7.00 on Fireworks, a gap larger than any measurement variance in this comparison. A deployment that specifically needs Fireworks' bring-your-own-cloud enterprise program or region-restricted hosting should weigh that feature against the price gap directly.

For latency-sensitive single-request workloads, this comparison calls it situational. Artificial Analysis' snapshot puts the fastest tracked model, Nemotron 3.5 Lightning, at 554 tokens per second on Fireworks against roughly 450 on DeepInfra, a real gap but one drawn from an undated rolling snapshot rather than a fixed benchmark run.

Is Fireworks AI or DeepInfra cheaper?

DeepInfra prices every shared model in this comparison lower than Fireworks: 34% lower on DeepSeek-V4-Pro-0813 and 267% lower on DeepSeek-V4-Flash-0731 output tokens.

Dedicated GPU pricing shows the widest gap, with DeepInfra's H100 80GB rate 3.18 times cheaper.

How many models do Fireworks AI and DeepInfra actually host?

Neither provider publishes an exact number. Fireworks describes "hundreds of models"; DeepInfra describes "100+ open source models." Third-party trackers count anywhere from 8 to 71, depending on whether they track a curated subset or attempt a fuller catalog.

Are Fireworks AI and DeepInfra OpenAI-compatible?

Yes. DeepInfra documents a drop-in OpenAI-compatible endpoint at api.deepinfra.com/v1/openai.

Fireworks documents OpenAI-compatible endpoints alongside its own structured-output and function-calling extensions.

Does either provider publish an uptime SLA?

No. Neither publishes a numeric uptime percentage on its public documentation. Fireworks publishes response-time support tiers instead; DeepInfra publishes only a 24-hour hardware replacement commitment for dedicated deployments.

Can prompts sent to Fireworks AI or DeepInfra be used to train a model?

Not by default on either provider. Fireworks documents that it does not store prompt data for open models without explicit opt-in. DeepInfra's terms state customer data is not used to train models except when a customer requests fine-tuning for its own use.

What happens when a rate limit is exceeded?

The two providers throttle differently. Fireworks documents a token-per-minute ceiling that scales with account spend tier. DeepInfra documents a concurrent-request-per-model ceiling, returning an HTTP 429 response once exceeded.

Published figures answer today's question, not next quarter's

A provider can change a price, a rate limit, or a certification list without a changelog entry. Every figure in this comparison carries its own retrieval date rather than one blanket timestamp for that reason. Check the linked source pages directly before committing budget to either provider.

Raw data: the complete sourced figure set behind every number in this piece, including models and metrics not covered in the prose above, is provided as a downloadable CSV alongside this article.

Changelog

2026-08-25: Published. Fireworks AI and DeepInfra compared on serverless pricing, dedicated GPU pricing, caching and batch discounts, rate limit design, compliance documentation, and catalog size.

Published: 2026-08-25 Last updated: 2026-08-25 Data collected: 2026-08-25 (UTC)

Sources. Provider pricing and docs: docs.fireworks.ai/serverless/pricing, fireworks.ai/pricing, fireworks.ai/models, fireworks.ai/enterprise, docs.fireworks.ai (rate limits, security, structured outputs, vision), fireworks.ai/blog; deepinfra.com/pricing, docs.deepinfra.com (models, chat, private-models, account/rate-limits), deepinfra.com/terms, deepinfra.com/blog. Third-party trackers: llm-stats.com/providers/fireworks, pricepertoken.com/endpoints/deepinfra, openrouter.ai/provider/deepinfra, artificialanalysis.ai/providers/fireworks, artificialanalysis.ai/providers/deepinfra. First-party site data: llmwaves.com/inference/fireworks and llmwaves.com/inference/deepinfra. All retrieved 2026-08-25.