This comparison covers 13 LLM API providers, dated 2026-08-26: Groq posts the fastest measured output speed, at 325.2 tokens per second on Llama 3.3 70B Instruct, while DeepInfra prices the same-class DeepSeek-V4-Pro-0813 model 34% below Together AI on output tokens.

Read the use-case verdicts below if your workload does not match either axis.

Provider

Representative model

Price in / out per 1M tokens

Context window

Speed signal

Free tier

Google

Gemini 2.5 Flash-Lite

$0.10 / $0.40

ND for this tier

ND

Yes, up to 500 req/day

AkashML

Llama 3.3 70B

$0.20 / $0.52

ND

ND

$100 credit

Groq

GPT-OSS 120B

$0.15 / $0.60

131,072 tokens

325.2 tokens/s measured, Llama 3.3 70B (Artificial Analysis); 280 tokens/s Groq's own figure, same model

Yes, rate-limited

Mistral

Mistral Large 3

$0.50 / $1.50

ND

ND

ND

DeepInfra

DeepSeek-V4-Pro-0813

$1.30 / $2.60

1,024K tokens (DeepInfra's own figure)

17.0 tokens/s measured, Llama 3.3 70B Turbo (Artificial Analysis)

No

Baseten

DeepSeek V4 Pro

$1.74 / $3.48

ND

ND

ND

Fireworks AI

DeepSeek-V4-Pro-0813

$1.74 / $3.48

1,048,576 tokens

ND

$1 credit

Together AI

DeepSeek-V4-Pro-0813

$1.32 / $3.96

ND

72.0 tokens/s measured, Turbo variant (Artificial Analysis)

ND

Anthropic

Claude Sonnet 5

$2.00 / $10.00

1,048,576 tokens

ND

No

OpenAI

GPT-5.6 Sol

$2.00 / $10.00

1,050,000 tokens (stated as "1.05M" on the models page)

ND

No

Cerebras

Llama 4 Scout

ND

ND

"Over 2,000 tokens/s" (Cerebras's own claim, not independently verified)

$5 credit

OpenRouter

Routes to hundreds of models

Pass-through, no markup (OpenRouter's own claim)

Varies by routed model

n/a, routing layer

Yes, 50 to 1,000 req/day

Amazon Bedrock

Claude Sonnet 5

E

1,048,576 tokens

153.7 tokens/s measured, Llama 3.3 70B (Artificial Analysis)

n/a

Table note: ND not disclosed on the page or pages checked · n/a field does not apply to this provider · E insufficient data, current-generation on-demand pricing not published in the source checked.

Rows are ordered by output price ascending for the 10 providers where a comparable output price was found; the remaining three rows are grouped below them. Lower output price is better on this axis.

Figures in this table are collected from the sources named in our methodology, retrieved 2026-08-26. DeepInfra prices DeepSeek-V4-Pro-0813 output tokens at $2.60 per million, 34% below Together AI's $3.96 and 25% below Fireworks AI's standard-tier $3.48 for the model under the same name, as reported on each provider's own pricing page on 2026-08-26.

Groq and DeepInfra sit at opposite ends of the one speed comparison we could run across four providers on identical hardware conditions, at 325.2 against 17.0 tokens per second for Llama 3.3 70B Instruct.

Basis and method

This comparison uses Basis 2, aggregated: every figure below was collected from a provider's own documentation, pricing page or model card, or from Artificial Analysis, an independent third-party benchmarking site, never from a test LLM Waves Research ran itself.

Each figure carries its source artifact, URL and retrieval date in the raw data file linked at the close of this article, and full sourcing detail lives on the methodology page. Percentage differences in this article are computed as (higher price minus lower price) divided by higher price, using each provider's own stated list price on 2026-08-26; list price is not always the price a given account pays after volume discounts or negotiated terms.

Speed and time-to-first-token (TTFT) figures attributed to Artificial Analysis come from its independent benchmark runs against each provider's first-party API, not from a provider's own marketing claim, and are labeled as such throughout. Methodology version 1.0, first publication.

This comparison ranks primarily on delivered output price per 1 million tokens for a shared model class where one exists, with independently measured or vendor-reported tokens per second and published context window reported alongside for every provider. The panel below pulls four of those figures out of the table for a faster read before the provider-by-provider detail starts.

None of these four figures describe every model a given provider hosts, and the full table above remains the source for any model this panel does not name.

Groq prices GPT-OSS 120B at $0.15 and posts the fastest measured speed

Groq lists GPT-OSS 120B at $0.15 per million input tokens and $0.60 per million output tokens, with a 131,072-token context window, on its model pricing page, retrieved 2026-08-26. Groq's own docs state 280 tokens per second for Llama 3.3 70B, on that same page.

Artificial Analysis measured Groq's Llama 3.3 70B Instruct endpoint at 325.2 tokens per second output and 0.91 seconds time to first token, the fastest and lowest-TTFT result among the four providers we could compare on that model, and a figure that differs from Groq's own published number rather than matching it.

The chart below sets Groq against the other three measured providers.

The gap holds throughout the measured set: each step down roughly halves the prior tokens-per-second figure.

A free tier exists, rate-limited to roughly 10 to 30 requests per minute depending on model, per Groq's rate limits documentation. Top pick for output speed among the providers this comparison could directly measure. Wrong pick if a workload needs a model family Groq does not host, since its catalog is narrower than Together AI's or Fireworks AI's.

Cerebras claims over 2,000 tokens per second, but its price list is not public

Cerebras states on its own inference page that it delivers "over 2,000 tokens per second" for Llama 4 Scout, a vendor claim LLM Waves Research has not independently verified.

A per-model dollar price table exists on Cerebras's pricing page but did not render across three fetch attempts on 2026-08-26; a free trial with $5 in credits is confirmed, alongside a $50-per-month Code Pro tier and a $200-per-month Code Max tier with expanded rate limits.

Situational: strong on Cerebras's own speed claim for buyers willing to create an account and check current pricing directly, but this comparison cannot rank Cerebras on price against the 12 other providers here.

DeepInfra prices output tokens lowest among three providers hosting the same model

DeepInfra lists DeepSeek-V4-Pro-0813 at $1.30 per million input tokens and $2.60 per million output tokens on its pricing page, the lowest output price of the three providers in this comparison hosting a model under that name, against $3.48 at Fireworks AI and $3.96 at Together AI.

DeepInfra confirmed no free tier is offered, stating plainly that an account needs a card on file or prepayment before use. Artificial Analysis measured DeepInfra's Turbo, FP8 variant of Llama 3.3 70B Instruct at 17.0 tokens per second output and 2.40 seconds TTFT, the slowest result in the four-provider speed comparison in this article.

Best value on disclosed output price for this model class. Wrong pick if time to first token matters more than price, or if a workload needs to start without a payment method on file.

Fireworks AI publishes context windows; Together AI mostly does not

Fireworks AI prices DeepSeek-V4-Pro-0813 at $1.74 per million input tokens and $3.48 per million output tokens on its standard tier, with a Priority tier at $2.61 and $5.22 marketed for lower time to first token, though Fireworks does not publish a number for that claim. A $1 free credit is offered on account creation.

Fireworks AI's models page lists context windows across most of its catalog, including 1,048,576 tokens for DeepSeek-V4-Pro-0813 and Kimi K3, where Together AI's own pages left context windows unconfirmed for the same models.

Close second to DeepInfra on output price for this model, ahead of Together AI, and the more completely documented of the two on context window.

Together AI lists the broadest open-weight catalog of the providers compared

Together AI's pricing page lists more than 20 chat models spanning DeepSeek, Qwen, Kimi, GLM, GPT-OSS, Llama, Gemma and MiniMax families, the widest catalog by model count among the specialist inference providers in this comparison.

DeepSeek-V4-Pro-0813 prices at $1.32 per million input tokens and $3.96 output, the highest output price of the three providers hosting that model here. Context windows were not found for most models on the pricing or models pages checked, and no free tier or first-party speed claim was found on either page.

Situational: the widest model selection in this comparison, but the least documented on context window and the most expensive of three providers on the one model priced identically elsewhere. The chart below compares all three providers hosting this model.

The $1.36 output-price spread is far wider than the $0.02 input-price spread between the same two providers.

OpenRouter passes through provider pricing and adds a payment fee, not an inference fee

OpenRouter states on its FAQ page that it passes through underlying provider pricing "without any markup on inference pricing," charging instead a 5.5% fee (minimum $0.80) on Stripe credit purchases.

Its free models are rate-limited to 50 requests per day without purchased credits, rising to 1,000 requests per day once an account holds $10 or more in credits, per the same page.

OpenRouter's public rankings page, retrieved 2026-08-26, ranks models by token volume actually routed through its API and states explicitly that this "does not rank models by accuracy, reasoning ability, or benchmark performance."

Best for reducing provider lock-in, since switching the underlying model does not require switching endpoints. Wrong pick as the sole provider for a single high-volume model, where a direct account with the hosting provider avoids the payment fee entirely.

AkashML prices Llama 3.3 70B on a decentralized GPU marketplace

AkashML, the managed inference layer built on the Akash Network's decentralized GPU marketplace, prices Llama 3.3 70B at $0.20 per million input tokens and $0.52 output, and GPT-OSS 20B at $0.03 and $0.13, on its own pricing page, retrieved 2026-08-26. New accounts receive $100 in credits.

AkashML states OpenAI-compatible drop-in migration, changing only the API base URL. Akash's own blog post announcing the product cites a different Llama 3.3 70B output price, $0.40 per million against akashml.com's $0.52; akashml.com is the more current of the two Akash-owned sources.

Situational: relevant for the self-hosting segment specifically, where a marketplace-priced backend is the appeal, not a fit for buyers who need a single named data center.

Native model APIs price highest and carry the largest published context windows

Anthropic, OpenAI, and Google each publish pricing directly for their own flagship models rather than routing to open-weight catalogs.

Anthropic's pricing page lists Claude Sonnet 5 at $2.00 input and $10.00 output per million tokens, with a 1,048,576-token context window confirmed on its models overview page.

OpenAI's pricing page lists GPT-5.6 Sol at the same $2.00 and $10.00 rates, with a 1.05-million-token context window per its models page.

Google's Gemini API pricing page lists Gemini 2.5 Flash-Lite at $0.10 input and $0.40 output, its cheapest tier with a confirmed free allowance of up to 500 requests per day.

Mistral's pricing page lists Mistral Large 3 at $0.50 and $1.50, the lowest native-provider output price in this group, though a confirmed context window was not found on Mistral's pricing or models pages. All three of Anthropic, OpenAI and Google state that paid-tier API inputs are not used to train their models, each on its own privacy or terms page.

Top pick for a single flagship model at native quality with the largest disclosed context ceiling; wrong pick for high-volume workloads where an open-weight specialist prices the same model class lower.

Baseten and Amazon Bedrock: two providers this comparison could not fully rank

Baseten prices DeepSeek V4 Pro at $1.74 input and $3.48 output per million tokens on its pricing page, identical to Fireworks AI's standard tier for the same model, and separately bills dedicated GPU deployments by the minute, from $0.01052 per minute for a T4 up to $0.16633 for a B200. No speed claim was found on the page checked.

Amazon Bedrock's pricing page surfaced only legacy Claude 3.5 Sonnet pricing in the fetch performed on 2026-08-26; current-generation pricing for Claude Sonnet 5 was not found there.

A Bedrock model card confirmed the model is available with a 1,048,576-token context window and routing options split between geography-scoped and global cross-region endpoints. Artificial Analysis measured Bedrock's Llama 3.3 70B endpoint at 153.7 tokens per second and 1.13 seconds TTFT.

Insufficient data for Amazon Bedrock on price; ranked on context window and third-party speed only.

Accented and non-English prompts are not a comparison axis on any of the 13 pricing pages

None of the 13 providers in this comparison publish accented-speech or non-English text accuracy figures on their pricing or documentation pages.

The performance difference a non-English workload experiences is driven mainly by which model is selected, not which API serves it, since the provider layer does not retrain the underlying weights; the exception is quantization, covered below, which can degrade multilingual output disproportionately to English output on the same model.

A buyer targeting non-English deployment should check the model card's own multilingual benchmark, where one is published, rather than the hosting provider's marketing page.

Cold-start latency is a different number from the time-to-first-token figure providers publish

Published time-to-first-token figures, including the Artificial Analysis numbers in this article, measure a warm request path, not a serverless container's first request after idle time or an entirely uncached prompt prefix.

Fireworks AI and DeepInfra both sell a "Priority" tier explicitly to reduce this delay, though neither publishes a number quantifying the improvement; Together AI offers dedicated deployments for the same reason.

A workload with bursty, unpredictable traffic should treat a provider's advertised TTFT as a floor, not an expectation, unless that provider names its own cold-start figure separately.

A stated context window is a ceiling, not a quality guarantee at that ceiling

Context window figures in the table above, from 131,072 tokens at Groq to 1,048,576 tokens at Fireworks AI, Anthropic, OpenAI and Amazon Bedrock, describe the maximum a provider accepts, not how the model performs as that maximum fills.

None of the 13 providers publish a long-context degradation curve on their own pricing or documentation pages.

The chart below plots every confirmed ceiling.

Claude Haiku 4.5 is a documented exception to the million-token pattern among Anthropic's tiers, capped at 200,000 tokens per its models overview page.

Output does not repeat identically across runs, and no provider here documents how much it varies

None of the 13 providers publish run-to-run output consistency data for identical prompts at fixed sampling settings. A seed parameter, where a provider's API supports one, narrows variance but does not guarantee identical output across requests, since backend routing and quantized serving can both introduce variation a seed does not control for.

This matters more on a quantized endpoint than on a full-precision one, which is the subject of the next section.

Model names carrying "Turbo" or "FP8" mark quantized serving, and not every provider discloses it the same way

Together AI and DeepInfra both label some model listings with suffixes such as "Turbo" and "FP8," for example DeepInfra's Llama-3.3-70B-Instruct-Turbo, that indicate a quantized serving path rather than the full-precision weights.

Fireworks AI's and Groq's listings in this comparison did not carry equivalent suffixes for the same model families.

A provider offering both a standard and a quantized listing of the same model, at different prices, is disclosing a tradeoff explicitly; a provider offering only one unlabeled listing is not confirming which precision it serves.

Decentralized, marketplace-priced compute is a fourth pricing model alongside flat, tiered and pass-through

AkashML's backend, the Akash Network, sets compute prices through a reverse-auction marketplace among independent infrastructure providers, a structurally different mechanism from the flat per-token rates at DeepInfra or Fireworks AI, the tiered rate-limit model at OpenAI and Anthropic, or OpenRouter's pass-through.

Akash's own blog cites $1.33 per hour for an H100 GPU against $3.93 on AWS. Given AkashML's own inconsistent published figures, this comparison treats it as a distinct, still-maturing category rather than folding it into the specialist-provider group above.

What we did not measure

LLM Waves Research ran no benchmark of its own for this comparison. Output quality, factual accuracy and reasoning performance were not evaluated for any model on any provider.

Speed and time-to-first-token figures come from Artificial Analysis for four providers on one model, Llama 3.3 70B Instruct, and from a vendor marketing claim for Cerebras; the remaining nine providers in the table have no independently measured speed figure in this article. Uptime, incident history and production behavior under sustained load were not evaluated for any provider.

Prices shown are list prices as published on 2026-08-26 and do not reflect volume discounts, enterprise contracts or promotional credits an individual account may receive.

Which provider fits which reader

For developers

OpenAI-compatible endpoints matter most for a drop-in migration; OpenRouter and Fireworks AI both document OpenAI-compatible routes, per their own pages cited above. Wrong pick if the target model is not in either catalog.

For startups

DeepInfra's disclosed output prices are the lowest of the three providers this comparison could compare on one model, but its lack of a free tier makes Groq's or Google's free allowance a better first step before committing budget.

For enterprise

Anthropic and Amazon Bedrock both publish region-routing detail, geography-scoped and global cross-region endpoints on Bedrock specifically; confirm current Bedrock pricing directly, since this comparison could not.

For high volume

DeepInfra's $2.60 output price for DeepSeek-V4-Pro-0813 is the lowest disclosed figure for that model in this comparison. Wrong pick if time to first token matters more than marginal cost, given DeepInfra's 17.0 tokens per second result above.

For low latency

Groq's 0.91-second measured TTFT on Llama 3.3 70B Instruct is the fastest in the four-provider comparison this article could run.

For on-device

None of the 13 providers compared here run on-device; every one is a hosted API. This segment does not apply to this comparison and belongs to a separate article on local inference.

For self-hosting

AkashML's decentralized marketplace backend and Baseten's per-minute dedicated GPU billing are the two closest fits among hosted options to a self-managed deployment model, though neither is literally self-hosting.

For non-English

Insufficient data across all 13 providers, none of which publish non-English or accented performance figures on their pricing or documentation pages, per the gap section above.

Which LLM API provider is the most affordable?

DeepInfra publishes the lowest output price this comparison found for a model shared across multiple providers, DeepSeek-V4-Pro-0813, though "most affordable" depends on which model a workload actually needs and whether a free tier matters more than the lowest per-token rate. See the comparison table above for the full price list by provider.

Which LLM API provider is the fastest?

Among the four providers this comparison could measure on the same model with the same independent source, Groq measured fastest on Llama 3.3 70B Instruct, per Artificial Analysis.

Cerebras states a higher figure for a different model on its own site, not independently verified here. See the comparison table above for every figure.

Is there an API for LLM?

Yes. Every provider in this comparison, from native model companies like Anthropic and OpenAI to specialist hosts like Groq and DeepInfra, exposes an HTTP API for sending prompts to a language model and receiving generated text back, most of them compatible with or adjacent to the OpenAI API format.

What is the most cost-effective LLM API?

Cost-effectiveness depends on the metric: DeepInfra's output price led this comparison's one shared-model test, while Groq's speed and Google's and OpenRouter's free-tier allowances change the calculation for a low-volume or early-stage workload. The use-case verdicts above break this down by reader.

Which LLM API provider is best for coding?

This comparison did not test coding accuracy for any provider. Fireworks AI, Together AI and DeepInfra all host coding-oriented open-weight models such as Kimi K2.7 Code, while OpenAI publishes a dedicated GPT-5.3 Codex tier; provider choice mainly determines price and speed for a coding workload, not the model's coding ability itself.

Is there a free LLM API provider with no credit card required?

Groq, Google and OpenRouter all confirmed free-tier access on the pages checked for this comparison, each with its own rate limits. DeepInfra confirmed the opposite, requiring a card or prepayment before any use. Check each provider's current terms directly, since free-tier availability changes without a changelog.

Prices change. This snapshot is from one afternoon.

Every dollar figure in this comparison is a snapshot from 2026-08-26, not a standing fact. Pricing pages change without notice, and the providers this comparison could not fully price, Cerebras and Amazon Bedrock among them, may publish the missing figures before this article's next refresh.