Groq's pricing page sells Llama 3.1 8B Instant at $0.05/$0.08 and Llama 3.3 70B Versatile at $0.59/$0.79 per million tokens, 36 days after Groq's own deprecation schedule retired both models on August 16, 2026.

Three more pricing rows carry the same contradiction. The table below has the rate for all ten models tracked.

Model

Sale status, Sept 21, 2026

Groq's own deprecation entry

Input $/1M

Output $/1M

Qwen/Qwen3.8-27B

Current

Not listed

$0.80

$4.00

Kimi K2 Instruct

Current

Not listed

$1.00

$3.00

Qwen 3.6 27B

Priced as current on three trackers

Retired 09/14/26 (7 days ago)

$0.60

$3.00

Llama 3.3 70B Versatile

Sold as self-serve on Groq's own page

Retired 08/16/26 (36 days ago)

$0.59

$0.79

Qwen3 32B

Priced as current on one tracker

Retired 07/17/26 (66 days ago)

$0.29

$0.59

GPT-OSS-120B

Current

Not listed

$0.15

$0.60

GPT-OSS-20B

Current

Not listed

$0.075

$0.30

Llama 3.1 8B Instant

Sold as self-serve on Groq's own page

Retired 08/16/26 (36 days ago)

$0.05

$0.08

MiniMax M2.7

Preview, contact sales

Not listed

†

†

DeepSeek R1 Distill Llama 70B

Removed from Groq; still priced live on one tracker (Sept 20, 2026)

Retired 10/02/25 (354 days ago)

n/r

n/r

Table note: † means Groq lists this model as contact sales rather than a self-serve rate. n/r means Groq no longer runs this model at all; the tracker figure comes from a third party, not from Groq. Rows are sorted by output price, highest first.

Four numbers carry the weight of that table. Qwen/Qwen3.8-27B costs $4.00 per million output tokens, the most expensive current row.

Llama 3.1 8B Instant costs $0.08, the cheapest row still selling on Groq's own page despite its retirement date.

GPT-OSS-20B costs $0.075 and $0.30. Groq's documentation never flags it as retired.

DeepSeek R1 Distill Llama 70B has carried no live Groq price for 354 days, yet a rate for it is still on sale elsewhere.

Basis and method

Measurement Basis 2 governs this article. Every rate below comes from Groq's own documentation, never a benchmark LLM Waves Research ran. We read console.groq.com's deprecations, models, rate limits, service tiers, batch, and compound tooling pages directly on September 21, 2026.

Groq's documentation states a model's price and its retirement date. It states nothing about whether a retired model ID still answers a live API call, and that question stays open below. This piece fills the Groq row of the LLM API pricing hub.

Full sourcing detail sits on the methodology page.

Two things determine the ranking here: the per-million-token rate Groq billed on September 21, 2026, and whether that rate matches Groq's own deprecation schedule for the same model ID.

No performance test accompanies this article: this article did not run a latency test, a throughput test, or an output-quality benchmark against any model discussed. Groq's documentation also doesn't say whether a request to a retired model ID still succeeds or gets rerouted, and this article doesn't close that gap.

Two of the four most searched Groq models are already retired

Llama 3.1 8B Instant and Llama 3.3 70B Versatile anchor two of the highest volume searches. Groq's deprecations page dates both retirements to August 16, 2026.

Neither model has been pulled from Groq's own pricing page; both still carry a live rate and no visible warning. Some of that same search volume targets models Groq never priced this way at all.

Llama 3.1 70B is not a Groq model ID; the current 70B class release carries the name Llama 3.3 70B Versatile above.

Llama 3.1 405B appears in neither the deprecations page nor the current model list checked for this piece.

The replacement chain gets worse on inspection. Groq's own entry for Llama 3.3 70B Versatile names two replacements: openai/gpt-oss-120b or qwen/qwen3.6-27b.

The second of those two options was itself retired on September 14, 2026. Groq's entry for DeepSeek R1 Distill Llama 70B recommends llama-3.3-70b-versatile or openai/gpt-oss-120b as a replacement; the first of those two is the same model this section opened with.

Following Groq's own migration advice can land a reader on a model Groq's own schedule has already retired.

Three more models are gone from Groq, and one price tracker has not noticed

Llama 4 Scout and Llama 4 Maverick, both full model IDs meta-llama/llama-4-scout-17b-16e-instruct and meta-llama/llama-4-maverick-17b-128e-instruct, carry retirement dates of July 17, 2026 and March 9, 2026 in Groq's own schedule.

DeepSeek R1 Distill Llama 70B was retired October 2, 2025, 354 days before this article, the oldest removal Groq's current documentation still lists.

AI Pricing Guru's Groq page, last synced September 20, 2026, one day before this article, still prices DeepSeek R1 Distill Llama 70B at $0.75 input and $0.99 output per million tokens.

That page updated one day before this audit. It still carried a rate for a model Groq removed nearly a year earlier.

A search for either Llama 4 model returns pages built around a model ID Groq's own infrastructure will reject.

The Qwen models Groq hosts climbed 6.8 times in output price across two generations

Qwen3 32B, retired July 17, 2026, billed $0.29 input and $0.59 output. Its successor, Qwen 3.6 27B, retired September 14, 2026 after seven weeks live, billed $0.60 and $3.00, a fivefold jump in output price on a model with fewer parameters.

The current replacement, Qwen/Qwen3.8-27B, bills $0.80 and $4.00, another 33 percent above the output rate for Qwen 3.6 27B and 6.8 times the rate for Qwen3 32B.

Seven weeks, start to end. Three Qwen generations, three price tiers.

Nothing in Groq's own documentation ties that price climb to a stated capability gain. The context window actually narrowed across the chain: Qwen/Qwen3.8-27B lists 131,042 tokens, smaller than the footprint Qwen3 32B carried before it left the lineup.

A reader pricing a Qwen model from an old bookmark can be planning against a number two generations and 6.8 times the output cost out of date.

Batch and prompt caching discounts do not stack, whatever one tracker says

Groq's batch documentation states the rule without qualification: all batch tokens bill at the 50% batch rate regardless of cache status.

Groq's prompt caching documentation confirms the same policy from the other side: the caching discount does not stack with the batch discount.

Both cut a rate in half alone. Run together, Groq's own pages say the bill still only halves once.

CloudZero's Groq guide, updated September 4, 2026, describes the opposite: batch and caching "can be stacked for an effective rate of roughly 25% of on demand pricing." That contradicts Groq's prompt caching documentation as directly as it contradicts the batch page, not a rounding difference.

A workload builder who priced a job at a quarter of the on demand rate, on that tracker's word, will see double the bill Groq's own pages describe.

Free tier limits sit per organization, and three current looking models are missing from the table entirely

Groq's rate limit documentation states limits apply at the organization level, not per user or per API key; running five keys against one account does not multiply the ceiling.

Live throughput and uptime for Groq's hosted models track separately on llmwaves.com's Groq inference page, outside the scope of this pricing. The current free tier table lists GPT-OSS-120B, GPT-OSS-20B, and Qwen/Qwen3.8-27B at 30 requests per minute, 1,000 per day, 8,000 tokens per minute, and 200,000 per day.

Groq's two Whisper models get 20 requests per minute and 2,000 per day. Llama 3.1 8B Instant, Llama 3.3 70B Versatile, and Kimi K2 Instruct appear nowhere in that table.

Groq's service tiers add a second axis on top: On Demand is the default, Flex trades reliability for throughput, Auto picks between them, and Performance is reserved for enterprise accounts.

None changes the per token rate. Nvidia's roughly $20 billion licensing deal with Groq, announced December 2025, left Groq an independent company after founder Jonathan Ross moved to Nvidia; Groq's cloud pricing has not moved because of it.

Tool fees on a single agentic workflow can outweigh its own token cost 9 to 1

Groq's compound tooling documentation prices four server-side tools separately from any model's per token rate: Basic Web Search at $5 per 1,000 requests, Advanced Web Search at $8, Visit Website at $1, and code execution at $0.18 per hour.

A fifth tool, Wolfram Alpha, bills through Wolfram's own key rather than Groq's invoice.

Here is the arithmetic, formula published: a workflow running 500 Advanced Web Search calls, 500 Visit Website calls, and two hours of code execution costs (500 x $0.008) + (500 x $0.001) + (2 x $0.18), or $4.86 in tool fees alone.

Add 2,000,000 input and 500,000 output tokens at the published rate for GPT-OSS-120B, (2 x $0.15) + (0.5 x $0.60), and the token line comes to $0.60. Total bill: $5.46.

A buyer who priced only the tokens undercounted the real bill by 9.1 times. Nine dollars in fees. One dollar in tokens, give or take. Groq API vs OpenAI API runs the same token math against a second vendor, without this tool fee layer.

Which Groq setup fits which reader

For developers

Best value: GPT-OSS-20B at $0.075 and $0.30, the cheapest model Groq's documentation does not flag as retired anywhere.

For startups

Situational: GPT-OSS-120B at $0.15 and $0.60 for anything needing more reasoning depth than the 20B tier, priced third from the bottom of the current lineup.

For enterprise

Situational: MiniMax M2.7 and the Performance service tier both require contacting Groq's sales team; neither publishes a self-serve number this article can carry.

For high volume

Best value: the batch discount, 50% off any current model's synchronous rate, run against GPT-OSS-20B for the lowest combined floor; it will not stack with cache savings.

For low latency

Insufficient data. This article priced tokens, not milliseconds, and did not measure Groq's LPU throughput against any other provider this session.

For self-hosting

Situational: on-premises Groq hardware is a contact sales conversation in every source checked for this piece; no self-serve rate exists to quote.

For on-device

Insufficient data. Groq publishes no on-device runtime or offline pricing this article could find.

For non-English

Situational: Orpheus Arabic Saudi, a preview text to speech model, bills $40 per million characters against $22 for Orpheus V1 English, the one language-specific price split in Groq's current lineup.

How much does Groq cost per million tokens?

Groq's current self-serve range runs from $0.05 input on Llama 3.1 8B Instant to $4.00 output on Qwen/Qwen3.8-27B, both confirmed live September 21, 2026. The table above carries all ten rates this article tracked.

Does Groq's free tier require a credit card?

Groq's own documentation describes free tier access without a card requirement stated anywhere in the rate limit or model pages checked for this piece.

Limits apply per organization, not per key, and several current-looking models carry no published free tier row at all.

Do Groq's batch discount and prompt caching discount combine?

No. Both of Groq's own pages agree: batch tokens bill at the 50% rate regardless of cache status, and caching does not stack on top of it.

One named tracker's claim that the two stack to a 25% effective rate matches neither page.

What does Kimi K2 cost on Groq?

Kimi K2 Instruct bills $1.00 input and $3.00 output per million tokens, a current row confirmed against three independent trackers and absent from Groq's own deprecations page. It does not appear in Groq's free tier rate limit table at all.

Is Llama 3.1 8B Instant still active on Groq, or was it deprecated?

Both. Groq's own deprecations page retired the model on August 16, 2026. Groq's own pricing page still bills it at $0.05 and $0.08 per million tokens, 36 days after that date, with no visible notice connecting the two pages.

Groq's own two pages do not agree with each other

One of them will have to change first.