Groq API and OpenAI API answer different questions. On output-token price, Groq's gpt-oss-20b lists at $0.30 per 1M tokens against $0.60 for OpenAI's gpt-5.6-luna, aggregated 2026-08-21. Groq wins on price and speed for open-weight chat; OpenAI wins on frontier reasoning. Stop here if that settles it.

Model

Provider

Input $/1M

Output $/1M

Context window

Reported speed†

gpt-oss-20b

Groq

$0.075

$0.30

131,072

1,000 tok/s

gpt-oss-120b

Groq

$0.15

$0.60

131,072

500 tok/s

qwen3.6-27b (preview)

Groq

$0.60

$3.00

131,072

500 tok/s

gpt-5.6-luna

OpenAI

$0.10

$0.60

short context

not published

gpt-5.6-terra

OpenAI

$1.00

$6.00

short context

not published

gpt-5.6-sol

OpenAI

$2.00

$10.00

short context

not published

Table note: † vendor-reported speed, not independently tested by LLM Waves Research. OpenAI's long-context input price runs exactly double the short-context rate shown for all three models; its long-context output price runs at 1.5 times the short-context rate. Groq's price stays flat regardless of context length used.

Groq's gpt-oss-20b prices output tokens at $0.30 per 1M against $0.60 for OpenAI'sgpt-5.6-luna, its cheapest current model, both aggregated 2026-08-21.

Groq's gpt-oss-120b input price sits at $0.15 per 1M, 25% below OpenAI's $0.20 long-context rate for gpt-5.6-luna. OpenAI's gpt-5.6-sol, the top of its lineup, has no Groq-hosted equivalent at any price.

Basis and method

Basis 2, aggregated data, governs every figure in this comparison. Every price and rate limit below comes from Groq's own documentation, retrieved 2026-08-21, or from OpenAI's own pricing and rate-limit pages, retrieved the same day. LLM Waves Research ran no prompts against either API for this article.

The daily-request-ceiling figures later in this article are original arithmetic performed on the vendors' published rate limits, with the formula shown in place. The full source list and version history sit on the methodology page.

Four axes rank Groq API against OpenAI API below: price per million tokens, rate-limit structure, model catalog freshness, and reported inference speed.

General chat models

Groq's four general chat models cluster below $1 per million tokens on both sides. OpenAI's three current models, priced on its pricing page retrieved 2026-08-21, range from $0.60 to $15.00 per million output tokens. gpt-oss-safeguard-20b, a moderation-focused preview model, matches the $0.075 input and $0.30 output pricing of gpt-oss-20b exactly.

Neither provider's cheapest model is its fastest. On Groq, gpt-oss-20b runs at a vendor-reported 1,000 tok/s against 500 tok/s for the pricier gpt-oss-120b.

Audio and moderation models

Groq prices Whisper-family transcription by the hour rather than by the token: whisper-large-v3 lists at $0.111 per hour and whisper-large-v3-turbo at $0.04 per hour, both retrieved 2026-08-21. Two preview voice models price by the character instead: canopylabs/orpheus-v1-english at $22 per 1M characters and canopylabs/orpheus-arabic-saudi at $40 per 1M characters.

Groq's supported-models documentation lists no embedding model, so a request for text-embedding functionality needs a different provider entirely.

Groq retired its two most searched models five days before this was published

Groq's deprecation log, retrieved 2026-08-21, lists llama-3.3-70b-versatile and llama-3.1-8b-instant as shut down on 2026-08-16. The replacement for llama-3.3-70b-versatile is openai/gpt-oss-120b or qwen/qwen3.6-27b; the replacement for llama-3.1-8b-instant is openai/gpt-oss-20b.

Any script, agent framework, or saved prompt template still calling either retired model ID against Groq's API started returning errors on 2026-08-16, five days before the data in this article was collected.

The free tier's advertised daily cap is not the real ceiling

Groq's rate-limit table, retrieved 2026-08-21, lists the free-tier limits for gpt-oss-120b as 30 requests per minute, 1,000 requests per day, 8,000 tokens per minute, and 200,000 tokens per day.

A 1,000-token request, a typical short chat completion, hits the 200,000 token-per-day cap at 200 requests, one fifth of the advertised 1,000-request figure. A 5,000-token request, closer to a long RAG prompt, hits it at 40 requests, one twenty-fifth of that figure.

Groq's own documentation notes that "higher limits are available for select workloads and enterprise use cases," and the published table lists identical numeric limits under both the Free Plan and Developer Plan columns for the models checked.

LLM Waves Research could not independently confirm whether every account sees this parity, since Groq directs users to their own account dashboard for the current exact figure.

Rate limits: fixed quotas against spend-gated tiers

Groq governs access with a fixed per-model quota, published in the table above, that does not change when a card is added to the account. OpenAI governs access with usage tiers instead: its rate-limit guide, retrieved 2026-08-21, states that Tier 1 access requires $5 paid into the account, and both the free tier and Tier 1 share a $100 monthly usage cap.

OpenAI's free tier additionally requires the account to be in an allowed geography, a gate Groq's documentation does not mention for any tier.

Groq is fast, but Cerebras is faster on the identical model

Third-party benchmarking answers the speed question Groq's own marketing does not. Artificial Analysis, retrieved 2026-08-21, recorded 474 output tokens per second for Groq's gpt-oss-120b endpoint, against 1,641 for Cerebras hosting the identical open-weight model, 702 for SambaNova, and 48 to 50 for the commodity GPU clouds CoreWeave and DeepInfra.

Groq's time to first token, 0.74 seconds, trails Cerebras at 0.46 seconds and beats SambaNova's 1.09 seconds.

CoreWeave's blended price, $0.04 per 1M tokens against Groq's $0.06, makes it the cheapest host for this model and the slowest of the five measured. A workload that tolerates 50 tokens per second finds a cheaper home than Groq; one that does not finds a faster one on Cerebras. This article's ranked comparison of LLM API providers covers that fuller field.

Getting a key and pointing existing OpenAI code at Groq

Groq's quickstart documentation, retrieved 2026-08-21, directs new accounts to console.groq.com/keys to generate an API key, with no separate approval step documented.

Groq's OpenAI-compatibility page lists the base URL as https://api.groq.com/openai/v1; swapping base_url and api_key in an existing OpenAI client library covers most chat completion requests.

Four parameter groups are not supported: logprobs, logit_bias, and top_logprobs; the messages[].name field; an n value other than 1; and the vtt and srt transcription output formats.

What this comparison did not measure

LLM Waves Research did not run prompts against either API for this article. This comparison did not test output quality on any task, did not measure latency from a specific region, and did not verify whether the Free Plan and Developer Plan rate limits diverge for any account beyond what the published table shows.

Which provider for which reader

Top pick for cheap, fast open-weight chat: Groq. gpt-oss-20b runs at a vendor-reported 1,000 tok/s and $0.30 per 1M output tokens, aggregated 2026-08-21. If the task needs frontier-level reasoning rather than a fast open-weight model, this pick is wrong; use OpenAI's gpt-5.6-sol instead.

Top pick for frontier reasoning: OpenAI. gpt-5.6-sol has no Groq-hosted equivalent in the current catalog. If cost per token matters more than reasoning depth, this pick is wrong; use Groq's qwen3.6-27b instead.

Situational for non-English voice: Groq. The Arabic-Saudi and English Orpheus models price at $40 and $22 per 1M characters, aggregated 2026-08-21, and OpenAI's pricing page lists no direct equivalent. Both sit in Groq's preview tier, not production; if a stable production audio endpoint matters more than price, this pick is wrong until they graduate.

Not recommended for either: on-device or self-hosted deployment. Both are hosted APIs; the open-weight gpt-oss models Groq serves are downloadable directly from their authors, making self-hosting a third option outside this comparison.

Is the Groq API free?

Groq's rate-limit table lists a Free Plan and a Developer Plan with identical published RPM, RPD, TPM, and TPD figures for the models checked, retrieved 2026-08-21.

Groq's own text notes that higher limits exist for select workloads and enterprise use, so a specific account may see more than the public table shows. The account dashboard's limits page carries the current exact figure.

What models does the Groq API support?

As of 2026-08-21, Groq's production catalog lists openai/gpt-oss-120b, openai/gpt-oss-20b, whisper-large-v3, and whisper-large-v3-turbo, plus groq/compound and groq/compound-mini as production systems. Five preview models sit alongside them, including qwen/qwen3.6-27b.

Meta's Llama models no longer appear in production; the two most-searched, llama-3.3-70b-versatile and llama-3.1-8b-instant, were both retired on 2026-08-16.

Does the Groq API support embeddings?

No. Groq's supported-models documentation, retrieved 2026-08-21, lists chat, audio transcription, and moderation models but no embedding model.

What is the Groq API base URL?

The base URL is https://api.groq.com/openai/v1, documented on Groq's OpenAI-compatibility page, retrieved 2026-08-21. Pointing an existing OpenAI client at this URL with a Groq API key covers most chat completion requests.

How do you get a Groq API key?

Groq's quickstart documentation, retrieved 2026-08-21, directs new accounts to the API Keys page at console.groq.com/keys to generate one, with no separate approval step documented beyond account creation.

Is the Groq API OpenAI compatible?

Mostly. Groq's compatibility page, retrieved 2026-08-21, confirms most OpenAI client parameters work after swapping base_url and api_key, though logprobs, logit_bias, top_logprobs, the messages[].name field, and the vtt and srt audio formats are not supported.

This comparison is already aging

Published 2026-08-21, this article reflects two vendors' pricing and rate-limit pages as they stood that day, and Groq changed its production model catalog five days earlier. Check console.groq.com/docs/models and OpenAI's pricing page directly before committing a production budget to either one. LLM Waves Research will re-check every figure here within 90 days.

Changelog: 2026-08-21: Published. Groq API and OpenAI API compared on price, rate limits, model catalog, and reported speed.

Published: 2026-08-21 Last updated: 2026-08-21 Data collected: 2026-08-21 (UTC)