Cerebras and Groq both sell inference for open-weight language models on custom silicon built to outrun general-purpose GPUs.
On gpt-oss-120b, the one model both platforms host on comparable terms, Cerebras posted 819 output tokens per second against Groq's 310, per OpenRouter's provider table retrieved 2026-08-25.
Groq priced input tokens 57% lower on the same model. Neither company leads on every axis measured here: match the row below to the workload, then stop reading.
Dimension | Cerebras | Groq | Source basis |
|---|---|---|---|
Output speed, gpt-oss-120b, OpenRouter | 819 tok/s | 310 tok/s | Reported † |
Output speed, gpt-oss-120b, Artificial Analysis | 1,714.6 tok/s | 476.7 tok/s | Reported † |
Time to first token, gpt-oss-120b, Artificial Analysis | 1.65 s | 4.93 s | Reported |
Price per 1M tokens, gpt-oss-120b (input / output) | $0.35 / $0.75 | $0.15 / $0.60 | Reported, vendor docs |
Context window, gpt-oss-120b (free tier / paid tier) | 65,000 / 131,072 tokens | 131,072 / 131,072 tokens | Reported, vendor docs |
Models with published pricing, Artificial Analysis | 4 | 11 | Reported |
Uptime, gpt-oss-120b, trailing window | 99.99% | 99.97% | Reported, OpenRouter |
Public company status | Nasdaq: CBRS, listed 2026-05-14 | Privately held | Reported |
Ties to Nvidia | None disclosed | Technology licensed, two executives hired, 2025-12-24 | Reported |
† Two independent providers measured the same model on the same day and disagree by roughly 2x. See "What we did not measure."
Basis and method
This comparison rests on Basis 2: figures collected from vendor documentation, Artificial Analysis, and OpenRouter, not from an inference run LLM Waves Research executed itself. Every figure carries its source and retrieval date, 2026-08-25 unless stated otherwise.
A fully controlled head-to-head was possible only on gpt-oss-120b, the one model both platforms publish under comparable defaults; other models carry wider sourcing caveats stated inline. Full methodology, including how the OpenRouter versus Artificial Analysis discrepancy was handled, is published on the LLM Waves methodology page.
Decision axis: this comparison ranks both providers on output speed, price per token, context window per pricing tier, and model catalog breadth, and treats company ownership as a fifth, non-technical axis that matters for procurement.
Output speed: Cerebras leads, but the size of the lead depends on who measured it
OpenRouter's live provider table for gpt-oss-120b, retrieved 2026-08-25, recorded Cerebras at 819 output tokens per second against Groq's 310, a 2.6x gap. Artificial Analysis's provider page for the same model, retrieved the same day, recorded Cerebras at 1,714.6 tokens per second against Groq's 476.7, a 3.6x gap.
Both services rank Cerebras first; neither figure should be read to one decimal place in a purchasing decision, since two independent benchmarks of an identical model on the same day disagree by more than 2x on Cerebras's own number.
The gap traces to a design choice, not a general capability edge. Cerebras's own comparison of the CS-3 to Groq's LPU describes 21 or more petabytes of aggregate on-chip SRAM against 230 megabytes per Groq LPU chip.
That figure is a vendor claim about Cerebras's own hardware and a named competitor, not an independent measurement, and is labeled that way here rather than repeated as fact.

Price per token: Groq is cheaper on every model both platforms publish
Groq lists gpt-oss-120b at $0.15 per million input tokens and $0.60 per million output tokens in its own documentation, retrieved 2026-08-25. Cerebras lists the same model at $0.35 per million input tokens and $0.75 per million output tokens in its own documentation, retrieved the same day. Groq's input price is 57% lower and its output price is 20% lower. OpenRouter's independent listing matches both vendors' figures exactly.
The pattern holds below the flagship model: Cerebras prices Llama 3.1 8B at $0.10 per million tokens, input and output alike, while Artificial Analysis lists Groq's Llama 3.1 8B at roughly $0.05 blended, half Cerebras's rate. On every model both providers publish, Groq costs less per token.

Context window: Cerebras's free tier is the tightest limit in this comparison
Context window is the dimension no competing article in the current search results measures. Cerebras caps gpt-oss-120b at 65,000 tokens on its free tier, rising to 131,072 tokens on paid tiers. For Llama 3.1 8B, the cap is tighter: Cerebras's own model page lists 8,000 tokens on the free tier and 32,000 tokens on paid tiers, against the model's native 128,000-token design.
Groq's documentation lists a flat 131,072-token window for gpt-oss-120b, Llama 3.1 8B, and Llama 3.3 70B, regardless of tier.
This is a mechanism, not an adjective: the wafer-scale SRAM design behind Cerebras's speed also narrows its context windows, since holding an entire model on-chip leaves less room for a long key-value cache than a memory-attached GPU affords. A long-document workload will not run at the native length of Llama 3.1 8B on Cerebras's free tier, or past a quarter of that length on a paid Cerebras plan.

Model catalog: Groq publishes roughly three times as many priced models
Artificial Analysis lists 4 priced models on Cerebras, two reasoning-effort variants of gpt-oss-120b and two of Gemma 4 31B. The same service lists 11 priced models on Groq, spanning Qwen3.6 27B, Qwen3 32B, gpt-oss-120b, gpt-oss-20b, Llama 4 Scout, Llama 3.3 70B, and Llama 3.1 8B, retrieved 2026-08-25.
A team standardizing on one vendor for more than a single model family has a narrower Cerebras catalog to work with.
What we did not measure
This comparison ran no inference requests of its own; every figure above is Reported or a labeled vendor claim, not measured by LLM Waves Research. It does not cover non-English prompts, streaming versus batch delivery, output consistency across repeated runs, or behavior under sustained concurrent load, since no source here published comparable figures for both platforms as of 2026-08-25.
It does not resolve why OpenRouter and Artificial Analysis disagree by roughly 2x on Cerebras's own throughput, beyond noting the two services use different prompt sets and sampling windows. A reader needing a number precise to the token should run the workload on both platforms directly.
Ownership, funding, and what changed since December 2025
Two events moved since most existing comparisons of these companies were published, and neither is optional context for a 2026 purchasing decision.
On 2025-12-24, Nvidia and Groq announced a non-exclusive inference technology licensing agreement. Groq founder Jonathan Ross and president Sunny Madra joined Nvidia under the deal, while Groq's own announcement states GroqCloud continued operating as an independent service. Reporting on the deal's value, including CNBC and Tom's Hardware, placed it near $20 billion.
Groq's next funding round priced the company well below what that deal implied. Crowdfund Insider reported that Groq raised $350 million at a $3.5 billion valuation in a round completed 2026-08-17, down from $6.9 billion in September 2025. Cerebras moved the opposite direction: it went public on Nasdaq as CBRS on 2026-05-14, priced its IPO at $185 per share, and closed its first day at $311.07, up 68.2%.
Cerebras is now a publicly traded, audited company; Groq is privately held, with its founding technical leadership now on Nvidia's payroll.

Use case verdicts
Raw single-model throughput: top pick, Cerebras. Both services agree Cerebras leads Groq by more than 2x on gpt-oss-120b.
Wrong for a workload that also needs Cerebras's free-tier context window, capped well below the paid-tier maximum.
Lowest cost per token: best value, Groq. Groq priced every model in this comparison lower than Cerebras's equivalent.
Wrong for a team already built around Cerebras's speed that cannot absorb Groq's roughly 2x to 3x lower throughput.
Wide model catalog on one vendor: top pick, Groq. Eleven priced models against Cerebras's four means fewer integrations for a team rotating between model families.
Wrong for a team standardizing on the single model Cerebras serves fastest.
Long-context workloads on a free or low tier: not recommended, Cerebras. An 8,000-token free-tier cap on Llama 3.1 8B, one-sixteenth the model's native window, disqualifies Cerebras's free tier for document-length prompts.
No Nvidia ownership stake or licensing tie: top pick, Cerebras. Cerebras discloses no Nvidia relationship; Groq licensed technology to Nvidia and lost its founder and president to Nvidia's payroll on 2025-12-24.
Situational: a buyer weighing only the API contract can treat this axis as irrelevant.
On-device and self-hosted deployment: insufficient data. Neither company publishes a self-hosted or on-device path for these models; both are cloud-API-only, so this segment carries no rank.
Is Groq faster than Cerebras?
No. On gpt-oss-120b, the model both platforms host under comparable terms, Cerebras posted higher output speed on both benchmarks checked: 819 against 310 tokens per second on OpenRouter, and 1,714.6 against 476.7 on Artificial Analysis, both retrieved 2026-08-25.
Cerebras also recorded a shorter time to first token, 1.65 seconds against 4.93. The gap size differs between the two services, but neither ranks Groq ahead of Cerebras on this model.
What is Cerebras Systems, and is it publicly traded?
Cerebras Systems builds the wafer-scale engine chip and CS-series systems behind Cerebras Inference.
It went public on 2026-05-14, listing on Nasdaq under the ticker CBRS at an IPO price of $185 per share and closing its first day at $311.07, a 68.2% gain, per reporting on the debut. Stock price analysis sits outside what this comparison covers.
Who competes with Groq, and does Nvidia now own it?
Nvidia does not own Groq. On 2025-12-24 the two companies entered a non-exclusive technology licensing agreement under which Groq's founder and president joined Nvidia, while GroqCloud, per Groq's own announcement, kept operating as an independent service.
Cerebras, SambaNova, and Nvidia's own GPU-based inference compete with GroqCloud on speed and price; OpenRouter lists SambaNova at $0.14 per million input tokens and 360 tokens per second on gpt-oss-120b.
Is Cerebras or Groq cheaper to run in production?
Groq, on every model this comparison could price on both platforms. Its input price for gpt-oss-120b, $0.15 per million tokens, is 57% below Cerebras's $0.35, and its output price, $0.60, is 20% below Cerebras's $0.75, both from vendor documentation and matched by OpenRouter, retrieved 2026-08-25.
A workload needing Cerebras's throughput may still cost less overall, but on listed price, Groq is cheaper.
Which models can you actually run on each platform?
Fewer on Cerebras than on Groq. Artificial Analysis lists 4 priced models on Cerebras against 11 on Groq, spanning Meta's Llama family, OpenAI's gpt-oss family, and Alibaba's Qwen3 and Qwen3.6 families, retrieved 2026-08-25.
Confirm a target model is hosted on both platforms before comparing speed or price at all; several third-party figures for Llama 3.3 70B could not be independently corroborated closely enough to report here with confidence.
The number that will move next
Every figure here is a snapshot of two companies mid-transition: one newly public, one newly reset in valuation after its own technology moved to the industry's largest GPU vendor. Check the live leaderboard for figures newer than this retrieval date, and the Groq API provider page and Cerebras provider page for model-by-model detail this comparison had no room for.
