xAI prices Grok 4.6 at $2 and $6 per million input and output tokens through 200,000 tokens, then $4 and $12 above it.

Anthropic prices Claude Opus 5 at a flat $5 and $25 across its full 1M window, both figures collected September 10, 2026.

Below 200,000 tokens, Grok is cheaper on both axes by 2.5x and 4.2x. Above it, the gap shrinks to 1.25x and 2.1x.

Model

Context window

Knowledge cutoff

Input $/M (≤200K)

Output $/M (≤200K)

Input $/M (>200K)

Output $/M (>200K)

Grok 4.6 (grok-4.6)

500K tokens

February 2026

$2.00

$6.00

$4.00

$12.00

Claude Opus 5 (claude-opus-5)

1M tokens

May 2026

$5.00‡

$25.00‡

$5.00‡

$25.00‡

Table note: bold marks the cheaper rate in each column. ‡ Claude Opus 5 carries no separate long-context tier; the same rate applies across the full 1,000,000-token window.

Basis and method

This is a Basis 2, aggregated comparison. Figures come from xAI's developer pricing documentation, xAI's Grok 4.6 model page, Anthropic's Claude Opus 5 announcement, and Anthropic's platform pricing and prompt-caching pages, cross-checked against two independent trackers, all collected September 10, 2026. No figure below was produced by testing.

The original contribution is the arithmetic connecting Grok's context-tier pricing to the exact point where its price advantage shrinks, a calculation not published elsewhere in this form. Every model name below carries its release date and configuration where one applies, per the locked naming table on the methodology page. Full sourcing, with URL and retrieval date per figure, sits on that same page.

This comparison ranks on three axes: cost per token across both context tiers, whether the two vendors' published coding benchmarks are actually comparable, and how far a weekly usage pool stretches under an agent workload.

xAI's price schedule turns on a single number, and it decides more of the total bill than either vendor's headline rate does.

Grok stays cheaper than Claude on every one of these four numbers, but the size of the saving depends entirely on which side of 200,000 tokens a request lands on, and that is the fact a flat headline ratio hides.

Grok 4.6: the sticker price hides a 200,000-token cliff

Grok 4.6 is xAI's newest model tier, released August 12, 2026, with a 500,000-token context window and a February 2026 knowledge cutoff.

Below 200,000 prompt tokens, xAI's developer pricing page lists Grok 4.6 at $2 per million input tokens and $6 per million output tokens.

Cross that threshold within a single request, and the same page lists $4 and $12, exactly double on both axes.

Claude Opus 5 carries no equivalent tier. A 900,000-token Opus 5 request bills at the same $5 and $25 rate as a 9,000-token one.

An agent session that starts under 200,000 tokens and grows past it mid-conversation moves into Grok's higher tier without any flag in the response payload.

What the sticker price doesn't include

Real-time access to X posts and web search is the advantage Grok is built around, and xAI meters it apart from the token price.

A single web search costs half a cent on top of whatever the underlying tokens already cost, and an agent that searches on every turn adds this line item on every turn too.

Pulling an X user's profile costs $10 per 1,000 calls, file attachments and code execution cost $10 and $5 per 1,000 calls, and collections search costs $2.50 per 1,000 calls.

xAI also charges $0.05 for each request its own systems block on usage-guideline grounds, plus storage fees of $0.025 per GiB per day for uploaded files, $0.10 per GiB per day for collections, and $0.20 per GiB on egress.

Cached input pricing moved in an unexpected direction between xAI's last two flagship releases.

Grok 4.6's cached-input rate is $0.50 per million tokens against Grok 4.5's $0.30, a 67% increase to reuse the same prompt on the newer model.

Claude Opus 5 uses a different mechanism entirely. Cached reads price at a fraction of the base input rate, and the minimum cacheable prompt on Opus 5 is 512 tokens, per Anthropic's prompt-caching documentation, covered in more depth in LLM Waves Research's Claude API pricing breakdown.

One weekly pool, and a governance stack worth knowing before procurement

xAI moved every paid Grok plan onto one shared weekly usage pool across Chat, Imagine, Voice, and Build in June 2026. An image-generation session and a coding session draw from the same weekly allowance under this structure, so a team running Grok Build against a repository can watch its budget drop from Imagine or Voice activity elsewhere on the same account.

Claude runs separate weekly caps per product instead, and Opus 5 draws its own cap down fastest of the three Claude tiers, per Anthropic's usage-limits documentation.

SpaceX completed its acquisition of xAI in February 2026, placing Grok inside a four-entity governance structure, per xAI's own announcement.

A January 2026 controversy over deepfake and CSAM-adjacent image generation on the platform, and a European Commission investigation opened afterward and still open as of September 10, 2026, are dated facts a procurement review would weigh before either model clears a security checklist.

This article states them as facts, not conclusions about either company's conduct, and xAI's full model roster sits on its creator page.

Claude Opus 5: one price for the whole window

Claude Opus 5 is Anthropic's newest model tier, released July 24, 2026, with a 1,000,000-token context window and a May 2026 knowledge cutoff. Anthropic's own pricing page lists Opus 5 at $5 per million input tokens and $25 per million output tokens, with no separate rate for requests that grow past any particular size. The Claude Opus 5 model page carries the full spec sheet.

Anthropic reports Claude Opus 5 at 96.0% on SWE-bench Verified and publishes OSWorld scores for computer-use tasks.

Neither benchmark has a Grok 4.6 equivalent on xAI's own model page, which is the mismatch the next section covers.

The coding benchmarks don't measure the same thing

xAI publishes Grok 4.6's coding performance on CursorBench 3.2, DeepSWE 1.1, Terminal-Bench 2.1, and GDPVal-AA, benchmarks Anthropic does not report for Claude Opus 5.

Anthropic publishes SWE-bench Verified and OSWorld, benchmarks xAI does not report for Grok 4.6. The one figure both sides get compared on is CursorBench 3.2.

xAI's own documentation lists Grok 4.6 at 70.8 on CursorBench 3.2 at maximum reasoning effort. Claude Opus 5 scores 70.0 on the same benchmark at Anthropic's standard effort setting, a figure reported by BenchLM.ai rather than published by Anthropic directly. Neither source discloses the number of tasks or runs behind either score.

A 0.8-point spread at two different effort settings, on a benchmark only one vendor grades, does not establish a coding winner between these two models.

On a shared, independently run benchmark, this article treats the coding comparison as insufficient data rather than a ranked result.

Three sources, three different Claude Opus 5 prices

Even a reader working from published trackers cannot get a consistent Claude Opus 5 rate. BenchLM.ai lists Claude Opus 5 at $3 per million input tokens and $15 per million output tokens. Layer3Labs.io publishes no per-token API rate for Opus 5 at all, substituting a $30 to $45 per seat, per month subscription figure in its place, a different product priced as if it were the same one. llm-stats.com and Ampere.sh both list $5 and $25, matching Anthropic's own pricing page directly.

A reader who checks two of these sources before sizing a workload budget has close to even odds of landing on the figure that does not match Anthropic's own page.

Grok 4.6 and Claude Opus 5 are the models actually being compared here

Grok 4.6 (grok-4.6) is the newest tier xAI sells. Claude Opus 5 (claude-opus-5) is the newest tier Anthropic sells. Earlier releases on both sides, Grok 4, 4.1, 4.3, and 4.5 from xAI, and Claude 3.5 Sonnet through Sonnet and Opus 4.8 from Anthropic, are prior generations neither company sells as its top tier as of September 10, 2026.

A price or benchmark figure attached to any of those older names describes a different product than the one covered in this article, and the two should not be paired as if they shipped at the same time.

Which model fits which job

The verdicts below apply LLM Waves Research's standard ladder, defined on the methodology page, to the eight buying situations this site tracks across every comparison.

Segment

Verdict

Why, and the disqualifier

Developers

Situational

Grok 4.6 is cheaper on every request that stays under 200,000 tokens; a codebase-sized prompt that regularly crosses it erases most of the saving.

Startups

Best value: Grok 4.6

The 2.5x and 4.2x gap below 200,000 tokens is the largest cost difference in this comparison; recompute before committing if requests will routinely exceed the threshold.

Enterprise

Situational

Run your own procurement review against the dated governance facts above before either vendor clears a security checklist; this article does not rank the two on governance.

High volume

Situational

Under 200,000 tokens, Grok's per-token rate wins outright; xAI's 20% batch discount applies only to Grok 4.3, not the 4.6 flagship this article prices.

Low latency

Insufficient data

This article did not run a latency test against either model.

On-device

Not recommended, either model

Both are hosted-API-only flagships; neither vendor publishes an on-device or locally run variant.

Self-hosting

Not recommended, either model

Both are closed-weight models; neither xAI nor Anthropic publishes downloadable weights for these tiers.

Non-English

Insufficient data

Neither vendor's pricing or model page sourced for this article publishes a non-English benchmark score.

What this article did not measure

This article did not run a latency test, a throughput test, or an output-quality benchmark against either model. Every figure above is collected from vendor documentation and published rate cards, not produced by testing this session.

Non-English performance, on-device deployment, and head-to-head coding accuracy on a shared, independently run benchmark are open questions this article leaves open rather than estimates it is willing to guess at.

Is Grok actually cheaper than Claude?

Below 200,000 prompt tokens, yes: Grok 4.6 lists at $2 and $6 per million input and output tokens against Claude Opus 5's flat $5 and $25, a 2.5x and 4.2x gap.

Cross 200,000 tokens in a single request and Grok's own rate doubles to $4 and $12, narrowing the gap to 1.25x and 2.1x. Add xAI's per-call tool fees, which Claude has no equivalent line item for, and a heavy agent workload closes more of that remaining gap than the headline numbers suggest.

Which is better for coding, Grok or Claude?

Neither model has a published score on a benchmark the other vendor also runs, so this question has insufficient data for a ranked answer.

The one figure both sides get compared on, a 70.8 to 70.0 result on xAI's own CursorBench 3.2, sets two different effort settings against each other on a benchmark only xAI grades, closer to a tie than a win for either side.

Which one holds up better in an agent workflow?

An agent workflow is exactly where Grok's 200,000-token pricing threshold and xAI's shared weekly usage pool both apply at once: a long-running session that crosses the threshold pays double on tokens for the rest of that request, and a team also running Grok Imagine or Voice on the same plan draws down the same weekly allowance.

Claude Opus 5's flat per-token rate and Anthropic's separate weekly caps per product remove both variables, at a higher starting price on every request.

What does Grok's real-time X data cost on top of tokens?

Web and X Search cost $5 per 1,000 calls, pulling an X user's profile costs $10 per 1,000 calls, and file attachments and code execution cost $10 and $5 per 1,000 calls, all billed apart from the token price on xAI's developer pricing page.

Claude Opus 5 has no equivalent real-time X access to price against these figures.

Is Grok safe for enterprise use?

That depends on what a buyer means by safe. On data handling, this article found no published difference material enough to rank the two vendors.

On governance, SpaceX's February 2026 acquisition of xAI, a January 2026 controversy over deepfake image generation on the platform, and an open European Commission investigation are dated facts a procurement review should weigh.

This article states them as facts, not a verdict on either vendor.

Both vendors will move at least one of these numbers before Grok 4.7 or Claude Opus 5.1 ships

Every dollar figure, context window, and benchmark score above reflects what xAI and Anthropic had published as of September 10, 2026. Grok's pricing has already restructured itself once around a context-length threshold this generation, and Claude's caching mechanics have changed across at least one prior release.