OpenAI's developer docs price GPT-5.6 Sol at $4 per million input tokens and $20 per million output tokens through November 21, 2026, reverting toward $5 and $30 after.

GPT-6 Astra, the top-capability tier, costs $10 and $50 short context, doubling past 272,000 input tokens.

The table below carries every current tier plus five costs a rate card will not show.

Model

Context tokens

Max output tokens

Input $/1M

Output $/1M

Cached input $/1M

Cache write $/1M

GPT-6 Astra, short context (up to 272,000 input tokens)

1,050,000

128,000

10.00

50.00

1.00

12.50

GPT-6 Astra, long context (above 272,000 input tokens)

1,050,000

128,000

20.00

75.00

2.00

25.00

GPT-5.6 Sol, short context, promotional through 2026-11-21

1,050,000

128,000

4.00

20.00

0.40

5.00

GPT-5.6 Sol, long context, promotional through 2026-11-21

1,050,000

128,000

8.00

30.00

0.80

10.00

GPT-5.6 Sol, short context, standard rate †

1,050,000

128,000

5.00 †

30.00 †

not published

not published

GPT-5.6 Terra, short context

1,050,000

128,000

2.00

12.00

0.20

2.50

GPT-5.6 Terra, long context

1,050,000

128,000

4.00

18.00

0.40

5.00

GPT-5.6 Luna, short context

1,050,000

128,000

0.20

1.20

0.02

0.25

GPT-5.6 Luna, long context

1,050,000

128,000

0.40

1.80

0.04

0.50

Table note: rows are grouped by model, short context above long context, not sorted by price. † vendor reported on OpenAI's separate marketing page, not yet confirmed as the developer-docs rate that takes effect once the Sol promotion ends. Not published: OpenAI's marketing page lists an input and output rate for Sol but no cached-input or cache-write figure.

GPT-6 Astra costs $10 per million input tokens and $50 per million output tokens inside its 272,000-token short-context tier, rising to $20 and $75 past that point. Astra costs more. It also does more.

GPT-5.6 Sol runs at $4 and $20 on the same axis through November 21, 2026, roughly 60% below the input and output rate of the top-capability tier. GPT-5.6 Luna, the cheapest current tier, costs $0.20 and $1.20, one-twentieth of the promotional input rate on Sol, inside the same 1,050,000-token context window every current tier shares.

Basis and method

This spoke fills the OpenAI row of LLM API Pricing Compared, the site's cross-vendor pricing hub, with the detail a single comparison table cannot carry.

The figures in this article are aggregated from OpenAI's own developer documentation. LLM Waves Research ran no benchmark and priced no request against a live endpoint. Every figure traces to developers.openai.com/api/docs/pricing, developers.openai.com/api/docs/guides/prompt-caching, developers.openai.com/api/docs/guides/spend-limits, and the model specification pages under developers.openai.com/api/docs/models, all collected September 14, 2026. Two figures are derived, not copied: the long-context multiplier (long-context input price divided by short-context input price, repeated for output) and the cache multiplier (cache-write price divided by uncached input price, and cached-read price divided by uncached input price).

Both hold at identical ratios across all four current tiers, computed from the published rate table rather than stated outright by OpenAI. A separate marketing page, openai.com/api/pricing, lists GPT-5.6 Sol at a different rate than the developer documentation on the same day; that conflict is reported as a finding here, not resolved by picking a winner.

This article ranks OpenAI's current model tiers on published price per million tokens, not on speed, accuracy, or any benchmark score. Full methodology: https://llmwaves.com/methodology

All four current tiers, from $0.20 to $10 per million short-context input tokens, exactly double their input rate once a request passes 272,000 tokens.

The promotional rate on GPT-5.6 Sol expires November 21, 2026

Four dollars per million input tokens and twenty dollars per million output tokens is not the permanent rate for GPT-5.6 Sol. OpenAI's developer pricing documentation labels this pricing promotional, effective through November 21, 2026. What follows is unstated on that page.

A separate page, openai.com's own marketing pricing table, already lists Sol at $5 input and $30 output with no promotional note attached, collected the same day as the developer-docs figures above. That is the strongest available signal for where the rate lands once the promotion ends, though OpenAI has not published a stated reversion date or rate, so this article treats $5 and $30 as reported, not confirmed.

A workload budgeted today at $4 and $20 underprices itself by 25% on input and 50% on output the moment the calendar turns, assuming the marketing page's number holds once the promotion lapses.

The expiry date sits on the developer-docs page in plain text. It does not sit on the invoice.

Input rises 25% and output rises 50% on the same model, based on the rate OpenAI's marketing page already shows.

A cache write costs 1.25 times more than a plain input token, not less

Prompt caching is sold as a discount. On GPT-5.6 and later, the first pass through a prompt prefix is not discounted at all. It costs 1.25 times the uncached input rate: $12.50 per million on GPT-6 Astra against a $10 standard input rate, $5.00 on Sol against $4.00, $2.50 on Terra against $2.00, $0.25 on Luna against $0.20.

The ratio holds exactly across all four tiers, a pattern OpenAI's prompt caching guide states as a multiplier but does not lay out side by side the way the numbers above do.

A cached read is the genuine discount: 0.1 times the uncached input rate, again identical across every tier.

The prefix has to be at least 1,024 visible input tokens to qualify for caching on GPT-5.6 and later, and the cache expires 30 minutes after its last write or reuse, configured throughprompt_cache_options.ttlwhich accepts only that one value.

Earlier model generations carry no separate cache-write charge at all and use a longer prompt_cache_retention window of either an in-memory few minutes or 24 hours.

A workload that writes a fresh prefix on every call and never reuses it inside that 30-minute window pays the write premium on every single request and collects none of the read discount back.

The write premium and the read discount hold at the same ratio on every tier, from GPT-6 Astra's $10 input down to GPT-5.6 Luna's $0.20.

Crossing 272,000 input tokens doubles the input price for the entire request

OpenAI's own model documentation for GPT-6 Astra states that prompts exceeding 272,000 input tokens are priced at 2x the standard input and cache rates.

The published rate table shows the output price moving too, at 1.5 times the standard rate, and the multiplier applies to the whole request once the threshold is crossed, not only to the tokens past it.

A 273,000-token prompt is billed as a long-context call from token one, at $20 per million on Astra instead of $10. One token over the line. The whole call reprices.

The 272,000-token threshold is the single largest line item most cost estimates miss. An agent that retrieves a large document, or a retrieval pipeline that stuffs several sources into one context window, can cross it without anyone deciding to.

The ratio, 2x on input and 1.5x on output, is stated nowhere on OpenAI's pricing page itself. It sits only on the model specification page, one click away from the table most readers actually budget from.

Short context

Below 272,000 input tokens, every current tier bills at its base rate: $10 and $50 on Astra, $4 and $20 on the promotional Sol rate, $2 and $12 on Terra, $0.20 and $1.20 on Luna.

Long context

Above that line, input exactly doubles and output rises by half on every tier: $20 and $75 on Astra, $8 and $30 on Sol, $4 and $18 on Terra, $0.40 and $1.80 on Luna.

The $10-per-1,000-calls web search rate is not the number on the invoice

The advertised rate

OpenAI's pricing documentation lists the web search tool at $10 per 1,000 calls, with content retrieved from the page billed separately at the model's standard input rate.

That second part is documented. The first part is not the whole story.

What developers report paying

A thread on OpenAI's own developer community, started by a user who found the visible charge running two to three times the advertised figure, drew replies describing the mechanism: one visible tool call can trigger several internal sub-searches, and each one is billed, invisibly, on top of the call a developer sees in their own logs.

One reply cites a bill of roughly $100 per 1,000 searches against the $10 sticker price. These figures are reported, not measured by LLM Waves Research, and OpenAI's pricing page does not disclose the sub-search behavior.

A workload that budgets $10 per 1,000 calls and gets billed for hidden multiples of that has no line item to point to that explains the gap.

A spend limit is a wall now, not an email

OpenAI's spend controls split into two kinds. A spend alert is a notification: traffic continues once the configured amount is reached. A hard spend limit blocks it: requests past the limit return a 429 error carrying organization_spend_limit_exceeded or project_spend_limit_exceeded.

OpenAI's own spend limits guide states enforcement is not instantaneous, so recorded spend can slightly exceed the configured number before the block takes effect.

An email alert used to be the only spend control OpenAI offered. Hard spend limits are a separate, newer mechanism layered on top, not a replacement for the alert.

A team that sets a hard limit expecting an exact ceiling should expect a small overshoot instead, not a precise stop.

What this article did not measure

This article did not run a latency test, a throughput test, or an output-quality benchmark against any model discussed.

Which OpenAI tier fits which reader

Reader

Verdict

Why

Number

Prototyping or low-stakes drafting

Best value

The rate on GPT-5.6 Luna, which carries no promotion, is already the floor of the current lineup

$0.20 / $1.20 per million

High-volume production text

Situational, if the workload fits under 272,000 tokens per call

The promotional rate on GPT-5.6 Sol is the cheapest mid-tier price through November 21, 2026, then rises

$4.00 / $20.00 per million, expiring

Long-context agents and document retrieval

Not recommended without a token budget check

Crossing 272,000 tokens doubles input on every tier, silently, from token one of the request

2x input, 1.5x output

Top-capability reasoning or coding workloads

Top pick

GPT-6 Astra is the only tier at this generation's top capability band, priced accordingly

$10.00 / $50.00 per million

Non-English workloads

Insufficient data

No OpenAI-published or independently sourced per-language cost or quality breakdown was found for any current tier

no rank assigned

Self-hosting or on-device deployment

Not applicable

OpenAI's current lineup ships as a hosted API only, with no open-weight release in this generation

n/a

GPT-6 Astra and GPT-5.6 Luna sit at opposite ends of the same 1,050,000-token context window and 128,000-token output ceiling, fifty times apart on input price. Claude Haiku 4.5 prices at $1.00 and $5.00 per million, Gemini 3.6 Flash at $0.75 and $3.75, and DeepSeek-V4-Flash at $0.22 and $0.66 off-peak, all cheaper on output than the promotional rate on GPT-5.6 Sol and all more expensive on input than GPT-5.6 Luna. None of the four vendors' cheapest tiers line up on both axes at once.

GPT-5.6 Luna is the cheapest input rate of the four, but the off-peak output rate on DeepSeek-V4-Flash undercuts it.

Is the OpenAI API included in a ChatGPT Plus or Pro subscription?

No. A ChatGPT Plus, Pro, or Business subscription pays for chat access on chatgpt.com and its apps. API access bills separately, per token, through a different account balance, with its own rate card.

A developer who already pays for ChatGPT Plus starts an API project at zero credit, not a pool to draw from.

Which current OpenAI model is the cheapest to run?

GPT-5.6 Luna, at $0.20 per million input tokens and $1.20 per million output tokens, inside the same 1,050,000-token context window as every other current tier.

Among the three rival vendors most often compared against it, the off-peak output rate on DeepSeek-V4-Flash, $0.66 per million, is cheaper than the $1.20 rate on Luna, though that rate only applies during a seven-hour daily window in UTC.

Does openai.com list the same GPT-5.6 Sol rate as the developer docs?

No, and that gap is the reason this article exists as a spoke rather than a repeat of the developer documentation. The developer docs, collected September 14, 2026, price the promotional rate on Sol at $4 and $20.

OpenAI's separate marketing pricing page, collected the same day, lists Sol at $5 and $30 with no promotional note. Both pages belong to OpenAI. Neither cites the other.

Is a failed or timed-out API request billed?

Not for the Batch API, according to OpenAI's own support team, quoted in a thread on OpenAI's developer community: a request that fails and returns no output appears only in the error file and is not charged, and completion tokens are billed only for output actually generated and recorded.

OpenAI has not published an equivalent statement for a rate-limited or timed-out request on the standard, synchronous API, and no published source resolves that narrower case either.

OpenAI's own two pages still do not match

The table above is built from developers.openai.com/api/docs/pricing, chosen because that page carries the promotional note the marketing page omits. That choice could look wrong within days if OpenAI updates either page. Recheck both before committing a budget to either number, and treat the raw data file below as the version this table is built from, not a permanent record.