Skip to content
llmwaves

Qwen3 Max Thinking

Qwen · released Feb 9, 2026

ProprietaryReasoningTool useStructured output

Qwen3-Max-Thinking is the flagship reasoning model in the Qwen3 series, designed for high-stakes cognitive tasks that require deep, multi-step reasoning. By significantly scaling model capacity and reinforcement learning compute, it...

Specification

Context window
262K
Max output
66K
Knowledge cutoff
Not stated
Parameters
Undisclosed
Licence
Proprietary
Serving providers
1
Moderated
No

Intelligence

32.5

78th percentile

Coding

Coding Index

Agentic

Agentic Index

Output speed

31 t/s

Median across providers

Latency

1.14s

Time to first token

Cost per task

$0.35

Estimated

Benchmarks

Where the score comes from

The Intelligence Index is a composite. These are the underlying evaluations this model was actually measured on.

Evaluation scores

Percentage correct · higher is better

  • GPQA Diamond
    86.1%
  • τ²-bench (Telecom)
    83.6%
  • IFBench
    70.7%
  • AA-LCR (long context)
    70.3%
  • SciCode
    43.1%
  • Humanity's Last Exam
    28.0%
  • Terminal-Bench Hard
    24.2%

An evaluation missing from this list was not run for this model — it is not a zero.

View as table
EvaluationScore
GPQA Diamond86.1%
τ²-bench (Telecom)83.6%
IFBench70.7%
AA-LCR (long context)70.3%
SciCode43.1%
Humanity's Last Exam28.0%
Terminal-Bench Hard24.2%

Against its peers

Intelligence Index · this model highlighted, nearest peers in grey

Peers are the models sitting closest on the Intelligence Index — the set you would realistically choose between.

View as table
ModelIntelligence
LongCat 2.033.9
Kimi K2 Thinking33.5
o3 Pro33.3
Qwen3.5-122B-A10B32.8
Qwen3 Max Thinking32.5
Qwen3.6 35B A3B32.1
M2.132.1
GPT-5.1-Codex-Mini31.3

Percentile among all indexed models

Intelligence78th
Terminal work64th

Pricing

What it costs to run

List prices per million tokens, plus what one representative task works out to.

List price

Input / 1M tokens
$0.78
Output / 1M tokens
$3.9
Cached input / 1M
Not offered
Blended 3:1
$1.56

One task, estimated

$0.35

Input tokens
50,000
Output tokens
80,000
Profile
Reasoning

Estimated from list pricing: 50K input tokens plus 80K output tokens for reasoning models (25K for non-reasoning).

Serving providers

Speed and latency figures are medians across these providers, so a widely-served model reports a blend rather than any single endpoint.