Skip to content
llmwaves

Leaderboards

One board per question

No single ranking answers every question. Pick the measurement that matches the work you are actually doing — a model that tops GPQA may be nowhere near the top on long-context retrieval.

Headline

The composite indices

Aggregate scores. Useful for a shortlist, misleading as a final answer.

Intelligence Index

index · higher is better

Composite of ten reasoning, knowledge, coding and agentic evaluations.

View as table
Modelindex
Claude Opus 5 (Anthropic)63.1
Claude Fable 5 (Anthropic)62.1
GPT-5.6 Sol (OpenAI)60.9
Kimi K3 (Moonshot AI)59.7
Qwen3.8 Max (Qwen)58.1
Claude Opus 4.8 (Anthropic)57.3
Muse Spark 1.2 (Meta)56.8
GPT-5.6 Terra (OpenAI)56.6
GPT-5.5 (OpenAI)56.3
Grok 4.5 (xAI)55.8

Coding Index

index · higher is better

SciCode, LiveCodeBench and Terminal-Bench Hard, combined.

View as table
Modelindex
Claude Opus 5 (Anthropic)78.0
GPT-5.6 Sol (OpenAI)77.4
GPT-5.6 Terra (OpenAI)76.7
Claude Fable 5 (Anthropic)76.5
Kimi K3 (Moonshot AI)76.2
GPT-5.5 (OpenAI)74.9
Claude Opus 4.8 (Anthropic)74.3
Claude Opus 4.7 (Anthropic)73.6
Grok 4.5 (xAI)72.4
Muse Spark 1.2 (Meta)72.2

Agentic Index

index · higher is better

Tool use, long-horizon planning and terminal work.

View as table
Modelindex
Claude Opus 5 (Anthropic)59.2
Qwen3.8 Max (Qwen)58.4
GPT-5.6 Sol (OpenAI)57.8
Claude Fable 5 (Anthropic)56.6
Kimi K3 (Moonshot AI)54.3
GPT-5.6 Terra (OpenAI)50.2
Claude Sonnet 5 (Anthropic)49.7
Claude Opus 4.8 (Anthropic)49.4
Grok 4.5 (xAI)48.9
V4 Flash 0731 (DeepSeek)48.4

Output speed

tokens/s · higher is better

Median output tokens per second across serving providers.

View as table
Modeltokens/s
V3 Fast (Morph)3,502 t/s
V3 Large (Morph)3,129 t/s
Apply 3 (Relace)3,080 t/s
gpt-oss-120b (OpenAI)773 t/s
GLM 4.7 (Z.ai)502 t/s
Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image) (Google)443 t/s
Qwen3 32B (Qwen)423 t/s
gpt-oss-safeguard-20b (OpenAI)422 t/s
Grok 4.20 Multi-Agent (xAI)337 t/s
M2.7 (MiniMax)307 t/s

Context window

tokens · higher is better

Maximum tokens the model accepts in one request.

View as table
Modeltokens
Grok 4.20 Multi-Agent (xAI)2M
Grok 4.20 (xAI)2M
Llama 4 Scout (Meta)1.31M
GPT-5.6 Luna Pro (OpenAI)1.05M
GPT-5.6 Luna (OpenAI)1.05M
GPT-5.6 Terra Pro (OpenAI)1.05M
GPT-5.6 Terra (OpenAI)1.05M
GPT-5.6 Sol Pro (OpenAI)1.05M
GPT-5.6 Sol (OpenAI)1.05M
GPT-5.5 Pro (OpenAI)1.05M

Arena Elo

Dot position on a 1,325–1,475 scale

1,3251,475

Head-to-head generation quality, judged pairwise. Elo has no meaningful zero, so bars from a zero baseline would flatten every gap.

View as table
ModelElo
Kimi K3 (Moonshot AI)1,455
Claude Fable 5 (Anthropic)1,397
Claude Opus 5 (Anthropic)1,393
GPT-5.6 Sol (OpenAI)1,380
GLM 5.2 (Z.ai)1,363
Gemini 3.6 Flash (Google)1,346
Claude Opus 4.7 (Anthropic)1,337
Claude Sonnet 5 (Anthropic)1,336
GPT-5.5 (OpenAI)1,336
Gemini 3.1 Pro Preview (Google)1,333

Individual evaluations

The measurements underneath

Each of these is a real test with its own failure modes. A model absent from a board was not run on it — that is not the same as scoring zero.

GPQA Diamond

Percentage correct · higher is better

Graduate-level science questions written to resist search.

View as table
ModelScore
GPT-5.6 Sol (OpenAI)94.1%
Gemini 3.1 Pro Preview (Google)94.1%
Kimi K3 (Moonshot AI)93.5%
GPT-5.5 (OpenAI)93.5%
Claude Opus 5 (Anthropic)93.2%
Grok 4.5 (xAI)93.1%
MiniMax M3 (MiniMax)92.9%
Gemini 3.6 Flash (Google)92.8%
Qwen3.8 Max (Qwen)92.7%
Claude Fable 5 (Anthropic)92.6%

Humanity's Last Exam

Percentage correct · higher is better

A deliberately brutal cross-disciplinary exam; scores stay low.

View as table
ModelScore
Claude Fable 5 (Anthropic)55.5%
Claude Opus 5 (Anthropic)54.9%
GPT-5.6 Sol (OpenAI)49.5%
Claude Opus 4.8 (Anthropic)48.7%
Gemini 3.1 Pro Preview (Google)47.0%
Kimi K3 (Moonshot AI)46.9%
Muse Spark 1.1 (Meta)46.2%
GPT-5.5 (OpenAI)45.8%
Muse Spark 1.2 (Meta)45.5%
GPT-5.4 (OpenAI)43.7%

SciCode

Percentage correct · higher is better

Research-grade scientific coding problems.

View as table
ModelScore
Claude Fable 5 (Anthropic)60.2%
Gemini 3.1 Pro Preview (Google)58.9%
Kimi K3 (Moonshot AI)58.7%
Muse Spark 1.1 (Meta)58.2%
GPT-5.4 (OpenAI)56.6%
Muse Spark 1.2 (Meta)56.4%
GPT-5.6 Sol (OpenAI)56.1%
GPT-5.5 (OpenAI)56.1%
Claude Opus 5 (Anthropic)55.7%
GPT-5.2-Codex (OpenAI)54.6%

τ²-bench (Telecom)

Percentage correct · higher is better

Tool-use and policy adherence in a simulated telecom support agent.

View as table
ModelScore
GLM 5.2 (Z.ai)99.1%
GLM 4.7 Flash (Z.ai)98.8%
Claude Fable 5 (Anthropic)98.5%
Step 3.7 Flash (StepFun)98.5%
GLM 5V Turbo (Z.ai)98.5%
GLM 5 Turbo (Z.ai)98.5%
GLM 5 (Z.ai)98.2%
Grok 4.3 (xAI)97.7%
GLM 5.1 (Z.ai)97.7%
Qwen3.6 Plus (Qwen)97.7%

Terminal-Bench Hard

Percentage correct · higher is better

Hard end-to-end tasks solved inside a real terminal.

View as table
ModelScore
GPT-5.6 Sol (OpenAI)65.9%
Claude Fable 5 (Anthropic)62.9%
GPT-5.5 (OpenAI)60.6%
Claude Opus 4.8 (Anthropic)58.3%
GPT-5.6 Terra (OpenAI)57.6%
GPT-5.4 (OpenAI)57.6%
Gemini 3.1 Pro Preview (Google)53.8%
GPT-5.3-Codex (OpenAI)53.0%
GPT-5.4 Mini (OpenAI)52.3%
Claude Opus 4.7 (Anthropic)51.5%

LiveCodeBench

Percentage correct · higher is better

Contest programming on problems released after training.

View as table
ModelScore
GLM 4.7 (Z.ai)89.4%
GPT-5.2 (OpenAI)88.9%
gpt-oss-120b (OpenAI)87.8%
GPT-5.1 (OpenAI)86.8%
o4 Mini High (OpenAI)85.9%
o4 Mini (OpenAI)85.9%
Kimi K2 Thinking (Moonshot AI)85.3%
GPT-5.1-Codex (OpenAI)84.9%
GPT-5 (OpenAI)84.6%
GPT-5 Mini (OpenAI)83.8%

AA-LCR (long context)

Percentage correct · higher is better

Long-context reasoning across large documents.

View as table
ModelScore
Muse Spark 1.2 (Meta)83.3%
Kimi K3 (Moonshot AI)82.7%
Muse Spark 1.1 (Meta)81.3%
Gemini 3.5 Flash (Google)81.0%
MiniMax M3 (MiniMax)80.3%
GPT-5.6 Terra (OpenAI)79.7%
GPT-5.2-Codex (OpenAI)79.3%
GPT-5.2 (OpenAI)79.3%
Gemini 3.6 Flash (Google)79.0%
GPT-5.5 (OpenAI)79.0%

AIME 2025

Percentage correct · higher is better

Competition mathematics.

View as table
ModelScore
GPT-5.2 (OpenAI)99.0%
GPT-5.1-Codex (OpenAI)95.7%
GLM 4.7 (Z.ai)95.0%
Kimi K2 Thinking (Moonshot AI)94.7%
GPT-5 (OpenAI)94.3%
GPT-5.1 (OpenAI)94.0%
gpt-oss-120b (OpenAI)93.4%
GPT-5.1-Codex-Mini (OpenAI)91.7%
GPT-5 Mini (OpenAI)90.7%
o4 Mini High (OpenAI)90.7%

MMLU-Pro

Percentage correct · higher is better

Broad multi-task knowledge, harder distractors than MMLU.

View as table
ModelScore
Claude Opus 4.5 (Anthropic)88.9%
M2.1 (MiniMax)87.5%
GPT-5.2 (OpenAI)87.4%
GPT-5 (OpenAI)87.1%
GPT-5.1 (OpenAI)87.0%
Gemini 2.5 Pro (Google)86.2%
GPT-5.1-Codex (OpenAI)86.0%
Claude Sonnet 4.5 (Anthropic)86.0%
Claude Opus 4 (Anthropic)86.0%
GLM 4.7 (Z.ai)85.6%

IFBench

Percentage correct · higher is better

Precise instruction following under explicit constraints.

View as table
ModelScore
MiniMax M3 (MiniMax)82.9%
Grok 4.3 (xAI)81.3%
Grok 4.20 (xAI)81.2%
Qwen3.7 Max (Qwen)80.5%
MiMo-V2.5-Pro (Xiaomi)79.9%
Qwen3.5 397B A17B (Qwen)78.8%
Qwen3.7 Plus (Qwen)78.0%
GPT-5.2-Codex (OpenAI)77.6%
Gemini 3.1 Flash Lite (Google)77.2%
Gemini 3.1 Flash Lite Preview (Google)77.2%

Cross-section

The frontier, benchmark by benchmark

The top twelve models on the Intelligence Index, and how each one actually did on the tests behind it.

Frontier models across six evaluations

Percentage correct · darker means stronger

ModelGPQA DiamondHumanity's Last ExamSciCodeτ²-benchTerminal-Bench HardAA-LCR
Claude Opus 5Anthropic93%55%56%42%76%
Claude Fable 5Anthropic93%56%60%99%63%77%
GPT-5.6 SolOpenAI94%50%56%85%66%78%
Kimi K3Moonshot AI94%47%59%46%83%
Qwen3.8 MaxQwen93%43%53%51%74%
Claude Opus 4.8Anthropic92%49%54%94%58%73%
Muse Spark 1.2Meta90%46%56%35%83%
GPT-5.6 TerraOpenAI93%43%54%86%58%80%
GPT-5.5OpenAI94%46%56%94%61%79%
Grok 4.5xAI93%43%54%42%74%
Claude Sonnet 5Anthropic91%41%54%37%77%
Claude Opus 4.7Anthropic91%42%55%89%52%75%
0%High

Read across a row to see a model's shape, down a column to see which test separates the field. Blank means not run.

View as table
ModelGPQA DiamondHumanity's Last ExamSciCodeτ²-bench (Telecom)Terminal-Bench HardAA-LCR (long context)
Claude Opus 593.2%54.9%55.7%42.1%75.7%
Claude Fable 592.6%55.5%60.2%98.5%62.9%76.7%
GPT-5.6 Sol94.1%49.5%56.1%85.1%65.9%77.7%
Kimi K393.5%46.9%58.7%46.0%82.7%
Qwen3.8 Max92.7%43.0%52.9%51.3%74.3%
Claude Opus 4.892.0%48.7%53.5%94.4%58.3%73.0%
Muse Spark 1.290.4%45.5%56.4%34.8%83.3%
GPT-5.6 Terra92.5%42.9%53.9%86.3%57.6%79.7%
GPT-5.593.5%45.8%56.1%93.9%60.6%79.0%
Grok 4.593.1%42.7%54.1%42.1%74.0%
Claude Sonnet 591.1%41.3%53.6%37.3%77.0%
Claude Opus 4.791.4%42.3%54.5%88.6%51.5%75.3%

Economics

Cost and responsiveness

The two boards where a low number wins.

Cheapest per task

Estimated USD · lower is better

Estimated from list pricing: 50K input tokens plus 80K output tokens for reasoning models (25K for non-reasoning). Limited to models scoring 30+ on the Intelligence Index.

View as table
ModelCost / taskIntelligence
Ling-3.0-flash (InclusionAI)$0.00637.8
V4 Flash 0423 (DeepSeek)$0.0251.8
V4 Flash 0731 (DeepSeek)$0.0251.8
Hy3 preview (Tencent)$0.0234.4
MiMo-V2.5 (Xiaomi)$0.0338.0
KAT-Coder-Pro V2 (KwaiPilot)$0.0433.9
Hy3 (Tencent)$0.0542.2
GPT-5.6 Luna (OpenAI)$0.0552.3
Ring-2.6-1T (InclusionAI)$0.0531.1
M2.5 (MiniMax)$0.0834.5

Lowest latency

Seconds to first token · lower is better

Median time to first token.

View as table
ModelLatency
Granite 4.1 8B (IBM Granite)118ms
Llama Guard 4 12B (Meta)138ms
Llama 3 8B Lunaris (Sao10K)158ms
gpt-oss-120b (OpenAI)205ms
Qwen3 32B (Qwen)205ms
Codestral 2508 (Mistral AI)208ms
Llama 3.2 3B Instruct (Meta)212ms
MythoMax 13B (Gryphe)222ms
Hermes 4 70B (Nous Research)229ms
gpt-oss-safeguard-20b (OpenAI)237ms

323 models considered for every board on this page.