Skip to content
llmwaves

Coding agents

Which models survive an agent loop

Writing a function is easy. Holding a plan across fifty tool calls, reading a failing test and fixing the right file is not. These are the measurements that separate the two.

Tool-capable models

260

Scored on agentic work

80

Best agentic score

59.2

Claude Opus 5

Best Terminal-Bench

66%

GPT-5.6 Sol

Indices

Coding and agentic ability

Two composites. Coding covers writing and fixing code; agentic covers planning, tool use and long-horizon work.

Coding Index

SciCode, LiveCodeBench and Terminal-Bench Hard combined · higher is better

View as table
ModelCoding Index
Claude Opus 5 (Anthropic)78.0
GPT-5.6 Sol (OpenAI)77.4
GPT-5.6 Terra (OpenAI)76.7
Claude Fable 5 (Anthropic)76.5
Kimi K3 (Moonshot AI)76.2
GPT-5.5 (OpenAI)74.9
Claude Opus 4.8 (Anthropic)74.3
Claude Opus 4.7 (Anthropic)73.6
Grok 4.5 (xAI)72.4
Muse Spark 1.2 (Meta)72.2

Agentic Index

Tool use, planning and terminal work · higher is better

View as table
ModelAgentic Index
Claude Opus 5 (Anthropic)59.2
Qwen3.8 Max (Qwen)58.4
GPT-5.6 Sol (OpenAI)57.8
Claude Fable 5 (Anthropic)56.6
Kimi K3 (Moonshot AI)54.3
GPT-5.6 Terra (OpenAI)50.2
Claude Sonnet 5 (Anthropic)49.7
Claude Opus 4.8 (Anthropic)49.4
Grok 4.5 (xAI)48.9
V4 Flash 0731 (DeepSeek)48.4

Economics

Agentic ability against what it costs

Agent loops burn output tokens. A model that scores two points higher and costs four times as much is not obviously the right call.

Agentic Index vs cost per task

Estimated USD per task on a log scale · up and to the right is the better deal

  • Proprietary
  • Open weights
01020304050$0.001$0.01$0.10$1.00$10.00Estimated cost per task, USD · better →Agentic IndexClaude Opus 5Qwen3.8 MaxGPT-5.6 TerraV4 Flash 0423gpt-oss-120bgpt-oss-20b

Labelled points are the Pareto frontier. Estimated from list pricing: 50K input tokens plus 80K output tokens for reasoning models (25K for non-reasoning).

View as table
ModelAgenticCost / task
Claude Opus 5 (Anthropic)59.2$2.25
Qwen3.8 Max (Qwen)58.4$0.58
GPT-5.6 Sol (OpenAI)57.8$2.65
Claude Fable 5 (Anthropic)56.6$4.50
Kimi K3 (Moonshot AI)54.3$1.35
GPT-5.6 Terra (OpenAI)50.2$0.53
Claude Sonnet 5 (Anthropic)49.7$0.90
Claude Opus 4.8 (Anthropic)49.4$2.25
Grok 4.5 (xAI)48.9$0.58
V4 Flash 0731 (DeepSeek)48.4$0.02
V4 Flash 0423 (DeepSeek)48.4$0.02
GPT-5.5 (OpenAI)47.4$2.65
GPT-5.6 Luna (OpenAI)46.9$0.05
Claude Opus 4.7 (Anthropic)46.3$2.25
GLM 5.2 (Z.ai)45.7$0.23
GPT-5.4 (OpenAI)44.2$1.32
Claude Sonnet 4.6 (Anthropic)42.1$1.35
Gemini 3.6 Flash (Google)40.5$0.68
Muse Spark 1.1 (Meta)39.7$0.40
Gemini 3.5 Flash (Google)39.7$0.79
V4 Pro (DeepSeek)37.8$0.09
MiniMax M3 (MiniMax)36.1$0.11
Inkling (Thinking Machines)34.1$0.37
Inkling Small (Thinking Machines)31.9$0.12
GPT-5.4 Mini (OpenAI)31.5$0.40

Underlying evaluations

The four that matter for agents

Terminal work, contest programming, research-grade scientific code, and tool-use discipline under a policy.

Terminal-Bench Hard

Percentage correct · higher is better

View as table
ModelScore
GPT-5.6 Sol (OpenAI)65.9%
Claude Fable 5 (Anthropic)62.9%
GPT-5.5 (OpenAI)60.6%
Claude Opus 4.8 (Anthropic)58.3%
GPT-5.6 Terra (OpenAI)57.6%
GPT-5.4 (OpenAI)57.6%
Gemini 3.1 Pro Preview (Google)53.8%
GPT-5.3-Codex (OpenAI)53.0%
GPT-5.4 Mini (OpenAI)52.3%
Claude Opus 4.7 (Anthropic)51.5%

τ²-bench (Telecom)

Percentage correct · higher is better

View as table
ModelScore
GLM 5.2 (Z.ai)99.1%
GLM 4.7 Flash (Z.ai)98.8%
Claude Fable 5 (Anthropic)98.5%
Step 3.7 Flash (StepFun)98.5%
GLM 5V Turbo (Z.ai)98.5%
GLM 5 Turbo (Z.ai)98.5%
GLM 5 (Z.ai)98.2%
Grok 4.3 (xAI)97.7%
GLM 5.1 (Z.ai)97.7%
Qwen3.6 Plus (Qwen)97.7%

LiveCodeBench

Percentage correct · higher is better

View as table
ModelScore
GLM 4.7 (Z.ai)89.4%
GPT-5.2 (OpenAI)88.9%
gpt-oss-120b (OpenAI)87.8%
GPT-5.1 (OpenAI)86.8%
o4 Mini High (OpenAI)85.9%
o4 Mini (OpenAI)85.9%
Kimi K2 Thinking (Moonshot AI)85.3%
GPT-5.1-Codex (OpenAI)84.9%
GPT-5 (OpenAI)84.6%
GPT-5 Mini (OpenAI)83.8%

SciCode

Percentage correct · higher is better

View as table
ModelScore
Claude Fable 5 (Anthropic)60.2%
Gemini 3.1 Pro Preview (Google)58.9%
Kimi K3 (Moonshot AI)58.7%
Muse Spark 1.1 (Meta)58.2%
GPT-5.4 (OpenAI)56.6%
Muse Spark 1.2 (Meta)56.4%
GPT-5.6 Sol (OpenAI)56.1%
GPT-5.5 (OpenAI)56.1%
Claude Opus 5 (Anthropic)55.7%
GPT-5.2-Codex (OpenAI)54.6%

Agent arena

Building whole applications

Elo from pairwise judgements on complete builds, not snippets — the closest proxy for what an agent is actually asked to do.

Full-stack apps

Arena Elo · dot plot on a stated scale, not bars from zero

1,2001,300
View as table
ModelEloWin rate
Claude Fable 5129563.7%
Claude Opus 4.8128562.2%
GLM 5.2127463.8%
Claude Sonnet 5127259.8%
Claude Opus 4.6126268.5%
Claude Sonnet 4.6125264.1%
Qwen3.7 Max124055.6%
Gemini 3.5 Flash123456.7%
MiniMax M3122653.0%
Kimi K2.7 Code121854.5%

Mobile apps

Arena Elo · dot plot on a stated scale, not bars from zero

1,2001,300
View as table
ModelEloWin rate
Claude Opus 4.6127964.8%
Claude Sonnet 4.6127163.4%
Claude Opus 4.8125557.5%
Claude Fable 5125457.5%
Claude Sonnet 5123755.1%
MiniMax M3123755.4%
Kimi K2.6122154.9%
GLM 5.2121954.0%
Gemini 3.5 Flash121853.5%
GLM 5.1121854.7%

Full list

260 tool-capable models

Every model that can call a tool, which is the floor for running an agent at all.

Filter by creator
260 of 260 models
Claude Opus 5
Anthropic
63.1$2.25
Qwen3.8 Max
Qwen
58.1$0.58
GPT-5.6 Sol
OpenAI
60.9$2.65
Claude Fable 5
Anthropic
62.1$4.50
Kimi K3
Moonshot AIopen
59.7$1.35
GPT-5.6 Terra
OpenAI
56.6$0.53
Claude Sonnet 5
Anthropic
55.3$0.90
Claude Opus 4.8
Anthropic
57.3$2.25
Grok 4.5
xAI
55.8$0.58
V4 Flash 0731
DeepSeekopen
51.8$0.02
V4 Flash 0423
DeepSeekopen
51.8$0.02
GPT-5.5
OpenAI
56.3$2.65
GPT-5.6 Luna
OpenAI
52.3$0.05
Claude Opus 4.7
Anthropic
55.0$2.25
GLM 5.2
Z.aiopen
52.6$0.23
GPT-5.4
OpenAI
53.1$1.32
Claude Sonnet 4.6
Anthropic
36.8$1.35
Gemini 3.6 Flash
Google
51.6$0.68
Muse Spark 1.1
Meta
53.2$0.40
Gemini 3.5 Flash
Google
52.0$0.79
V4 Pro
DeepSeekopen
45.3$0.09
MiniMax M3
MiniMaxopen
45.4$0.11
Inkling
Thinking Machinesopen
42.3$0.37
Inkling Small
Thinking Machinesopen
41.2$0.12
GPT-5.4 Mini
OpenAI
40.9$0.40
Hy3
Tencentopen
42.2$0.05
Hy3 preview
Tencentopen
34.4$0.02
Kimi K2.6
Moonshot AIopen
45.1$0.23
Qwen3.7 Max
Qwen
46.7$0.43
GLM 5.1
Z.aiopen
41.0$0.29
Kimi K2.7 Code
Moonshot AIopen
43.0$0.32
GPT-5.4 Nano
OpenAI
39.7$0.11
MiMo-V2.5-Pro
Xiaomiopen
42.9$0.09
Qwen3.6 Plus
Qwen
40.5$0.17
Qwen3.6 27B
Qwenopen
37.7$0.32
Gemini 3.5 Flash Lite
Google
37.4$0.22
GPT-5
OpenAI
35.3$0.86
Claude Sonnet 4.5
Anthropic
29.9$1.35
GLM 4.7
Z.aiopen
34.5$0.16
M2.7
MiniMaxopen
38.9$0.10