Coding agents
Which models survive an agent loop
Writing a function is easy. Holding a plan across fifty tool calls, reading a failing test and fixing the right file is not. These are the measurements that separate the two.
Tool-capable models
260
Scored on agentic work
80
Best agentic score
59.2
Claude Opus 5
Best Terminal-Bench
66%
GPT-5.6 Sol
Indices
Coding and agentic ability
Two composites. Coding covers writing and fixing code; agentic covers planning, tool use and long-horizon work.
Coding Index
SciCode, LiveCodeBench and Terminal-Bench Hard combined · higher is better
- Claude Opus 5Anthropic78.0
- GPT-5.6 SolOpenAI77.4
- GPT-5.6 TerraOpenAI76.7
- Claude Fable 5Anthropic76.5
- Kimi K3Moonshot AI76.2
- GPT-5.5OpenAI74.9
- Claude Opus 4.8Anthropic74.3
- Claude Opus 4.7Anthropic73.6
- Grok 4.5xAI72.4
- Muse Spark 1.2Meta72.2
View as table
| Model | Coding Index |
|---|---|
| Claude Opus 5 (Anthropic) | 78.0 |
| GPT-5.6 Sol (OpenAI) | 77.4 |
| GPT-5.6 Terra (OpenAI) | 76.7 |
| Claude Fable 5 (Anthropic) | 76.5 |
| Kimi K3 (Moonshot AI) | 76.2 |
| GPT-5.5 (OpenAI) | 74.9 |
| Claude Opus 4.8 (Anthropic) | 74.3 |
| Claude Opus 4.7 (Anthropic) | 73.6 |
| Grok 4.5 (xAI) | 72.4 |
| Muse Spark 1.2 (Meta) | 72.2 |
Agentic Index
Tool use, planning and terminal work · higher is better
- Claude Opus 5Anthropic59.2
- Qwen3.8 MaxQwen58.4
- GPT-5.6 SolOpenAI57.8
- Claude Fable 5Anthropic56.6
- Kimi K3Moonshot AI54.3
- GPT-5.6 TerraOpenAI50.2
- Claude Sonnet 5Anthropic49.7
- Claude Opus 4.8Anthropic49.4
- Grok 4.5xAI48.9
- V4 Flash 0731DeepSeek48.4
View as table
| Model | Agentic Index |
|---|---|
| Claude Opus 5 (Anthropic) | 59.2 |
| Qwen3.8 Max (Qwen) | 58.4 |
| GPT-5.6 Sol (OpenAI) | 57.8 |
| Claude Fable 5 (Anthropic) | 56.6 |
| Kimi K3 (Moonshot AI) | 54.3 |
| GPT-5.6 Terra (OpenAI) | 50.2 |
| Claude Sonnet 5 (Anthropic) | 49.7 |
| Claude Opus 4.8 (Anthropic) | 49.4 |
| Grok 4.5 (xAI) | 48.9 |
| V4 Flash 0731 (DeepSeek) | 48.4 |
Economics
Agentic ability against what it costs
Agent loops burn output tokens. A model that scores two points higher and costs four times as much is not obviously the right call.
Agentic Index vs cost per task
Estimated USD per task on a log scale · up and to the right is the better deal
- Proprietary
- Open weights
Labelled points are the Pareto frontier. Estimated from list pricing: 50K input tokens plus 80K output tokens for reasoning models (25K for non-reasoning).
View as table
| Model | Agentic | Cost / task |
|---|---|---|
| Claude Opus 5 (Anthropic) | 59.2 | $2.25 |
| Qwen3.8 Max (Qwen) | 58.4 | $0.58 |
| GPT-5.6 Sol (OpenAI) | 57.8 | $2.65 |
| Claude Fable 5 (Anthropic) | 56.6 | $4.50 |
| Kimi K3 (Moonshot AI) | 54.3 | $1.35 |
| GPT-5.6 Terra (OpenAI) | 50.2 | $0.53 |
| Claude Sonnet 5 (Anthropic) | 49.7 | $0.90 |
| Claude Opus 4.8 (Anthropic) | 49.4 | $2.25 |
| Grok 4.5 (xAI) | 48.9 | $0.58 |
| V4 Flash 0731 (DeepSeek) | 48.4 | $0.02 |
| V4 Flash 0423 (DeepSeek) | 48.4 | $0.02 |
| GPT-5.5 (OpenAI) | 47.4 | $2.65 |
| GPT-5.6 Luna (OpenAI) | 46.9 | $0.05 |
| Claude Opus 4.7 (Anthropic) | 46.3 | $2.25 |
| GLM 5.2 (Z.ai) | 45.7 | $0.23 |
| GPT-5.4 (OpenAI) | 44.2 | $1.32 |
| Claude Sonnet 4.6 (Anthropic) | 42.1 | $1.35 |
| Gemini 3.6 Flash (Google) | 40.5 | $0.68 |
| Muse Spark 1.1 (Meta) | 39.7 | $0.40 |
| Gemini 3.5 Flash (Google) | 39.7 | $0.79 |
| V4 Pro (DeepSeek) | 37.8 | $0.09 |
| MiniMax M3 (MiniMax) | 36.1 | $0.11 |
| Inkling (Thinking Machines) | 34.1 | $0.37 |
| Inkling Small (Thinking Machines) | 31.9 | $0.12 |
| GPT-5.4 Mini (OpenAI) | 31.5 | $0.40 |
Underlying evaluations
The four that matter for agents
Terminal work, contest programming, research-grade scientific code, and tool-use discipline under a policy.
Terminal-Bench Hard
Percentage correct · higher is better
- GPT-5.6 SolOpenAI65.9%
- Claude Fable 5Anthropic62.9%
- GPT-5.5OpenAI60.6%
- Claude Opus 4.8Anthropic58.3%
- GPT-5.6 TerraOpenAI57.6%
- GPT-5.4OpenAI57.6%
- Gemini 3.1 Pro PreviewGoogle53.8%
- GPT-5.3-CodexOpenAI53.0%
- GPT-5.4 MiniOpenAI52.3%
- Claude Opus 4.7Anthropic51.5%
View as table
| Model | Score |
|---|---|
| GPT-5.6 Sol (OpenAI) | 65.9% |
| Claude Fable 5 (Anthropic) | 62.9% |
| GPT-5.5 (OpenAI) | 60.6% |
| Claude Opus 4.8 (Anthropic) | 58.3% |
| GPT-5.6 Terra (OpenAI) | 57.6% |
| GPT-5.4 (OpenAI) | 57.6% |
| Gemini 3.1 Pro Preview (Google) | 53.8% |
| GPT-5.3-Codex (OpenAI) | 53.0% |
| GPT-5.4 Mini (OpenAI) | 52.3% |
| Claude Opus 4.7 (Anthropic) | 51.5% |
τ²-bench (Telecom)
Percentage correct · higher is better
- GLM 5.2Z.ai99.1%
- GLM 4.7 FlashZ.ai98.8%
- Claude Fable 5Anthropic98.5%
- Step 3.7 FlashStepFun98.5%
- GLM 5V TurboZ.ai98.5%
- GLM 5 TurboZ.ai98.5%
- GLM 5Z.ai98.2%
- Grok 4.3xAI97.7%
- GLM 5.1Z.ai97.7%
- Qwen3.6 PlusQwen97.7%
View as table
| Model | Score |
|---|---|
| GLM 5.2 (Z.ai) | 99.1% |
| GLM 4.7 Flash (Z.ai) | 98.8% |
| Claude Fable 5 (Anthropic) | 98.5% |
| Step 3.7 Flash (StepFun) | 98.5% |
| GLM 5V Turbo (Z.ai) | 98.5% |
| GLM 5 Turbo (Z.ai) | 98.5% |
| GLM 5 (Z.ai) | 98.2% |
| Grok 4.3 (xAI) | 97.7% |
| GLM 5.1 (Z.ai) | 97.7% |
| Qwen3.6 Plus (Qwen) | 97.7% |
LiveCodeBench
Percentage correct · higher is better
- GLM 4.7Z.ai89.4%
- GPT-5.2OpenAI88.9%
- gpt-oss-120bOpenAI87.8%
- GPT-5.1OpenAI86.8%
- o4 Mini HighOpenAI85.9%
- o4 MiniOpenAI85.9%
- Kimi K2 ThinkingMoonshot AI85.3%
- GPT-5.1-CodexOpenAI84.9%
- GPT-5OpenAI84.6%
- GPT-5 MiniOpenAI83.8%
View as table
| Model | Score |
|---|---|
| GLM 4.7 (Z.ai) | 89.4% |
| GPT-5.2 (OpenAI) | 88.9% |
| gpt-oss-120b (OpenAI) | 87.8% |
| GPT-5.1 (OpenAI) | 86.8% |
| o4 Mini High (OpenAI) | 85.9% |
| o4 Mini (OpenAI) | 85.9% |
| Kimi K2 Thinking (Moonshot AI) | 85.3% |
| GPT-5.1-Codex (OpenAI) | 84.9% |
| GPT-5 (OpenAI) | 84.6% |
| GPT-5 Mini (OpenAI) | 83.8% |
SciCode
Percentage correct · higher is better
- Claude Fable 5Anthropic60.2%
- Gemini 3.1 Pro PreviewGoogle58.9%
- Kimi K3Moonshot AI58.7%
- Muse Spark 1.1Meta58.2%
- GPT-5.4OpenAI56.6%
- Muse Spark 1.2Meta56.4%
- GPT-5.6 SolOpenAI56.1%
- GPT-5.5OpenAI56.1%
- Claude Opus 5Anthropic55.7%
- GPT-5.2-CodexOpenAI54.6%
View as table
| Model | Score |
|---|---|
| Claude Fable 5 (Anthropic) | 60.2% |
| Gemini 3.1 Pro Preview (Google) | 58.9% |
| Kimi K3 (Moonshot AI) | 58.7% |
| Muse Spark 1.1 (Meta) | 58.2% |
| GPT-5.4 (OpenAI) | 56.6% |
| Muse Spark 1.2 (Meta) | 56.4% |
| GPT-5.6 Sol (OpenAI) | 56.1% |
| GPT-5.5 (OpenAI) | 56.1% |
| Claude Opus 5 (Anthropic) | 55.7% |
| GPT-5.2-Codex (OpenAI) | 54.6% |
Agent arena
Building whole applications
Elo from pairwise judgements on complete builds, not snippets — the closest proxy for what an agent is actually asked to do.
Full-stack apps
Arena Elo · dot plot on a stated scale, not bars from zero
- Claude Fable 5Anthropic1,295
- Claude Opus 4.8Anthropic1,285
- GLM 5.2Z.ai1,274
- Claude Sonnet 5Anthropic1,272
- Claude Opus 4.6Anthropic1,262
- Claude Sonnet 4.6Anthropic1,252
- Qwen3.7 MaxQwen1,240
- Gemini 3.5 FlashGoogle1,234
- MiniMax M3MiniMax1,226
- Kimi K2.7 CodeMoonshot AI1,218
View as table
| Model | Elo | Win rate |
|---|---|---|
| Claude Fable 5 | 1295 | 63.7% |
| Claude Opus 4.8 | 1285 | 62.2% |
| GLM 5.2 | 1274 | 63.8% |
| Claude Sonnet 5 | 1272 | 59.8% |
| Claude Opus 4.6 | 1262 | 68.5% |
| Claude Sonnet 4.6 | 1252 | 64.1% |
| Qwen3.7 Max | 1240 | 55.6% |
| Gemini 3.5 Flash | 1234 | 56.7% |
| MiniMax M3 | 1226 | 53.0% |
| Kimi K2.7 Code | 1218 | 54.5% |
Mobile apps
Arena Elo · dot plot on a stated scale, not bars from zero
- Claude Opus 4.6Anthropic1,279
- Claude Sonnet 4.6Anthropic1,271
- Claude Opus 4.8Anthropic1,255
- Claude Fable 5Anthropic1,254
- Claude Sonnet 5Anthropic1,237
- MiniMax M3MiniMax1,237
- Kimi K2.6Moonshot AI1,221
- GLM 5.2Z.ai1,219
- Gemini 3.5 FlashGoogle1,218
- GLM 5.1Z.ai1,218
View as table
| Model | Elo | Win rate |
|---|---|---|
| Claude Opus 4.6 | 1279 | 64.8% |
| Claude Sonnet 4.6 | 1271 | 63.4% |
| Claude Opus 4.8 | 1255 | 57.5% |
| Claude Fable 5 | 1254 | 57.5% |
| Claude Sonnet 5 | 1237 | 55.1% |
| MiniMax M3 | 1237 | 55.4% |
| Kimi K2.6 | 1221 | 54.9% |
| GLM 5.2 | 1219 | 54.0% |
| Gemini 3.5 Flash | 1218 | 53.5% |
| GLM 5.1 | 1218 | 54.7% |
Full list
260 tool-capable models
Every model that can call a tool, which is the floor for running an agent at all.
| Claude Opus 5 Anthropic | 63.1 | $2.25 |
| Qwen3.8 Max Qwen | 58.1 | $0.58 |
| GPT-5.6 Sol OpenAI | 60.9 | $2.65 |
| Claude Fable 5 Anthropic | 62.1 | $4.50 |
| Kimi K3 Moonshot AIopen | 59.7 | $1.35 |
| GPT-5.6 Terra OpenAI | 56.6 | $0.53 |
| Claude Sonnet 5 Anthropic | 55.3 | $0.90 |
| Claude Opus 4.8 Anthropic | 57.3 | $2.25 |
| Grok 4.5 xAI | 55.8 | $0.58 |
| V4 Flash 0731 DeepSeekopen | 51.8 | $0.02 |
| V4 Flash 0423 DeepSeekopen | 51.8 | $0.02 |
| GPT-5.5 OpenAI | 56.3 | $2.65 |
| GPT-5.6 Luna OpenAI | 52.3 | $0.05 |
| Claude Opus 4.7 Anthropic | 55.0 | $2.25 |
| GLM 5.2 Z.aiopen | 52.6 | $0.23 |
| GPT-5.4 OpenAI | 53.1 | $1.32 |
| Claude Sonnet 4.6 Anthropic | 36.8 | $1.35 |
| Gemini 3.6 Flash Google | 51.6 | $0.68 |
| Muse Spark 1.1 Meta | 53.2 | $0.40 |
| Gemini 3.5 Flash Google | 52.0 | $0.79 |
| V4 Pro DeepSeekopen | 45.3 | $0.09 |
| MiniMax M3 MiniMaxopen | 45.4 | $0.11 |
| Inkling Thinking Machinesopen | 42.3 | $0.37 |
| Inkling Small Thinking Machinesopen | 41.2 | $0.12 |
| GPT-5.4 Mini OpenAI | 40.9 | $0.40 |
| Hy3 Tencentopen | 42.2 | $0.05 |
| Hy3 preview Tencentopen | 34.4 | $0.02 |
| Kimi K2.6 Moonshot AIopen | 45.1 | $0.23 |
| Qwen3.7 Max Qwen | 46.7 | $0.43 |
| GLM 5.1 Z.aiopen | 41.0 | $0.29 |
| Kimi K2.7 Code Moonshot AIopen | 43.0 | $0.32 |
| GPT-5.4 Nano OpenAI | 39.7 | $0.11 |
| MiMo-V2.5-Pro Xiaomiopen | 42.9 | $0.09 |
| Qwen3.6 Plus Qwen | 40.5 | $0.17 |
| Qwen3.6 27B Qwenopen | 37.7 | $0.32 |
| Gemini 3.5 Flash Lite Google | 37.4 | $0.22 |
| GPT-5 OpenAI | 35.3 | $0.86 |
| Claude Sonnet 4.5 Anthropic | 29.9 | $1.35 |
| GLM 4.7 Z.aiopen | 34.5 | $0.16 |
| M2.7 MiniMaxopen | 38.9 | $0.10 |