Skip to content
llmwaves

Arena

Judged head to head, not scored on a test

Benchmarks measure whether an answer is correct. The arena measures whether a person prefers it — two models get the same brief, a judge picks a winner, and Elo does the rest.

Models with Elo

100

Categories

11

Highest Elo

1,455

Kimi K3

Tournaments run

663,461

Overall

The standing order

Elo pooled across every category a model competed in, alongside its raw win rate.

Overall Elo

Higher is better · scale runs 1,300–1,475, not from zero

1,3001,475

Elo is relative: a 100-point gap means the higher model wins roughly two thirds of head-to-heads. Zero Elo is meaningless, so this is a dot plot on a stated scale rather than bars from a zero baseline.

View as table
ModelEloWin rate
Kimi K3 (Moonshot AI)145569.4%
Claude Fable 5 (Anthropic)139765.4%
Claude Opus 5 (Anthropic)139365.2%
GPT-5.6 Sol (OpenAI)138062.7%
GLM 5.2 (Z.ai)136359.7%
Gemini 3.6 Flash (Google)134658.0%
Claude Opus 4.7 (Anthropic)133760.0%
Claude Sonnet 5 (Anthropic)133655.5%
GPT-5.5 (OpenAI)133660.7%
Gemini 3.1 Pro Preview (Google)133368.8%
Kimi K2.6 (Moonshot AI)132960.3%
Claude Opus 4.6 (Anthropic)132962.5%
Muse Spark 1.2 (Meta)132255.9%
Qwen3.7 Max (Qwen)132256.8%
MiMo-V2.5-Pro (Xiaomi)131458.7%

Win rate

Share of head-to-heads won · higher is better

40even odds80

Win rate and Elo disagree when a model faced an unusually strong or weak field. The tick at 50% is even odds — anything left of it loses more than it wins.

View as table
ModelWin rate
Gemini 2.5 Pro (Google)71.8%
Kimi K3 (Moonshot AI)69.4%
Gemini 3.1 Pro Preview (Google)68.8%
Nano Banana Pro (Gemini 3 Pro Image Preview) (Google)65.8%
Claude Fable 5 (Anthropic)65.4%
Claude Opus 5 (Anthropic)65.2%
Nano Banana 2 (Gemini 3.1 Flash Image Preview) (Google)64.4%
GPT-5.4 (OpenAI)63.6%
GPT-5 (OpenAI)62.9%
GPT-5.6 Sol (OpenAI)62.7%
Gemini 3 Flash Preview (Google)62.7%
Claude Opus 4.6 (Anthropic)62.5%
M2.1 (MiniMax)60.9%
GPT-5.5 (OpenAI)60.7%
Kimi K2.6 (Moonshot AI)60.3%

Cross-section

Nobody wins everywhere

The same twelve models across every well-covered category. Read down a column to see which brief separates the field.

Elo by model and category

Darker means stronger · blank means the model did not compete

Model3D scenesData visualisationFull-stack appsGame developmentMobile appsSVGUI componentsWebsites
Kimi K3Moonshot AI1,4551,3811,3941,378
Claude Fable 5Anthropic1,3761,3441,2951,3971,2541,3521,3551,327
Claude Opus 5Anthropic1,3901,3931,3841,339
GPT-5.6 SolOpenAI1,3641,3431,3801,3651,3651,344
GLM 5.2Z.ai1,3631,3251,2741,3371,2191,2601,3391,337
Gemini 3.6 FlashGoogle1,3321,3461,2881,3411,322
Claude Opus 4.7Anthropic1,3051,3111,3241,2701,3371,317
Claude Sonnet 5Anthropic1,3081,2641,2721,3361,2371,2381,3121,296
GPT-5.5OpenAI1,2501,2831,1281,3361,2091,2781,2921,280
Gemini 3.1 Pro PreviewGoogle1,2881,2381,1081,2411,1511,3331,2621,255
Kimi K2.6Moonshot AI1,3291,2861,1981,2901,2211,2251,2991,297
Claude Opus 4.6Anthropic1,3291,3101,2621,3201,2791,2751,3201,315
1,108High

Only categories with at least eight competing models are shown, so a single lucky matchup cannot define a column.

View as table
Model3D scenesData visualisationFull-stack appsGame developmentMobile appsSVGUI componentsWebsites
Kimi K31455138113941378
Claude Fable 513761344129513971254135213551327
Claude Opus 51390139313841339
GPT-5.6 Sol136413431380136513651344
GLM 5.213631325127413371219126013391337
Gemini 3.6 Flash13321346128813411322
Claude Opus 4.7130513111324127013371317
Claude Sonnet 513081264127213361237123813121296
GPT-5.512501283112813361209127812921280
Gemini 3.1 Pro Preview12881238110812411151133312621255
Kimi K2.613291286119812901221122512991297
Claude Opus 4.613291310126213201279127513201315

By category

One board per brief

Data visualisation rewards different things than game development. These are the boards that matter if you know what you are building.

3D scenes

Arena Elo · dot plot on a stated scale, not bars from zero

1,3001,475
View as table
ModelEloWin rate
Kimi K3145569.4%
Claude Opus 5139064.7%
Claude Fable 5137662.9%
GPT-5.6 Sol136459.3%
GLM 5.2136359.7%
Gemini 3.6 Flash133254.2%
Kimi K2.6132960.3%
Claude Opus 4.6132962.5%
Qwen3.7 Max132256.8%
V4 Pro131358.6%

Data visualisation

Arena Elo · dot plot on a stated scale, not bars from zero

1,3001,400
View as table
ModelEloWin rate
Claude Opus 5139365.2%
Kimi K3138165.4%
Gemini 3.6 Flash134658.0%
Claude Fable 5134459.3%
GPT-5.6 Sol134357.9%
GLM 5.2132557.1%
Qwen3.7 Max131858.1%
Claude Opus 4.7131158.2%
Claude Opus 4.6131058.7%
Claude Sonnet 4.6130858.2%

Full-stack apps

Arena Elo · dot plot on a stated scale, not bars from zero

1,2001,300
View as table
ModelEloWin rate
Claude Fable 5129563.7%
Claude Opus 4.8128562.2%
GLM 5.2127463.8%
Claude Sonnet 5127259.8%
Claude Opus 4.6126268.5%
Claude Sonnet 4.6125264.1%
Qwen3.7 Max124055.6%
Gemini 3.5 Flash123456.7%
MiniMax M3122653.0%
Kimi K2.7 Code121854.5%

Game development

Arena Elo · dot plot on a stated scale, not bars from zero

1,3001,400
View as table
ModelEloWin rate
Claude Fable 5139765.4%
GPT-5.6 Sol138062.7%
GLM 5.2133759.1%
Claude Sonnet 5133655.5%
GPT-5.5133660.7%
Claude Opus 4.7132460.0%
Muse Spark 1.2132255.9%
Claude Opus 4.6132061.2%
MiMo-V2.5-Pro131458.7%
Qwen3.7 Max131157.6%

Graphic design

Arena Elo · dot plot on a stated scale, not bars from zero

View as table
ModelEloWin rate
Nano Banana 2 (Gemini 3.1 Flash Image Preview)127765.0%
Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image)126855.5%
Nano Banana Pro (Gemini 3 Pro Image Preview)125465.8%
Nano Banana (Gemini 2.5 Flash Image)117952.5%

Image

Arena Elo · dot plot on a stated scale, not bars from zero

View as table
ModelEloWin rate
Nano Banana 2 (Gemini 3.1 Flash Image Preview)129164.4%
Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image)126053.9%
Nano Banana Pro (Gemini 3 Pro Image Preview)123960.4%
Nano Banana (Gemini 2.5 Flash Image)120153.4%

Mobile apps

Arena Elo · dot plot on a stated scale, not bars from zero

1,2001,300
View as table
ModelEloWin rate
Claude Opus 4.6127964.8%
Claude Sonnet 4.6127163.4%
Claude Opus 4.8125557.5%
Claude Fable 5125457.5%
Claude Sonnet 5123755.1%
MiniMax M3123755.4%
Kimi K2.6122154.9%
GLM 5.2121954.0%
Gemini 3.5 Flash121853.5%
GLM 5.1121854.7%

SVG

Arena Elo · dot plot on a stated scale, not bars from zero

1,2501,375
View as table
ModelEloWin rate
GPT-5.6 Sol136564.7%
Claude Fable 5135266.1%
Gemini 3.1 Pro Preview133368.8%
Gemini 3.5 Flash129660.9%
GPT-5.5127858.1%
Claude Opus 4.6127561.1%
Claude Opus 4.7127059.2%
Qwen3.7 Max126559.4%
GLM 5.2126055.5%
GPT-5.6 Terra125451.8%

UI components

Arena Elo · dot plot on a stated scale, not bars from zero

1,3001,400
View as table
ModelEloWin rate
Kimi K3139464.5%
Claude Opus 5138462.0%
GPT-5.6 Sol136559.6%
Claude Fable 5135559.1%
Gemini 3.6 Flash134155.9%
GLM 5.2133958.3%
Claude Opus 4.7133760.0%
Claude Opus 4.6132060.0%
Claude Sonnet 5131255.1%
Qwen3.7 Max130955.5%

Websites

Arena Elo · dot plot on a stated scale, not bars from zero

1,3001,400
View as table
ModelEloWin rate
Kimi K3137863.5%
GPT-5.6 Sol134459.8%
Claude Opus 5133958.9%
GLM 5.2133760.1%
Claude Fable 5132759.8%
Gemini 3.6 Flash132257.8%
Muse Spark 1.2131855.7%
Claude Opus 4.7131759.0%
Claude Opus 4.6131561.4%
Claude Sonnet 4.6130860.1%