Skip to content
llmwaves

Independent model analysis

Which model, and what does it cost you?

323 models from 51 creators, measured on intelligence, output speed, latency and the price of getting real work done. No vendor benchmarks, no marketing numbers.

Top Intelligence Index

63.1

Claude Opus 5 · Anthropic

Fastest output

3,502 t/s

V3 Fast

Cheapest benchmarked

$0.015

Ling-2.6-flash · per 1M tokens

Models benchmarked

161

of 323 indexed

Highlights

Intelligence, speed and cost

The three numbers that decide most model choices. Every chart on this site ships a table view; hover a bar for the rest of a model's profile.

Intelligence

Intelligence Index · higher is better

Composite of ten reasoning, knowledge, coding and agentic evaluations.

View as table
ModelIndex
Claude Opus 5 (Anthropic)63.1
Claude Fable 5 (Anthropic)62.1
GPT-5.6 Sol (OpenAI)60.9
Kimi K3 (Moonshot AI)59.7
Qwen3.8 Max (Qwen)58.1
Claude Opus 4.8 (Anthropic)57.3
Muse Spark 1.2 (Meta)56.8
GPT-5.6 Terra (OpenAI)56.6
GPT-5.5 (OpenAI)56.3
Grok 4.5 (xAI)55.8

Speed

Output tokens per second · higher is better

Median output tokens per second across serving providers.

View as table
ModelTokens/s
V3 Fast (Morph)3502
V3 Large (Morph)3129
Apply 3 (Relace)3080
gpt-oss-120b (OpenAI)773
GLM 4.7 (Z.ai)502
Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image) (Google)443
Qwen3 32B (Qwen)423
gpt-oss-safeguard-20b (OpenAI)422
Grok 4.20 Multi-Agent (xAI)337
M2.7 (MiniMax)307

Cost per task

Estimated USD per task · lower is better

Estimated from list pricing: 50K input tokens plus 80K output tokens for reasoning models (25K for non-reasoning). Limited to models scoring 30+ on the Intelligence Index, so the ranking compares like with like.

View as table
ModelUSD
Ling-3.0-flash (InclusionAI)$0.006
V4 Flash 0423 (DeepSeek)$0.02
V4 Flash 0731 (DeepSeek)$0.02
Hy3 preview (Tencent)$0.02
MiMo-V2.5 (Xiaomi)$0.03
KAT-Coder-Pro V2 (KwaiPilot)$0.04
Hy3 (Tencent)$0.05
GPT-5.6 Luna (OpenAI)$0.05
Ring-2.6-1T (InclusionAI)$0.05
M2.5 (MiniMax)$0.08

Latency

Seconds to first token · lower is better

Median time to first token.

View as table
ModelSeconds
Granite 4.1 8B (IBM Granite)118ms
Llama Guard 4 12B (Meta)138ms
Llama 3 8B Lunaris (Sao10K)158ms
gpt-oss-120b (OpenAI)205ms
Qwen3 32B (Qwen)205ms
Codestral 2508 (Mistral AI)208ms
Llama 3.2 3B Instruct (Meta)212ms
MythoMax 13B (Gryphe)222ms
Hermes 4 70B (Nous Research)229ms
gpt-oss-safeguard-20b (OpenAI)237ms

Blended price

USD per 1M tokens, 3:1 input:output · lower is better

3:1 weighted blend of input and output list price. Free endpoints are excluded; only benchmarked models appear.

View as table
ModelUSD / 1M
Ling-2.6-flash (InclusionAI)$0.015
Ling-3.0-flash (InclusionAI)$0.032
gpt-oss-20b (OpenAI)$0.055
Granite 4.1 8B (IBM Granite)$0.063
gpt-oss-120b (OpenAI)$0.07
Nemotron 3 Nano 30B A3B (NVIDIA)$0.088
Phi 4 (Microsoft)$0.088
Hy3 preview (Tencent)$0.1
V4 Flash 0423 (DeepSeek)$0.11
V4 Flash 0731 (DeepSeek)$0.113

Trade-offs

Nothing is free

Intelligence, speed and price pull against each other. These two plots are where the actual decision gets made.

Intelligence vs price

Blended USD per 1M tokens on a log scale · up and to the right is a better deal

  • Proprietary
  • Open weights
0102030405060$0.01$0.1$1$10$100$1000Blended price, USD / 1M tokens · better →Intelligence IndexClaude Opus 5Kimi K3Qwen3.8 MaxMuse Spark 1.2GLM 5.2GPT-5.6 Luna

Price axis is inverted so cheaper sits right. Labelled points are the Pareto frontier — nothing in the index is both smarter and cheaper. Only models with both a published price and an Intelligence Index score appear.

View as table
ModelIntelligenceBlended / 1M
Claude Opus 5 (Anthropic)63.1$10
Claude Fable 5 (Anthropic)62.1$20
GPT-5.6 Sol (OpenAI)60.9$11.25
Kimi K3 (Moonshot AI)59.7$6
Qwen3.8 Max (Qwen)58.1$3
Claude Opus 4.8 (Anthropic)57.3$10
Muse Spark 1.2 (Meta)56.8$2
GPT-5.6 Terra (OpenAI)56.6$2.25
GPT-5.5 (OpenAI)56.3$11.25
Grok 4.5 (xAI)55.8$3
Claude Sonnet 5 (Anthropic)55.3$4
Claude Opus 4.7 (Anthropic)55.0$10
Muse Spark 1.1 (Meta)53.2$2
GPT-5.4 (OpenAI)53.1$5.63
GLM 5.2 (Z.ai)52.6$1.18
GPT-5.6 Luna (OpenAI)52.3$0.225
Gemini 3.5 Flash (Google)52.0$3.38
V4 Flash 0731 (DeepSeek)51.8$0.113
V4 Flash 0423 (DeepSeek)51.8$0.11
Gemini 3.6 Flash (Google)51.6$3
Gemini 3.1 Pro Preview (Google)47.7$4.5
Qwen3.7 Max (Qwen)46.7$2.21
GPT-5.3-Codex (OpenAI)45.5$4.81
MiniMax M3 (MiniMax)45.4$0.525
V4 Pro (DeepSeek)45.3$0.544

Intelligence vs output speed

Median tokens per second · up and to the right is both smart and fast

  • Proprietary
  • Open weights
01020304050600100200300400500600700Output speed, tokens/sIntelligence IndexClaude Opus 5Muse Spark 1.2Muse Spark 1.1Gemini 3.5 FlashM2.7GLM 4.7

Labelled points are the Pareto frontier — nothing is both smarter and faster. Output speed is the median across serving providers, so a model served by many reports a blended figure.

View as table
ModelIntelligenceTokens/s
Claude Opus 5 (Anthropic)63.187
Claude Fable 5 (Anthropic)62.154
GPT-5.6 Sol (OpenAI)60.946
Kimi K3 (Moonshot AI)59.771
Qwen3.8 Max (Qwen)58.138
Claude Opus 4.8 (Anthropic)57.364
Muse Spark 1.2 (Meta)56.8125
GPT-5.6 Terra (OpenAI)56.656
GPT-5.5 (OpenAI)56.385
Grok 4.5 (xAI)55.852
Claude Sonnet 5 (Anthropic)55.3102
Claude Opus 4.7 (Anthropic)55.077
Muse Spark 1.1 (Meta)53.2159
GPT-5.4 (OpenAI)53.1120
GLM 5.2 (Z.ai)52.6142
GPT-5.6 Luna (OpenAI)52.388
Gemini 3.5 Flash (Google)52.0182
V4 Flash 0731 (DeepSeek)51.8148
V4 Flash 0423 (DeepSeek)51.887
Gemini 3.6 Flash (Google)51.6108
Gemini 3.1 Pro Preview (Google)47.7110
Qwen3.7 Max (Qwen)46.747
GPT-5.3-Codex (OpenAI)45.561
MiniMax M3 (MiniMax)45.484
V4 Pro (DeepSeek)45.371

AI trends

The frontier, month by month

Best Intelligence Index score available at any point in time. The line steps only when a release actually beats the record.

Frontier Intelligence Index

Best score available on any given date · one series, so no legend is needed

0102030405060202420252026Intelligence Index

A step means a new record holder; a flat run means nobody beat it.

View as table
DateModelIndex
May 2023GPT-46.8
Apr 2024GPT-4 Turbo7.7
May 2024GPT-4o11.1
Dec 2024o123.9
Apr 2025o4 Mini26.1
Apr 2025o331.1
Jun 2025o3 Pro33.3
Aug 2025GPT-535.3
Nov 2025GPT-5.1-Codex35.6
Nov 2025GPT-5.137.5
Dec 2025GPT-5.243.3
Feb 2026Gemini 3.1 Pro Preview47.7
Mar 2026GPT-5.453.1
Apr 2026Claude Opus 4.755.0
Apr 2026GPT-5.556.3
May 2026Claude Opus 4.857.3
Jun 2026Claude Fable 562.1
Jul 2026Claude Opus 563.1

Start from the number that matters to you

Cheapest model that clears a quality bar, fastest model that can still use tools, or the single best score on one benchmark — every ranking is one click away.

Cost per task figures are estimates derived from list pricing — see methodology. Example: $2.25 for Claude Opus 5.