Independent model analysis
Which model, and what does it cost you?
323 models from 51 creators, measured on intelligence, output speed, latency and the price of getting real work done. No vendor benchmarks, no marketing numbers.
Top Intelligence Index
63.1
Claude Opus 5 · Anthropic
Fastest output
3,502 t/s
V3 Fast
Cheapest benchmarked
$0.015
Ling-2.6-flash · per 1M tokens
Models benchmarked
161
of 323 indexed
Highlights
Intelligence, speed and cost
The three numbers that decide most model choices. Every chart on this site ships a table view; hover a bar for the rest of a model's profile.
Intelligence
Intelligence Index · higher is better
- Claude Opus 5Anthropic63.1
- Claude Fable 5Anthropic62.1
- GPT-5.6 SolOpenAI60.9
- Kimi K3Moonshot AI59.7
- Qwen3.8 MaxQwen58.1
- Claude Opus 4.8Anthropic57.3
- Muse Spark 1.2Meta56.8
- GPT-5.6 TerraOpenAI56.6
- GPT-5.5OpenAI56.3
- Grok 4.5xAI55.8
Composite of ten reasoning, knowledge, coding and agentic evaluations.
View as table
| Model | Index |
|---|---|
| Claude Opus 5 (Anthropic) | 63.1 |
| Claude Fable 5 (Anthropic) | 62.1 |
| GPT-5.6 Sol (OpenAI) | 60.9 |
| Kimi K3 (Moonshot AI) | 59.7 |
| Qwen3.8 Max (Qwen) | 58.1 |
| Claude Opus 4.8 (Anthropic) | 57.3 |
| Muse Spark 1.2 (Meta) | 56.8 |
| GPT-5.6 Terra (OpenAI) | 56.6 |
| GPT-5.5 (OpenAI) | 56.3 |
| Grok 4.5 (xAI) | 55.8 |
Speed
Output tokens per second · higher is better
- V3 FastMorph3,502 t/s
- V3 LargeMorph3,129 t/s
- Apply 3Relace3,080 t/s
- gpt-oss-120bOpenAI773 t/s
- GLM 4.7Z.ai502 t/s
- 443 t/s
- Qwen3 32BQwen423 t/s
- gpt-oss-safeguard-20bOpenAI422 t/s
- 337 t/s
- M2.7MiniMax307 t/s
Median output tokens per second across serving providers.
View as table
| Model | Tokens/s |
|---|---|
| V3 Fast (Morph) | 3502 |
| V3 Large (Morph) | 3129 |
| Apply 3 (Relace) | 3080 |
| gpt-oss-120b (OpenAI) | 773 |
| GLM 4.7 (Z.ai) | 502 |
| Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image) (Google) | 443 |
| Qwen3 32B (Qwen) | 423 |
| gpt-oss-safeguard-20b (OpenAI) | 422 |
| Grok 4.20 Multi-Agent (xAI) | 337 |
| M2.7 (MiniMax) | 307 |
Cost per task
Estimated USD per task · lower is better
- Ling-3.0-flashInclusionAI$0.006
- V4 Flash 0423DeepSeek$0.02
- V4 Flash 0731DeepSeek$0.02
- Hy3 previewTencent$0.02
- MiMo-V2.5Xiaomi$0.03
- KAT-Coder-Pro V2KwaiPilot$0.04
- Hy3Tencent$0.05
- GPT-5.6 LunaOpenAI$0.05
- Ring-2.6-1TInclusionAI$0.05
- M2.5MiniMax$0.08
Estimated from list pricing: 50K input tokens plus 80K output tokens for reasoning models (25K for non-reasoning). Limited to models scoring 30+ on the Intelligence Index, so the ranking compares like with like.
View as table
| Model | USD |
|---|---|
| Ling-3.0-flash (InclusionAI) | $0.006 |
| V4 Flash 0423 (DeepSeek) | $0.02 |
| V4 Flash 0731 (DeepSeek) | $0.02 |
| Hy3 preview (Tencent) | $0.02 |
| MiMo-V2.5 (Xiaomi) | $0.03 |
| KAT-Coder-Pro V2 (KwaiPilot) | $0.04 |
| Hy3 (Tencent) | $0.05 |
| GPT-5.6 Luna (OpenAI) | $0.05 |
| Ring-2.6-1T (InclusionAI) | $0.05 |
| M2.5 (MiniMax) | $0.08 |
Latency
Seconds to first token · lower is better
- Granite 4.1 8BIBM Granite118ms
- 138ms
- Llama 3 8B LunarisSao10K158ms
- gpt-oss-120bOpenAI205ms
- Qwen3 32BQwen205ms
- Codestral 2508Mistral AI208ms
- 212ms
- MythoMax 13BGryphe222ms
- Hermes 4 70BNous Research229ms
- gpt-oss-safeguard-20bOpenAI237ms
Median time to first token.
View as table
| Model | Seconds |
|---|---|
| Granite 4.1 8B (IBM Granite) | 118ms |
| Llama Guard 4 12B (Meta) | 138ms |
| Llama 3 8B Lunaris (Sao10K) | 158ms |
| gpt-oss-120b (OpenAI) | 205ms |
| Qwen3 32B (Qwen) | 205ms |
| Codestral 2508 (Mistral AI) | 208ms |
| Llama 3.2 3B Instruct (Meta) | 212ms |
| MythoMax 13B (Gryphe) | 222ms |
| Hermes 4 70B (Nous Research) | 229ms |
| gpt-oss-safeguard-20b (OpenAI) | 237ms |
Blended price
USD per 1M tokens, 3:1 input:output · lower is better
- Ling-2.6-flashInclusionAI$0.015
- Ling-3.0-flashInclusionAI$0.032
- gpt-oss-20bOpenAI$0.055
- Granite 4.1 8BIBM Granite$0.063
- gpt-oss-120bOpenAI$0.07
- Nemotron 3 Nano 30B A3BNVIDIA$0.088
- Phi 4Microsoft$0.088
- Hy3 previewTencent$0.1
- V4 Flash 0423DeepSeek$0.11
- V4 Flash 0731DeepSeek$0.113
3:1 weighted blend of input and output list price. Free endpoints are excluded; only benchmarked models appear.
View as table
| Model | USD / 1M |
|---|---|
| Ling-2.6-flash (InclusionAI) | $0.015 |
| Ling-3.0-flash (InclusionAI) | $0.032 |
| gpt-oss-20b (OpenAI) | $0.055 |
| Granite 4.1 8B (IBM Granite) | $0.063 |
| gpt-oss-120b (OpenAI) | $0.07 |
| Nemotron 3 Nano 30B A3B (NVIDIA) | $0.088 |
| Phi 4 (Microsoft) | $0.088 |
| Hy3 preview (Tencent) | $0.1 |
| V4 Flash 0423 (DeepSeek) | $0.11 |
| V4 Flash 0731 (DeepSeek) | $0.113 |
Trade-offs
Nothing is free
Intelligence, speed and price pull against each other. These two plots are where the actual decision gets made.
Intelligence vs price
Blended USD per 1M tokens on a log scale · up and to the right is a better deal
- Proprietary
- Open weights
Price axis is inverted so cheaper sits right. Labelled points are the Pareto frontier — nothing in the index is both smarter and cheaper. Only models with both a published price and an Intelligence Index score appear.
View as table
| Model | Intelligence | Blended / 1M |
|---|---|---|
| Claude Opus 5 (Anthropic) | 63.1 | $10 |
| Claude Fable 5 (Anthropic) | 62.1 | $20 |
| GPT-5.6 Sol (OpenAI) | 60.9 | $11.25 |
| Kimi K3 (Moonshot AI) | 59.7 | $6 |
| Qwen3.8 Max (Qwen) | 58.1 | $3 |
| Claude Opus 4.8 (Anthropic) | 57.3 | $10 |
| Muse Spark 1.2 (Meta) | 56.8 | $2 |
| GPT-5.6 Terra (OpenAI) | 56.6 | $2.25 |
| GPT-5.5 (OpenAI) | 56.3 | $11.25 |
| Grok 4.5 (xAI) | 55.8 | $3 |
| Claude Sonnet 5 (Anthropic) | 55.3 | $4 |
| Claude Opus 4.7 (Anthropic) | 55.0 | $10 |
| Muse Spark 1.1 (Meta) | 53.2 | $2 |
| GPT-5.4 (OpenAI) | 53.1 | $5.63 |
| GLM 5.2 (Z.ai) | 52.6 | $1.18 |
| GPT-5.6 Luna (OpenAI) | 52.3 | $0.225 |
| Gemini 3.5 Flash (Google) | 52.0 | $3.38 |
| V4 Flash 0731 (DeepSeek) | 51.8 | $0.113 |
| V4 Flash 0423 (DeepSeek) | 51.8 | $0.11 |
| Gemini 3.6 Flash (Google) | 51.6 | $3 |
| Gemini 3.1 Pro Preview (Google) | 47.7 | $4.5 |
| Qwen3.7 Max (Qwen) | 46.7 | $2.21 |
| GPT-5.3-Codex (OpenAI) | 45.5 | $4.81 |
| MiniMax M3 (MiniMax) | 45.4 | $0.525 |
| V4 Pro (DeepSeek) | 45.3 | $0.544 |
Intelligence vs output speed
Median tokens per second · up and to the right is both smart and fast
- Proprietary
- Open weights
Labelled points are the Pareto frontier — nothing is both smarter and faster. Output speed is the median across serving providers, so a model served by many reports a blended figure.
View as table
| Model | Intelligence | Tokens/s |
|---|---|---|
| Claude Opus 5 (Anthropic) | 63.1 | 87 |
| Claude Fable 5 (Anthropic) | 62.1 | 54 |
| GPT-5.6 Sol (OpenAI) | 60.9 | 46 |
| Kimi K3 (Moonshot AI) | 59.7 | 71 |
| Qwen3.8 Max (Qwen) | 58.1 | 38 |
| Claude Opus 4.8 (Anthropic) | 57.3 | 64 |
| Muse Spark 1.2 (Meta) | 56.8 | 125 |
| GPT-5.6 Terra (OpenAI) | 56.6 | 56 |
| GPT-5.5 (OpenAI) | 56.3 | 85 |
| Grok 4.5 (xAI) | 55.8 | 52 |
| Claude Sonnet 5 (Anthropic) | 55.3 | 102 |
| Claude Opus 4.7 (Anthropic) | 55.0 | 77 |
| Muse Spark 1.1 (Meta) | 53.2 | 159 |
| GPT-5.4 (OpenAI) | 53.1 | 120 |
| GLM 5.2 (Z.ai) | 52.6 | 142 |
| GPT-5.6 Luna (OpenAI) | 52.3 | 88 |
| Gemini 3.5 Flash (Google) | 52.0 | 182 |
| V4 Flash 0731 (DeepSeek) | 51.8 | 148 |
| V4 Flash 0423 (DeepSeek) | 51.8 | 87 |
| Gemini 3.6 Flash (Google) | 51.6 | 108 |
| Gemini 3.1 Pro Preview (Google) | 47.7 | 110 |
| Qwen3.7 Max (Qwen) | 46.7 | 47 |
| GPT-5.3-Codex (OpenAI) | 45.5 | 61 |
| MiniMax M3 (MiniMax) | 45.4 | 84 |
| V4 Pro (DeepSeek) | 45.3 | 71 |
AI trends
The frontier, month by month
Best Intelligence Index score available at any point in time. The line steps only when a release actually beats the record.
Frontier Intelligence Index
Best score available on any given date · one series, so no legend is needed
A step means a new record holder; a flat run means nobody beat it.
View as table
| Date | Model | Index |
|---|---|---|
| May 2023 | GPT-4 | 6.8 |
| Apr 2024 | GPT-4 Turbo | 7.7 |
| May 2024 | GPT-4o | 11.1 |
| Dec 2024 | o1 | 23.9 |
| Apr 2025 | o4 Mini | 26.1 |
| Apr 2025 | o3 | 31.1 |
| Jun 2025 | o3 Pro | 33.3 |
| Aug 2025 | GPT-5 | 35.3 |
| Nov 2025 | GPT-5.1-Codex | 35.6 |
| Nov 2025 | GPT-5.1 | 37.5 |
| Dec 2025 | GPT-5.2 | 43.3 |
| Feb 2026 | Gemini 3.1 Pro Preview | 47.7 |
| Mar 2026 | GPT-5.4 | 53.1 |
| Apr 2026 | Claude Opus 4.7 | 55.0 |
| Apr 2026 | GPT-5.5 | 56.3 |
| May 2026 | Claude Opus 4.8 | 57.3 |
| Jun 2026 | Claude Fable 5 | 62.1 |
| Jul 2026 | Claude Opus 5 | 63.1 |
Explore
Seven ways into the data
Models
Every model in the index, filterable by capability, creator and price, with the full metric table.
OpenCoding Agents
Agentic and coding indices, terminal work, tool use, and head-to-head full-stack build quality.
OpenSpeech / Image / Video
Which models hear audio, read video, and emit images — and what that capability costs.
OpenInference
Serving providers ranked on output speed, time to first token and blended price.
OpenLeaderboards
One ranked board per benchmark: GPQA, Humanity's Last Exam, SciCode, τ²-bench, long context.
OpenAI Trends
How the frontier moved, and how fast the price of a given intelligence level fell.
OpenArena
Pairwise Elo across websites, UI components, data visualisation, 3D and game development.
OpenStart from the number that matters to you
Cheapest model that clears a quality bar, fastest model that can still use tools, or the single best score on one benchmark — every ranking is one click away.
Cost per task figures are estimates derived from list pricing — see methodology. Example: $2.25 for Claude Opus 5.