MiMo-V2.5
Xiaomi · released Apr 22, 2026
MiMo-V2.5 is a native omnimodal model by Xiaomi. It delivers Pro-level agentic performance at roughly half the inference cost, while surpassing MiMo-V2-Omni in multimodal perception across image and video understanding...
Specification
- Context window
- 1.05M
- Max output
- 131K
- Knowledge cutoff
- Not stated
- Parameters
- 310.8B
- Licence
- mit
- Serving providers
- 6
- Moderated
- No
- Uptime
- 99.7%
Intelligence
38.0
86th percentile
Coding
56.8
Coding Index
Agentic
24.4
Agentic Index
Output speed
55 t/s
Median across providers
Latency
846ms
Time to first token
Cost per task
$0.03
Estimated
Benchmarks
Where the score comes from
The Intelligence Index is a composite. These are the underlying evaluations this model was actually measured on.
Evaluation scores
Percentage correct · higher is better
- τ²-bench (Telecom)90.6%
- GPQA Diamond84.9%
- AA-LCR (long context)68.3%
- IFBench67.1%
- SciCode43.1%
- Terminal-Bench Hard41.7%
- Humanity's Last Exam27.2%
An evaluation missing from this list was not run for this model — it is not a zero.
View as table
| Evaluation | Score |
|---|---|
| τ²-bench (Telecom) | 90.6% |
| GPQA Diamond | 84.9% |
| AA-LCR (long context) | 68.3% |
| IFBench | 67.1% |
| SciCode | 43.1% |
| Terminal-Bench Hard | 41.7% |
| Humanity's Last Exam | 27.2% |
Against its peers
Intelligence Index · this model highlighted, nearest peers in grey
- Claude Opus 4.6Anthropic38.8
- MiMo-V2.5Xiaomi38.0
- Grok 4.20xAI38.0
- Grok 4.3xAI37.9
- Ling-3.0-flashInclusionAI37.8
- Qwen3.6 27BQwen37.7
- GPT-5.1OpenAI37.5
- Gemini 3.5 Flash LiteGoogle37.4
Peers are the models sitting closest on the Intelligence Index — the set you would realistically choose between.
View as table
| Model | Intelligence |
|---|---|
| Claude Opus 4.6 | 38.8 |
| MiMo-V2.5 | 38.0 |
| Grok 4.20 | 38.0 |
| Grok 4.3 | 37.9 |
| Ling-3.0-flash | 37.8 |
| Qwen3.6 27B | 37.7 |
| GPT-5.1 | 37.5 |
| Gemini 3.5 Flash Lite | 37.4 |
Percentile among all indexed models
Pricing
What it costs to run
List prices per million tokens, plus what one representative task works out to.
List price
- Input / 1M tokens
- $0.14
- Output / 1M tokens
- $0.28
- Cached input / 1M
- $0.003
- Blended 3:1
- $0.175
One task, estimated
$0.03
- Input tokens
- 50,000
- Output tokens
- 80,000
- Profile
- Reasoning
Estimated from list pricing: 50K input tokens plus 80K output tokens for reasoning models (25K for non-reasoning).
Arena
Head-to-head generation quality
Elo from pairwise judgements, broken out by the kind of thing the model was asked to build.
- Overall Elo
- 1,292
- Win rate
- 55.5%
- Strongest at
- Websites
- Tournaments
- 17,542
Elo by category
Dot position on a 1,200–1,300 scale · Elo has no meaningful zero
- 1,292
- 1,292
- 1,281
- 1,281
- 1,270
- 1,217
Agent categories (full-stack apps, mobile apps) are only scored for models tested in agent harnesses.
View as table
| Category | Elo | Win rate |
|---|---|---|
| Websites | 1292 | 55.5% |
| UI components | 1292 | 55.3% |
| Data visualisation | 1281 | 55.4% |
| Game development | 1281 | 55.7% |
| 3D scenes | 1270 | 51.9% |
| SVG | 1217 | 52.5% |