Compare
Four models, one page
Pick the models you are actually choosing between. Every row is scaled within itself, so a bar means something next to its neighbours and nothing across rows.
| Specification | Qwen3 VL 30B A3B InstructQwen | Claude Opus 5Anthropic |
|---|---|---|
| Released | Oct 6, 2025 | Jul 24, 2026 |
| Context window | 262K | 1M |
| Max output | 16K | 128K |
| Input / 1M | $0.15 | $5 |
| Output / 1M | $0.6 | $25 |
| Cost per task | $0.02 | $2.25 |
| Arena Elo | — | 1,393 |
| Serving providers | 4 | 5 |
| Parameters | 31.1B | Undisclosed |
| Licence | apache-2.0 | Proprietary |
| Capabilities |
|
|
Metrics side by side
Each row is scaled to the largest value in that row — bars compare within a row, never across rows
- Qwen3 VL 30B A3B Instruct
- Claude Opus 5
Intelligence Index
Qwen3 VL 30B A3B Instruct9.9Claude Opus 563.1Coding Index
Qwen3 VL 30B A3B Instructnot measuredClaude Opus 578.0Agentic Index
Qwen3 VL 30B A3B Instructnot measuredClaude Opus 559.2Output speed
Qwen3 VL 30B A3B Instruct16 t/sClaude Opus 587 t/sContext window
Qwen3 VL 30B A3B Instruct262KClaude Opus 51MLatency · lower is better
Qwen3 VL 30B A3B Instruct417msClaude Opus 51.12sBlended price / 1M · lower is better
Qwen3 VL 30B A3B Instruct$0.262Claude Opus 5$10Cost per task · lower is better
Qwen3 VL 30B A3B Instruct$0.02Claude Opus 5$2.25
Rows marked “lower is better” still draw a longer bar for a larger number — read the value, not just the length. Arena Elo is in the specification table above instead: it has no meaningful zero, so a bar would flatten the gaps.
View as table
| Metric | Qwen3 VL 30B A3B Instruct | Claude Opus 5 |
|---|---|---|
| Intelligence Index | 9.9 | 63.1 |
| Coding Index | — | 78.0 |
| Agentic Index | — | 59.2 |
| Output speed | 16 t/s | 87 t/s |
| Context window | 262K | 1M |
| Latency · lower is better | 417ms | 1.12s |
| Blended price / 1M · lower is better | $0.262 | $10 |
| Cost per task · lower is better | $0.02 | $2.25 |
Evaluation scores
Percentage correct on a common 0–100% scale
- Qwen3 VL 30B A3B Instruct
- Claude Opus 5
GPQA Diamond
Qwen3 VL 30B A3B Instruct69.5%Claude Opus 593.2%Humanity's Last Exam
Qwen3 VL 30B A3B Instruct6.3%Claude Opus 554.9%SciCode
Qwen3 VL 30B A3B Instruct30.8%Claude Opus 555.7%τ²-bench
Qwen3 VL 30B A3B Instruct19.0%Claude Opus 542.1%Terminal-Bench Hard
Qwen3 VL 30B A3B Instruct6.1%Claude Opus 5not measuredLiveCodeBench
Qwen3 VL 30B A3B Instruct47.6%Claude Opus 5not measuredAA-LCR long context
Qwen3 VL 30B A3B Instruct27.3%Claude Opus 575.7%AIME 2025
Qwen3 VL 30B A3B Instruct72.3%Claude Opus 5not measured
A missing bar means that evaluation was not run for that model — it is not a zero.
View as table
| Evaluation | Qwen3 VL 30B A3B Instruct | Claude Opus 5 |
|---|---|---|
| GPQA Diamond | 69.5% | 93.2% |
| Humanity's Last Exam | 6.3% | 54.9% |
| SciCode | 30.8% | 55.7% |
| τ²-bench | 19.0% | 42.1% |
| Terminal-Bench Hard | 6.1% | — |
| LiveCodeBench | 47.6% | — |
| AA-LCR long context | 27.3% | 75.7% |
| AIME 2025 | 72.3% | — |
Estimated from list pricing: 50K input tokens plus 80K output tokens for reasoning models (25K for non-reasoning).