Skip to content
llmwaves

Mixtral 8x22B Instruct

Mistral AI · released Apr 17, 2024

Open weightsTool useStructured output

Mistral's official instruct fine-tuned version of [Mixtral 8x22B](/models/mistralai/mixtral-8x22b). It uses 39B active parameters out of 141B, offering unparalleled cost efficiency for its size. Its strengths include: - strong math, coding,...

Specification

Context window
66K
Max output
Knowledge cutoff
Jan 31, 2024
Parameters
140.6B
Licence
apache-2.0
Serving providers
1
Moderated
No

Intelligence

4.0

13th percentile

Coding

Coding Index

Agentic

Agentic Index

Output speed

101 t/s

Median across providers

Latency

543ms

Time to first token

Cost per task

$0.25

Estimated

Benchmarks

Where the score comes from

The Intelligence Index is a composite. These are the underlying evaluations this model was actually measured on.

Evaluation scores

Percentage correct · higher is better

  • MMLU-Pro
    53.7%
  • GPQA Diamond
    33.2%
  • SciCode
    18.8%
  • LiveCodeBench
    14.8%
  • Humanity's Last Exam
    4.0%
  • AIME 2025
    0.0%

An evaluation missing from this list was not run for this model — it is not a zero.

View as table
EvaluationScore
MMLU-Pro53.7%
GPQA Diamond33.2%
SciCode18.8%
LiveCodeBench14.8%
Humanity's Last Exam4.0%
AIME 20250.0%

Against its peers

Intelligence Index · this model highlighted, nearest peers in grey

Peers are the models sitting closest on the Intelligence Index — the set you would realistically choose between.

View as table
ModelIntelligence
Olmo 3 32B Think6.1
Hermes 3 70B Instruct4.8
Phi 44.6
Mistral Large4.1
Mixtral 8x22B Instruct4.0
Reka Flash 33.7
Claude 3 Haiku3.5
GPT-3.5 Turbo3.2

Percentile among all indexed models

Intelligence13th

Pricing

What it costs to run

List prices per million tokens, plus what one representative task works out to.

List price

Input / 1M tokens
$2
Output / 1M tokens
$6
Cached input / 1M
$0.2
Blended 3:1
$3

One task, estimated

$0.25

Input tokens
50,000
Output tokens
25,000
Profile
Standard

Estimated from list pricing: 50K input tokens plus 80K output tokens for reasoning models (25K for non-reasoning).

Serving providers

Speed and latency figures are medians across these providers, so a widely-served model reports a blend rather than any single endpoint.