Skip to content
llmwaves

Mistral Large

Mistral AI · released Feb 26, 2024

ProprietaryTool useStructured output

This is Mistral AI's flagship model, Mistral Large 2 (version `mistral-large-2407`). It's a proprietary weights-available model and excels at reasoning, code, JSON, chat, and more. Read the launch announcement [here](https://mistral.ai/news/mistral-large-2407/)....

Specification

Context window
128K
Max output
Knowledge cutoff
Nov 30, 2024
Parameters
Undisclosed
Licence
Proprietary
Serving providers
1
Moderated
No
Uptime
100.0%

Intelligence

4.1

13th percentile

Coding

Coding Index

Agentic

Agentic Index

Output speed

35 t/s

Median across providers

Latency

656ms

Time to first token

Cost per task

$0.25

Estimated

Benchmarks

Where the score comes from

The Intelligence Index is a composite. These are the underlying evaluations this model was actually measured on.

Evaluation scores

Percentage correct · higher is better

  • MMLU-Pro
    51.5%
  • GPQA Diamond
    35.1%
  • SciCode
    20.8%
  • LiveCodeBench
    17.8%
  • Humanity's Last Exam
    3.5%
  • AIME 2025
    0.0%

An evaluation missing from this list was not run for this model — it is not a zero.

View as table
EvaluationScore
MMLU-Pro51.5%
GPQA Diamond35.1%
SciCode20.8%
LiveCodeBench17.8%
Humanity's Last Exam3.5%
AIME 20250.0%

Against its peers

Intelligence Index · this model highlighted, nearest peers in grey

Peers are the models sitting closest on the Intelligence Index — the set you would realistically choose between.

View as table
ModelIntelligence
Olmo 3 32B Think6.1
Hermes 3 70B Instruct4.8
Phi 44.6
Mistral Large4.1
Mixtral 8x22B Instruct4.0
Reka Flash 33.7
Claude 3 Haiku3.5
GPT-3.5 Turbo3.2

Percentile among all indexed models

Intelligence13th

Pricing

What it costs to run

List prices per million tokens, plus what one representative task works out to.

List price

Input / 1M tokens
$2
Output / 1M tokens
$6
Cached input / 1M
$0.2
Blended 3:1
$3

One task, estimated

$0.25

Input tokens
50,000
Output tokens
25,000
Profile
Standard

Estimated from list pricing: 50K input tokens plus 80K output tokens for reasoning models (25K for non-reasoning).

Serving providers

Speed and latency figures are medians across these providers, so a widely-served model reports a blend rather than any single endpoint.