Skip to content
llmwaves

Mistral Large 2407

Mistral AI · released Nov 19, 2024

ProprietaryTool useStructured output

This is Mistral AI's flagship model, Mistral Large 2 (version mistral-large-2407). It's a proprietary weights-available model and excels at reasoning, code, JSON, chat, and more. Read the launch announcement [here](https://mistral.ai/news/mistral-large-2407/)....

Specification

Context window
131K
Max output
Knowledge cutoff
Mar 31, 2024
Parameters
Undisclosed
Licence
Proprietary
Serving providers
1
Moderated
No

Intelligence

7.0

24th percentile

Coding

Coding Index

Agentic

Agentic Index

Output speed

47 t/s

Median across providers

Latency

751ms

Time to first token

Cost per task

$0.25

Estimated

Benchmarks

Where the score comes from

The Intelligence Index is a composite. These are the underlying evaluations this model was actually measured on.

Evaluation scores

Percentage correct · higher is better

  • MMLU-Pro
    68.3%
  • GPQA Diamond
    47.2%
  • τ²-bench (Telecom)
    33.0%
  • IFBench
    31.6%
  • SciCode
    27.1%
  • LiveCodeBench
    26.7%
  • Humanity's Last Exam
    2.9%
  • AA-LCR (long context)
    1.7%
  • AIME 2025
    0.0%

An evaluation missing from this list was not run for this model — it is not a zero.

View as table
EvaluationScore
MMLU-Pro68.3%
GPQA Diamond47.2%
τ²-bench (Telecom)33.0%
IFBench31.6%
SciCode27.1%
LiveCodeBench26.7%
Humanity's Last Exam2.9%
AA-LCR (long context)1.7%
AIME 20250.0%

Against its peers

Intelligence Index · this model highlighted, nearest peers in grey

Peers are the models sitting closest on the Intelligence Index — the set you would realistically choose between.

View as table
ModelIntelligence
Nemotron 3 Nano 30B A3B7.2
Mistral Large 24077.0
Qwen2.5 Coder 32B Instruct6.9
GLM 4.5V6.8
GPT-46.8
Gemini 2.5 Flash Lite6.7
GPT-4o-mini6.7
GPT-4o-mini (2024-07-18)6.7

Percentile among all indexed models

Intelligence24th
Mathematics0th

Pricing

What it costs to run

List prices per million tokens, plus what one representative task works out to.

List price

Input / 1M tokens
$2
Output / 1M tokens
$6
Cached input / 1M
$0.2
Blended 3:1
$3

One task, estimated

$0.25

Input tokens
50,000
Output tokens
25,000
Profile
Standard

Estimated from list pricing: 50K input tokens plus 80K output tokens for reasoning models (25K for non-reasoning).

Serving providers

Speed and latency figures are medians across these providers, so a widely-served model reports a blend rather than any single endpoint.