Skip to content
llmwaves

Claude 3 Haiku

Anthropic · released Mar 13, 2024

ProprietaryTool useVision

Claude 3 Haiku is Anthropic's fastest and most compact model for near-instant responsiveness. Quick and accurate targeted performance. See the launch announcement and benchmark results [here](https://www.anthropic.com/news/claude-3-haiku) #multimodal

Specification

Context window
200K
Max output
4K
Knowledge cutoff
Aug 31, 2023
Parameters
Undisclosed
Licence
Proprietary
Serving providers
1
Moderated
Yes
Uptime
100.0%

Intelligence

3.5

11th percentile

Coding

Coding Index

Agentic

Agentic Index

Output speed

79 t/s

Median across providers

Latency

423ms

Time to first token

Cost per task

$0.04

Estimated

Benchmarks

Where the score comes from

The Intelligence Index is a composite. These are the underlying evaluations this model was actually measured on.

Evaluation scores

Percentage correct · higher is better

  • GPQA Diamond
    37.4%
  • IFBench
    36.1%
  • AA-LCR (long context)
    25.0%
  • τ²-bench (Telecom)
    21.1%
  • SciCode
    18.6%
  • LiveCodeBench
    15.4%
  • Humanity's Last Exam
    4.1%
  • AIME 2025
    1.0%
  • Terminal-Bench Hard
    0.8%

An evaluation missing from this list was not run for this model — it is not a zero.

View as table
EvaluationScore
GPQA Diamond37.4%
IFBench36.1%
AA-LCR (long context)25.0%
τ²-bench (Telecom)21.1%
SciCode18.6%
LiveCodeBench15.4%
Humanity's Last Exam4.1%
AIME 20251.0%
Terminal-Bench Hard0.8%

Against its peers

Intelligence Index · this model highlighted, nearest peers in grey

Peers are the models sitting closest on the Intelligence Index — the set you would realistically choose between.

View as table
ModelIntelligence
Olmo 3 32B Think6.1
Hermes 3 70B Instruct4.8
Phi 44.6
Mistral Large4.1
Mixtral 8x22B Instruct4.0
Reka Flash 33.7
Claude 3 Haiku3.5
GPT-3.5 Turbo3.2

Percentile among all indexed models

Intelligence11th
Terminal work11th

Pricing

What it costs to run

List prices per million tokens, plus what one representative task works out to.

List price

Input / 1M tokens
$0.25
Output / 1M tokens
$1.25
Cached input / 1M
$0.03
Blended 3:1
$0.5

One task, estimated

$0.04

Input tokens
50,000
Output tokens
25,000
Profile
Standard

Estimated from list pricing: 50K input tokens plus 80K output tokens for reasoning models (25K for non-reasoning).

Serving providers

Speed and latency figures are medians across these providers, so a widely-served model reports a blend rather than any single endpoint.