Skip to content
llmwaves

Hermes 3 70B Instruct

Nous Research · released Aug 18, 2024

Open weightsStructured output

Hermes 3 is a generalist language model with many improvements over [Hermes 2](/models/nousresearch/nous-hermes-2-mistral-7b-dpo), including advanced agentic capabilities, much better roleplaying, reasoning, multi-turn conversation, long context coherence, and improvements across the...

Specification

Context window
131K
Max output
16K
Knowledge cutoff
Dec 31, 2023
Parameters
70.6B
Licence
llama3
Serving providers
1
Moderated
No
Uptime
100.0%

Intelligence

4.8

16th percentile

Coding

Coding Index

Agentic

Agentic Index

Output speed

34 t/s

Median across providers

Latency

344ms

Time to first token

Cost per task

$0.05

Estimated

Benchmarks

Where the score comes from

The Intelligence Index is a composite. These are the underlying evaluations this model was actually measured on.

Evaluation scores

Percentage correct · higher is better

  • MMLU-Pro
    57.1%
  • GPQA Diamond
    40.1%
  • SciCode
    23.1%
  • LiveCodeBench
    18.8%
  • Humanity's Last Exam
    4.0%
  • AIME 2025
    2.3%

An evaluation missing from this list was not run for this model — it is not a zero.

View as table
EvaluationScore
MMLU-Pro57.1%
GPQA Diamond40.1%
SciCode23.1%
LiveCodeBench18.8%
Humanity's Last Exam4.0%
AIME 20252.3%

Against its peers

Intelligence Index · this model highlighted, nearest peers in grey

Peers are the models sitting closest on the Intelligence Index — the set you would realistically choose between.

View as table
ModelIntelligence
Saba6.2
Olmo 3 32B Think6.1
Hermes 3 70B Instruct4.8
Phi 44.6
Mistral Large4.1
Mixtral 8x22B Instruct4.0
Reka Flash 33.7
Claude 3 Haiku3.5

Percentile among all indexed models

Intelligence16th

Pricing

What it costs to run

List prices per million tokens, plus what one representative task works out to.

List price

Input / 1M tokens
$0.7
Output / 1M tokens
$0.7
Cached input / 1M
Not offered
Blended 3:1
$0.7

One task, estimated

$0.05

Input tokens
50,000
Output tokens
25,000
Profile
Standard

Estimated from list pricing: 50K input tokens plus 80K output tokens for reasoning models (25K for non-reasoning).

Serving providers

Speed and latency figures are medians across these providers, so a widely-served model reports a blend rather than any single endpoint.