Skip to content
llmwaves

Inference provider

Groq

8 models served. Speed, latency and price below are medians across this provider's catalogue, not per-endpoint measurements.

Models served

8

Median output speed

280 t/s

Median latency

261ms

Median price / 1M

$0.131

Catalogue

Strongest and fastest on this provider

What the catalogue looks like at the top end, on the two axes that usually decide the choice.

Highest Intelligence Index

Among models this provider serves

View as table
ModelIntelligence
M2.7 (MiniMax)38.9
gpt-oss-120b (OpenAI)24.1
gpt-oss-20b (OpenAI)15.2
Llama 4 Scout (Meta)10.3

Fastest output

Median tokens per second

View as table
ModelTokens/s
gpt-oss-120b (OpenAI)773
Qwen3 32B (Qwen)423
gpt-oss-safeguard-20b (OpenAI)422
M2.7 (MiniMax)307
gpt-oss-20b (OpenAI)252
Llama 3.3 70B Instruct (Meta)234
Llama 4 Scout (Meta)152
Llama 3.1 8B Instruct (Meta)119

Full catalogue

All 8 models on Groq

Filter by creator
8 of 8 models
M2.7
MiniMaxopen
38.9$0.10
gpt-oss-120b
OpenAIopen
24.1$0.02
gpt-oss-20b
OpenAIopen
15.2$0.01
Llama 4 Scout
Metaopen
10.3$0.01
gpt-oss-safeguard-20b
OpenAIopen
$0.03
Qwen3 32B
Qwenopen
$0.03
Llama 3.3 70B Instruct
Metaopen
$0.01
Llama 3.1 8B Instruct
Metaopen
$0.005