Inference provider
Groq
8 models served. Speed, latency and price below are medians across this provider's catalogue, not per-endpoint measurements.
Models served
8
Median output speed
280 t/s
Median latency
261ms
Median price / 1M
$0.131
Catalogue
Strongest and fastest on this provider
What the catalogue looks like at the top end, on the two axes that usually decide the choice.
Highest Intelligence Index
Among models this provider serves
- M2.7MiniMax38.9
- gpt-oss-120bOpenAI24.1
- gpt-oss-20bOpenAI15.2
- Llama 4 ScoutMeta10.3
View as table
| Model | Intelligence |
|---|---|
| M2.7 (MiniMax) | 38.9 |
| gpt-oss-120b (OpenAI) | 24.1 |
| gpt-oss-20b (OpenAI) | 15.2 |
| Llama 4 Scout (Meta) | 10.3 |
Fastest output
Median tokens per second
- gpt-oss-120bOpenAI773 t/s
- Qwen3 32BQwen423 t/s
- gpt-oss-safeguard-20bOpenAI422 t/s
- M2.7MiniMax307 t/s
- gpt-oss-20bOpenAI252 t/s
- 234 t/s
- Llama 4 ScoutMeta152 t/s
- 119 t/s
View as table
| Model | Tokens/s |
|---|---|
| gpt-oss-120b (OpenAI) | 773 |
| Qwen3 32B (Qwen) | 423 |
| gpt-oss-safeguard-20b (OpenAI) | 422 |
| M2.7 (MiniMax) | 307 |
| gpt-oss-20b (OpenAI) | 252 |
| Llama 3.3 70B Instruct (Meta) | 234 |
| Llama 4 Scout (Meta) | 152 |
| Llama 3.1 8B Instruct (Meta) | 119 |
Full catalogue
All 8 models on Groq
Filter by creator
8 of 8 models| M2.7 MiniMaxopen | 38.9 | $0.10 |
| gpt-oss-120b OpenAIopen | 24.1 | $0.02 |
| gpt-oss-20b OpenAIopen | 15.2 | $0.01 |
| Llama 4 Scout Metaopen | 10.3 | $0.01 |
| gpt-oss-safeguard-20b OpenAIopen | — | $0.03 |
| Qwen3 32B Qwenopen | — | $0.03 |
| Llama 3.3 70B Instruct Metaopen | — | $0.01 |
| Llama 3.1 8B Instruct Metaopen | — | $0.005 |