Inference provider
43 models served. Speed, latency and price below are medians across this provider's catalogue, not per-endpoint measurements.
Models served
43
Median output speed
102 t/s
Median latency
843ms
Median price / 1M
$1.13
Catalogue
Strongest and fastest on this provider
What the catalogue looks like at the top end, on the two axes that usually decide the choice.
Highest Intelligence Index
Among models this provider serves
- Claude Opus 5Anthropic63.1
- Claude Fable 5Anthropic62.1
- Claude Opus 4.8Anthropic57.3
- Claude Sonnet 5Anthropic55.3
- Claude Opus 4.7Anthropic55.0
- Gemini 3.5 FlashGoogle52.0
- Gemini 3.6 FlashGoogle51.6
- Gemini 3.1 Pro PreviewGoogle47.7
- Claude Opus 4.6Anthropic38.8
- Gemini 3.5 Flash LiteGoogle37.4
View as table
| Model | Intelligence |
|---|---|
| Claude Opus 5 (Anthropic) | 63.1 |
| Claude Fable 5 (Anthropic) | 62.1 |
| Claude Opus 4.8 (Anthropic) | 57.3 |
| Claude Sonnet 5 (Anthropic) | 55.3 |
| Claude Opus 4.7 (Anthropic) | 55.0 |
| Gemini 3.5 Flash (Google) | 52.0 |
| Gemini 3.6 Flash (Google) | 51.6 |
| Gemini 3.1 Pro Preview (Google) | 47.7 |
| Claude Opus 4.6 (Anthropic) | 38.8 |
| Gemini 3.5 Flash Lite (Google) | 37.4 |
Fastest output
Median tokens per second
- gpt-oss-120bOpenAI773 t/s
- GLM 4.7Z.ai502 t/s
- 443 t/s
- gpt-oss-20bOpenAI252 t/s
- 242 t/s
- 234 t/s
- Gemini 3.1 Flash LiteGoogle183 t/s
- Gemini 3.5 FlashGoogle182 t/s
- 180 t/s
- 166 t/s
View as table
| Model | Tokens/s |
|---|---|
| gpt-oss-120b (OpenAI) | 773 |
| GLM 4.7 (Z.ai) | 502 |
| Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image) (Google) | 443 |
| gpt-oss-20b (OpenAI) | 252 |
| Qwen3 Next 80B A3B Thinking (Qwen) | 242 |
| Llama 3.3 70B Instruct (Meta) | 234 |
| Gemini 3.1 Flash Lite (Google) | 183 |
| Gemini 3.5 Flash (Google) | 182 |
| Nano Banana (Gemini 2.5 Flash Image) (Google) | 180 |
| Nano Banana 2 (Gemini 3.1 Flash Image) (Google) | 166 |
Full catalogue
All 43 models on Google
Filter by creator
43 of 43 models