Inference provider
Nebius
12 models served. Speed, latency and price below are medians across this provider's catalogue, not per-endpoint measurements.
Models served
12
Median output speed
72 t/s
Median latency
313ms
Median price / 1M
$0.201
Catalogue
Strongest and fastest on this provider
What the catalogue looks like at the top end, on the two axes that usually decide the choice.
Highest Intelligence Index
Among models this provider serves
- GLM 5.1Z.ai41.0
- gpt-oss-120bOpenAI24.1
View as table
| Model | Intelligence |
|---|---|
| GLM 5.1 (Z.ai) | 41.0 |
| gpt-oss-120b (OpenAI) | 24.1 |
Fastest output
Median tokens per second
- gpt-oss-120bOpenAI773 t/s
- Qwen3 32BQwen423 t/s
- 242 t/s
- 234 t/s
- Nemotron 3 SuperNVIDIA164 t/s
- Hermes 4 70BNous Research76 t/s
- GLM 5.1Z.ai67 t/s
- 60 t/s
- 57 t/s
- Gemma 3 27BGoogle46 t/s
View as table
| Model | Tokens/s |
|---|---|
| gpt-oss-120b (OpenAI) | 773 |
| Qwen3 32B (Qwen) | 423 |
| Qwen3 Next 80B A3B Thinking (Qwen) | 242 |
| Llama 3.3 70B Instruct (Meta) | 234 |
| Nemotron 3 Super (NVIDIA) | 164 |
| Hermes 4 70B (Nous Research) | 76 |
| GLM 5.1 (Z.ai) | 67 |
| Qwen3 235B A22B Instruct 2507 (Qwen) | 60 |
| Qwen3 30B A3B Instruct 2507 (Qwen) | 57 |
| Gemma 3 27B (Google) | 46 |
Full catalogue
All 12 models on Nebius
Filter by creator
12 of 12 models| GLM 5.1 Z.aiopen | 41.0 | $0.29 |
| gpt-oss-120b OpenAIopen | 24.1 | $0.02 |
| Nemotron 3 Super NVIDIAopen | — | $0.09 |
| Qwen3 Next 80B A3B Thinking Qwenopen | — | $0.10 |
| Hermes 4 70B Nous Researchopen | — | $0.04 |
| Hermes 4 405B Nous Researchopen | — | $0.29 |
| Qwen3 30B A3B Instruct 2507 Qwenopen | — | $0.007 |
| Qwen3 235B A22B Instruct 2507 Qwenopen | — | $0.02 |
| Qwen3 32B Qwenopen | — | $0.03 |
| Gemma 3 27B Googleopen | — | $0.02 |
| Qwen2.5 VL 72B Instruct Qwenopen | — | $0.03 |
| Llama 3.3 70B Instruct Metaopen | — | $0.01 |