Inference provider
Crusoe
8 models served. Speed, latency and price below are medians across this provider's catalogue, not per-endpoint measurements.
Models served
8
Median output speed
114 t/s
Median latency
359ms
Median price / 1M
$0.344
Catalogue
Strongest and fastest on this provider
What the catalogue looks like at the top end, on the two axes that usually decide the choice.
Highest Intelligence Index
Among models this provider serves
- GLM 5.2Z.ai52.6
- Kimi K2.6Moonshot AI45.1
- GLM 5.1Z.ai41.0
- Nemotron 3 Nano 30B A3BNVIDIA7.2
View as table
| Model | Intelligence |
|---|---|
| GLM 5.2 (Z.ai) | 52.6 |
| Kimi K2.6 (Moonshot AI) | 45.1 |
| GLM 5.1 (Z.ai) | 41.0 |
| Nemotron 3 Nano 30B A3B (NVIDIA) | 7.2 |
Fastest output
Median tokens per second
- 234 t/s
- Gemma 4 31BGoogle232 t/s
- GLM 5.2Z.ai142 t/s
- Nemotron 3 Nano 30B A3BNVIDIA138 t/s
- Kimi K2.6Moonshot AI89 t/s
- GLM 5.1Z.ai67 t/s
- 60 t/s
- V3 0324DeepSeek32 t/s
View as table
| Model | Tokens/s |
|---|---|
| Llama 3.3 70B Instruct (Meta) | 234 |
| Gemma 4 31B (Google) | 232 |
| GLM 5.2 (Z.ai) | 142 |
| Nemotron 3 Nano 30B A3B (NVIDIA) | 138 |
| Kimi K2.6 (Moonshot AI) | 89 |
| GLM 5.1 (Z.ai) | 67 |
| Qwen3 235B A22B Instruct 2507 (Qwen) | 60 |
| V3 0324 (DeepSeek) | 32 |
Full catalogue
All 8 models on Crusoe
Filter by creator
8 of 8 models| GLM 5.2 Z.aiopen | 52.6 | $0.23 |
| Kimi K2.6 Moonshot AIopen | 45.1 | $0.23 |
| GLM 5.1 Z.aiopen | 41.0 | $0.29 |
| Nemotron 3 Nano 30B A3B NVIDIAopen | 7.2 | $0.02 |
| Gemma 4 31B Googleopen | — | $0.03 |
| Qwen3 235B A22B Instruct 2507 Qwenopen | — | $0.02 |
| V3 0324 DeepSeekopen | — | $0.04 |
| Llama 3.3 70B Instruct Metaopen | — | $0.01 |