Inference provider
Chutes
6 models served. Speed, latency and price below are medians across this provider's catalogue, not per-endpoint measurements.
Models served
6
Median output speed
83 t/s
Median latency
430ms
Median price / 1M
$1.21
Catalogue
Strongest and fastest on this provider
What the catalogue looks like at the top end, on the two axes that usually decide the choice.
Highest Intelligence Index
Among models this provider serves
- Kimi K3Moonshot AI59.7
- Kimi K2.6Moonshot AI45.1
- GLM 5.1Z.ai41.0
- Qwen3.6 27BQwen37.7
- 34.3
View as table
| Model | Intelligence |
|---|---|
| Kimi K3 (Moonshot AI) | 59.7 |
| Kimi K2.6 (Moonshot AI) | 45.1 |
| GLM 5.1 (Z.ai) | 41.0 |
| Qwen3.6 27B (Qwen) | 37.7 |
| Qwen3.5 397B A17B (Qwen) | 34.3 |
Fastest output
Median tokens per second
- Gemma 4 31BGoogle232 t/s
- 116 t/s
- Kimi K2.6Moonshot AI89 t/s
- Qwen3.6 27BQwen77 t/s
- Kimi K3Moonshot AI71 t/s
- GLM 5.1Z.ai67 t/s
View as table
| Model | Tokens/s |
|---|---|
| Gemma 4 31B (Google) | 232 |
| Qwen3.5 397B A17B (Qwen) | 116 |
| Kimi K2.6 (Moonshot AI) | 89 |
| Qwen3.6 27B (Qwen) | 77 |
| Kimi K3 (Moonshot AI) | 71 |
| GLM 5.1 (Z.ai) | 67 |
Full catalogue
All 6 models on Chutes
Filter by creator
6 of 6 models| Kimi K3 Moonshot AIopen | 59.7 | $1.35 |
| Kimi K2.6 Moonshot AIopen | 45.1 | $0.23 |
| GLM 5.1 Z.aiopen | 41.0 | $0.29 |
| Qwen3.6 27B Qwenopen | 37.7 | $0.32 |
| Qwen3.5 397B A17B Qwenopen | 34.3 | $0.21 |
| Gemma 4 31B Googleopen | — | $0.03 |