Inference provider
BaseTen
8 models served. Speed, latency and price below are medians across this provider's catalogue, not per-endpoint measurements.
Models served
8
Median output speed
138 t/s
Median latency
427ms
Median price / 1M
$1.12
Catalogue
Strongest and fastest on this provider
What the catalogue looks like at the top end, on the two axes that usually decide the choice.
Highest Intelligence Index
Among models this provider serves
- Kimi K3Moonshot AI59.7
- GLM 5.2Z.ai52.6
- V4 Flash 0731DeepSeek51.8
- V4 ProDeepSeek45.3
- Kimi K2.6Moonshot AI45.1
- InklingThinking Machines42.3
- gpt-oss-120bOpenAI24.1
View as table
| Model | Intelligence |
|---|---|
| Kimi K3 (Moonshot AI) | 59.7 |
| GLM 5.2 (Z.ai) | 52.6 |
| V4 Flash 0731 (DeepSeek) | 51.8 |
| V4 Pro (DeepSeek) | 45.3 |
| Kimi K2.6 (Moonshot AI) | 45.1 |
| Inkling (Thinking Machines) | 42.3 |
| gpt-oss-120b (OpenAI) | 24.1 |
Fastest output
Median tokens per second
- gpt-oss-120bOpenAI773 t/s
- V4 Flash 0731DeepSeek148 t/s
- InklingThinking Machines147 t/s
- GLM 5.2Z.ai142 t/s
- Nemotron 3 UltraNVIDIA133 t/s
- Kimi K2.6Moonshot AI89 t/s
- Kimi K3Moonshot AI71 t/s
- V4 ProDeepSeek71 t/s
View as table
| Model | Tokens/s |
|---|---|
| gpt-oss-120b (OpenAI) | 773 |
| V4 Flash 0731 (DeepSeek) | 148 |
| Inkling (Thinking Machines) | 147 |
| GLM 5.2 (Z.ai) | 142 |
| Nemotron 3 Ultra (NVIDIA) | 133 |
| Kimi K2.6 (Moonshot AI) | 89 |
| Kimi K3 (Moonshot AI) | 71 |
| V4 Pro (DeepSeek) | 71 |
Full catalogue
All 8 models on BaseTen
Filter by creator
8 of 8 models| Kimi K3 Moonshot AIopen | 59.7 | $1.35 |
| GLM 5.2 Z.aiopen | 52.6 | $0.23 |
| V4 Flash 0731 DeepSeekopen | 51.8 | $0.02 |
| V4 Pro DeepSeekopen | 45.3 | $0.09 |
| Kimi K2.6 Moonshot AIopen | 45.1 | $0.23 |
| Inkling Thinking Machinesopen | 42.3 | $0.37 |
| gpt-oss-120b OpenAIopen | 24.1 | $0.02 |
| Nemotron 3 Ultra NVIDIAopen | — | $0.32 |