Inference provider
Together
20 models served. Speed, latency and price below are medians across this provider's catalogue, not per-endpoint measurements.
Models served
20
Median output speed
89 t/s
Median latency
446ms
Median price / 1M
$0.534
Catalogue
Strongest and fastest on this provider
What the catalogue looks like at the top end, on the two axes that usually decide the choice.
Highest Intelligence Index
Among models this provider serves
- Kimi K3Moonshot AI59.7
- GLM 5.2Z.ai52.6
- V4 Flash 0731DeepSeek51.8
- MiniMax M3MiniMax45.4
- V4 ProDeepSeek45.3
- Kimi K2.6Moonshot AI45.1
- Kimi K2.7 CodeMoonshot AI43.0
- InklingThinking Machines42.3
- Inkling SmallThinking Machines41.2
- gpt-oss-120bOpenAI24.1
View as table
| Model | Intelligence |
|---|---|
| Kimi K3 (Moonshot AI) | 59.7 |
| GLM 5.2 (Z.ai) | 52.6 |
| V4 Flash 0731 (DeepSeek) | 51.8 |
| MiniMax M3 (MiniMax) | 45.4 |
| V4 Pro (DeepSeek) | 45.3 |
| Kimi K2.6 (Moonshot AI) | 45.1 |
| Kimi K2.7 Code (Moonshot AI) | 43.0 |
| Inkling (Thinking Machines) | 42.3 |
| Inkling Small (Thinking Machines) | 41.2 |
| gpt-oss-120b (OpenAI) | 24.1 |
Fastest output
Median tokens per second
- gpt-oss-120bOpenAI773 t/s
- gpt-oss-20bOpenAI252 t/s
- 234 t/s
- Gemma 4 31BGoogle232 t/s
- V4 Flash 0731DeepSeek148 t/s
- InklingThinking Machines147 t/s
- GLM 5.2Z.ai142 t/s
- Kimi K2.7 CodeMoonshot AI140 t/s
- Nemotron 3 UltraNVIDIA133 t/s
- Kimi K2.6Moonshot AI89 t/s
View as table
| Model | Tokens/s |
|---|---|
| gpt-oss-120b (OpenAI) | 773 |
| gpt-oss-20b (OpenAI) | 252 |
| Llama 3.3 70B Instruct (Meta) | 234 |
| Gemma 4 31B (Google) | 232 |
| V4 Flash 0731 (DeepSeek) | 148 |
| Inkling (Thinking Machines) | 147 |
| GLM 5.2 (Z.ai) | 142 |
| Kimi K2.7 Code (Moonshot AI) | 140 |
| Nemotron 3 Ultra (NVIDIA) | 133 |
| Kimi K2.6 (Moonshot AI) | 89 |
Full catalogue
All 20 models on Together
Filter by creator
20 of 20 models| Kimi K3 Moonshot AIopen | 59.7 | $1.35 |
| GLM 5.2 Z.aiopen | 52.6 | $0.23 |
| V4 Flash 0731 DeepSeekopen | 51.8 | $0.02 |
| MiniMax M3 MiniMaxopen | 45.4 | $0.11 |
| V4 Pro DeepSeekopen | 45.3 | $0.09 |
| Kimi K2.6 Moonshot AIopen | 45.1 | $0.23 |
| Kimi K2.7 Code Moonshot AIopen | 43.0 | $0.32 |
| Inkling Thinking Machinesopen | 42.3 | $0.37 |
| Inkling Small Thinking Machinesopen | 41.2 | $0.12 |
| gpt-oss-120b OpenAIopen | 24.1 | $0.02 |
| Qwen3.5-9B Qwenopen | 21.8 | $0.02 |
| gpt-oss-20b OpenAIopen | 15.2 | $0.01 |
| Nemotron 3 Ultra NVIDIAopen | — | $0.32 |
| Gemma 4 31B Googleopen | — | $0.03 |
| Cogito v2.1 671B Deep Cogito | — | $0.16 |
| Gemma 3n 4B Googleopen | — | $0.006 |
| Virtuoso Large Arcee AI | — | $0.07 |
| Llama Guard 4 12B Metaopen | — | $0.01 |
| Llama 3.3 70B Instruct Metaopen | — | $0.01 |
| Qwen2.5 7B Instruct Qwenopen | — | $0.01 |