Skip to content
llmwaves

Inference provider

Together

20 models served. Speed, latency and price below are medians across this provider's catalogue, not per-endpoint measurements.

Models served

20

Median output speed

89 t/s

Median latency

446ms

Median price / 1M

$0.534

Catalogue

Strongest and fastest on this provider

What the catalogue looks like at the top end, on the two axes that usually decide the choice.

Highest Intelligence Index

Among models this provider serves

View as table
ModelIntelligence
Kimi K3 (Moonshot AI)59.7
GLM 5.2 (Z.ai)52.6
V4 Flash 0731 (DeepSeek)51.8
MiniMax M3 (MiniMax)45.4
V4 Pro (DeepSeek)45.3
Kimi K2.6 (Moonshot AI)45.1
Kimi K2.7 Code (Moonshot AI)43.0
Inkling (Thinking Machines)42.3
Inkling Small (Thinking Machines)41.2
gpt-oss-120b (OpenAI)24.1

Fastest output

Median tokens per second

View as table
ModelTokens/s
gpt-oss-120b (OpenAI)773
gpt-oss-20b (OpenAI)252
Llama 3.3 70B Instruct (Meta)234
Gemma 4 31B (Google)232
V4 Flash 0731 (DeepSeek)148
Inkling (Thinking Machines)147
GLM 5.2 (Z.ai)142
Kimi K2.7 Code (Moonshot AI)140
Nemotron 3 Ultra (NVIDIA)133
Kimi K2.6 (Moonshot AI)89

Full catalogue

All 20 models on Together

Filter by creator
20 of 20 models
Kimi K3
Moonshot AIopen
59.7$1.35
GLM 5.2
Z.aiopen
52.6$0.23
V4 Flash 0731
DeepSeekopen
51.8$0.02
MiniMax M3
MiniMaxopen
45.4$0.11
V4 Pro
DeepSeekopen
45.3$0.09
Kimi K2.6
Moonshot AIopen
45.1$0.23
Kimi K2.7 Code
Moonshot AIopen
43.0$0.32
Inkling
Thinking Machinesopen
42.3$0.37
Inkling Small
Thinking Machinesopen
41.2$0.12
gpt-oss-120b
OpenAIopen
24.1$0.02
Qwen3.5-9B
Qwenopen
21.8$0.02
gpt-oss-20b
OpenAIopen
15.2$0.01
Nemotron 3 Ultra
NVIDIAopen
$0.32
Gemma 4 31B
Googleopen
$0.03
Cogito v2.1 671B
Deep Cogito
$0.16
Gemma 3n 4B
Googleopen
$0.006
Virtuoso Large
Arcee AI
$0.07
Llama Guard 4 12B
Metaopen
$0.01
Llama 3.3 70B Instruct
Metaopen
$0.01
Qwen2.5 7B Instruct
Qwenopen
$0.01