llmproviders.ai
← All providers

Together AI pricing

verified 2026-09-01

Broad catalog of open-weight models on serverless endpoints — a one-stop shop for Llama, DeepSeek, and Qwen inference.

current models · $ per 1M tokens
ModelInput /1MCached inOutput /1MContext
GPT-OSS 120B
GPT-OSS
$0.15$0.60128K
GPT-OSS 20B
GPT-OSS
$0.05$0.20128K
GLM-5.3
GLM
$1.40$4.401M
GLM-5.3 Flash
GLM
$0.15$0.501M
Gemma 4 31B Instruct
Gemma 4
$0.39$0.97262K
DeepSeek V4 Flash 0731
DeepSeek V4
$0.14$0.281M
DeepSeek V4 Pro 0813
DeepSeek V4
$1.32$3.961.0M
Qwen 3.8-2.4T-A95B
Qwen 3.8
$2$6
Qwen 3.8 Flash
Qwen 3.8
$0.15$0.471M
Qwen 3.7 Max
Qwen 3.7
$1.25$3.75
Qwen 3.7 Plus
Qwen 3.7
$0.32$1.281M
Qwen 3.5 9B
Qwen 3.5
$0.17$0.25262K
Llama 3.3 70B Instruct Turbo
Llama 3.3
$1.04$1.04131K
Kimi K3
Kimi
$3$151.0M
MiniMax M3
MiniMax
$0.30$1.20524K
Inkling
Inkling
$1$4.05524K
Together AI totals · 1M in + 1M out
GPT-OSS 20B
$0.2500
DeepSeek V4 Flash 0731
$0.4200
Qwen 3.5 9B
$0.4200
Qwen 3.8 Flash
$0.6200
GLM-5.3 Flash
$0.6500
GPT-OSS 120B
$0.7500
Gemma 4 31B Instruct
$1.36
Qwen 3.7 Plus
$1.60
Qwen 3.7 Max
$5.00
DeepSeek V4 Pro 0813
$5.28
GLM-5.3
$5.80
Qwen 3.8-2.4T-A95B
$8.00

Cost of 1M input + 1M output tokens. Bar length is square-root scaled so cheap models stay visible next to premium ones.