← All providers
Together AI pricing
verified 2026-09-01Broad catalog of open-weight models on serverless endpoints — a one-stop shop for Llama, DeepSeek, and Qwen inference.
current models · $ per 1M tokens
| Model | Input /1M | Cached in | Output /1M | Context |
|---|---|---|---|---|
GPT-OSS 120B GPT-OSS | $0.15 | — | $0.60 | 128K |
GPT-OSS 20B GPT-OSS | $0.05 | — | $0.20 | 128K |
GLM-5.3 GLM | $1.40 | — | $4.40 | 1M |
GLM-5.3 Flash GLM | $0.15 | — | $0.50 | 1M |
Gemma 4 31B Instruct Gemma 4 | $0.39 | — | $0.97 | 262K |
DeepSeek V4 Flash 0731 DeepSeek V4 | $0.14 | — | $0.28 | 1M |
DeepSeek V4 Pro 0813 DeepSeek V4 | $1.32 | — | $3.96 | 1.0M |
Qwen 3.8-2.4T-A95B Qwen 3.8 | $2 | — | $6 | — |
Qwen 3.8 Flash Qwen 3.8 | $0.15 | — | $0.47 | 1M |
Qwen 3.7 Max Qwen 3.7 | $1.25 | — | $3.75 | — |
Qwen 3.7 Plus Qwen 3.7 | $0.32 | — | $1.28 | 1M |
Qwen 3.5 9B Qwen 3.5 | $0.17 | — | $0.25 | 262K |
Llama 3.3 70B Instruct Turbo Llama 3.3 | $1.04 | — | $1.04 | 131K |
Kimi K3 Kimi | $3 | — | $15 | 1.0M |
MiniMax M3 MiniMax | $0.30 | — | $1.20 | 524K |
Inkling Inkling | $1 | — | $4.05 | 524K |
Together AI totals · 1M in + 1M out
GPT-OSS 20B
$0.2500
DeepSeek V4 Flash 0731
$0.4200
Qwen 3.5 9B
$0.4200
Qwen 3.8 Flash
$0.6200
GLM-5.3 Flash
$0.6500
GPT-OSS 120B
$0.7500
Gemma 4 31B Instruct
$1.36
Qwen 3.7 Plus
$1.60
Qwen 3.7 Max
$5.00
DeepSeek V4 Pro 0813
$5.28
GLM-5.3
$5.80
Qwen 3.8-2.4T-A95B
$8.00
Cost of 1M input + 1M output tokens. Bar length is square-root scaled so cheap models stay visible next to premium ones.