Pricing change·Jun 1 → Jun 8, 2026·Week 24, 2026·AI & Machine Learning
Together AI raises serverless inference prices across Llama, Qwen models
What changed
Price increase
Multiple serverless inference models increased: Qwen3.5 9B ($0.10→$0.17 in, $0.15→$0.25 out), Llama 3.3 70B ($0.88→$1.04), Llama 3 8B Instruct Lite ($0.10→$0.14) per 1M tokens.
Feature added
NVIDIA Nemotron 3 Ultra added to serverless inference chat models at $0.60/1M input ($0.20 cached) and $3.60/1M output tokens.
Before and after
Jun 1, 2026 vs Jun 8, 2026Pulse.
100%
Change 1 of 2 · price increased
Multiple serverless inference models increased: Qwen3.5 9B ($0.10→$0.17 in, $0.15→$0.25 out), Llama 3.3 70B ($0.88→$1.04), Llama 3 8B Instruct Lite ($0.10→$0.14) per 1M tokens.
Sign in to view the full comparison
See the before/after screenshots and every marked-up change.
Week 23, 2026 · Together AI cuts Qwen3 inference prices 50%, adds cachingTogether AI reveals cached input pricing for GLM-5.1 and Qwen3.5 · Week 25, 2026
More from Together AI
Full history →- Week 41, 2026·improvedTogether AI adds Tev1 4B Experimental at $0.04 input with free output
- Week 39, 2026·increaseTogether AI: Qwen3.7-Max inference price increased 25% ($2.00→$2.50)
- Week 38, 2026·improvedTogether AI: Qwen3.7-Max price hike + new DeepSeek V4.1 Flash model
- Week 34, 2026·improvedTogether AI unveils Qwen3.8-2.4T-A95B, Muse Glimmer, DeepSeek V4 Pro pricing on ...
- Week 32, 2026·increaseTogether AI: HGX H100 GPU price increase + Kimi K3 & Inkling Small models added
