Together AI hikes reserved GPU rates up to 30% and drops three AI models
What changed
Three models removed from Together AI serverless inference pricing list: Llama 4 Maverick ($0.27/$0.85/1M tokens), Qwen3-Next-80B-A3B-Instruct ($0.15/$1.50/1M tokens), and GLM-4.5-Air ($0.20/$1.10/1M tokens).
Reserved GPU cluster prices increased across all duration tiers: NVIDIA HGX H100 from $2.69-$2.25/hr to $2.99-$2.55/hr; H200 from $3.19-$2.59/hr to $3.49-$2.89/hr; B200 from $5.49-$4.49/hr to $7.15-$6.39/hr (30% increase).
Before and after
Mar 30, 2026 vs Apr 6, 2026Three models removed from Together AI serverless inference pricing list: Llama 4 Maverick ($0.27/$0.85/1M tokens), Qwen3-Next-80B-A3B-Instruct ($0.15/$1.50/1M tokens), and GLM-4.5-Air ($0.20/$1.10/1M tokens).
Sign in to view the full comparison
See the before/after screenshots and every marked-up change.
More from Together AI
Full history →- Week 41, 2026·improvedTogether AI adds Tev1 4B Experimental at $0.04 input with free output
- Week 39, 2026·increaseTogether AI: Qwen3.7-Max inference price increased 25% ($2.00→$2.50)
- Week 38, 2026·improvedTogether AI: Qwen3.7-Max price hike + new DeepSeek V4.1 Flash model
- Week 34, 2026·improvedTogether AI unveils Qwen3.8-2.4T-A95B, Muse Glimmer, DeepSeek V4 Pro pricing on ...
- Week 32, 2026·increaseTogether AI: HGX H100 GPU price increase + Kimi K3 & Inkling Small models added
