Together AI: HGX H100 GPU price increase + Kimi K3 & Inkling Small models added
NVIDIA HGX H100 GPU cluster reserved pricing increased across all commitment tiers (31-90 days $3.59→$3.69/GPU/hr, 91-180 days $3.29→$3.45/GPU/hr, 181+ days $3.09→$3.19/GPU/hr; on-demand $3.99/hr unchanged) Added Kimi K3 ($3.00/$15.00 per 1M input/output tokens, $0.30 cached) and Inkling Small ($0.50/$1.20 per 1M tokens, $0.10 cached) to the serverless inference model catalog
What changed
Added two new serverless inference models to the model catalog: Kimi K3 ($3.00/$15.00 per 1M input/output tokens, $0.30 cached input) and Inkling Small ($0.50/$1.20 per 1M input/output tokens, $0.10 cached input)
NVIDIA HGX H100 GPU cluster reserved pricing increased across all commitment tiers: 31-90 days $3.59→$3.69/GPU/hr, 91-180 days $3.29→$3.45/GPU/hr, 181+ days $3.09→$3.19/GPU/hr (on-demand rate $3.99/hr unchanged)
Before and after
Jul 27, 2026 vs Aug 3, 2026Added two new serverless inference models to the model catalog: Kimi K3 ($3.00/$15.00 per 1M input/output tokens, $0.30 cached input) and Inkling Small ($0.50/$1.20 per 1M input/output tokens, $0.10 cached input)
Sign in to view the full comparison
See the before/after screenshots and every marked-up change.
More from Together AI
Full history →- Week 41, 2026·improvedTogether AI adds Tev1 4B Experimental at $0.04 input with free output
- Week 39, 2026·increaseTogether AI: Qwen3.7-Max inference price increased 25% ($2.00→$2.50)
- Week 38, 2026·improvedTogether AI: Qwen3.7-Max price hike + new DeepSeek V4.1 Flash model
- Week 34, 2026·improvedTogether AI unveils Qwen3.8-2.4T-A95B, Muse Glimmer, DeepSeek V4 Pro pricing on ...
- Week 31, 2026·packagingTogether AI: LFM2.5-8B-A1B model removed from Serverless Inference
