
Together reverses course, cuts Qwen inference prices 40% two days after a hike
Together cut serverless inference pricing on two Qwen models: Qwen3.8 Flash fell from $0.15/$0.47 to $0.09/$0.28 per 1M input/output tokens, and Qwen3.7-Max fell from $2.50/$7.50 to $1.50/$4.50 — both roughly 40% reductions.
Together’s pricing page · Change Detected
WHY IT MATTERS
Inference hosting for popular open-weight models like Qwen, DeepSeek, and Kimi is looking increasingly commodity-priced, with providers competing on thin, fast-moving margins rather than sticky list prices. Watch whether this cut holds or gets reversed again within days, and whether the quiet rise in cached-token pricing — up even as headline prices fell — becomes where Together rebuilds margin.
BY THE NUMBERS
- 36 pricing/catalog changes in 365 days; ~7-day average cadence
- Qwen3.7-Max cut 40% just 2 days after a 25% hike (Sep 21, 2026)
- 20 changes in the last 90 days alone
- Cached-token price for Qwen3.7-Max rose $0.25→$0.30 despite the cut
THEIR RATIONALE
The cut lands just two days after Together raised Qwen3.7-Max pricing 25% to $2.50/$7.50 on Sep 21, itself one of several hikes to that model through early September. Together frames this as a straightforward price cut, but the timing — a reversal within 48 hours of the last increase — reads as reactive, most plausibly a response to competitive inference pricing elsewhere or a recalibration of GPU costs rather than a planned discount cycle.
THE SIGNAL
This fits Together's overall rhythm: 36 tracked pricing and catalog changes in the past year, one roughly every 7 days, with 20 of those in the last 90 days alone — by far the fastest-moving page in the credits-and-usage-billing pattern that has touched 348 companies industry-wide in six months. Unlike most companies in that pattern, which are restructuring plans or credit packaging, Together is adjusting raw per-token unit prices near-weekly, treating inference cost like a spot market tracking GPU economics and rival hosts of the same open-weight models.
IN CONTEXT
This cut lands just two days after Together raised Qwen3.7-Max pricing 25% to $2.50/$7.50 on Sep 21, the tail end of a stretch that saw the same model hiked repeatedly through early September before today's reversal. That whiplash fits Together's overall cadence — 36 changes in the past year, roughly one every 7 days — making it the fastest-moving page in the credits-and-usage-billing pattern that touched 348 companies industry-wide over six months. Unlike peers such as Snyk restructuring Enterprise onto credit-based pricing or Notion ending its free Workers beta ahead of October credit billing, Together isn't repackaging — it's adjusting raw per-token unit prices near-weekly, treating inference cost like a spot market tied to GPU economics and rival hosts of the same open-weight models.