Fireworks overhauls AI model lineup with GLM-5 and MiniMax M2 additions
What changed
GLM-4.6 removed from serverless inference pricing [was $0.55/1M input, $2.19/1M output tokens]
Qwen3 Coder 480B removed from serverless inference pricing [was $0.45/1M input, $1.80/1M output tokens]
MiniMax M2 family cached input pricing added at $0.03/1M tokens [previously no cached pricing; model listing consolidated from M2/M2.1 to M2 family]
Qwen3 235B Family removed from serverless inference pricing [was $0.22/1M input, $0.88/1M output tokens]
GLM-5 added to serverless inference pricing at $1.00/1M input, $0.20 cached input, $3.20/1M output tokens
DeepSeek R1 0528 removed from serverless inference pricing [was $1.35/1M input, $5.40/1M output tokens]
Before and after
Feb 9, 2026 vs Feb 18, 2026GLM-4.6 removed from serverless inference pricing [was $0.55/1M input, $2.19/1M output tokens]
Sign in to view the full comparison
See the before/after screenshots and every marked-up change.
More from Fireworks
Full history →- Week 40, 2026·packagingFireworks adds GLM 5.3 Flash with 200K context, per-token training pricing
- Week 37, 2026·improvedFireworks rebuilds Serverless Training API catalog: drops Qwen 3.5 9B, adds DeepSeek V4 Flash and Muse Glimmer 30B
- Week 34, 2026·increaseFireworks: On-Demand GPU pricing up 11-30% from Sep 1; region surcharge to 1.5x
- Week 33, 2026·improvedFireworks: New GB300 GPU tier ($18/hr) + region premium pricing
- Week 32, 2026·packagingFireworks: New Serverless Training API added (Qwen 3.5 9B, 3.6 27B, Kimi K3)
