Fireworks hikes H100 GPU pricing 50% while trimming AI model lineup
What changed
Streaming ASR v1/v2 removed from STT pricing. DeepSeek V3, Kimi K2 Instruct/Thinking, and gpt-oss-20b lost cached pricing. Caching footnote updated to 50% universal rate.
DeepSeek R1 0528 ($1.35 input/$0.68 cached/$5.4 output), Qwen3 235B Family ($0.22 input/$0.88 output), and Qwen3 Coder 480B ($0.45 input/$0.23 cached/$1.80 output) removed from serverless pricing table.
H100 80 GB GPU on-demand price increased from $4.00/hr to $6.00/hr (50% increase). MiniMax M2 family cached token price decreased from $0.15 to $0.03 per 1M tokens.
Before and after
Dec 29, 2025 vs Mar 30, 2026Streaming ASR v1/v2 removed from STT pricing. DeepSeek V3, Kimi K2 Instruct/Thinking, and gpt-oss-20b lost cached pricing. Caching footnote updated to 50% universal rate.
Sign in to view the full comparison
See the before/after screenshots and every marked-up change.
More from Fireworks
Full history →- Week 40, 2026·packagingFireworks adds GLM 5.3 Flash with 200K context, per-token training pricing
- Week 37, 2026·improvedFireworks rebuilds Serverless Training API catalog: drops Qwen 3.5 9B, adds DeepSeek V4 Flash and Muse Glimmer 30B
- Week 34, 2026·increaseFireworks: On-Demand GPU pricing up 11-30% from Sep 1; region surcharge to 1.5x
- Week 33, 2026·improvedFireworks: New GB300 GPU tier ($18/hr) + region premium pricing
- Week 32, 2026·packagingFireworks: New Serverless Training API added (Qwen 3.5 9B, 3.6 27B, Kimi K3)
