Fireworks Unveils Major 10x Increase in Serverless Inference RPM for Developer Plan, Boosting Capacity to 6,000
What changed
Meta Llama 4 Scout (Basic) model added with pricing $0.15 input, $0.60 output per 1M tokens
DeepSeek R1 (Fast) model added with pricing $3.00 input, $8.00 output per 1M tokens
Streaming transcription service added at $0.0032 per audio minute
Spending tier qualification thresholds reduced - Tier 2 from $100+ to $50+, Tier 3 from $1,000+ to $500+, Tier 4 from $10,000+ to $5,000+
Meta Llama 4 Maverick (Basic) model added with pricing $0.22 input, $0.88 output per 1M tokens
Developer plan serverless inference daily token limit added: 2.5 billion tokens/day
Developer plan monthly GPU hours limit added: 2,000 GPU hours/month
Developer plan on-demand GPU deployment reduced from 16 GPUs to 8 GPUs
Developer plan serverless inference RPM increased from 600 to 6,000 (10x increase)
DeepSeek R1 (Basic) model added with pricing $0.55 input, $2.19 output per 1M tokens
Before and after
Jan 3, 2025 vs Apr 9, 2025Meta Llama 4 Scout (Basic) model added with pricing $0.15 input, $0.60 output per 1M tokens
Sign in to view the full comparison
See the before/after screenshots and every marked-up change.
More from Fireworks
Full history →- Week 40, 2026·packagingFireworks adds GLM 5.3 Flash with 200K context, per-token training pricing
- Week 37, 2026·improvedFireworks rebuilds Serverless Training API catalog: drops Qwen 3.5 9B, adds DeepSeek V4 Flash and Muse Glimmer 30B
- Week 34, 2026·increaseFireworks: On-Demand GPU pricing up 11-30% from Sep 1; region surcharge to 1.5x
- Week 33, 2026·improvedFireworks: New GB300 GPU tier ($18/hr) + region premium pricing
- Week 32, 2026·packagingFireworks: New Serverless Training API added (Qwen 3.5 9B, 3.6 27B, Kimi K3)
