Fireworks Pricing Restructured: DeepSeek R1 0528 Sees 55% Input and 32.5% Output Price Reduction
What changed
DeepSeek R1 (Fast) model removed from pricing
Prompt caching introduced with 50% discount for all serverless text and vision models (except Qwen3 VL series)
Meta Llama 4 Maverick (Basic) model removed from pricing
DeepSeek R1 0528 pricing changed from $3.00 input/$8.00 output to $1.35 input/$0.68 cached/$5.4 output (55% price reduction on input, 32.5% on output)
Meta Llama 3.1 405B model removed from pricing
Fine-tuning restructured into Supervised Fine Tuning (SFT) and Direct Preference Optimization (DPO) with separate pricing columns
Qwen3 235B Family and GLM-4.5 Air renamed to Qwen3 235B Family (GLM-4.5 Air component removed)
Kimi K2 Thinking model added alongside existing Kimi K2 Instruct, both with same pricing $0.60 input, $0.30 cached, $2.50 output
DeepSeek V3 pricing changed from flat $0.90/1M tokens to separate input ($0.56), cached ($0.28), and output ($1.68) pricing
Qwen3 8B embeddings model added with pricing $0.1 per 1M input tokens
Qwen3 VL 30B A3B model added with pricing $0.15 input, $0.60 output
AMD MI300X GPU ($4.99/hour) removed from on-demand deployments
Reinforcement Fine Tuning (RFT) section added, priced per GPU hour at same rates as on-demand deployment
DPO pricing for Models >300B introduced at $20.00 (vs $10.00 for SFT)
DPO pricing for Models 80B-300B introduced at $12.00 (vs $6.00 for SFT)
DPO pricing for Models up to 16B introduced at $1.00 (vs $0.50 for SFT)
Meta Llama 4 Scout (Basic) model removed from pricing
MiniMax M2 model added with pricing $0.30 input, $0.15 cached, $1.20 output
Qwen3 30B and Qwen Coder Flash model removed from pricing
DPO pricing for Models 16.1B-80B introduced at $6.00 (vs $3.00 for SFT)
Prompt caching added for Kimi K2 models at $0.30 cached (50% discount from $0.60 input)
GLM-4.5 and DeepSeek R1 (Basic) renamed to GLM-4.5, GLM-4.6 (DeepSeek R1 Basic component removed)
Before and after
Oct 5, 2025 vs Jan 1, 2026DeepSeek R1 (Fast) model removed from pricing
Sign in to view the full comparison
See the before/after screenshots and every marked-up change.
More from Fireworks
Full history →- Week 40, 2026·packagingFireworks adds GLM 5.3 Flash with 200K context, per-token training pricing
- Week 37, 2026·improvedFireworks rebuilds Serverless Training API catalog: drops Qwen 3.5 9B, adds DeepSeek V4 Flash and Muse Glimmer 30B
- Week 34, 2026·increaseFireworks: On-Demand GPU pricing up 11-30% from Sep 1; region surcharge to 1.5x
- Week 33, 2026·improvedFireworks: New GB300 GPU tier ($18/hr) + region premium pricing
- Week 32, 2026·packagingFireworks: New Serverless Training API added (Qwen 3.5 9B, 3.6 27B, Kimi K3)
