Skip to content

PricingSaaS joins Willingness to PayRead the announcement →

Pricing change·Oct 5 → Jan 1, 2026·Q4 2025·Marketing & Content

Fireworks Pricing Restructured: DeepSeek R1 0528 Sees 55% Input and 32.5% Output Price Reduction

FireworksPricingPlan removedFeature addedPackage changePlan renamedNew planPrice increase

What changed

Plan removed

DeepSeek R1 (Fast) model removed from pricing

Feature added

Prompt caching introduced with 50% discount for all serverless text and vision models (except Qwen3 VL series)

Plan removed

Meta Llama 4 Maverick (Basic) model removed from pricing

Package change

DeepSeek R1 0528 pricing changed from $3.00 input/$8.00 output to $1.35 input/$0.68 cached/$5.4 output (55% price reduction on input, 32.5% on output)

Plan removed

Meta Llama 3.1 405B model removed from pricing

Package change

Fine-tuning restructured into Supervised Fine Tuning (SFT) and Direct Preference Optimization (DPO) with separate pricing columns

Plan renamed

Qwen3 235B Family and GLM-4.5 Air renamed to Qwen3 235B Family (GLM-4.5 Air component removed)

New plan

Kimi K2 Thinking model added alongside existing Kimi K2 Instruct, both with same pricing $0.60 input, $0.30 cached, $2.50 output

Package change

DeepSeek V3 pricing changed from flat $0.90/1M tokens to separate input ($0.56), cached ($0.28), and output ($1.68) pricing

New plan

Qwen3 8B embeddings model added with pricing $0.1 per 1M input tokens

New plan

Qwen3 VL 30B A3B model added with pricing $0.15 input, $0.60 output

Plan removed

AMD MI300X GPU ($4.99/hour) removed from on-demand deployments

Feature added

Reinforcement Fine Tuning (RFT) section added, priced per GPU hour at same rates as on-demand deployment

Price increase

DPO pricing for Models >300B introduced at $20.00 (vs $10.00 for SFT)

Price increase

DPO pricing for Models 80B-300B introduced at $12.00 (vs $6.00 for SFT)

Price increase

DPO pricing for Models up to 16B introduced at $1.00 (vs $0.50 for SFT)

Plan removed

Meta Llama 4 Scout (Basic) model removed from pricing

New plan

MiniMax M2 model added with pricing $0.30 input, $0.15 cached, $1.20 output

Plan removed

Qwen3 30B and Qwen Coder Flash model removed from pricing

Price increase

DPO pricing for Models 16.1B-80B introduced at $6.00 (vs $3.00 for SFT)

Feature added

Prompt caching added for Kimi K2 models at $0.30 cached (50% discount from $0.60 input)

Plan renamed

GLM-4.5 and DeepSeek R1 (Basic) renamed to GLM-4.5, GLM-4.6 (DeepSeek R1 Basic component removed)

Before and after

Oct 5, 2025 vs Jan 1, 2026
Pulse.
100%
Change 1 of 22 · plan removed

DeepSeek R1 (Fast) model removed from pricing

Sign in to view the full comparison

See the before/after screenshots and every marked-up change.