Skip to content

PricingSaaS joins Willingness to PayRead the announcement →

Pricing change·Jul 6 → Oct 5, 2025·Q3 2025·Marketing & Content

Fireworks Announces Major Price Decrease on H100 GPU, Now Just $4.00/hour (31% Off)

FireworksPricingFeature addedPrice decreaseDiscount addedPackage change

What changed

Feature added

Prompt caching discounts now available for enterprise deployments

Feature added

FLUX.1 Kontext Max image model added at $0.08 per image

Feature added

OpenAI gpt OSS 20b model added at $0.07 input, $0.30 output per 1M tokens

Price decrease

B200 180 GB GPU price decreased from $11.99/hour to $9.00/hour (25% decrease)

Feature added

OpenAI gpt OSS 120b model added at $0.15 input, $0.60 output per 1M tokens

Feature added

FLUX.1 Kontext Pro image model added at $0.04 per image

Discount added

Batch inference pricing added at 50% discount of serverless pricing for input and output tokens

Feature added

VLM supervised fine-tuning (fine-tuning with images) added, billed per token

Feature added

Streaming ASR v2 model added at $0.0035 per audio minute

Feature added

Kimi K2 Instruct model added at $0.60 input, $2.50 output per 1M tokens

Feature added

Qwen3 Coder 480B model added at $0.45 input, $1.80 output per 1M tokens

Package change

Fine tuning pricing expanded from 3 tiers to 4 tiers: split DeepSeek R1/V3 ($10.00) into Models 80B-300B ($6.00) and Models >300B ($10.00)

Price decrease

H200 141 GB GPU price decreased from $6.99/hour to $6.00/hour (14% decrease)

Price decrease

H100 80 GB GPU price decreased from $5.80/hour to $4.00/hour (31% decrease)

Before and after

Jul 6, 2025 vs Oct 5, 2025
Pulse.
100%
Change 1 of 14 · feature added

Prompt caching discounts now available for enterprise deployments

Sign in to view the full comparison

See the before/after screenshots and every marked-up change.