Skip to content

PricingSaaS joins Willingness to PayRead the announcement →

Pricing change·Jul 4 → Oct 5, 2025·Q3 2025·AI & Machine Learning

Groq Removes Multiple Models While Adding High-Performance GPT OSS 20B and 120B with New Prompt Caching Feature

GroqPricingPlan removedNew planFeature addedDiscount change

What changed

Plan removed

Gemma 2 9B 8k model removed (previously $0.20 input/$0.20 output per million tokens)

Plan removed

Llama Guard 3 8B 8k model removed (previously $0.20 input/$0.20 output per million tokens)

New plan

GPT OSS 20B 128k model added at $0.10 input/$0.50 output per million tokens with 1,000 TPS

New plan

Kimi K2-0905 1T 256k model added at $1.00 input/$3.00 output per million tokens with 200 TPS

New plan

GPT OSS 120B 128k model added at $0.15 input/$0.75 output per million tokens with 500 TPS

Plan removed

DeepSeek R1 Distill Llama 70B 128k model removed (previously $0.75 input/$0.99 output per million tokens)

Feature added

Prompt Caching feature added with 50% discount on cached input tokens for select models (Kimi K2, GPT OSS 20B, GPT OSS 120B)

Plan removed

Mistral Saba 24B 32k model removed (previously $0.79 input/$0.79 output per million tokens)

Plan removed

Llama 3 8B 8k model removed (previously $0.05 input/$0.08 output per million tokens)

Plan removed

Distil-Whisper ASR model removed (previously $0.02 per hour transcribed with 250x speed factor)

Plan removed

Llama 3 70B 8k model removed (previously $0.59 input/$0.79 output per million tokens)

Discount change

Batch API discount increased from 25% to 50%

Plan removed

Qwen QwQ 32B (Preview) 128k model removed (previously $0.29 input/$0.39 output per million tokens)

Before and after

Jul 4, 2025 vs Oct 5, 2025
Pulse.
100%
Change 1 of 13 · plan removed

Gemma 2 9B 8k model removed (previously $0.20 input/$0.20 output per million tokens)

Sign in to view the full comparison

See the before/after screenshots and every marked-up change.