Groq Introduces New High-Performance Models and 50% Discount on Batch API for Paid GroqCloud Customers
What changed
Llama 3 Groq 8B Tool Use Preview 8k model removed (was $0.19/M tokens for both input and output at 1250 tokens/second)
Qwen QwQ 32B (Preview) 128k model added with input price $0.29/M tokens and output price $0.39/M tokens at 400 tokens/second
Llama 3 Groq 70B Tool Use Preview 8k model removed (was $0.89/M tokens for both input and output at 335 tokens/second)
Qwen 2.5 Coder 32B Instruct 128k model added with input price $0.79/M tokens and output price $0.79/M tokens at 390 tokens/second
DeepSeek R1 Distill Llama 70B model added with input price $0.75/M tokens and output price $0.99/M tokens at 275 tokens/second
Qwen 2.5 32B Instruct 128k model added with input price $0.79/M tokens and output price $0.79/M tokens at 200 tokens/second
Mistral Saba 24B model added with input price $0.79/M tokens and output price $0.79/M tokens at 330 tokens/second
Batch API introduced with 25% discount rate, doubled to 50% off through end of April 2025 for paid GroqCloud customers
Text-to-Speech (TTS) Models section introduced with PlayAI Dialog v1.0 model priced at $50.00 per million characters at 140 characters/second
Mixtral 8x7B Instruct 32k model removed (was $0.24/M tokens for both input and output at 575 tokens/second)
DeepSeek R1 Distill Qwen 32B 128k model added with input price $0.69/M tokens and output price $0.69/M tokens at 140 tokens/second
Gemma 7B 8k Instruct model removed (was $0.07/M tokens for both input and output at 950 tokens/second)
Before and after
Jan 3, 2025 vs Apr 1, 2025Llama 3 Groq 8B Tool Use Preview 8k model removed (was $0.19/M tokens for both input and output at 1250 tokens/second)
Sign in to view the full comparison
See the before/after screenshots and every marked-up change.
More from Groq
Full history →- Week 30, 2026·packagingGroq drops Llama 4 Scout and Qwen3 32B, upgrades Minimax to M2.7
- Q2 2026·packagingGroq: Qwen 3.6 27B added, new Enterprise-only LLM tier + 1 more change
- Q4 2025·decreaseGroq Reduces GPT OSS 20B Input and Output Token Prices by 25% and 40%, Enhancing Cost Efficiency for Users
- Week 1, 2026·packagingGroq swaps PlayAI for Orpheus TTS, cuts text-to-speech price 56%
- Q3 2025·improvedGroq Removes Multiple Models While Adding High-Performance GPT OSS 20B and 120B with New Prompt Caching Feature
