Groq Introduces Llama 4 Maverick and Qwen3 Models Amid Removal of Multiple High-Cost Plans and End of 50% Batch Processing Discount
What changed
Llama 3.2 11B Vision 8k (Preview) model removed (was $0.18 input/$0.18 output)
DeepSeek R1 Distill Qwen 32B model removed (was $0.69 input/$0.69 output)
Qwen 2.5 32B Instruct model removed (was $0.79 input/$0.79 output)
Llama 4 Maverick (17Bx128E) model added with 562 tokens/sec speed, $0.20 input/$0.60 output per M tokens
Llama 3.2 3B (Preview) model removed (was $0.06 input/$0.06 output)
Qwen3 32B model added with 662 tokens/sec speed, $0.29 input/$0.59 output per M tokens
Qwen 2.5 Coder 32B Instruct model removed (was $0.79 input/$0.79 output)
Llama 3.3 70B SpecDec model removed (was $0.59 input/$0.99 output)
Whisper Large v3 Turbo speed factor increased from 216x to 228x
Llama 3.2 90B Vision 8k (Preview) model removed (was $0.90 input/$0.90 output)
Llama 3.1 8B Instant speed increased from 750 to 840 tokens/sec
Llama 3.2 1B (Preview) model removed (was $0.04 input/$0.04 output)
Llama 4 Scout (17Bx16E) model added with 594 tokens/sec speed, $0.11 input/$0.34 output per M tokens
Llama Guard 4 12B model added with 325 tokens/sec speed, $0.20 input/$0.20 output per M tokens
Whisper V3 Large speed factor increased from 189x to 217x
Llama 3 8B speed increased from 1250 to 1345 tokens/sec
DeepSeek R1 Distill Llama 70B speed increased from 275 to 400 tokens/sec
50% batch processing discount promotion ended (was doubled from 25% through end of April 2025)
Llama 3.3 70B Versatile speed increased from 275 to 394 tokens/sec
Before and after
Apr 1, 2025 vs Jul 4, 2025Llama 3.2 11B Vision 8k (Preview) model removed (was $0.18 input/$0.18 output)
Sign in to view the full comparison
See the before/after screenshots and every marked-up change.
More from Groq
Full history →- Week 30, 2026·packagingGroq drops Llama 4 Scout and Qwen3 32B, upgrades Minimax to M2.7
- Q2 2026·packagingGroq: Qwen 3.6 27B added, new Enterprise-only LLM tier + 1 more change
- Q4 2025·decreaseGroq Reduces GPT OSS 20B Input and Output Token Prices by 25% and 40%, Enhancing Cost Efficiency for Users
- Week 1, 2026·packagingGroq swaps PlayAI for Orpheus TTS, cuts text-to-speech price 56%
- Q3 2025·improvedGroq Removes Multiple Models While Adding High-Performance GPT OSS 20B and 120B with New Prompt Caching Feature
