Skip to content

PricingSaaS joins Willingness to PayRead the announcement →

AI Brief·Sep 21 – Sep 27·6 leaders · 56 sources

Cheaper flagships, metered everything else: this week in AI pricing

Anthropic cut Opus pricing 20% and moved Opus 5 to legacy, Fireworks and Baseten added GLM 5.3 rows while retiring older models, Fastly put a per-request price on AI routing, and Notion and Aircall reworked how AI work is billed. Our capture logged 107 AI-relevant changes.

The moves

Anthropic replaced its flagship. Our capture shows Opus 5.5 entering the API price list at $4 per million input tokens and $20 per million output, with Opus 5 moved to the legacy section at its original $5/$25 [44]. Anthropic's own announcement states the cut plainly: "Input and output tokens are $4 and $20 per million, 20% less than Opus 5" [1], and cache reads fall to $0.20 per million, which it calls 60% less than Opus 5 [2]. The page, which renders its price table client-side, carries the same "Daily driver for agentic coding and enterprise work" framing our capture recorded [4]. The old rate survives only as a legacy line.

Fireworks added GLM 5.3 Flash to its Serverless Training API table with a 200K context window and per-token prices for prefill, cached prefill, sample and train [45]. The page lists the row at $2.96, $0.593, $7.41 and $8.89 per million tokens respectively, under a header that says "You pay only for the tokens you prefill, sample, and train" [8]. The same week, Fireworks' changelog put a serverless deprecation into effect [9], pointing Kimi K2.6 and GLM 5.2 users at GLM 5.3 [10]. Training priced like inference, on a shared always-on pool, is the more interesting half of this move: fine-tuning becomes a metered line item rather than a GPU reservation [11].

Baseten did similar housekeeping in one edit. Our capture shows GLM-5.3 Fast entering the Model APIs price list at $2.10 input, $0.21 cached input and $6.60 output per million tokens, DeepSeek V4.1 Flash's cached-input price cut from $0.03 to $0.007, and six models, including GLM 4.7, Kimi K2.6 and DeepSeek V4 Pro, removed [46]. The page lagged the changelog, which announced GLM 5.3 Fast on Sep 3 as "the same 753B-parameter model as GLM 5.3 on dedicated Fast capacity" [14] and scheduled the retirements for "5pm PT September 25th" [15]. The live page carries the new rows [16], and Baseten's $0.30 and $1.20 for DeepSeek V4.1 Flash match DeepSeek's own peak rates [20].

Fastly created an AI Platform category on its price page with two products. AI Runtime Control is metered per 10,000 requests, with 100,000 free each month, $2.50 per block up to a million requests and $1.75 beyond a hundred million [47][37][38]; AI Firewall is contact-sales [38]. The press release frames the pair as "real-time visibility and control across their AI systems" [35], and the routing product carries "visibility into token spend, rate limiting, budget controls, and failover" [39]. What matters for pricing is the unit: Fastly is charging for governed model calls, not tokens, which puts an edge vendor's meter in front of the model vendors' meters.

Notion removed the trial notice from its Workers (Beta) add-on. Our capture had the page stating Workers were free to try with credit billing starting October 15; that sentence is gone [48]. The live page says only that Workers "Requires Notion credits" and pitches them as a way to "Extend Notion with custom code to build agent tools, sync external data, and trigger Notion workflows from anywhere" [24]. Notion's help centre still calls Workers "free to try on Business and Enterprise plans" [22] and prices a run at $0.0023, or "~4,348 runs per 1,000 monthly Notion credits ($10)" [21]. Third-party guides still carry the October 15 date [25][27]; the web now says more about the deadline than Notion does.

Aircall redesigned its page around AI. Our capture shows the Custom plan renamed Enterprise with HIPAA, GDPR, SOC2 and SCIM listed, a standalone AI Voice Agents plan at a $100 per month minimum, Branded caller ID and Spam Management add-ons at $0.08 per call and $0.03 per check, and AI Messaging Agents simplified to a flat $0.15–$0.35 per conversation [49]. The page sells the voice plan as "Natural-sounding AI that answers and resolves calls, with or without a phone system" [29] and the messaging add-on as resolving "SMS and WhatsApp 24/7 on your existing numbers" [31]. The FAQ still says "A custom plan is available for teams of 25+ licenses" [30]. Pre-change reviews still describe a quote-only Custom tier [32][33][34].

What they said

The official framing this week is efficiency, not discounting. Anthropic says Opus 5.5 "performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5" [3], which turns a 20% list-price cut into a claimed 40% cost reduction once token efficiency is included [1]. Fireworks describes serverless training as a shared pool with no provisioning and no idle cost, where customers pay only for the tokens they prefill, sample and train [8]; Baseten describes GLM 5.3 Fast as the same model "on dedicated Fast capacity" and "distinct from GLM 5.3 Flash" [14], which is a speed tier sold at a higher per-token rate rather than a different model. Fireworks, for its part, positions GLM 5.3 Flash as "the first natively multimodal model in the GLM-5 series" [12].

Fastly's Kelly Shortridge argues that enterprises "need control in production, at runtime, without delay, friction, or disruption" [36], and the launch coverage ties the product to a finding that 93% of organizations are exceeding their AI budgets [40]. Notion's help centre keeps the economics concrete: Workers "typically cost $0.0023 per run" [21] and are "cheaper compared to Custom Agents because pricing is based on how much work they do when they run" [23]. Aircall sells the voice plan as working "with or without a phone system" [29], which is the tell: the AI agent is now the product, and the phone system is optional.

What the market said

Coverage confirmed the Anthropic numbers within hours: TechCrunch reported that "Output tokens will be charged at $20 per million tokens for Opus 5.5, compared to $25 for the previous model" [5], and TestingCatalog gave the full pair against Opus 5 [6]. On Hacker News the reaction went straight to margins: one commenter noted Opus 5 was the highest-spend model on OpenRouter and argued that being "forced to reduce price despite raising capabilities" says something about the market [7], while replies countered that cost to serve had simply fallen. The GLM-5.3-Flash thread reads the inference providers the same way: unlike the labs they are not heavily leveraged, so "the more hardware they can bring online, the cheaper they can serve tokens" [13]. For Fastly, the trade press repeated the feature list [42][43] and Futurum supplied the commercial logic, arguing that "a single platform that simultaneously protects AI applications and enforces cost policies is a materially stronger procurement argument than two separate point solutions" [41]; we found no community thread on the pricing itself.

The web lags our capture in three places. Aircall guides from April, July and August still describe a quote-only Custom tier and per-minute voice pricing [33][32][34], and Aircall's own FAQ has not caught up with the Enterprise rename [30]. Notion's October 15 deadline now lives in third-party guides [25][26][27] and in the analyst question of "Whether Workers billing on October 15 creates new friction" [28], not on Notion's page [24]. Baseten is the reverse case: the changelog [14][15] and coverage of the DeepSeek V4.1 Flash addition [17] ran weeks ahead of the price list, and a LiteLLM pull request had to add the GLM-5.3-Fast entry so requests stopped logging zero spend [19]; Vercel's gateway lists Baseten as a provider at the same $2.10 and $6.60 [18].

The pattern

The week's volume was the highest in the Monitor's trailing window: 173 AI-related changes in W40 against 94 at the low point in W36 [50], and 107 AI-relevant diffs among 276 published changes across 4,388 tracked pages [51]. Of the changes in the last window, 51% bundled AI into existing tiers, 31% priced it by usage, credits or tokens, 10% created or restructured a tier and 8% sold it as a separate paid add-on [52]. This week's leads sit almost entirely in the usage and add-on buckets: Fireworks, Baseten, Fastly and Notion are all metering, and Aircall and Oneflow are packaging AI as add-ons.

Credits are the fastest-growing theme: 338 credit-related changes in the last window against 208 in the prior one, from 183 companies, with 145 introducing credits for the first time; 94 expanded an allowance and 50 shrank one [53]. Free AI moved the same way, 131 changes against 80, with 32 expansions and 24 tightenings [54]. Notion's Workers deadline and Fireworks' token-priced training are both credit and usage stories; Anthropic's cut is the exception that proves the metering rule, since a lower per-token rate only matters when the meter is running.

Direction still points up. Across the window, 70 AI-related price increases were logged against 27 decreases, with a median quantified increase of about 15% [55]. Agent pricing remains vague: of 293 agent-related changes, only 16 stated a per-seat meter and 1 a per-conversation meter, with 276 unspecified [56]. Aircall's flat per-conversation rate and per-month voice minimum are therefore unusual for stating a unit at all, and Fastly's per-request tiers are unusual for stating one below the token.

What to watch

  • Whether Notion restores a billing date for Workers before October 15, or lets the help centre's "eventually" stand while credits start to burn.

  • Whether Anthropic's legacy-tier pattern spreads: Opus 5 kept at $5/$25 rather than repriced, which rewards migration without punishing incumbents.

  • Whether Fireworks' per-token training table grows beyond its six launch models, and whether Together or Baseten answer with a metered training line.

  • Whether Aircall's FAQ catches up with the Enterprise rename, and whether the $100 voice-agent minimum shows up in competitor pricing guides.

  • Whether Fastly's per-request AI Runtime Control meter becomes the template for other edge vendors pricing AI governance.

Sources

  1. [1]Anthropic — Introducing Claude Opus 5.5verified“Input and output tokens are $4 and $20 per million, 20% less than Opus 5.”
  2. [2]Anthropic — Introducing Claude Opus 5.5 (cache pricing)verified“Cache reads (which make up the majority of agentic and coding work costs) are $0.20 per million tokens, 60% less than Opus 5.”
  3. [3]Anthropic — Introducing Claude Opus 5.5 (cost claim)verified“It performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5.”
  4. [4]Claude — Pricing (live page)unverified“Daily driver for agentic coding and enterprise work”
  5. [5]TechCrunch — Anthropic releases Opus 5.5 with lower prices and Fable-level performanceverified“Output tokens will be charged at $20 per million tokens for Opus 5.5, compared to $25 for the previous model.”
  6. [6]TestingCatalog — Anthropic launches Claude Opus 5.5 with lower API costsverified“API pricing is $4 per million input tokens and $20 per million output tokens, compared with $5 and $25 for Opus 5.”
  7. [7]Hacker News — Claude Opus 5.5 (comment on the price drop)verified“If you are forced to reduce price despite raising capabilities, that certainly tells something about the market, and potentially about Anthropic future profitability too”
  8. [8]Fireworks — Pricing (Serverless Training API)verified“You pay only for the tokens you prefill, sample, and train. Base model Context Prefill / 1M Cached prefill / 1M Sample / 1M Train / 1M GLM 5.3 Flash 200K $2.96 $0.593 $7.41 $8.89”
  9. [9]Fireworks — Changelog (serverless deprecation)verified“The serverless deprecation announced for September 25, 2026 is now in effect. The models below are no longer available on public serverless”
  10. [10]Fireworks — Changelog (recommended migrations)verified“Kimi K2.6 — migrate to GLM 5.3 or Kimi K3 Kimi K2.7 Code — migrate to GLM 5.3 or Kimi K3 GLM 5.2 , including GLM 5.2 Fast , GLM 5.2 Fast US , and GLM 5.2 US — migrate to GLM 5.3”
  11. [11]Unite.AI — Fireworks AI Makes Training API Generally Availableverified“Serverless training lets customers train LoRA adapters on shared infrastructure with per-token billing, with sampling running in the same session”
  12. [12]Fireworks — GLM 5.3 Flash model pageverified“GLM-5.3-Flash is the first natively multimodal model in the GLM-5 series with 320B total parameters and 18B active parameters.”
  13. [13]Hacker News — GLM-5.3-Flashverified“the more hardware they can bring online, the cheaper they can serve tokens”
  14. [14]Baseten — Changelog (GLM 5.3 Fast available)verified“Sep 3, 2026 GLM 5.3 Fast available on Baseten GLM 5.3 Fast is now available through Baseten Model APIs. It serves the same 753B-parameter model as GLM 5.3 on dedicated Fast capacity and is distinct from GLM 5.3 Flash.”
  15. [15]Baseten — Changelog (Model API deprecation)verified“GLM 4.7, Kimi K2.7, Kimi K2.6, Inkling, Inkling Small, and DeepSeek V4 Pro will be deprecated at 5pm PT September 25th.”
  16. [16]Baseten — Pricing (Model APIs, live page)verified“DeepSeek V4.1 Flash $0.30 $0.30 $0.007 $0.007 $1.20”
  17. [17]Unite.AI — Baseten Adds DeepSeek-V4.1-Flash to Model APIsverified“DeepSeek-V4.1-Flash is available now on Baseten Model APIs, Baseten announced on September 11, 2026”
  18. [18]Vercel AI Gateway — GLM 5.3 Fastverified“Speed-optimized version of Z.AI's GLM-5.3 agentic coding model, built for responsive, real-time workloads.”
  19. [19]GitHub — LiteLLM PR: add baseten/zai-org/GLM-5.3-Fast pricingverified“Adds the GLM-5.3-Fast entry next to GLM-5.3 and GLM-5.3-Flash”
  20. [20]DEV Community — AI Weekly: DeepSeek Cuts Prices as Agents Go Hostedverified“peak rates are $0.30 per million input tokens and $1.20 per million output tokens, down from $0.44 and $1.32 for V4 Flash. The cache read dropped from $0.014 to $0.006, a 57 percent cut.”
  21. [21]Notion Help — Understand pricing for Workers (per-run cost)verified“Workers typically cost $0.0023 per run, which works out to ~4,348 runs per 1,000 monthly Notion credits ($10).”
  22. [22]Notion Help — Understand pricing for Workers (beta status)verified“During the beta, Workers are free to try on Business and Enterprise plans (including Business trials).”
  23. [23]Notion Help — Understand pricing for Workers (vs Custom Agents)verified“cheaper compared to Custom Agents because pricing is based on how much work they do when they run”
  24. [24]Notion — Pricing (live page)verified“Extend Notion with custom code to build agent tools, sync external data, and trigger Notion workflows from anywhere.”
  25. [25]Coworker.ai — Notion AI Pricing in 2026verified“free to try now and starts consuming credits on October 15, per the same page.”
  26. [26]Breeze — Notion AI pricing in 2026: plans, credits and Workersverified“Workers are free to try on Business and Enterprise during the beta, including Business trials, and will eventually require Notion credits.”
  27. [27]Matthias Frank — Notion Workers: The Complete Beginner's Guideverified“Notion Workers are currently in public beta on Business and Enterprise plans. Free during the beta, then on Notion credits from 15 October 2026.”
  28. [28]Pricing Innovation — Notion's AI pricing journey: from flat fee to tier gate to creditsverified“Whether Workers billing on October 15 creates new friction.”
  29. [29]Aircall — Pricing (AI Voice Agents plan)verified“Natural-sounding AI that answers and resolves calls, with or without a phone system. $100 per month minimum”
  30. [30]Aircall — Pricing (FAQ still names the Custom plan)verified“A custom plan is available for teams of 25+ licenses with SLA, SSO, and dedicated onboarding.”
  31. [31]Aircall — Pricing (AI Messaging Agents add-on)verified“Resolve SMS and WhatsApp 24/7 on your existing numbers, with full-context handoff. $0.15–$0.35 per conversation”
  32. [32]Aloware — Aircall pricing 2026: Plans, AI costs, and hidden feesverified“AI Voice Agents cost $0.49 per minute, dropping to $0.39 above 2,551 minutes”
  33. [33]CloudTalk — Aircall Pricing: Breakdown of Plans, Costs, and Extra Feesverified“Aircall pricing runs from $30 to $50 per license/month on annual billing across two published plans—Essentials and Professional—plus a quote-based Custom tier.”
  34. [34]JustCall — Aircall Review 2026verified“a single-user Starter AI plan also exists, built around AI Voice Agents rather than as a team phone system”
  35. [35]Fastly — Press release: Fastly Launches AI Firewall and AI Runtime Controlverified“today announced AI Runtime Control, AI Firewall, and new API Security capabilities designed to give organizations real-time visibility and control across their AI systems.”
  36. [36]Fastly — Press release (Kelly Shortridge quote)verified“enterprises need control in production, at runtime, without delay, friction, or disruption”
  37. [37]Fastly — Pricing (AI Runtime Control tiers)verified“Requests 100,000 free Billed per 10,000 requests Requests (per month) Monthly price (per 10,000 requests) 0 - 100,000 Free 100,000 - 1 million $2.50”
  38. [38]Fastly — Pricing (top tier and AI Firewall)verified“25 million - 100 million $2.00 Beyond 100 million $1.75 AI Firewall Protect AI-enabled applications from LLM-based threats Contact sales for pricing”
  39. [39]Help Net Security — Fastly gives enterprises real-time control over AI models and agentsverified“With real-time visibility into token spend, rate limiting, budget controls, and failover, organizations maintain model choice while enforcing policy centrally.”
  40. [40]Help Net Security — Fastly AI Runtime Control (budget context)verified“McKinsey finding that 93% of organizations are exceeding their AI budgets”
  41. [41]Futurum Group — Fastly bets the edge on AI governance as machine traffic tops 50%verified“a single platform that simultaneously protects AI applications and enforces cost policies is a materially stronger procurement argument than two separate point solutions.”
  42. [42]Channel Insider — Fastly Launches AI Runtime Controls for Enterprise Agentsverified“All three capabilities are now available, according to Fastly.”
  43. [43]Security Today — Fastly Launches AI Firewall and Runtime Controlverified“AI Runtime Control provides a central endpoint for requests to public and self-hosted providers.”
  44. [44]Pulse capture — Anthropic API 2026W40“Anthropic API: Opus 5.5 launches, replaces Opus 5 as flagship”
  45. [45]Pulse capture — Fireworks 2026W40“Fireworks adds GLM 5.3 Flash with 200K context, per-token training pricing”
  46. [46]Pulse capture — Baseten 2026W40“Baseten adds GLM-5.3 Fast, retires 6 models, cuts DeepSeek V4.1 Flash cache price 77%”
  47. [47]Pulse capture — Fastly 2026W40“Fastly.default launches AI Platform: usage-based Runtime Control from $1.75”
  48. [48]Pulse capture — Notion 2026W40“Notion silently drops Workers trial notice ahead of Oct 15 billing”
  49. [49]Pulse capture — Aircall 2026W40“Aircall: Custom→Enterprise rename + AI Voice Agents plan ($100/mo min)”
  50. [50]AI Monetization Monitor — weekly counts“weekly counts W28–W40: 119, 134, 102, 111, 99, 98, 155, 137, 94, 141, 121, 128, 173”
  51. [51]AI Monetization Monitor — this week's totals“ai relevant diffs 107; week changes 276; pages tracked 4388”
  52. [52]AI Monetization Monitor — mix“mix: Bundled into existing tiers 782 (51%); Usage / credits / tokens 474 (31%); New / restructured tier 160 (10%); Separate paid add-on 115 (8%)”
  53. [53]AI Monetization Monitor — credits theme“themes.credits: now 338, prior 208, companies 183, first time 145, shrank 50, expanded 94”
  54. [54]AI Monetization Monitor — free AI theme“themes.free ai: now 131, prior 80, companies 84, tightened 24, expanded 32”
  55. [55]AI Monetization Monitor — direction“direction: inc 70, dec 27, inc median 15.2”
  56. [56]AI Monetization Monitor — agent metering“themes.agents: now 293, prior 281; metering Per seat 16, Per conversation / message 1, Other / unspecified 276”

Questions about this brief?

The Pulse Agent has the moves, the monitor numbers and every source loaded.

Ask the agent

By the numbers

AI pricing moves, last 13 weeks· +21% vs prior 13 weeks
1,612
Credit-priced moves· 145 companies priced in credits for the first time
338
App / infra price index· spread 0
100 / 100
AI moves this week· 173 at the window's peak
173