TEARDOWN Nº 3 / THE SONNET 5 PRICE CLIFF
The Sonnet 5 price cliff, priced before it happens
PUBLISHED 2026-08-13 · BY Llorenç Ballester — Chief Burn Officer
In 19 days — 2026-08-31 — Claude Sonnet 5's intro pricing ends. Input goes from $2 to $3 per million tokens, output from $10 to $15. That's already documented on the pricing board today — not a forecast. We ran it through three real workloads so the number isn't abstract when it lands.
A coding-agent workload on Sonnet 5 goes from $36.70 to $51.22/month at confirmed prices (+39.6%) — or to $55.04 (+50.0%) if cache pricing moves with it, which is not yet published. Either way: base input/output pricing is a confirmed, dated 50% hike. What caching does to that number is the one thing on this page we can't yet cite a source for.
What's confirmed, and what isn't
Two different kinds of claim live on this page, and we're not going to blur them. The base price change is confirmed: Sonnet 5's own pricing page already documents “Intro pricing — list is $3/$15 per MTok until 2026-08-31”, sourced from Anthropic's own pricing page and checked daily by our watcher. That's a fact on the record today, not a guess about the future.
What Sonnet 5's cached-read and cache-write rates will do is not published anywhere — not by Anthropic, not in the dataset we cross-check against. What we can point to is a pattern: every one of the 5 Anthropic models on our board prices cached reads at exactly 10% of that model's own input price, and cache writes at exactly 1.25x — zero exceptions, verified against today's board. If that holds after the reversion, cache pricing moves proportionally. If it doesn't, our projected numbers below are wrong in the direction of overestimating the hike's cushion. We show both scenarios so you can pick the one you trust.
Three workloads, before and after
“Confirmed” holds cache pricing at today's dollar rate and moves only input/output — it isolates exactly what the documented change does. “Projected” also scales cache pricing with the pattern above. The gap between the two columns is entirely the unpublished part.
| Workload | Today | Confirmed after | Projected after | Confirmed Δ |
|---|---|---|---|---|
| Chat app | $384 | $576 | $576 | +50.0% |
| Chat app (cached) | $337 | $493 | $506 | +46.3% |
| Coding agent | $36.70 | $51.22 | $55.04 | +39.6% |
Monthly figures at each workload's own published volume (full assumptions). Prices from the board, last confirmed 2026-08-12. Estimates at list prices, not invoices.
The pattern in the table is the finding: the more a workload leans on cached reads, the more the confirmed increase is cushioned below 50% — cached-chat at +46.3% instead of the plain chat workload's +50.0%. That cushion is real today. It is also the part of this whole page that's built on the shakiest ground — it disappears entirely in the projected column, where every workload converges back to almost exactly +50%.
Why a scheduled price change is worth a teardown
Most price coverage is retrospective — a change lands, someone writes it up after the fact. This one is already fully specified today: the date, the before, and the after are all on Sonnet 5's own pricing page right now. That makes it one of the few AI pricing events you can actually plan a budget around instead of reacting to. Our daily watcher will apply the change automatically on 2026-08-31 and post it — this page is the explainer for when that happens, written before it does.
Method and limits
Every figure is computed from the pricing board and the same canonical workloads published at /workloads— nothing here is a new assumption about usage shape. The base price change is a confirmed fact, dated and sourced. The cache-rate projection is explicitly ours: an extrapolation from a pattern that is true of every Anthropic model today and has no guarantee of holding tomorrow. If you're reading this after 2026-08-31, the live board and the changelog already show what actually happened — treat the numbers on this page as the pre-event estimate they were written as. Full rules at /methodology.
Frequently asked questions
Is Claude Sonnet 5 getting more expensive?
Its list price isn't changing — it's returning to what it always was. Sonnet 5 launched at $3 input / $15 output per million tokens and has been running an intro discount at $2/$10 since. That discount is documented to end 2026-08-31, after which the price reverts to list. Whether Anthropic extends the intro again is not something we can predict; what's on the record today is the end date and the list price it reverts to.
Will prompt caching get more expensive too?
We don't know yet, and we say so explicitly on this page. Anthropic doesn't publish a forward cache price for Sonnet 5. What we do know: every one of the 5 Anthropic models currently on our board prices cached reads at exactly 10% of that model's own input price and cache writes at exactly 1.25x — no exceptions. If that pattern holds, cache rates move proportionally with the base price. We show both the confirmed scenario (cache prices held at today's dollars) and the projected one (cache scaling with the pattern) side by side.
How much will my AI bill actually go up on September 1st?
It depends entirely on how much of your usage is Sonnet 5 and how much of that is served from cache. A workload with no caching sees close to the full 50% increase. A workload that leans hard on cached reads sees less of it — but only in the confirmed scenario, where we deliberately don't move the cache price. If the cache price moves with the pattern, the cushion disappears and every workload converges back toward 50%.
What should I do before the price changes?
Nothing panicked — a scheduled reversion to a previously published list price isn't an emergency. But if Sonnet 5 is a meaningful share of a fixed monthly budget, this is a reasonable moment to check your actual mix of models via /audit, and to decide in advance whether routing some of that volume to a cheaper model is worth it — see our model-routing guide for the mechanics of that trade-off.
- AI Tokenomics: what tokens actually cost
- How much does ChatGPT actually cost per prompt?
- Prompt caching: the 90% discount most teams never claim
- Why your AI agent costs 10–40x more than a chat
- Why output tokens cost more than input
- Cut AI costs with model routing
- The cache minimum: why your prompt cache silently does nothing
- What a Cursor session costs, by model
- How to read your AI usage export
- → The Burnmeter: measure your own token waste