Systems & AdoptionAugust 30, 2026·5 min read

Your AI Costs Keep Blowing Past Budget. It's the Same Mistake Four Times.

A silent 40% token hike, a subsidized subscription, a budget built on sticker price, and an agent deployed where a prompt would do — four unrelated cost overruns from this year share one root cause. Here's the check that catches it before the invoice does.

By Patin Team · Examples are illustrative composites

If your AI tools bill has surprised you at least once this year, it wasn't one unlucky invoice. It's been the same mistake, showing up four different ways, in four stories that never referenced each other.

Four stories, one mechanism

In April, Claude's Opus 4.7 upgrade quietly used about 40% more tokens per task than the version before it — no pricing announcement, no line on the plans page, just a bigger number on anyone billed by usage.

In May, Simon Willison disclosed that his $200/month Claude subscription was generating $2,180 in actual compute. A 10x subsidy, paid to build market share, not a stable price.

In June, the same pattern hit budgets directly: Uber blew through its annual AI budget in months, because the number finance had approved was built from subscription sticker price, not from what the workflows actually consumed once they scaled.

That same month, a third version of it showed up as a tooling choice: a marketer's competitive-intelligence agent tripled her bill running continuous monitoring for work her team only read once a day. A scheduled prompt would have done the job at roughly a fortieth of the cost.

What's actually happening

None of these four teams did anything reckless. They upgraded when prompted, subscribed at the advertised price, budgeted from the invoice in front of them, and reached for the most capable tool available. Each decision was reasonable in isolation. What none of them did was price the task before choosing how to run it — what it actually costs in compute, at the tier of capability it actually needs, independent of what the subscription happens to charge this quarter.

That's the one mechanism under all four stories: the price on the invoice and the cost of the work are two different numbers, and only one of them is currently visible to you.

What to do differently Monday morning

Before your next AI tool renewal, upgrade, or deployment decision, run one number: what would this specific task cost if you paid for the compute it actually uses, at the simplest tier — single prompt, bounded workflow, or always-on agent — that would still do the job well? Not the subscription price. Not what a vendor charges today while it's still subsidizing adoption. The real number, for the real tier.

You don't need to change anything based on that number today. You need to know it before the subsidy narrows, the model upgrades again, or someone reaches for an agent because it's available rather than because the task needs it.

Raj: setting a real budget instead of a subscription budget

Raj runs operations for a 40-person logistics brokerage. Four dispatchers use Claude Pro at $20/month each; the team budgeted $960 for the year and called it done. When he ran Willison's ratio against that spend, the real compute those four accounts were likely drawing came out closer to $8,000–10,000 annually — a number his current budget doesn't survive if pricing normalizes even partway.

He didn't cut anything. He set a task-level budget instead: for each recurring AI use in his team, what's the maximum this specific task is worth paying for, regardless of what the subscription currently charges. Two uses cleared that bar easily. One — a nightly shipment-status summary nobody had ever timed — didn't, and he moved it to a once-a-day scheduled prompt instead of leaving it on an always-on connector.

Elena: matching tier to task before deploying anything new

Elena leads content ops at a 90-person media company. Her team wanted a research agent that would monitor competitor publishing and flag gaps continuously. Before deploying it, she asked the tier question first: does this need to watch continuously, or does it need to answer once a day?

The answer was once a day. She built a scheduled prompt that runs each morning instead of an always-on agent, and the team gets the same gap-flagging with a bill that scales with days run, not with how much the underlying model happens to think per query.

The one thing

Four teams hit the same wall this year for the same underlying reason: none of them had priced the work before the tool told them what it cost. The subsidy will narrow, the models will keep upgrading their own token use, and agents will keep being the most capable option in the room. The number worth knowing before any of that happens isn't what the invoice says today — it's what the task is actually worth running at the tier it actually needs.

Reading about it only gets you so far

Patin turns this into five-minute drills that score what you write and tell you why. It's in closed beta — join the waitlist and we'll email you when your cohort opens.

Just want the writing? .