Measuring Whether AI Is Actually Working
Licences issued and hours saved are the two most common AI metrics and neither survives contact with reality. Four measures that tell you something, and one number worth having.
By Patin Team · Examples are illustrative composites
Two numbers dominate AI reporting, and neither means anything.
Licences issued measures procurement. It's reported as adoption because it's the only figure available on day one, and it stays flat while the interesting things happen underneath it.
Hours saved is worse, because it's usually an estimate multiplied by a headcount. Ask where the figure came from and it's typically self-reported time savings on individual tasks, extrapolated. It doesn't survive the obvious follow-up: what happened in the saved hours? If the answer is "more of the same work at the same headcount", nothing was saved — the standard moved.
What's worth measuring is harder to game and easier to collect than either.
Persistence: is anyone still doing it in month three?
The single best indicator, and nearly free.
Take the workflows people adopted in month one. In month three, ask which are still running unprompted. Enthusiasm at week two tells you about novelty; something still in use at week twelve is genuinely better than what it replaced, because nobody is sustaining a habit that isn't paying.
Expect attrition and don't treat it as failure. Three of eight surviving is a good outcome — you've found three real things and stopped paying attention to five that weren't.
Cycle time on a specific deliverable
Not "productivity". One named thing your team produces regularly: the monthly report, the tender response, the client onboarding pack.
Measure it end to end, including review and rework. This is where the honest picture appears — the drafting got much faster and the reviewing got slower, and the net is sometimes smaller than the drafting number suggests, occasionally negative.
Whichever it is, it's a real number about a real thing, and you can act on it. "23% more productive" isn't actionable in any direction.
Rework rate
The number nobody collects and the one that catches the failures.
What proportion of AI-assisted work comes back — corrections, complaints, revisions, things spotted at review? You want it trending flat or down. Rising rework alongside rising output means the quality cost is being paid downstream by whoever reviews, and that person is usually not the one reporting the time saving.
Even a rough count kept for a quarter tells you more than any survey.
Where the freed time went
If a workflow saved real hours, something must have absorbed them. Name it.
Legitimate answers: more of a valuable activity that was previously squeezed, faster turnaround for a client, a backlog cleared, work brought back in-house. Also legitimate: people are less exhausted, which is a real return even if it never shows in a metric.
The answer that should worry you is "nobody knows". Unaccounted time usually means the volume expectation rose to fill it, which is a workload change described as an efficiency gain.
The one number worth having
If you keep only one thing, keep a quality sample.
For one repeated AI-assisted output, take a fixed number — ten, twenty — at random each month and assess them against your own standard. Record the figure.
A single month's number is nearly meaningless. The series is the point: it's the only instrument that detects drift, which is the failure mode that matters most and announces itself least. Everything else on this list tells you about speed; this tells you whether the thing is still working.
What not to measure
Prompts per person. Rewards activity over outcome, and the people getting most value often use it least often on bigger tasks.
Anything self-reported about time. People are poor at estimating time saved and there's mild social pressure toward optimism. It's not dishonesty, it's measurement error, and it compounds when multiplied by headcount.
Tool-reported engagement. Vendor dashboards measure engagement with the vendor's product. Reasonable for them, not evidence for you.
Rahul — the report that took the same time
Rahul runs an analytics team. Everyone agreed AI had made reporting dramatically faster, and his measurement of the monthly report end to end found the total unchanged.
Drafting had gone from three days to half a day. Review had gone from half a day to two and a half, because the drafts needed a different kind of checking — plausible and occasionally wrong beats obviously incomplete for readability and loses badly for reviewability.
Once he could see it, he fixed the actual problem: better criteria in the brief, and a source requirement on every figure. Review came back down. His comment is that the team's perception had been accurate about their own experience and wrong about the deliverable.
Katya — the series she started
Katya is head of operations at a logistics firm. Her only AI metric is a monthly sample: twenty AI-assisted customer responses, scored against the team's own criteria.
The first month's number told her nothing. The value appeared at month five, when the score dropped and she could see it was a drop rather than an impression.
Her view is that every other metric she'd tried measured how much AI was being used, and this is the only one that measures whether it's still working.
The one thing
Licences and estimated hours saved measure procurement and optimism. Measure persistence at month three, cycle time on a named deliverable including review, rework rate, and where the freed time went.
If you keep one thing, keep a monthly quality sample. It's the only measure that catches drift, and drift is the failure you won't otherwise see coming.
Put this into practice
Reading is a start — but skill comes from doing. Try these drills now.
Reading about it only gets you so far
Patin turns this into five-minute drills that score what you write and tell you why. It's in closed beta — join the waitlist and we'll email you when your cohort opens.
Just want the writing? .
Keep reading on this
Your AI Costs Blew Up Because You Deployed an Agent. A Prompt Would Have Done It.
Uber exhausted its 2026 AI budget in four months. Sam Altman says enterprise cost complaints are now 'a meme.' The root cause isn't reckless spend — it's deploying always-on agents for tasks a single prompt could handle. Here's the three-tier framework that fixes it.
5 min readYour AI Budget Is Already Wrong. Here's How to Fix It.
Uber blew its annual AI budget in months. Simon Willison's $200 subscription runs $2,180 in compute. Anthropic's revenue went 5x in five months. If your team budgeted for AI tools in Q4, those numbers are already wrong — here's how to reset them.
4 min readAI Subscription Prices Are Subsidised. Here's What Happens When They're Not.
Claude Pro subscribers consume $200–400 of compute per $20/month. OpenAI loses $1.22 for every $1 earned. GitHub Copilot switches to usage-based billing June 1. The price you pay for AI today is not the price it costs to run. Here's how to plan for when that changes.
5 min read