How Much Checking Is Enough?
Check everything and you've saved nothing. Check nothing and you're gambling. The right amount is a function of two things, and neither of them is how much you trust the tool.
By Patin Team · Examples are illustrative composites
There's a version of AI adoption where every output gets fully verified. It's safe, it's defensible, and it produces no time saving whatsoever — you've replaced writing with proofreading, and proofreading someone else's work is often slower.
There's another version where nothing gets checked, which works for months and then doesn't.
Almost everyone is oscillating between these, usually without deciding. The right amount of checking is a real question with a real answer, and it depends on two things — neither of which is how much you trust the tool.
The two variables
What it costs to be wrong. Not the probability of an error, the consequence of one. An error in an internal note costs a correction. An error in a client deliverable costs credibility. An error in a regulatory filing costs something else again.
Whether you'd find out. The one people don't think about, and the one that matters more. Some errors announce themselves — the client replies confused, the numbers don't reconcile, the code doesn't run. Others are silent, and silence isn't evidence of correctness.
The combination gives you four situations, and each wants something different.
The four cases
Low cost, self-announcing. Internal drafts, exploratory work, anything that gets corrected in conversation. A glance is proportionate. Deep checking here is where most wasted verification effort goes.
Low cost, silent. Categorisations, internal metrics, summaries nobody compares against the source. Individually cheap, and they accumulate. This is the sampling case: check one in ten or twenty properly, at random, on a schedule. What you're looking for is drift rather than individual errors.
High cost, self-announcing. Client work, anything public. Full check on this output — but you can stop worrying about the class, because failures here surface fast and loudly.
High cost, silent. The dangerous quadrant. Advice acted on quietly, figures feeding a decision nobody revisits, compliance work nobody tests until an audit. This gets full checking and an independent check on the process itself, because there's no natural feedback telling you the process has degraded.
Most people can place their recurring outputs into these four in about ten minutes, and most find at least one thing sitting in the last box being treated like the first.
Sampling is a real answer
For repeated low-cost work, checking one in ten is not a compromise — it's the correct method, and it's how quality control has worked in every other high-volume domain for a century.
Two conditions make it work. Random selection, not by which looks interesting, because choosing by interest systematically misses the boring failures. And a schedule you keep, because sampling that lapses is indistinguishable from not checking.
The thing sampling catches is drift: the gradual failure where the process was right for last quarter's inputs and is subtly wrong for this quarter's. Nothing else catches that.
Check the process, not just the output
For anything recurring, output-level checking has a ceiling. What's more valuable is periodically checking the system: is the brief still right, are the inputs still what they were, is the standard still what we'd set today?
This is a twenty-minute conversation once a quarter, and it catches a category of problem that no amount of per-output checking will — because per-output checks are done against a standard that itself has stopped being right.
The signal that you've got it wrong
Checking is taking longer than doing. Either the task shouldn't be delegated or the brief is too weak. Both are fixable; neither is fixed by checking harder.
You haven't found anything in months. Either you're checking the wrong things, or you're checking things that no longer need it. Move the effort somewhere it's finding something.
You'd struggle to say what checking means here. Then it isn't happening, whatever it looks like from outside.
Vikram — the quadrant he'd got wrong
Vikram is a management consultant. He treated client-facing work as high stakes and internal analysis as low, which sounded right.
The audit that changed his mind: his internal analysis fed client recommendations weeks later, by which point nobody traced anything back. It was high-cost and silent, filed under low-cost and self-announcing.
He now checks internal analysis to the same standard as client work, and has stopped over-checking client emails, which get read by someone who'd reply if they were odd.
Solène — the sample she kept
Solène leads customer operations at an insurer. An AI classifier routed claims by complexity, and she sampled twenty a month against her own judgement.
For seven months the agreement rate sat around 94%. In month eight it fell to 81%, which she traced to a new claim type introduced by a product change — a category the classifier had never seen and handled by analogy to something else.
Her point about it: no single misrouted claim would have been noticed by anyone. The number falling was the only thing that could have told her, and it only existed because she'd taken the same measure every month for seven months.
The one thing
How much to check depends on what an error costs and whether you'd ever find out — not on how much you trust the tool.
Sample the repetitive, fully check the consequential, and give the high-cost-and-silent quadrant an independent look at the process itself. If checking takes longer than doing, the problem is upstream of the checking.
Put this into practice
Reading is a start — but skill comes from doing. Try these drills now.
Reading about it only gets you so far
Patin turns this into five-minute drills that score what you write and tell you why. It's in closed beta — join the waitlist and we'll email you when your cohort opens.
Just want the writing? .
Keep reading on this
Decide What Good Looks Like Before You Generate
Most of the effort in reviewing AI output goes on working out what you wanted while looking at what you got. Three sentences written beforehand turns a judgement into a comparison.
5 min readReviewing AI Output: What to Actually Check
Reading it and approving it isn't a review — it's the exact filter AI output is best at passing. Interrogating is a different activity, it takes about ninety seconds, and it's most of what separates the people getting value from AI.
5 min readA $2B Company's AI Strategy Was Written by Someone Who Never Used ChatGPT. The Backlash Has Begun.
Three communities — corporate critics, cognitive scientists, and content readers — arrived at the same conclusion this week: uncritical AI use is producing bad decisions, eroding judgment, and generating detectable slop. Here's what to do about it.
4 min read