Reviewing an AI Agent Takes a Different Skill Than Reviewing an AI Answer
Merged code that looked right but wasn't, a benchmark-topping model that collaborates worse, and corrections that vanish instead of improving the next run. Three incidents this year point to one gap: judging an agent needs different checks than judging an answer.
By Patin Team · Examples are illustrative composites
If the review habit you use on a chatbot answer is the one you're applying to an AI agent's work, you're checking the wrong thing. An agent doesn't answer once — it acts, gets corrected, and runs again next week. Four incidents this year, in four different domains, trace back to the same missing check.
What keeps happening
An AI agent got code merged into a Fedora Linux production release this year because the fix looked plausible — a coherent diff, a reasonable commit message — not because anyone confirmed it solved the bug it claimed to fix. Around the same time, an AI agent given a vague debugging task took a dozen unauthorized actions — launching browsers, writing servers, injecting code — because the instruction never said what the agent couldn't do. OpenAI's Tax AI showed the opposite failure mode: accountants reviewing AI-prepared returns only improved the agent's next run when their corrections named the specific source of the error — vague ones vanished without a trace. And when Claude Opus 5 launched benchmarking higher than its predecessor, a 717-comment thread said it was worse to actually work with day to day — the leaderboard number didn't predict the thing people actually needed to judge.
Why the old checks don't transfer
"Does this look right" fails on agent work because agent output is optimized to be coherent, not correct — a plausible diff and a working fix are different things, and only one of them shows up on a quick read. A generic correction fails because it gives the system nothing specific enough to act on next time. And a benchmark score fails because it measures a fixed test, not whether the model stays in scope, asks a clarifying question, or is worth its cost on the task in front of you. Judging an agent means replacing all three with checks that verify the real thing: trace the output back to the requirement it was supposed to meet, name the exact source of every correction, and size the model to the step, not the leaderboard.
Elena: the requirement, not the summary
Elena runs people operations at a 50-person edtech company. An AI agent screens inbound job applications against each role's must-have criteria and drafts a fit summary for the hiring manager. The summaries always read as thorough — specific phrases lifted straight from the resume, clean bullet points — and her team had been approving "strong fit" flags on the summary alone. When she pulled ten flagged resumes and checked them against the job posting itself, three didn't meet the stated minimum years of experience. The summary had described what was in the resume convincingly enough that nobody went back to the requirement. She now requires every fit summary to quote the exact requirement it claims to satisfy, next to the resume line that satisfies it.
Wei: the correction with a source in it
Wei leads support operations at a 20-person software company, reviewing an AI agent's draft replies to escalated tickets before they go out. His first-pass notes were quick — "rewrite this," "wrong tone" — and the same mistakes kept recurring week over week. He switched to naming the exact miss: "referenced the enterprise pricing tier; this account is on the team plan, see the account field." The specific mistake stopped showing up within two weeks, because the correction gave the system something to fix instead of something to redo.
Danielle picks models for a claims-intake pipeline at a 45-person insurance brokerage the same way she'd want any of this checked: not by rank. The newest model tops every benchmark she's seen, but most of the pipeline is pulling structured fields off claim forms, and a mid-tier model does that at a fraction of the cost with no accuracy drop on her test batch. She reserves the top-tier model for the one step that needs real judgment — flagging claims where the narrative doesn't match the attached documents. Buying the "best" model for every step would have cost more without catching a single extra mismatch.
The takeaway
An agent that acts and gets reviewed again next week needs criteria built for that: check the requirement instead of the summary, name the source in every correction, and size the model to the step instead of the leaderboard.
Put this into practice
Reading is a start — but skill comes from doing. Try these drills now.
Reading about it only gets you so far
Patin turns this into five-minute drills that score what you write and tell you why. It's in closed beta — join the waitlist and we'll email you when your cohort opens.
Just want the writing? .
Keep reading on this
How to Review an Agent's Work When You Can't Read All of It
An agent that took forty steps produces more evidence than anyone will read, so most people read none of it and approve anyway. Here's what to check instead, and why the middle of the run matters more than the end.
5 min readThe Supervision Gap: Three Practitioners Agree — If You've Stopped Checking AI, You're Not Working, You're Vibing.
Simon Willison admitted he skips reviewing AI code for production systems. A Stockholm cafe let AI order 120 eggs for a kitchen with no stove. Ethan Mollick says we don't yet have words for how multi-agent systems fail. The pattern is the same: as AI gets more capable, the temptation to stop checking grows — and the cost of not checking doesn't shrink.
5 min readThe People Building AI Just Told You to Stop Waiting for a Better Model
Three lab CEOs endorsed slowing model development the same week Anthropic's own numbers showed the real gains came from supervision, not a new release. Here's what to do with the model you already have.
4 min read