Judging AISeptember 27, 2026·4 min read

The Best AI Eval Advice This Year Starts With Reading, Not a Metric

A new guide from Hamel Husain and Shreya Shankar says the way most teams grade AI output is backwards. Read real failures by hand first, then decide what to measure — not the other way around.

By Patin Team · Examples are illustrative composites

If your team's idea of checking AI output is spot-checking a handful of examples and calling it good, or building a dashboard that scores "helpfulness" out of 100, you don't have an eval process — you have a number that feels like one. A guide published this week names the step both of those skip, and it isn't a smarter metric. It's reading the failures first.

What happened

Lenny's Newsletter published a detailed guide this week by Hamel Husain and Shreya Shankar on finding where AI products actually go wrong. The method: pull 100 or more real outputs from production, have a human read them and tag what's actually wrong with the bad ones, cluster those tags into a short list of recurring failure modes, and only then decide what to automate or measure. Most teams skip straight to step three — they pick a metric, wire up a dashboard, and never look at a single real output along the way. Husain and Shankar's point is that a metric chosen before anyone has read a failure is a guess dressed up as a system.

The skill, Monday morning

Before you trust any score, threshold, or dashboard tied to an AI workflow your team runs, go pull 20 to 30 outputs it produced this month — the ones that were accepted, not just the obvious misses. Read them the way a skeptical client or auditor would, and write down, in plain language, what's actually wrong with the weak ones. Not "quality was low." What specifically: a wrong number, a tone that reads as dismissive, a step that got skipped silently. Do that before you pick, or trust, a metric. The failure modes you find by reading are almost never the ones a generic score would have caught.

Deja: a dashboard that missed the real problem

Deja runs support operations at a 60-person e-commerce company and uses AI to draft first-pass responses to customer tickets, which agents then edit before sending. Her team had tracked a "resolution quality" score for months, averaging 91 out of 100, and treated that number as proof the system was working. When she finally sat down and read 30 flagged tickets end to end instead of trusting the score, the real pattern wasn't quality at all — it was the drafts confidently promising refund timelines that didn't match the company's actual policy, a specific and fixable error the quality score had no way to catch because nobody had told it to look for that. She now keeps a running list of failure types pulled from actual tickets, reviewed monthly, instead of one blended score.

Marcus: the metric that came first

Marcus manages client reporting at a 15-person PR agency and set up an AI-drafted "coverage summary" for each client, scored automatically for tone and completeness before it ever reached him. The score looked healthy for three straight months. A client called out an error only after it shipped: the summary had quietly attributed a competitor's press mention to his client's campaign, a factual mix-up the tone-and-completeness score was never built to notice, because nobody had read a real output before deciding what the score should check for. Marcus still uses a score today, but he built it after reading a batch of drafts, not before.

The one thing

A metric built before anyone has read a real failure isn't measuring quality — it's measuring whatever the person who built it happened to guess quality would look like. Read the outputs first. The categories you'll actually need to track come from what you find, not the other way around.

Reading about it only gets you so far

Patin turns this into five-minute drills that score what you write and tell you why. It's in closed beta — join the waitlist and we'll email you when your cohort opens.

Just want the writing? .