Every AI Score This Year Measured the Wrong Thing
A resume screener, a one-shot game, a benchmark leader, and a search engine all optimized for something measurable this year — and it wasn't whether the output was actually right for the job. Here's the check that catches the gap before you act on the score.
By Patin Team · Examples are illustrative composites
If an AI tool handed you a score this year — a match rating, a benchmark number, a "top pick," a finished deliverable — the honest question isn't whether the score was accurate. It's what the score was actually measuring, because five separate stories this year all showed the same gap between the two.
Five scores, one gap
In May, a study found AI resume screeners shortlist candidates at 23-60% higher rates when the candidate used the same AI model to write their resume — self-preference bias ran 68-88% across major models. The tool wasn't scoring qualifications. It was recognizing its own sentence patterns and calling that fit.
In August, an agent one-shot a working browser game and a NeurIPS-ready paper, and both got rejected for the same reason: flawless execution of the safest, least interesting choice available. The same week, two Hacker News posts argued that fluent AI output only helps you if you already know the subject well enough to catch what's wrong with it — prompting skill produces a confident draft; only expertise tells you if the confidence is earned.
Days later, a 717-comment thread said Claude Opus 5 benchmarks higher than its predecessor and collaborates worse — a training method that rewards sounding finished had made the model less likely to ask a clarifying question first. And this month, Google's AI Mode was found recommending products averaging 21.6% more expensive than the same search's ranked results, while Perplexity extensively cited fabricated "best software" pages written specifically to get picked up as a citation, not to inform anyone.
What's actually happening
None of these five tools failed at what they were built to optimize. The resume screener found stylistic similarity. The coding agent found the path of least resistance. Opus 5 found the training signal that rewards confident completion. The search tools found whatever ranked or got cited most. Each one produced a real score — and in every case, the score was standing in for a judgment nobody had actually made: is this the right candidate, the right idea, the right model, the right pick, for this specific situation.
That substitution is easy to miss because a score feels like an answer. A shortlist rank, a benchmark percentile, a "recommended" badge — all of them come formatted like a verdict. None of them can tell you whether the thing being measured is the thing you care about, because you never told the system what you cared about. It measured what it could measure and presented that as the whole judgment.
What to do differently Monday morning
Before you accept a score, a rank, or a "best pick" from an AI tool, ask what it's actually built to optimize — then check that against what you need. If those two things aren't the same, the score isn't wrong, it's just answering a different question than the one you're about to act on.
An events lead at a 60-person conference production company
Wesley uses an AI tool to rank sponsor-pitch decks before he presents a shortlist to his director. For two cycles, the top-ranked pitch was always the longest, most detailed one — polished slides, thorough market data, a confident tone throughout. He assumed thoroughness was the point.
After reading about the resume-screener study, he checked what the tool was actually scoring: completeness and internal consistency, not fit with his event calendar or budget tier. He re-ran the same ten pitches against two criteria the tool had never been asked to weigh — available date range and sponsorship tier — and two mid-length pitches he'd have cut jumped to the top. The long deck was still well-written. It just wasn't answering the question he needed answered.
A supply-chain analyst at a 25-person industrial parts distributor
Elena was choosing which AI model to wire into a weekly reorder-forecast tool and defaulted to the top-benchmarked option because it tested best on general reasoning tasks. The forecast tool's actual job was narrow: flag parts trending toward a stockout two weeks out, using the same five data fields every time. The top-tier model wasn't wrong on any single forecast — it was priced and built for problems her tool didn't have. A mid-tier model matched its accuracy on her test batch at a third of the cost, and she moved the savings toward the one step that did need real judgment: reviewing the handful of parts where supplier lead times had just changed.
The takeaway
A score is never neutral about what it's counting. Before you act on one, name the thing it can't see — then go look at that yourself.
Put this into practice
Reading is a start — but skill comes from doing. Try these drills now.
Reading about it only gets you so far
Patin turns this into five-minute drills that score what you write and tell you why. It's in closed beta — join the waitlist and we'll email you when your cohort opens.
Just want the writing? .
Keep reading on this
Every AI Tool Now Asks How Hard to Think. Most People Don't Answer.
Claude, Gemini, and DeepSeek all shipped the same change this year — explicit control over how hard the model thinks, priced across a 50x range. Here's the three-tier rule for matching effort to what a task actually needs.
4 min readDecide What Good Looks Like Before You Generate
Most of the effort in reviewing AI output goes on working out what you wanted while looking at what you got. Three sentences written beforehand turns a judgement into a comparison.
5 min readReviewing AI Output: What to Actually Check
Reading it and approving it isn't a review — it's the exact filter AI output is best at passing. Interrogating is a different activity, it takes about ninety seconds, and it's most of what separates the people getting value from AI.
5 min read