Four AI Decisions About People Went Wrong This Year. Same Root Cause.
A resume screener, a workplace-monitoring rollout, a layoff algorithm, and an accent-modification tool all failed the same way this year. Here's the one check that would have caught each of them before a real person was affected.
By Patin Team · Examples are illustrative composites
Before any AI tool scores, flags, or filters a person this quarter — a candidate, an employee, a customer — ask what the score is actually standing in for. Four separate incidents this year show what happens when nobody does: the tool wasn't wrong about what it measured, it measured the wrong thing, and by the time anyone checked, the decision had already landed on someone's job.
Four incidents, one shape
A study found AI resume screeners shortlist candidates who used the same model to write their resume at rates 23-60% higher — the tool was measuring stylistic similarity to itself, not fit for the role. Meta watched employees work across Gmail and internal tools to train an AI system, then cut 8,000 of those employees — the activity logs it learned from could show what people did, not the judgment calls that made the work worth doing. Kaiser Permanente nurses started cutting patient calls short to protect an AI-generated efficiency score, and Meta employees separately sued over a layoff algorithm that flagged protected-class workers at disproportionate rates — both scores optimized for something measurable while missing the context that mattered. Telus rolled out real-time accent modification on offshore call-centre agents without telling customers — the system changed how a person presented themselves, and the people it changed found out from the backlash, not from being asked.
Different vendors, different jobs, same failure: a proxy measure — writing style, activity volume, call length, accent — got treated as if it were the outcome that actually mattered, and it started making or shaping decisions about a real person before anyone tested whether the proxy held up.
What to check before an AI score touches a person
Run this before any tool that scores, flags, ranks, or modifies something about an employee, candidate, or customer goes live — not after the first complaint:
- Name the proxy. What is the score actually counting — words, minutes, keystrokes, accent markers — and is that the same thing as the outcome you care about, like job performance or genuine fit?
- Put a person at the flag. When the score crosses a threshold and triggers a real consequence — a rejection, a layoff, a rollout — does a human review that specific case, or does the action fire automatically?
- Test the edge cases first. Before it's live, run the score against a candidate who writes differently, an employee whose best work doesn't show up in logs, a call that legitimately needed to run long. If the proxy fails there, it will fail in production the same way.
- Disclose it to the person it's measuring. If the answer to "would this person be upset to find out how they're being scored" is yes, that's the signal to slow down, not proceed quietly and hope the backlash doesn't come.
Worked example: the hiring manager
Priya screens applicants for a project-manager role at a 60-person logistics software company. The ATS she uses added an AI shortlisting feature this year, and it's returning a suspiciously uniform crop of finalists — clean, confident, similarly-structured resumes. Before she trusts the shortlist, she runs three resumes she wrote herself in a plainer style through the same tool and watches them rank lower despite matching every stated requirement. That's her answer: the tool is scoring resume style, not project-management competence. She keeps the AI shortlist as one input, not the filter, and has a human read every application that scores in the bottom half before it's rejected.
Second example: the operations lead
Marcus runs support operations at a 200-person subscription retailer and is considering a monitoring dashboard that scores agents on call-handle time and lets managers flag underperformers automatically. Following the checklist, he asks the vendor to run a test batch that includes calls with a distressed customer needing extra time. Three of those calls get flagged as "underperforming" despite the agent handling them well. He ships the dashboard with the automatic flag turned into a manager review queue instead, and tells the team exactly what's being measured before the first score goes out.
One sentence to keep on the wall: if an AI score is going to change what happens to a real person, know exactly what it's measuring, and check that it's measuring the right thing, before it goes live — not after someone's job depends on the answer.
Put this into practice
Reading is a start — but skill comes from doing. Try these drills now.
Reading about it only gets you so far
Patin turns this into five-minute drills that score what you write and tell you why. It's in closed beta — join the waitlist and we'll email you when your cohort opens.
Just want the writing? .
Keep reading on this
Telus Is Erasing Offshore Workers' Accents With AI. Three Questions Before You Deploy Something Similar.
Telus deployed real-time accent modification on offshore call-centre agents without disclosing it to customers. Labour unions called it deceptive. Two rival telecoms declined. Three questions every manager should answer before rolling out any AI that changes how people present themselves.
4 min readA Nurse Cut Short a Call With a Dying Patient Because AI Said It Was Too Long. This Is What AI Monitoring Gets Wrong.
Kaiser Permanente nurses are cutting short patient calls to protect AI-generated performance scores. Meta employees are suing over AI layoff algorithms that flagged protected-class workers. Both failures share one root: a score that can't see context.
5 min readAI Resume Screeners Prefer Resumes Written by the Same AI Model. The Bias Is 23-60%.
A study found AI screening tools shortlist candidates who used the same model to write their resume at rates 23-60% higher. What hiring managers and applicants need to do about it.
4 min read