Systems & AdoptionJuly 21, 2026·5 min read

A Nurse Cut Short a Call With a Dying Patient Because AI Said It Was Too Long. This Is What AI Monitoring Gets Wrong.

Kaiser Permanente nurses are cutting short patient calls to protect AI-generated performance scores. Meta employees are suing over AI layoff algorithms that flagged protected-class workers. Both failures share one root: a score that can't see context.

By Patin Team · Examples are illustrative composites

A Kaiser Permanente nurse ended a call with a terminally ill patient this week because her AI-generated performance score tracks call length. The patient needed more time. The nurse needed to protect her rating. The algorithm didn't know the difference — and that's not a bug. It's how the system was designed.

What happened

A Markup/CalMatters investigation published July 18 found that Kaiser Permanente call centre nurses face AI-powered surveillance that monitors keystrokes, time on electronic health records, tone, and empathy — and penalises calls that run beyond target length. Multiple nurses described shortening conversations they knew needed more time because the algorithm was watching and the score would drop.

The same week, twenty-six Meta employees filed a federal lawsuit (July 15) alleging the company used an internal AI scoring system to generate layoff lists that disproportionately flagged workers on medical, parental, or family leave. A federal judge denied an emergency injunction on July 17 but acknowledged the case raises serious questions about AI in workforce decisions. This is the first major AI employment discrimination suit against a Big Tech company.

Neither story involves a model that malfunctioned. Both systems were almost certainly producing scores exactly as designed. That's the point.

Three questions for Monday

If you manage a team, approve new tools, or work in HR or operations, you are now operating in an environment where AI-generated scores affect employment outcomes and are being contested in court. Three questions to run before deploying or approving any AI monitoring or evaluation tool:

What is this score blind to? Every metric captures one thing and ignores everything adjacent to it. A call-length score is blind to the reason the call ran long. A productivity score is blind to the difference between shallow tasks that close fast and deep work that takes time because it's hard, not because someone is slow. Name the blindspot before you use the score.

What behaviour does this incentivise? Nurses are cutting short calls with dying patients. If employees start structuring their leave timing, project choices, or work habits around an AI score, that's the system working — just not in the direction you intended. Before you deploy a metric, think through the version of this where people optimise for it. Is that the behaviour you want?

Who reviews a score before it affects someone? A human checkpoint before any AI evaluation influences employment, compensation, or access is not optional overhead. It's what makes the evaluation defensible — legally and professionally.

A people operations lead at a 200-person fintech

Her company is evaluating an AI tool that scores customer support calls — sentiment, resolution rate, and handle time. The vendor demo shows clean dashboards and a tidy ranking.

Before signing, she runs a calibration test: she picks three calls her team leads rated as excellent, and three they flagged as poor. She runs them through the vendor's tool.

Two "excellent" calls score below median. Both involved complex situations that took 20 minutes to resolve properly. One "poor" call sits in the top quartile — short, efficient, and the customer called back the next day with the same unresolved issue.

She doesn't reject the tool. She adds one contract requirement: the tool flags calls for review. It does not influence performance ratings without a team lead listening to the flagged call first.

A team lead at a 70-person professional services firm

She receives a weekly AI-generated productivity report for her six-person team. The report scores output volume: documents reviewed, emails sent, tasks closed.

One team member sits at the bottom of the volume ranking three weeks in a row. Before acting on it, she checks what that person has been doing. The answer: leading two large client proposals that require deep review of existing work — the kind of work that doesn't produce many "tasks closed" events.

The score was measuring the wrong thing for the work that person was doing. She adjusts her process: the AI score surfaces people to check in with. Not conclusions to act on.

The one thing

When an AI scores a person's work, someone decided what to measure. You own what happens when the measurement turns out to be wrong.

<BlogPracticeSection />

Reading about it only gets you so far

Patin turns this into five-minute drills that score what you write and tell you why. It's in closed beta — join the waitlist and we'll email you when your cohort opens.

Just want the writing? .