AI Skills for Government: Practical AI for the Public Sector
Public sector AI use carries constraints the private sector doesn't: decisions must be explainable, records must be defensible, and the people affected can't take their custom elsewhere. Here's what that changes.
By Forge Team · Examples are illustrative composites
Public sector AI use runs into three constraints most private sector guidance ignores.
Decisions have to be explainable — to the person affected, to an appeal, sometimes to a court or a committee. Records have to be defensible years later, when nobody remembers the context. And the people on the receiving end usually can't go elsewhere, which means an error isn't a lost customer, it's someone who's stuck with it.
None of that makes AI unusable. It makes the boundary between assistance and decision the thing you have to get right, and get right in writing.
The line that matters: drafting versus deciding
The workable rule is that AI can help produce material that a person then owns, and cannot produce an outcome that affects someone's entitlements, liberty, or record.
Summarising a case file so an officer reads it faster: fine. Drafting standard correspondence: fine. Producing a first-pass analysis a policy officer then interrogates: fine.
Determining eligibility, scoring an applicant, flagging someone for attention, or generating the reasoning that goes on the record as the basis for a decision: not without a human who has independently reached the same conclusion and can say why in their own words.
The test is not whether a human clicked approve. It's whether the human could explain the decision if the tool vanished.
Skill one: interrogate the output, don't review it
Reviewing means reading and looking for something obviously wrong. Interrogating means checking specific things: where did this figure come from, what's the source for this claim, what does this conclusion assume, what would change it.
Well-formatted output invites reviewing. Public sector consequences require interrogating. The difference is a habit, and it's the one worth training.
Skill two: place the checkpoints where reversal is impossible
Any process where AI touches a citizen-facing outcome needs an explicit map of where a human must intervene.
Put the checkpoints where an error can't be undone: before anything is communicated to a member of the public, before anything enters a permanent record, before anything triggers an entitlement change or an enforcement step.
Write this down. An unwritten checkpoint isn't a control — it's a habit that survives exactly as long as the person who has it.
Skill three: keep the record intact
Anything that might be reviewed later needs to show what was considered and why. If AI produced part of the reasoning, the record should reflect what it produced and what the officer did with it.
This is more than diligence. A file that can't distinguish officer judgement from generated text is a file that can't be defended, and the people who'll need to defend it won't be the ones who created it.
Skill four: watch for consistent unfairness
A human making inconsistent judgements produces scattered errors. An automated process applying a flawed rule produces the same error every time, at scale, in the same direction — and consistency is easily mistaken for fairness.
That pattern is harder to notice than random error precisely because it looks orderly. Anything applied at volume needs someone periodically asking not "is this consistent" but "is this consistently wrong about a particular group".
Aoife — the summary that lost the caveat
Aoife is a senior caseworker at a housing authority. Her team used AI to summarise lengthy application files so officers could triage faster.
One summary described an applicant's circumstances accurately but dropped a caveat from a medical letter — a consultant's note that a condition was expected to deteriorate. The summary said the condition was managed. The letter said it was managed for now.
The application was correctly assessed in the end, because the officer opened the source document. But it was close, and the summary was not wrong so much as flattened.
Her team's rule now: summaries are a navigation aid, not evidence. Any factor that materially affects a decision gets read in the original before it's relied on.
Ben — the checkpoint that wasn't written down
Ben runs a licensing team in a local authority. They'd built a workflow where AI drafted decision letters and an officer approved each one before sending. It worked well for eight months.
Then the officer who'd designed it moved teams. Her replacement, reasonably, treated approval as a formatting check — nobody had written down that the approval step existed to independently verify the reasoning, because to her it had been obvious.
Two letters went out with conditions that didn't match the application. Both were caught on appeal.
The control had never been a control. It was one person's understanding of why the step existed, and it left when she did.
The one thing
In the private sector, an AI error costs a customer. In the public sector it costs someone who often has no alternative and limited means to challenge it.
That doesn't mean using less AI. It means the boundary between assistance and decision has to be explicit, written down, and understood by whoever holds the job next — not just by the person who set it up.
Put this into practice
Reading is a start — but skill comes from doing. Try these drills now.
Reading about it only gets you so far
Forge turns this into five-minute drills that score what you write and tell you why. It's in closed beta — join the waitlist and we'll email you when your cohort opens.
Just want the writing? .
Keep reading on this
The Supervision Gap: Three Practitioners Agree — If You've Stopped Checking AI, You're Not Working, You're Vibing.
Simon Willison admitted he skips reviewing AI code for production systems. A Stockholm cafe let AI order 120 eggs for a kitchen with no stove. Ethan Mollick says we don't yet have words for how multi-agent systems fail. The pattern is the same: as AI gets more capable, the temptation to stop checking grows — and the cost of not checking doesn't shrink.
5 min readAI Found 18 Rare Diagnoses Specialists Had Missed. The Workflow Is Worth Stealing.
A NEJM AI study of 376 unresolved cases shows what the best AI collaboration pattern looks like in practice: AI generates hypotheses, humans decide what to do with them. Here is the structure worth applying to any analysis-heavy work.
4 min readTwo Viral Posts Prove the AI Bottleneck Isn't the Technology — It's You.
Two Hacker News posts hit the front page the same week with a combined 1,374 points. One named the upstream problem — vague briefs going in. One named the downstream problem — raw output going out. The model was fine in both cases.
4 min read