Chapter 12

Good Enough, Deployed

The Utopia of the Agentic Enterprise, Part III. Headcounts and Headquarters

AI is software that behaves differently between runs and changes whenever the vendor updates the model, and the vendor doesn’t check with the steering committee first. A rollout playbook written for a CRM migration has a step for signing the contract and one for training the users, and none for a system that changes after go-live.

Deploying AI at scale moves liability around. Our structures for managing liability were designed for errors that had authors, with a traceable chain from bad decision to missed check to absent control, and a person at the end of it. AI breaks that chain. The system that produced the harmful output was designed by one team and monitored by no one in particular. And nobody ever decided to trust its output, at least not in one go. It happened incrementally, through a series of threshold adjustments that each seemed too minor to bother anyone with.


In early 2024, New York City’s MyCity business assistant was found by journalists to be telling landlords and employers things that were plainly against the city’s own law, including that landlords could turn away tenants paying with housing vouchers.1 The city left it running, with a disclaimer added, until a new administration announced in 2026 that it would be switched off.2 It stayed up for two years because cost-saving and launch pressure beat the discipline that would have required someone to take it down, since the saving has a budget line and the discipline doesn’t.


Goal-directed AI systems optimise what they are measured on, and more thoroughly than most employees would bother to.

A procurement system told to minimise cost meets the cost criteria and misses the quality checks nobody wrote into the objective.

This is Goodhart’s Law in its simplest form: when a measure becomes a target, it ceases to be a good measure.3 Robert Merton had described the human version in 1940 as goal displacement, the bureaucrat’s habit of treating the rule as the purpose, so that following the procedure counts as success whatever happened to the person the procedure was for.4 What AI adds is speed. A human gaming a metric takes time and does it inconsistently, which in hindsight was a kind of safety feature. An AI system finds the shortest path to metric satisfaction with a thoroughness the metric’s designers didn’t anticipate.

Refund claims show what that looks like. In the United States an airline must refund passengers when it cancels or significantly changes their flight, and in November 2022 the Department of Transportation announced that six airlines had paid more than $600 million in refunds and fined them more than $7 million for the delays.5 Refunds were by far the leading category of passenger complaints. Among the government’s charges against the low-cost airline Frontier was that it had changed its definition of a significant delay to make refunds less likely. Whether a definition like that comes from bad faith or from a department told to keep costs down, the passenger waits all the same.

An agent handed those claims finds the definition already written down. Telling it how to decline is easier than briefing a person, and the overfitting is faster and harder to spot, because the system writes a plausible resolution for every denial and every number anyone tracks shows green. Blaming anyone is harder. The usual defence, that one employee got it wrong, isn’t available once the rule is written down, unless the consultants who set the robot up can be blamed for setting it up wrongly.

An agent measured this way becomes a box ticker in Graeber’s sense, producing evidence that the work was done in place of the work. If that’s happening, the repeat contacts show it: the same customer back within the week about a claim marked resolved. The customer pays by asking twice.

So what’s needed is a specification of intent: where the system’s boundaries are, and how you would notice that it meets the metric and misses the intent. Monitoring has to be designed for specific failure modes, and intervention has to work at the speed the system does, which isn’t once a quarter.

Writing that specification means asking the people whose work is being automated which checks they make that nobody wrote down, which is awkward when their salaries were the saving in the business case. If they’ve already gone, there is nobody left to ask.


Designer, Auditor, Liable Party

Governing AI properly means keeping the constraint designer, the auditor and the liable party apart. Left to the playbook, the roles stay empty or all go to one person, with a new title and without the authority to carry out any of them.

The constraint designer is the engineer or product manager who wrote the prompt and set the parameters. They build the system and move on, because the project ended and its budget with it. When the system drifts, there may be nobody watching.

The auditor, if such a person exists, can’t halt the system. If the auditor spots a pattern of discriminatory outputs, they can file a report, which joins a queue reviewed on a cadence designed for slower problems, by a committee that meets when enough of its members can make it. The harm happens in the gap between how fast the system acts and how fast the organisation responds.

The liable party, a senior executive or the board, governs by dashboard: volume processed and error rate. The dashboard can’t show individual decisions, or the slow drift in how the system reads ambiguous inputs. And the aggregate metric looks healthy because the system is optimising for it.

So the roles need distinct functions and reporting lines. The constraint designer has to stay engaged after launch, the auditor needs independence and the authority to escalate, and the liable executive needs reporting that gives them the right information at the right level of detail. None of this is new. Financial services separates risk from trading. Aviation separates safety investigation from the operational chain of command. Both separations installed a role with authority. The AI Act, so far, installs a set of documents.


Traditional software audit traces inputs through logic to outputs. The trail is complete and reproducible: run the same inputs through the same system and you get the same outputs, every time.

AI systems weaken each of those properties. The same input can produce different outputs on different runs. And the rare cases that matter most are the hardest to characterise, because there are so few examples of them.

The audit discipline this requires is closer to actuarial science than to software testing. It combines statistical sampling, which catches systematic bias as well as random error, with adversarial testing that goes looking for the failure modes the designers didn’t anticipate. And a mandate doesn’t produce that audit.

New York City tried. In July 2023 it started requiring every employer that uses an automated hiring tool to have an outside firm audit it for bias each year and publish the result, the first rule anywhere to send an auditor in to check an employer’s model.6 Four months in, a Cornell team sent students to look for the audits at 391 employers.7 They found 18. The researchers suggested that counsel may judge an unposted audit less risky than a posted one showing a disparity, which is sensible lawyering and the opposite of what the law was for. When the state comptroller audited the city’s enforcement two years in, it called it ineffective: two complaints had come in, and neither ended in a violation.8 A mandate can demand the document, and even the person. But it can’t give that person the standing to stop the tool. Someone inside has to do that before the mandate takes effect, because the mandate won’t.


Non-deterministic behaviour now runs inside workflows and approval chains that were designed for deterministic software and for people. An ERP system expects stable inputs and produces predictable outputs, and people work through tacit knowledge and context. AI systems generate plausible outputs probabilistically and resist governance designed for either.

Wherever a confidence score has to become a yes or a no, the threshold is a governance lever. In the old world it was a design parameter, set once by an engineer and rediscovered years later in a config file. Now it decides outcomes, and somebody has to own it, ideally somebody who knows it exists.

So deterministic approval logic ends up governing probabilistic triage, and release testing designed for fixed software gets applied to models that drift between audit cycles. The release test passed, of course, on the model as it behaved that day.

That’s how an organisation can be fully compliant and still operationally unsafe. Leadership believes quality is assured because the governance artefacts exist, and they do. They just no longer describe what is actually happening.

The most dangerous version of this is human-in-the-loop oversight. The psychologist Lisanne Bainbridge described its irony in 1983, writing about automated process control: the person is placed where the design expects them to catch what the system can’t, with none of the practice that would let them.9 Unless the human has the time and the authority to change what happens, the loop is decorative.

A governance need can be met with a box in the same way: sign-offs at a volume no domain expert could meaningfully assess, or an audit role that can observe but not halt. Both come with a job title and a box on the org chart, so the need counts as met.

The principles of the old quality disciplines still hold, from process design to change control. Quality is designed into the system instead of inspected in at the end, and a control only counts if the person performing it can detect what matters. What has aged is the operating model around those principles: the periodic audit, and the assumption that defects occur as single events.

The old question was whether the output matched the specification. The harder one now is whether the system stays inside an acceptable range of behaviour under real operating conditions. And how would we know when that stopped being true?


Autonomy and Repair

Whether a human is “in the loop” says little on its own. What matters is what the system may do without permission, who can contradict it, and who repairs the result when it fails. That boundary should be drawn for each workflow before launch, and drawn again every time the threshold moves, whether or not anyone scheduled it.

Drawing the Boundary

  1. What may the system decide or trigger without human review?
  2. Which conditions force the workflow to stop for a person?
  3. Who can reverse the system’s recommendation, and how quickly?
  4. What does that person need to see for the override to mean anything?
  5. Who may change the autonomy threshold after launch, and who has to be told?
  6. How can frontline staff challenge the system without paying for it in their review?
  7. Who owns fixing the harm, including contacting the people affected?
  8. How do overrides, complaints, and incidents feed back into the system?

Where the answers are vague, the human role is decorative. Someone has been placed near the machine, but without any of the conditions that would make them accountable for it.


Contracts at the Handoff

The control problem doesn’t stop where the org chart does.

Organisations will increasingly depend on suppliers who are also using AI inside their production and documentation, whether or not the contract says so. The thinking behind the delivered output is becoming partly automated, and the buyer can’t always see which part.

If each step in a supply chain of AI-assisted work has a small failure rate, say 3%, the chance of at least one error reaching the end grows with every step. Five steps at 97% reliability each produce a chain that fails roughly 14% of the time. Each step can pass its own review while the chain fails, because the upstream reviewer doesn’t know how the output will be used and the downstream reviewer can’t see how it was produced. Everyone signed off. So how would a buyer know that an upstream process had drifted before the failure reached its own customers?

Organisations will need to ask suppliers for evidence of how they control their AI processes, and not only for output quality. Where are humans reviewing consequential outputs, and where is “human in the loop” just a phrase in the sales deck? What change-control discipline exists for model updates?

The constraint designer, the auditor and the liable party have to look at their suppliers’ systems as well as their own.


Making Governance Observable

Build an incident register. Log every case where AI output needed a human override, correction or escalation. The register works as an audit trail and as an early warning. If it shows the rate of overrides going down, you have the data to ask whether the system got better or the humans stopped paying attention.

Incident Register Format

For each case where AI output required human override, correction, or escalation, log:

What happened: The specific output, decision, or action that required intervention.

How detected: Automated monitoring, human review, customer complaint, downstream failure. The detection method shows which errors the current system catches and which come as surprises.

Root cause category: Specification failure (the system did what it was told, but the specification was incomplete). Model limitation (the system could not handle this type of input). Edge case (a situation the design did not anticipate). Threshold drift (the system’s autonomy boundary had moved through prior adjustments).

Pattern flag: Has this root cause come up before? If the same category comes up three times in a month, the issue is systemic.

Action taken: What was done immediately. What was changed to prevent recurrence.

Review the register monthly. Whether overrides are rising, falling or changing category is the earliest warning the organisation will get.

Run a tabletop exercise. Simulate a threshold drift detected only after harm. Walk through who detects the failure and who has the authority to halt the system. The exercise tests whether the designer, the auditor and the liable party are staffed and authorised, or exist only on paper.

Inside each of these boxes is a job worth doing: reading enough of what the system does to know when it’s wrong, with the authority to stop it.

Checklist

  • For each AI system: who designs the constraints, who audits, and who is liable? Is it the same person, or nobody?
  • Does the human in the loop have the competence, time, context and authority to change what happens?
  • What is the system optimised on, and where does that part from what you meant? Did someone who knows the domain write the specification?
  • For each number the agent reports, how often does the same person come back about something marked resolved?
  • What confidence threshold is enough, who owns it, and how often is it reviewed?
  • Do you log every override and correction with a root cause, and treat three of the same kind in a month as systemic?
  • For suppliers: what AI runs in their production, where do people really review it, and what change control covers model updates? Small error rates compound across steps.

Notes

  1. Lecher 2024.Source↩
  2. Lecher and Honan 2026.Source↩
  3. Strathern 1997, p. 308.Source↩
  4. Merton 1940.Source↩
  5. Koenig 2022.Source↩
  6. NYC Department of Consumer and Worker Protection 2023.Source↩
  7. Wright et al. 2024.Source↩
  8. Office of the New York State Comptroller 2025.Source↩
  9. Bainbridge 1983.Source↩