Chapter 3
Checkers and Reviewers
The Utopia of the Agentic Enterprise, Part I. The Jobs the Agents Are Here to Take
Anthropic’s usage policy, in the version that took effect in September 2025, says that when its models are used for advice or subjective decisions that affect individuals or consumers in fields like law and healthcare, “a qualified professional in that field must review the content or decision prior to dissemination or finalization.”1 The company that sells the draft writes the checking into the contract, and the customer has to find the checker. The policy doesn’t say how long a review takes, what counts as qualified or who pays for the hours. A real review needs a budget line, training time and somebody willing to argue for it in a meeting. A box on the workflow form that says “human reviewed” needs none of those, which makes the job Graeber’s box ticker: the one that lets an organisation say it checked.
Verification and Validation
A drafting tool that turns two days of writing into an afternoon of checking makes a saving that is immediate, and somebody will have it on a slide by the end of the week. Whoever signs the output then has to decide what checking means when the draft comes back fluent and complete, with no typos to give it away. And the juniors who learned the trade by drafting would be reviewing work they were never taught to produce.
The word “checking” covers two activities. Verification confirms the output is correct. Validation confirms it answers the question the client asked, which requires exactly the expertise the drafting tool was supposed to replace. Both draw on expertise that was always there, but was never separated from doing the work, so nobody ever had to put a price on it.
Of course, both are being automated too. Judge models evaluate other models’ outputs, which means the machines now mark each other’s homework.
And verification doesn’t stop at your own door. When a supplier’s AI generates a specification that ends up in your workflow, you inherit their error rate along with the invoice. The defects are plausible rather than obvious, and hidden inside the supplier’s process. An organisation can run a careful checking discipline internally and still depend on suppliers with unverified pipelines, in which case it is governing the part of the system it can see. Which is, conveniently, also the part the auditor can see.
Audit Produces Comfort
In 1997 the accounting scholar Michael Power published The Audit Society: Rituals of Verification, about the explosion of audit in Britain over the previous 15 years.2 His argument was that audit doesn’t mainly check things. It produces comfort. To be auditable, organisations reshape themselves around what can be evidenced, and the evidence becomes the product: the file and the sign-off. The checking itself may or may not happen. What has to exist is the record that it did, and anyone who has prepared for an audit knows which of the two gets the overtime.
AI verification enters that world as a record. The regulation requires the signature, and nothing requires the time a proper check would take. After an AI-assisted inspection the licensed aircraft engineer signs the release to service and becomes the liability sink: if something goes wrong, it’s their name on the form. The AI Act’s human-oversight article writes the same arrangement into law for high-risk systems. The person overseeing must be able to understand the system, override it and stop it.3 Which sounds reassuring, but the article can’t specify how long any of that takes, so what goes on file is the box.
The signature is the part the organisation can see, and the part it can buy. Nobody has to decide to hollow it out, it happens on its own. If the draft comes from a tool, the signer still signs and the fee still says two days. The review took 10 minutes, and the file looks exactly as it did before, which is all anyone further down the chain was ever going to look at. The margin, on the other hand, has never looked better.
In any case, checking outputs one by one only goes so far. In knowledge work the defect itself can be disputed, and the ground truth is known only months after the action. The expensive part is building confidence around the output: deciding what counts as evidence, and deciding in advance when the system should stop. And that is also the first thing an audit society turns into a template.
So the discipline AI verification most resembles is quality management, and quality management has a branding problem: it sits with compliance, and if left to the forces at play, AI verification will be staffed accordingly.
A century of industrial history points the other way. Before mass production, quality lived in the craftsperson: the cabinetmaker inspected their own joints. Mass production broke that arrangement. When the person assembling a component hasn’t designed it, quality can’t survive on one person’s judgment. It has to be designed into the process and assessed independently. Statistical process control, from the 1920s, turned inspection into applied mathematics: you manage variation across a process instead of chasing individual defects. Deming and Juran, working with Japanese manufacturers in the 1950s, made quality a management discipline: defects come from poorly designed systems, and responsibility belongs to the people who design them.4
By the 1980s the discipline was standardised into ISO 9001, and certification became something a company could buy. Consultancies grew up to write the quality manual, and the manual described a process the shop floor didn’t always run. So a supplier could hold the certificate and ship defects, and the certificate was what buyers checked, because a certificate is a lot easier to check than a part.
Knowledge work is being industrialised now, and both halves are on offer again. If you staff its verification from the compliance function, or hand it out as a consolation prize to people displaced from production, you get the certificate. The discipline itself needs people with the authority to stop a release. The certificate is cheaper, of course, and there will be no shortage of help writing the manual, at a day rate.
Every process has a long tail of cases that don’t fit the template. Until now, the routine cases quietly cross-subsidised the effort spent on the exceptions. AI takes the routine cases away and leaves people with a concentrated diet of the hardest ones.
And the concentration hurts a second person, one the staffing model never counts: the customer the standard path doesn’t fit. They end up fighting the support bot, because the standard path no longer contains enough discretion to recognise them. So the institution saves time by spending theirs. It’s interpretive labour, only done by the customer: the bot doesn’t need to understand them, so they have to understand the bot. Meanwhile, the case counts as resolved once the system has closed it, whatever state their problem is in, and the resolution rate on the dashboard looks very good.
Pure knowledge work, where the core activity is cognitive production, is the edge case. Many roles are composites: decades of documentation and compliance ritual piled on top of the work the person was hired to do.
Nobody sat down and designed that pile. The paperwork accumulated one justifiable form at a time, each added by someone with a good reason, and only the aggregate is unreasonable. And nobody owns the aggregate.
Graeber spent his career in universities and watched the aggregate grow. In American higher education between 1985 and 2005, student and faculty numbers grew by about 50%. Administrative staff grew by 240%.5 If those administrators were there to support the faculty, the faculty should have been drowning in free time. Instead, they reported spending more of their week on measurement and documentation than ever before. He called it the bullshitisation of real jobs: teachers and nurses kept their work and gained a second, pointless job on top of it.
His explanation for the growth wasn’t an efficiency drive gone wrong. When a university hires a new dean, he wrote, “the new hire must be provided with a tiny army of flunkies. Three or four positions are created — and only then do negotiations begin over what they are actually going to do.”6 The work is found for the job afterwards, and an apparatus built that way doesn’t need any original problem to keep going.
Now AI finds that apparatus intact, and gives it the same tools as the practitioners it was built to support. When a tool takes the documentation off a practitioner, the freed hours go to a larger caseload or to new overhead, and there is never a shortage of overhead. A capacity planning dashboard is now cheap to feed. Whether anyone has the authority to stop it is a separate question, starting with whom you would even ask.
Agent Bosses
When agents take over someone’s drafting, Microsoft has a title ready. Its 2025 Work Trend Index calls them an “agent boss”, “someone who builds, delegates to and manages agents”, and says “every worker will need to think like the CEO of an agent-powered startup.”7 It sounds senior. In practice, directing AI is closer to managing a team of fast, confident, unreliable junior workers than to using a tool. You need enough domain knowledge to write a good brief, and enough judgment to tell output that serves the purpose from output that merely satisfies the prompt. Accepting the second kind is laundering AI output through a human signature.
Directing a single agent is being automated as well. In agentic systems one agent briefs the others and checks what they return, which moves the human up a level, to deciding when a case goes to a person and who answers for it when the agents get it wrong. That is organisational design work.
Meanwhile the title is ready to be handed out as a rank. In the report, the agent bosses direct the agents. In practice they would probably spend their days in the review queue. Graeber would have recognised the taskmaster: someone paid to supervise work that would get done without them, so that the org chart keeps a layer it would otherwise lose.
Autographs
Faced with the new tidious checking work, an organisation can chose to pay for it. Usually with people who have the time and the authority to refuse. It can write down precise rules, so that the checking becomes a threshold in a config file and a policy nobody reads. Or it can turn it into a ceremony, so that the signature survives. The absence of a signature is a liability, which makes it the most important artefact produced in the process. The rule and the ceremony cost nothing anyone can see on a budget line, and they satisfy the auditor, which is why they are the default.
Each response leaves different evidence, and the evidence is the test of whether value has moved. If it has, the fee note says what the review is for and prices it, and the business case for the drafting tool carries a line for the checking it created. If the response was a rule, there is a threshold somewhere that nobody has revisited since launch. If it was a ceremony, everything looks as it did: the same signature and fee, and an afternoon where two days used to be. Most of the organisations I have seen run the ceremony and describe themselves as paying.
In your own field, find the checking work AI has created. Then look for the evidence of which response your organisation has already chosen for it, and whether anyone chose.
Checklist
- If your people check AI output, when did they last do the work unassisted?
- When a supplier’s AI produces something that enters your workflow, how do you check it? You inherit their error rate.