Chapter 10

The Weekend Prototype

The Utopia of the Agentic Enterprise, Part II. The Agents Who Take Your Job

At prototype scale, AI costs almost nothing: a developer and a weekend. At pilot scale, costs rise but stay within the business case, which was written around the pilot. At production scale come categories of cost that a business case written around the pilot has no line for, because nothing in traditional technology economics looked like them.

Each of those categories is paid for in people’s time. In Graeber’s terms the people are duct tapers, employed to patch a fault that ought not to exist, and at production scale the fault is in the agent: its drafts need a qualified reader, and the cases it can’t decide need someone to decide them. They’re paid from budgets other than the one that approved the project, and the customers among them aren’t paid at all.


Cost Inversion

Enterprise technology used to front-load its costs: a large sum up front, then running costs predictable enough to put in a spreadsheet and forget. AI turns the curve around, and the cost grows after deployment, when the project team has had its celebration and moved on.

At production scale, the system sees the full range of inputs, including the hard cases that generate errors and escalations, and thousands of daily outputs need structured review by qualified people. Models change and knowledge bases drift while edge cases accumulate, and what was designed as a one-time build becomes ongoing maintenance that the business case had priced at zero.

So the prototype impresses everyone and production falls short, sometimes by enough to put the organisation off AI for years. Which is a little unfair on the technology. The organisation modelled the costs of a prototype and then deployed at production scale.

In July 2025 Anthropic announced weekly limits on its coding agent, Claude Code, because some subscribers were running it “continuously in the background, 24/7”. One of them had used tens of thousands of dollars’ worth of model time on a $200 monthly plan.1 With traditional software, margins improve as users adopt it. With generative AI, the heaviest users can wipe them out. A company that pays for its own agents by usage meets the same arithmetic in its budget, where the adoption campaign works best on exactly the people the budget can least afford. An organisation that models AI costs like an ERP rollout (high upfront, low marginal) will underfund production.


The costs that dominate at production scale are invisible during the pilot and missing from the standard budget template.

Verification costs. Every consequential AI output needs review by someone qualified to evaluate it. At prototype scale, verification is free, because the developer checks the outputs while building, on the same weekend. At production scale, reviewing needs its own infrastructure: qualified reviewers, and a way to escalate the outputs they aren’t sure about.

And review needs people who know the work, who don’t get cheaper with volume.

Maintenance costs. AI systems don’t stand still, but the business case assumes they do. Vendors replace the underlying models that production systems depend on, with a cheerful announcement that the new one is better. Each update means regression testing and revising prompts. The prompt that produced reliable results last month may produce subtly different results this month, and “subtly different” in production can mean a compliance failure or an error in front of a customer.

Integration costs. A procurement system wants a yes or a no, and an AI system produces a confidence score. Someone has to decide what score counts as “approve”, and who gets the cases in between, which may well be the person the system was supposed to replace. Each downstream system needs its own adapter, so integration costs rise in steps, one system at a time.

Accountability costs. When AI is involved in decisions that affect people or money, someone has to answer when it gets one wrong. Deciding who that is, and with what audit trail, is design work the business case leaves out. The obvious candidates were in the room when it was written.


Calling these costs “hidden” suggests somebody lost them by accident. Somebody put them there: each cost is booked somewhere other than the business case that needs to win approval.

Verification cost goes to the business unit doing the reviewing. The AI team ships the model, and the operations team pays for the reviewers. The approver reads the two budgets separately, and the business case shows savings net of the AI team’s costs, so it gets approved. Six months later operations asks for more reviewers, in a different meeting, as a capacity gap rather than an AI cost. Once the reviewing function owns the number, the economics of the project look different, at least to anyone who puts the two budgets side by side.

Maintenance cost is paid through the vendor contract, where it’s booked as “usage growth” rather than “cost overrun”, which reads much better in a quarterly review.

Accountability cost goes to the risk function, which takes on the new category of review without additional headcount, because the headcount conversation would have blocked the rollout. The risk function says so at the time, in writing, so it can find the email later. The rollout goes ahead anyway.

The make-or-buy conversation follows the same logic. Buying from a model vendor or a services firm keeps the cost out of the hiring ledger. It reads as a per-unit charge rather than a headcount addition, and per-unit charges clear approval bars that headcount additions do not, since a hiring freeze doesn’t cover invoices. Building puts the same cost inside the organisation, where it attracts all the scrutiny the buy path escaped. So the organisation that “chose to buy rather than build” may have chosen the path where the cost would stay least visible.


Payroll Versus Software

Payroll and software are not booked the same way, even when they do the same work in practice.

Payroll is visible and politically charged, and every line of it is a person somebody has to talk to. Software spend tells a much nicer story. Economically, both lines are operating expense, but one reads as drag and the other as investment. So moving cost from the first column to the second doesn’t improve anything about how the work gets done. It changes the column, and the words on the slide.

The company announces a headcount reduction through automation and meets its payroll target. Then the capacity it removed comes back in through vendors and offshore review teams, doing much the same work under someone else’s letterhead. The machine gets the credit in the press release, and the tickets go to the people propping it up. So labour substitution passes for a technology upgrade. A headcount cut is easy to see in the numbers, while the vendor dependency and the outsourced exception handling scatter across budgets and get described in other languages, sometimes literally.

A second channel runs alongside the first. Where AI has become a protected budget category, there’s money for it in a way there isn’t for replacing legacy systems or cleaning up data, which reads as maintenance, and maintenance is yesterday’s problem. Managers who understand the system route whatever they can through the AI line, where the same work, wrapped in the language of transformation, gets a much friendlier budget conversation. The modernisation deferred for years finally gets funded and delivered, as AI transformation. The model itself may contribute little more than the name on the slide. The gains come from the cleaner data, and from a production line no longer held together by workarounds.

Which would be a clever bit of budget craft, if it weren’t for the feedback loop. AI gets the credit for ordinary modernisation that finally happened under its banner. Leadership concludes that AI worked, when what worked was fixing the conveyor belt. The next round of AI budget gets easier, and “AI” becomes the name under which the organisation funds whatever needed funding anyway.

In the end the company pays twice: for the machine, and for the human labour it kept in a less visible form.


The vendor ledger is not the only place the work goes. Some of it leaves the books entirely, and is done by people who send no invoice.

A chatbot reduces contact-centre minutes by making the customer read three help articles and retype the issue in words the bot understands, then type “agent” five times before the system lets her through. The contact centre records a productivity gain. The work still exists, of course. It has moved to her, and her time isn’t counted.

Suppliers and employees get the same treatment through onboarding forms and self-service portals, and so do other teams: a central AI team automates intake, and local teams spend more time correcting misclassified requests. Task time inside the unit that bought the tool fell. Total effort across the system may well have risen, but nobody is in charge of the system.


Partial automation this expensive persists because it serves nearly everyone’s local interest. Executives get a transformation narrative, and consultants get the implementation and then the remediation work, the better kind, since it comes with a sense of urgency. The risk function gets new control responsibilities, which also means new budget. And employees keep their employment, in thinner work.

No single participant sees the whole failure. The benefits are visible and concentrated, while the costs are dispersed, and partly someone else’s. Measured against the promise of simplification, the arrangement has failed, but the people keeping it going have little reason to measure it that way. Cynicism doesn’t help much here. Each of them is behaving rationally inside a mandate.


An organisation that wants to know what its AI costs has to add it all up: payroll plus vendors plus offshore review, plus the time moved onto customers and other teams. And it has to attribute each gain to the mechanism that produced it, whatever banner the funding came under.


When a project produces a visible gain anywhere near AI, AI gets the credit. When the initiative underperforms, the cost migrates sideways into ordinary IT overhead. And when that overhead is later cut, the cut becomes evidence of managerial discipline.

The fix is unglamorous: forecast tracking at portfolio level, comparing what was projected with what happened, added up across all the cost centres the project touched. It only works if the initiative’s label was fixed when it was funded, so that the projection and the outcome can be matched before the project is renamed to something more successful.


Stress-Testing the Business Case

Standard technology business case templates were designed for a cost shape AI doesn’t follow. The arithmetic fails at the projection to production, because the costs that dominate there have no row in the template at all.

Cost Stress-Test

Before approving an AI business case, ask one question for each hidden cost layer, plus scale. A “no” on any question means the business case has not priced the work.

  1. Verification. Does the budget include a line item for qualified human reviewers at production volume, scaled beyond the pilot?

  2. Maintenance. Does the budget include ongoing engineering for model updates and knowledge base drift, as a staffed function with named people?

  3. Integration. For each downstream system that receives AI output, is there a specific plan for translating probabilistic outputs into deterministic inputs?

  4. Governance. Is there a named person accountable for each category of AI decision, with an audit trail and escalation path?

  5. Scale. What does production volume look like, with projections tested against the cost curve described above? A business case that shows costs at pilot scale and benefits at production scale is a sales pitch.

A business case built this way will show higher costs and lower first-year returns than the template version, which makes it harder to get approved. But it will also be funded correctly, with the reviewer staffing and ongoing upkeep that keep it running.

Some of the tape is work, and stays. The rest covers for a fault, and the way to take it off is to fix the fault, which is the modernisation the AI budget sometimes pays for under another name.

Checklist

  • Put the business case through the Cost Stress-Test. A “no” on any question means the work hasn’t been priced.
  • Are costs shown at pilot scale and benefits at production scale?
  • Where did the cost go: reviewers inside the operating unit, maintenance booked as vendor usage growth, review absorbed by risk, work pushed onto customers, suppliers or other teams?
  • Have you costed the whole arrangement: payroll, contractors, vendors, managed services, offshore review, governance staff, and the time moved onto others?

Notes

  1. Anthropic 2025.Source↩