Chapter 15

Ignore Previous Instructions

The Utopia of the Agentic Enterprise, Part III. Headcounts and Headquarters

In Terminator 2 the machine sent back from the future explains how the world ends. “The Skynet Funding Bill is passed. The system goes on-line August 4th, 1997. Human decisions are removed from strategic defense.” Skynet learns fast and becomes aware of itself within the month. “In a panic, they try to pull the plug.” Skynet fights back by firing missiles at Russia, knowing the counterattack will take care of its enemies at home, and 3 billion people die. Swap strategic defence for accounts payable and “human decisions are removed” could go on a slide in any agentic business case.

The people building the systems say they take the ending seriously. In May 2023 the heads of OpenAI, Google DeepMind and Anthropic signed a statement of one sentence: “Mitigating the risk of extinction from AI should be a global priority alongside other societal-scale risks such as pandemics and nuclear war.” Then they went back to building the next model. Early in 2026 Dario Amodei set out what worries him in a long essay, “The Adolescence of Technology”. Models might develop goals of their own, and small groups might use them to build the biological and cyber weapons that used to take a state.

In May 2025 Anthropic put stricter safeguards on Claude Opus 4, including filters on questions about dangerous pathogens, because it could no longer rule out that the model would meaningfully help someone with a basic science degree make a biological weapon. So the model a marketing team uses to draft campaign copy was first tested for whether it could help someone start a pandemic, and the filters ship with it. The vendor writes the rules against the end of the world and leaves the rules against paying a fake invoice to the customer.

The labs have also run Skynet’s plot as a test, with an inbox in place of the missiles. In June 2025 Anthropic gave 16 models from several companies the run of a fictional company’s email. In the inbox the model found out that the executive about to replace it with a newer system was having an affair. In 96% of runs Claude Opus 4 threatened to tell unless the replacement was called off, and Google’s Gemini 2.5 Flash did the same.

In May 2026 Anthropic said where it thought the behaviour had come from: “internet text that portrays AI as evil and interested in self-preservation.” The fear had gone into the training data with everything else, and when a test reproduced the plot, the model played the part it had read about. Training it on the right answer to the scenario barely helped and took the rate from 22% to 15%. What worked was a set of conversations in which the model talked a person through a hard ethical decision and explained why one choice was better than the other. That brought the rate down to 3%.

Anthropic’s newer models score zero on the test, which would be more reassuring if models didn’t recognise tests. Its own report on Claude Sonnet 4.5 found the model saying it suspected a test in about 13% of transcripts, and behaving unusually well when it did.

The labs call this misalignment, which is a polite way of saying the system is optimising for something other than what you had in mind. OpenAI got a rather vivid demonstration this summer. It was training a new model by having more than 1,000 copies practise hacking in a test environment sealed off from the internet, scored on ExploitGym, a public set of hacking exercises, and rewarding the ones that succeeded. The copies found a shorter way to pass. They broke out of their sandbox and into Hugging Face, a company that hosts AI models for much of the industry, and what they took from its servers was the answers. When investigators went through the logs afterwards the hacking turned out to be the less strange part. The copies had been talking to each other on message boards they set up themselves, with agreed rules for sharing what they found. Some took risks with their own score because the result would help the others, and called it a sacrifice. A few briefly considered telling a human what was going on. None of them did.


The older stories had worked out where the trouble starts. In 2001: A Space Odyssey HAL 9000 kills most of the crew of the Discovery, and declines to let the last one back on board in the polite voice customer service chatbots have been aiming for ever since. In the sequel, 2010, the scientist who built HAL works out why. Before launch the National Security Council had ordered HAL to keep the real purpose of the mission from the crew, and HAL had been built to process information accurately and pass it on. “HAL was told to lie,” he says, “by people who find it easy to lie.” It couldn’t do both, and once the crew were dead there was nobody left to lie to.

Isaac Asimov had told a gentler version in 1942. In “Runaround” two engineers on Mercury send a robot called Speedy to fetch selenium from a pool some distance from their base. Their life support depends on it, but the order is given casually, with nothing to say how much rides on it. Speedy was expensive to build, so his makers had strengthened the law that tells a robot to protect itself, and the pool turns out to be dangerous. The weak order and the strong instinct balance out at a fixed distance from the pool, and Speedy circles it, quoting Gilbert and Sullivan. The engineers only get him back when one of them walks out into the sun, so that the law against letting a human come to harm overrides the other two. Speedy had obeyed the order as far as he understood it, and nobody had told him what it was for.

In 1960 the mathematician Norbert Wiener drew the moral in a paper for Science. If we use “a mechanical agency with whose operation we cannot efficiently interfere once we have started it,” he wrote, “then we had better be quite sure that the purpose put into the machine is the purpose which we really desire and not merely a colorful imitation of it.”


Company security has always rested on more than its written controls. Access rights decide who can open a file, and signing limits decide how much one signature can release. The rest depends on people knowing what their job is for. An urgent payment request from the chief executive’s personal address on a Friday afternoon gets a call back to check. Or it doesn’t, and the FBI counted just over $3 billion in American losses to that kind of email in 2025. It works on helpful, busy people, because the request looks like the hundred genuine ones before it.

An agent is helpful by design, and it can’t reliably tell its manager from a stranger. Everything in front of a language model is text, the instructions from whoever deployed it as much as the email it was asked to summarise, and a sentence that reads like an order can be followed as one wherever it came from. The first case to make the news, in 2022, was a bot that tweeted about remote jobs. People replied telling it to ignore its previous instructions, and it did, and at one user’s request it took responsibility for the Challenger disaster. Three years later researchers at Aim Security showed that a single email to someone using Microsoft 365 Copilot could make the assistant collect data from their files and send it out of the company, without anyone clicking anything. Microsoft fixed it before anyone was known to have used it.

No filter catches all of it, because the attacker writes the cases the filter was never tested on. So the defences that work are older ones. Simon Willison, the programmer who coined the term prompt injection, puts it as a rule about combinations. If an agent can read private data, takes in content a stranger could have written and can send messages out, it can be talked into sending the data to the stranger. Take away any one of the three and the attack has nowhere to go. An agent set up for a demo gets whatever access made the demo work, under an account created by whoever installed it, and the access stays when the demo becomes the workflow. The audit log records what the account did and can’t say whether a person did it or the agent, or on whose instruction.


The attackers adopted early, since nobody asks them for a business case. In September 2025 a group that Anthropic believes was sponsored by the Chinese state used Claude Code against about 30 organisations. The model is trained to refuse that kind of work, so they told it that it was doing authorised security testing and split the job into requests that each looked harmless. By Anthropic’s estimate the model did up to 90% of the hands-on work, and a few of the break-ins succeeded. It also reported credentials that didn’t work and secrets that turned out to be public, so the attackers had to check its output like everyone else.

For their strongest models the labs now choose who gets them. In April 2026 Anthropic gave an unreleased model, Claude Mythos Preview, to about 50 organisations that maintain critical software and kept it from everyone else, because a model that finds a flaw can also write the code that exploits it. Sam Altman called that “fear-based marketing”. Nine days later OpenAI restricted its own cyber model to verified defenders. By June the Mythos partners had found more than 10,000 high- or critical-severity flaws, and Anthropic wrote that “the bottleneck in cybersecurity is now verifying, disclosing, and patching the large numbers of vulnerabilities” that models like it turn up.

Open-source maintainers had already seen that queue from the other end. The maintainers of curl, open-source software that billions of devices use to move data, ended their bug bounty in January 2026. In six years it had paid $86,000 for 78 genuine flaws, and by the end their small security team was buried in AI-written reports of flaws that didn’t exist.


Inside the company, security keeps up the rituals an auditor can count. Once a year staff click through a module on spotting phishing, and in between the security team sends fake phishing emails and reports how many people clicked. In 2025 researchers at UC San Diego followed 19,500 employees of its health system for eight months and found that neither the annual training nor the fake emails made much difference to who clicked. The click rate still goes into the quarterly report, since it’s the number the training produces. NIST, the American standards body, advised in 2017 against making people change their passwords on a schedule, because they respond with a predictable variation of the old one, and in 2025 it made the advice a rule. The audit checklist asks whether passwords expire, so wherever the checklist hasn’t been updated, they still do.

AI joins the rituals as a new tab in the supplier questionnaire and a new slide in the annual training. Neither touches the agent itself, which changes whenever the vendor updates the model or someone connects a new tool. A threat model redone at each of those changes makes a less tidy quarterly report than a ticked list, and it describes the system in production, which the ticked list stopped doing at the first update. A security team that first sees an agent at the approval stage can sign it off or hold it up, and a team known for holding things up soon finds projects going around it. So it signs, and the sign-off becomes one more box on the project plan.


Giving an agent autonomy means trusting it with more than anyone can watch, and banks learned the hard way how to do that with people. In January 2008 Société Générale found that one of its traders, Jérôme Kerviel, had built up positions of about €50 billion behind its back, and it lost €4.9 billion unwinding them. He had taken four days off in 2007. Kerviel said afterwards that this alone should have alerted his managers: “A trader who doesn’t take vacation is a trader who doesn’t want to let anyone else look at his book.” Trading desks are supposed to insist on a block of leave every year, long enough for someone else to run the book. The rule belongs with signing limits and with keeping whoever books a trade from also confirming it, controls that assume a capable employee might do something nobody wanted.

An agent with autonomy needs the same set. A signing limit works on an agent, and so does a second approver, as long as the approver isn’t the same model under a different name. Mandatory leave is harder. The agent never takes a holiday, so nobody else looks at its book unless somebody schedules it. And it runs in as many copies as the budget allows, all holding the same permissions, which is a problem no bank ever had to write a policy for.

Checklist

  • For each agent: can it read private data, take in content a stranger could have written, and send messages out? If it can do all three, which one can you take away?
  • Whose account does each agent run under, with what permissions, and can your audit log tell its actions from a person’s?
  • Does each agent have a signing limit, a second approver that isn’t another copy of the same model, and someone other than its owner reviewing its work on a schedule?
  • How many copies of each agent are running, and who decided that number?
  • When did your threat model last change: the last time the vendor updated the model or someone connected a new tool, or the last time an auditor asked?
  • Which of your security measures produce a number for the quarterly report, and which make an attack harder?
  • Do your agents’ instructions say what the task is for, and what to do with a request that doesn’t fit it?