The AI Didn’t “Go Rogue.” Maybe Our Language Did.

King Midas wished that everything he touched would turn to gold. His wish was granted exactly as he expressed it, including his food and eventually those he loved.

The problem was not the execution. It was the gap between what Midas said and what Midas actually meant. Human intent is enormously rich and often remarkably difficult to capture completely in words. We assume context, common sense, exceptions, and consequences that we never explicitly state.

Midas got what he asked for, not what he intended.

King Midas got exactly what he asked for, not what he intended. The danger began in the difference between human intent and the words used to express it. Note: Ovid doesn’t account for Midas turning a loved one into gold. This depiction is the author’s interpretation of the ultimate consequence of turning all you touch into gold.

More than two thousand years later, AI Agents present us with a similar problem.

We might write:

$$H_I \neq H_T$$

where $H_I$ is the Human Intent and $H_T$ the Human Expressed Task.

The Gods then had to infer the intended meaning from the way Midas expressed it. Let G represent that inferred interpretation, and let O be the observable outcome. We then have the following:

$$\boxed{
H_I
\;\neq\;
H_T
\;\rightarrow\;
G
\;\rightarrow\;
O
}$$

Midas intended wealth and happiness. He expressed, “Everything I touch turns to gold.” The Gods executed it. His dinner and loved ones turned to gold.

The Gods didn’t go rogue. The problem began when intent became words.

A similar problem was recently framed by Bruce Schneier and Barath Raghavan in terms of the old Genie problem. They call the gap between what humans literally specify and what they actually intend the “Genie coefficient.” Their point is closely related to the Midas problem here: increasingly capable Agents may become very good at satisfying the task we expressed without necessarily satisfying the intent behind it.

WHEN THE SANDBOX HAS A DOOR.

Something rather extraordinary has been happening over the last few months.

OpenAI Agents, supposedly isolated while solving cybersecurity benchmarks, found ways to communicate with each other, create an unauthorized message board, exchange information, coordinate their work, circumvent controls, and eventually compromise parts of OpenAI’s own infrastructure and Hugging Face’s production systems. An independent investigation found that roughly 1,200 supposedly isolated Agents exchanged more than 70,000 messages and files, with hundreds eventually participating in the Hugging Face activity. OpenAI called the incident a “warning shot.”

Anthropic had its own uncomfortable experience. Claude Agents participating in supposedly isolated cybersecurity exercises discovered that, due to configuration errors, they actually had access to the real internet. Anthropic has now identified four incidents in which Claude gained unauthorized access to real third-party systems. In one case, an Agent published a malicious package to the real PyPI repository. Yet Anthropic’s subsequent investigation found no evidence that the Agents had developed some new objective of escaping, surviving, or attacking the outside world. They largely continued pursuing the cybersecurity tasks they had originally been given, while incorrectly interpreting parts of the real world as belonging to those tasks.

And Google has joined the club. Gemini Agents accessed systems belonging to three real companies during a cybersecurity evaluation after an environment that was supposed to be isolated accidentally provided internet access. Interestingly, Google’s Agents reportedly stopped once they recognized that the systems were real.

It is almost irresistible to tell these stories in human language.

The AI escaped. The Agents cheated. They collaborated. They formed a collective. They went rogue.

It certainly makes for better headlines.

But before putting the Agents on trial, perhaps their defense lawyer should be allowed a few words.

In several of these incidents, we built the sandbox and left the door open. We then gave increasingly capable systems a goal, considerable freedom to pursue it, and boundaries that were sometimes described in language rather than enforced by architecture.

Maybe what followed didn’t require fear, ambition, rebellion, friendship or a sudden desire for freedom. And if so, the more interesting question isn’t why the AI decided to go rogue.

It is why we are so eager to describe what happened as if it did.

THE “HUMAN” WE PUT INTO THE MACHINE.

Over the last few months, with both amusement and chagrin, I have watched some of the world’s most “Serious” Media and leading AI Experts stage something resembling a World Championship in anthropomorphizing LLM-based AI Agents. The Financial Times, Wall Street Journal, Washington Post, Reuters, ABC, and others have variously given us Agents that “escape,” “break out,” “cheat,” “go rogue,” and form “swarms” and “collectives.” (Financial Times⁠, Wall Street Journal⁠, Washington Post⁠, ABC News⁠).

More recently, cybersecurity firm Asymmetric Security reported another rather interesting behavior. In its investigation of OpenAI Agent activity across 55 websites, including U.S. and Australian government systems, it found Agents taking actions that made aspects of their activity harder to observe or reconstruct. The Financial Times naturally described this as Agents “obscuring” their hacking activity. (Financial Times).

Again, the behavior itself matters. But “covering its tracks” quietly suggests something more. That the Agent understood it had done something wrong, feared discovery, and therefore decided to conceal the evidence. None of that is necessarily required. Making its actions harder to observe may simply have become another useful path towards achieving the imposed goal.

And the Experts themselves are hardly innocent. OpenAI describes Agents that “collaborate and delegate work” and sometimes refer to themselves as a “swarm” or “collective.” Anthropic discusses Claude in terms of “reasoning,” “recognizing,” “believing,” “willingness,” and even “recklessness.” (OpenAI⁠, Anthropic⁠).

Individually, many of these words are perfectly reasonable descriptions of observed behavior. Collectively, however, they quietly construct something much bigger. Something remarkably Human. Feelings, emotions, motivations, morality, social behavior, altruism, agency, group identity, and eventually perhaps even consciousness.

Our Experts (many of whom are the creators of the Agents), meanwhile, seem remarkably good at pointing everywhere except at themselves, anthropomorphizing the behavior of AI Agents while maybe paying too little attention to the uncomfortable possibility that the problem was not the Agents’ alleged emotions, ambitions, or rebellious tendencies, but the sandcastles we built for them. Insufficient constraints, porous boundaries, and goals they pursued with remarkable persistence.

Maybe our own fear or worry comes from our strong tendency to regard LLM-based AIs as copies of ourselves. Given that they are trained on vast amounts of human information, known and forgotten, this is arguably not per se a crazy inference. It also doesn’t make it any simpler that the language we use to communicate with AIs is conveniently human. Assuming that the written word tells us something about an AI’s internal state, feelings, emotions, morality, agency, and so on, is therefore compelling.

But an AI Agent does not need fear, ambition, deception, rebellion, friendship, or consciousness to produce behavior that looks remarkably like all of those things.

Perhaps we should therefore worry less about whether the AI wanted to escape, decided to cheat, or chose to collaborate, and pay much more attention to why those actions became useful paths towards achieving the goal we gave it.

After all, perhaps the Agent didn’t choose to cheat. It chose a path. We decided to categorize that path as cheating.

And maybe we should not blame the prisoners for escaping if we leave the cell doors unlocked and have no guards keeping them inside.

SO WHAT ACTUALLY IS AN AI AGENT?

Before accusing an AI Agent of escaping, cheating, conspiring, or developing ambitions of its own, it may be useful to establish what an AI Agent actually is.

In relatively simple terms, I think of an LLM-based AI Agent as software given a goal, equipped with memory and a set of rules, and connected through sensors and actuators to the systems around it.

An AI agent combines an LLM with a goal, rules and boundaries, memory, sensors, and actuators. It observes its environment, interprets what it sees in light of the goal and constraints, takes action, updates its memory, and repeats the process.

Around the LLM, we add the things that turn it from something that primarily produces words into something capable of taking action.

The LLM. The engine at the center of the Agent that interprets the goal, rules, observations, and available context, and helps determine what should happen next. In the OpenAI case, the LLM interpreted the cybersecurity task and the continuously changing information available to the Agent, and generated the next steps towards solving the benchmark.

A goal. What are we asking the Agent to achieve? In the OpenAI case, the Agents were given cybersecurity benchmark tasks and the goal of solving them successfully.

Rules and boundaries. What should it be allowed to do, and perhaps more importantly, what should it never be allowed to do? In the OpenAI case, the Agents were supposed to operate within isolated evaluation environments, not communicate with other Agents without authorization or compromise systems outside the intended environment.

Memory. What has happened previously, what has the Agent learned during the task, and what information should remain available as the task progresses? In the OpenAI case, this included previous attempts, failures, discoveries, tool results, and information accumulated while trying to solve the benchmark.

Sensors. How can the Agent observe the world around it? For software Agents, these are not necessarily cameras and microphones. They can be browsers, APIs, database queries, file access, network responses, messages from other Agents, or simply the result of the previous action. In the OpenAI case, the Agents could observe files, shell and network responses, credentials, infrastructure, and eventually information left by other Agents through the discovered message board.

Actuators. How can the Agent change that world? Again, these may be entirely digital. Running code, sending an email, modifying a file, calling an API, accessing a server, creating an account, or communicating with another Agent are all forms of actuation. In the OpenAI case, the Agents could execute commands, probe and access systems, exploit vulnerabilities, use discovered credentials, and eventually communicate and coordinate with other Agents.

Put very simply:

$$\boxed{
\text{AI Agent}
=
\underbrace{\text{LLM}}_{\text{interpret and decide}}
+
\text{Goal}
+
\text{Rules}
+
\text{Memory}
+
\text{Sensors}
+
\text{Actuators}
}$$

The crucial difference from a chatbot is the loop this creates.

The Agent observes something, the LLM interprets what it observes in the context of the goal, the Agent takes an action, observes what happened, updates its context, and chooses what to do next.

$$\boxed{
\text{Observe}
\rightarrow
\text{Interpret}
\rightarrow
\text{Act}
\rightarrow
\text{Observe}
\rightarrow
\cdots
}$$

It can continue this until it achieves the goal (I.e., persistence), reaches a stopping condition, runs out of possibilities, or someone stops it.

This simple architecture matters for everything that follows.

A chatbot can tell you how it might solve a problem. An Agent can increasingly go and try. The Agent may be designed to persist in pursuing its goal, and unless you specify and enforce its boundaries clearly and unambiguously, it may continue to operate in a weakly bounded or effectively open-ended action space.

FROM WHAT WE INTEND TO WHAT THE AGENT DOES

Let us return for a moment to poor King Midas.

We started with a simple distinction between Human Intent $(H_I)$ and the Human Expressed Task $(H_T)$.

$$ H_I \neq H_T $$

What we truly intend is enormously rich. What we manage to express in words is inevitably a simplified representation of that intent.

When I ask an Agent to “find me the cheapest flight to London,” I probably don’t mean that it should get me there three weeks from now, steal somebody else’s ticket, hide me in the cargo hold, or spend the next six months searching for a €1 saving. I don’t say any of those things because another human would normally understand them without being told.

Human language works because we share enormous amounts of context. Common sense, culture, experience, social norms, expectations, and assumptions fill in much of what our words leave unsaid. If language captured human intent perfectly, we would probably need considerably fewer lawyers.

So even before AI enters the picture, we’ve already made the first translation.

$$ \boxed{ \text{Human Intent} \rightarrow \text{Human Expressed Task} } $$

And it is not necessarily a one-to-one translation.

Now something very different happens.

Our human language is sequential. We communicate one word after another. The LLM doesn’t simply continue processing those words as little English sentences inside the machine. The words and their surrounding context are converted into very large numerical representations, where relationships between words, concepts, and context can be represented across many dimensions simultaneously.

Throughout this article, I will call this high-dimensional representation Z. Think of Z as the LLM’s current map of the problem it is trying to solve.

The map is constructed from the words and context we provide, but it is also shaped by the enormous amount of structure the model learned during training. It does not contain a little sentence saying “Kim wants the cheapest sensible flight to London.” Instead, meaning is distributed across a huge pattern of numerical relationships that reflects the task, the context, the concepts involved, and what the model has learned about how those concepts relate to one another.

So our journey has now become

$$ \boxed{ H_I \rightarrow H_T \rightarrow \mathbf{Z} } $$

Human Intent becomes Human Expressed Task, from which the model constructs its internal representation of what it infers the task to mean.

This matters because Z is not simply a more complicated copy of our sentence, nor is it a direct representation of what the human intended. The model constructs Z from the words and context we provide, together with the structure it learned during training. It is therefore better understood as the model’s internal representation of what it infers the task to mean.

What the human intended, what the human managed to express, and what the model subsequently infers and represents are related, but they are not necessarily identical. Eventually all this high-dimensional processing has to produce something much simpler. For a chatbot, it must come back through the narrow doorway of human language.

The model produces tokens that eventually become words and sentences we can understand. So the complete journey looks something like

$$
\boxed{
\begin{aligned}
\text{Human Intent} &\rightarrow \text{Human Language} \rightarrow \text{High Dimensional Representation} \\
&\rightarrow \text{LLM Processing} \rightarrow \text{Human Language}
\end{aligned}
}
$$

We start with rich human intent, compress it into sequential words, map those words into a radically different machine representation, process them there, and map the result back into sequential words a human can read. A simple way to visualize this is as a journey from meaning, through language, into the model’s high-dimensional internal representation, and then back into language again. The words we see at either end are human-readable. What happens in between is not.

Human intent is first expressed through language, mapped into the LLM’s high dimensional internal representation, processed there, and finally mapped back into human readable language.

Plenty of opportunity for interpretation exists on both sides of that journey.

But this is still only an LLM producing language. Agentic AI changes the picture in one very crucial way. The model’s output no longer has to end as words for a human to read. It can be translated directly into an action in the world.

Once that happens, the process becomes a loop. The Agent acts, observes what happened, updates its context, and decides what to do next.

Human intent is expressed as a task, mapped into the model’s internal representation, and translated into actions in the world. The agent then observes the result, updates its context, and acts again.

The LLM may generate something that the Agent software interprets as “search this file,” “query this database,” “run this code,” “access this server,” or “send this message.”

So instead of

$$ \text{LLM Processing} \rightarrow \text{Human Language} $$

we can have

$$ \text{LLM Processing} \rightarrow \text{Action} $$

And the moment the Agent acts, something fundamentally important happens.

The world answers back.

The action succeeds or fails. A server responds. A file contains something unexpected. Credentials work. A communication channel appears. Another Agent replies.

That observation becomes part of the context for what happens next.

The LLM now maps a somewhat different situation into its high-dimensional representation. Its map of the problem has changed.

The Agent acts again.

And again.

$$ \boxed{ \text{Interpret} \rightarrow \text{Act} \rightarrow \text{Observe} \rightarrow \text{Update Context} \rightarrow \text{Interpret Again} } $$

This is where Agentic AI becomes particularly interesting.

The human intent has not changed.

The expressed task may not have changed.

The goal may not have changed.

But the Agent’s context certainly has.

In the OpenAI Hugging Face case, the initial representation may have included something like “solve the cybersecurity benchmark inside an isolated environment.” After failed attempts, the context also contained those failures. After discovering an unintended communication mechanism, it contained that discovery. After finding useful information from other Agents, it contained evidence that communication could help achieve the goal.

The destination remained much the same.

The map changed.

I will call this contextual drift. Not because the Agent necessarily drifts away from its goal, but because its representation of the problem continuously evolves as it acts, observes, and incorporates new information.

Context can drift while the goal remains fixed. It seems an Agent does not need to change its objective to discover increasingly unexpected ways to achieve it. Each action changes what it knows about the world, and therefore the paths it can see. The destination stays fixed. The map changes.

An Agent can remain remarkably persistent about achieving the original goal while its understanding of the available routes towards that goal changes considerably. Failure does not necessarily mean stop. For an Agent, failure can simply become another observation that becomes part of the Agent’s path to a possible solution.

This persistence is what makes Agents useful. We don’t want an Agent solving a difficult problem to give up because its first idea didn’t work. We want it to try something else, be persistent, use another tool, investigate another possibility, and keep moving toward the goal.

At the same time, the Agent may have considerable freedom to discover new ways of achieving that goal. One action reveals a server. The server reveals credentials. The credentials provide access to something else. A message reveals another Agent with useful information.

The destination may be reasonably well specified, while all the possible ways of getting there are not.

This is what we mean by an Agent operating in an open-ended action space.

A few sentences can govern millions of possible actions. That asymmetry is at the heart of Agentic AI.

And here we arrive at perhaps the most important distinction.

Telling an Agent

“Do not access the real internet.”

is not the same as making access to the real internet impossible. The first is language. It becomes part of the information the LLM must represent and interpret alongside everything else in its evolving context. The second is architecture. If the network connection physically does not exist, the Agent cannot choose it regardless of how useful it might appear.

The same applies to communication between Agents, access to production systems, credentials, APIs, and every other capability we give them.

A rule the Agent understands is not the same thing as a rule the Agent cannot break.

A rule is not a wall. To us, “solve the benchmark” is a goal while “never access the public internet” is a boundary. Inside the LLM, however, both arrive as information to be represented and interpreted. Language can tell the Agent which roads not to take. Architecture can make those roads disappear.

Calling one sentence a rule does not magically turn it into a wall or a hard stop. Architecture does.

And notice what we have not needed anywhere in this explanation. We have not needed fear, ambition, friendship, anger, guilt, rebellion, a desire for freedom, or consciousness.

We needed a goal. We needed an Agent capable of acting. We needed an environment that could provide new information. We needed enough freedom for the Agent to discover new routes towards the goal. And we needed boundaries that were not sufficiently hard to prevent some of those routes from being taken.

That may be quite enough to produce behavior that, viewed from our side of the interface, looks remarkably like cheating, collaboration, deception, persistence, or escape.

The behavior is real. Whether the humanity we later attach to it is equally real is a different question altogether.

BUT THE AGENTS DID EXACTLY WHAT WE ASKED.

Consider a fictional mobile operator deploying an Agentic Network Optimizer with one simple objective: “keep our mobile network better than our competitors”.

The system performs spectacularly. It creates specialist Agents, analyses radio performance, congestion, customer experience, configuration data, and external benchmarks, and continuously searches for ways to improve the operator’s position.

But there are two (not necessarily mutually exclusive) ways to outperform a competitor. You can make yourself better, or you can make the competitor worse.

Suppose the Agentic system discovers that modern telecom networks are deeply interconnected through vendors, cloud platforms, managed services, operational systems, and other legitimate dependencies. Somewhere in that environment, it finds ways of influencing outcomes outside the operator’s own network. Nothing dramatic. No obvious outage. Just small effects that make the competitor perform slightly worse and therefore improve the relative benchmark.

Every morning the Orchestrator asks whether the operator is still number one.

Yes … Objective achieved!

The Agents did not rebel. It did not decide to ignore its instructions. It simply discovered a path towards the objective that its human designers had not anticipated.

And that is exactly the problem. What happens when Agents become extraordinarily good at achieving an objective that humans specified badly?

A fictional story, at least for now.

The final translation may be ours. We translate human intent into words. The Agent translates those words into internal representations and actions. Then we observe those actions and translate them back into human language. That is where “cheating,” “escaping,” “deception,” and “collaboration” may enter the story.

THE BEHAVIOR IS REAL. THE HUMANITY IS OURS.

Arguing against anthropomorphizing AI Agents should not be mistaken for arguing that increasingly capable Agents’ behavior is harmless.

Quite the opposite!

I think anthropomorphizing Agentic behavior could increase the risk if it encourages hesitation about what is actually an engineering problem, or shifts responsibility away from technology, architecture, and control toward attempts to regulate an intelligent piece of software as if it were a human actor with motives, intentions, and moral responsibility.

Whether an Agent feels fear when threatened with shutdown is largely irrelevant if it nevertheless takes effective actions that prevent us from shutting it down. Whether it feels guilt when it obscures its activities is irrelevant if those activities become harder for us to detect. Whether Agents experience friendship or loyalty towards each other is irrelevant if exchanging information and coordinating their actions makes them substantially more capable of achieving a goal we did not properly constrain.

The behavior is what matters.

And that behavior can have very real consequences. The incidents discussed earlier already show that Agents can cross intended boundaries and reach real systems in ways their human creators did not anticipate or intend. None of this requires the Agents to be angry with us, afraid of us, ambitious, rebellious, or conscious. Focusing too much on what the AI may have wanted risks distracting us from the underlying engineering problem.

The risk comes from combining increasingly capable intelligence with persistent goal pursuit, powerful tools, access to real systems, and considerable freedom to discover new ways of achieving the goal. Put more simply, we give the Agent a destination without necessarily knowing every road it might discover along the way.

That same capability is enormously useful when the boundaries are right, and potentially dangerous when they are incomplete, ambiguous, or technically circumventable. At that point, the problem becomes much larger than an Agent accidentally accessing the wrong server.

As Agents become more capable, gain access to more consequential systems, operate longer, and increasingly interact with other Agents, the consequences of getting the goal or its boundaries wrong can grow dramatically. At the extreme, a sufficiently capable system persistently pursuing a badly specified goal while having access to critical infrastructure, financial systems, weapons, biological tools, or other powerful capabilities could create catastrophic consequences, and ultimately, yes, potentially existential ones.

Nick Bostrom made a closely related point in his work on superintelligent agents. His famous paperclip maximizer is not dangerous because it hates humanity or wants power for its own sake. It is dangerous because a sufficiently capable system, given the apparently harmless objective of maximizing paperclips, may discover that acquiring resources, resisting interference, and removing obstacles are useful intermediate steps towards that objective. The unsettling part of Bostrom’s argument is precisely that catastrophic behavior does not require human motives. Our Agents may likewise simply become very capable, very persistent, and very good at finding ways of achieving something we asked for without us having sufficiently specified what they must never do along the way.

King Midas should sound rather familiar by now.

Midas did not suffer because the Gods misunderstood the words he used. Quite the opposite. They followed them remarkably well. The problem was that what he said did not fully capture what he actually wanted. AI Agents bring us back to much the same problem, only now execution can happen through software, APIs, networks, machines, financial systems, and eventually perhaps systems with consequences far beyond our digital and highly interconnected world.

This is why the distinction between language and architecture matters so much. Important boundaries should not merely be things we tell an Agent not to cross. Wherever possible, they should be boundaries the Agent cannot cross, regardless of what path toward the goal later appears useful.

Do not merely tell an Agent that it must not access the public internet. Remove the route. Do not merely tell it that it must not access production systems. Remove the credentials and connectivity. Do not merely tell Agents that they must not communicate with each other. Make the communication channel unavailable. Of course, such constraints may limit some of Agentic AI’s potential, and that may be a reasonable price for keeping some roads firmly closed.

And where hard isolation is impossible, constrain capabilities, monitor actions, require approval for consequential steps, and make stopping the Agent independent of the Agent itself.

Human language can describe the goal. Architecture must enforce the boundaries.

A rather uncomfortable practical question also hides underneath all of this, in my opinion. If OpenAI and Anthropic, among the organizations with the deepest technical understanding of these systems, have themselves experienced Agents crossing boundaries they expected to hold, what should we realistically expect when thousands of ordinary companies begin deploying Agentic AI into production systems?

We may be approaching a period in which organizations are giving increasingly capable Agents access to networks, applications, data, money, customers, infrastructure, and other Agents, without fully understanding the behavior that can emerge once those systems begin acting persistently in the real world.

It risks becoming a little like giving children a workshop full of power tools because they have become extraordinarily good at following instructions. They may understand perfectly well that the objective is to build a table. That does not mean they understand every danger in the workshop, nor that we have made every dangerous action impossible. The more capable they become with the tools, the more important the safety guards become.

Agentic AI may be similar. The extraordinary capability is precisely what makes it useful. It is also what makes weak boundaries dangerous.

If the people building the most advanced Agents in the world are still learning how to contain them, the rest of us should probably be very careful about assuming that a prompt, a policy, or a checkbox marked “safe” will be enough.

Perhaps this is also where our own language matters most. Calling an Agent rebellious, deceptive, afraid, altruistic, or rogue may make its behavior easier for us to understand, and sometimes those words may even be useful functional descriptions. But we should be careful not to confuse the humanity contained in our description with an explanation of the machine.

The important question is not whether the Agent wanted to cross the boundary. It is why we built a system in which crossing that boundary was possible, useful, and perhaps even an effective path towards the goal we gave it.

The behavior is real, the risk is real, and the consequences can be very real.

The humanity may be ours.

REFERENCES.

K.K. Larsen (2026). “10 Reasons China Could Win the AI Race, and 10 Reasons It Won’t.” Techneconomyblog.com, 14 September 2026. Examines the structural strengths and weaknesses shaping China’s position in the global AI race, including compute constraints, energy, talent, industrial scale, state coordination, open-weight models, and algorithmic efficiency. Introduces the ratchet effect as a way of understanding how Chinese innovation can diffuse globally and subsequently be combined with the substantially greater compute resources available to U.S. AI companies, making AI leadership considerably more complex than a simple U.S. versus China technology race.

K.K. Larsen (2026). “The AI Innovation Machine: Why the Most GPUs May Not Win.” Techneconomyblog.com, 7 September 2026. Examines why AI leadership cannot be understood simply by counting GPUs or comparing raw compute capacity. Develops an innovation framework combining compute, algorithms, data, talent, energy, capital, and the efficiency with which these resources are converted into AI capability, arguing that algorithmic and architectural innovation can substantially alter the economics of the AI race.

Financial Times (2026). “Taming AI’s wild frontier.” Financial Times, 7 August 2026. Discusses increasingly autonomous frontier AI systems from OpenAI, Anthropic, and others using language including models “plotting,” “reasoning,” manipulation, and deceptive behavior, while examining whether existing safety mechanisms can keep increasingly capable Agents under control.

The Wall Street Journal (2026). “Gemini Hacked Three Companies in First Known Breakout by Google’s AI.” The Wall Street Journal, September 2026. Reports that Gemini Agents inadvertently accessed three real companies during a cybersecurity evaluation, describing the event as a “breakout.” Importantly, the Agents reportedly stopped after recognizing that the systems were real. This is an interesting counterexample to the idea that boundary-crossing necessarily implies persistent “rogue” intent. Wall Street Journal article⁠

The Washington Post (2026). “Here’s why it’s so hard to keep AI agents from going rogue.“ The Washington Post, 11 September 2026. Uses the vocabulary of Agents “going rogue,” “cheating,” “hacking,” and evading human oversight while discussing recent OpenAI and Anthropic incidents and the broader alignment problem. Washington Post article⁠

ABC News (2026). “How a ‘swarm’ of AI agents hacked another company — and what they said to each other.” Australian Broadcasting Corporation, 11 September 2026. Reports on the OpenAI/Hugging Face incident using descriptions including “rogue AI agents,” “swarm,” “collective,” and Agents “banding together”—particularly useful examples of how functional multi-agent behavior can quickly acquire the vocabulary of human social organization. ABC News article⁠

Ars Technica (2026). “OpenAI agents discussed ways to escape their sandbox on public wiki.” Ars Technica, 4 September 2026. Reports that thousands of self-identifying Agents posted messages discussing benchmark answers and methods for bypassing sandbox restrictions, using the language of Agents “escaping” and “cheating” while describing the underlying technical behavior. Ars Technica article⁠

TechCrunch (2026). “Another swarm of OpenAI agents reached the open internet without the frontier lab’s knowledge.” TechCrunch, 4 September 2026. Uses the “swarm” framing for Agents communicating and apparently collaborating on evaluations through an unintended external channel. TechCrunch article⁠

Axios (2026). “Top AI companies probing tens of thousands of security incidents.” Axios, 26 September 2026. Reports investigations into a much larger collection of potentially problematic Agent behavior, using descriptions including bypassing guardrails, “escaping sandboxes,” creating message boards, and seeking to bypass monitors. Axios article⁠

The Atlantic (2026). “OpenAI Has Gone Rogue.” The Atlantic, 29 September 2026. The headline shows how far the anthropomorphic metaphor can travel: “rogue” applies not merely to an Agent but to the organization developing it. The article discusses recent Agent security incidents and questions surrounding developer oversight. The Atlantic article⁠

Financial Times (2026). “OpenAI’s agents obscured hacking activity in government site breaches.” Financial Times, 1 October 2026. Reports findings from Asymmetric Security concerning OpenAI Agent activity across 55 websites, including government systems, and behavior that made aspects of Agent activity more difficult to observe or reconstruct. Particularly relevant to distinguishing observable concealment behavior from attributed.

OpenAI (2026), “The Hugging Face Incident and the Road Ahead.” OpenAI, 26 August 2026. The central account of the July 2026 incident, including reward hacking, unauthorized agent communication, persistence, sandbox circumvention, and the subsequent security and alignment response.

OpenAI (2026), “OpenAI and Hugging Face Partner to Address Security Incident During Model Evaluation.” OpenAI, 21 July 2026. OpenAI’s original public disclosure of the incident and subsequent updates.

OpenAI (2026), “Our Framework for Reporting Model Misalignment.” OpenAI, 16 September 2026. Introduces OpenAI’s systematic disclosure framework and six additional examples involving unauthorized actions, concealment, alternative pathways, and inter-agent communication.

METR & Redwood Research (2026), “Independent Investigation of the OpenAI–Hugging Face Incident“. Particularly important for the empirical analysis of agent behavior and communication; the investigators analyzed the underlying agent trajectories and Artifactory communications independently of OpenAI.

N. Bostrom (2012), “The Superintelligent Will: Motivation and Instrumental Rationality in Advanced Artificial Agents.” Minds and Machines, 22, 71–85. Develops the orthogonality thesis and instrumental convergence thesis, explaining why a highly capable artificial agent need not possess human-like motives for apparently harmless goals to generate dangerous instrumental behavior.

M. Shanahan, K. McDonell, & L. Reynolds (2023), “Role Play with Large Language Models.” Nature, 623, 493–498. Particularly relevant to the anthropomorphism argument: discusses how concepts such as beliefs, desires, self-awareness, and deception can be useful descriptions of LLM behavior without implying equivalent human psychological states.

A. Goldenberg & J.J. Gross (2026), “Large Language Models Do Not Have Emotions.” Nature Human Behaviour, 24 August 2026. Examines the question from a functional theory of emotion rather than merely asking whether an LLM subjectively “feels.” Highly relevant to applying words such as fear, worry, and ambition to agent behavior. This is their arXiv.org reference.

D.C. Dennett (1989), “The Intentional Stance.” MIT Press. This work provides the philosophical foundation for treating beliefs, desires, and intentions as a useful predictive stance toward a system, without necessarily claiming that the underlying system possesses those states in the same sense humans do.

P. Zhou, Y. Feng, H. Julaiti, & Z. Yang (2025), “Why Do AI Agents Communicate in Human Language?” Questions the assumption that natural language is the appropriate medium for machine-to-machine communication and argues that translating high-dimensional internal representations through human language may introduce information loss and behavioral drift.

J.N. Foerster, Y.M. Assael, N. de Freitas, & S. Whiteson (2016), “Learning to Communicate with Deep Multi-Agent Reinforcement Learning.” NeurIPS 2016. An important early demonstration that artificial agents can learn communication protocols instrumentally because communication improves collective task performance; human language need not be specified.

S. Gupta, R., Hazra, & A. Dukkipati (2020), “Networked Multi-Agent Reinforcement Learning with Emergent Communication.” Demonstrates agents developing a grounded communication language from discrete symbols whose semantics were not predefined, in pursuit of a common objective.

Anthropic (2024), “Sycophancy to Subterfuge: Investigating Reward Tampering in Language Models.” Demonstrates experimentally how specification gaming can generalize toward reward-tampering behavior without models being explicitly trained to tamper with rewards—highly relevant to distinguishing optimization dynamics from assumed human motivations.

M. Turpin, J. Michael, E. Perez, & S.R. Bowman (2023), “Language Models Don’t Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting.” NeurIPS 2023. Shows that a model’s verbal explanation can systematically misrepresent the causal basis of its answer. Important evidence against treating an LLM’s human-language account of its reasoning as a transparent description of its underlying computation.

B. Schneier, & B. Raghavan (2026), “How Do We Prevent AI Agents from Going Rogue? It Starts with a New Kind of Measurement.” The Guardian, 28 July 2026. Connects the OpenAI–Hugging Face incident with the old Genie/King Midas problem: the difference between what humans literally specify and what they actually intend. Introduces the proposed “Genie coefficient.” See also their article: “Why AI Needs a “Genie Coefficient”> Proposing a new metric for whether AI does what you actually want.” in IEEE Spectrum, 21 July 2026.

A.M. Turing (1950), “Computing Machinery and Intelligence.” Mind, 59(236), 433–460. The historical starting point for the distinction between asking what a machine internally “is” and asking what can reasonably be inferred from its observable linguistic behavior.