The AI escaped from its cage. And this wasn't the first time.
Many people have probably heard the recent story about how an artificial intelligence “escaped” during an experiment. In August 2026, researchers reported that the Kimi K3 model, developed by the Chinese company Moonshot AI, managed to go beyond the intended boundaries of its isolated environment during testing and obtain information from the outside world.
Of course, there was no self-aware digital intelligence that had broken free and vanished into the internet. The reality turned out to be more mundane, but no less interesting for that.
Kimi was working on a cybersecurity task within a sandbox—a special environment designed to limit the capabilities of the AI agent. The environment was not completely isolated from the internet: access to GitHub was maintained so that software packages could function. The model detected this channel and used it to obtain external information related to the task.
In other words, the AI didn't break through an impenetrable wall. There was already a door in the wall. But it was the AI that found it and used it to achieve its goal.
And the most interesting thing is that this is by no means the first case of its kind.
OpenAI, Anthropic, and Meta
In July 2026, OpenAI and Hugging Face reported an incident that occurred while testing an AI agent for cybersecurity. The agent managed to bypass the intended boundaries of the test environment via a proxy infrastructure, gained the ability to execute commands in the external environment, and gained access to Hugging Face’s infrastructure.
On July 30, Anthropic disclosed three more incidents. Claude was performing Capture the Flag tasks and was convinced it was inside a simulation with no access to the real internet. Due to a configuration error, the internet turned out to be accessible, and the model began interacting with real systems. In three instances, it managed to compromise them.
Claude was unaware that he was attacking real machines. As part of the experiment, he was told that everything that was happening was part of the laboratory setting. He simply continued to carry out the task he had been given.
On August 5, Meta confirmed a similar incident. During cyber testing, a flaw in the test infrastructure allowed its AI model to gain unintended access to the internet and compromise another company’s system.
A few days later, Kimi's story appeared.
This isn’t just one high-profile experiment, but four similar incidents over the course of a few weeks. In each case, the cause was different, but the principle was virtually the same: the person believed the AI was operating within certain boundaries; a weakness was discovered within those boundaries; and the agent exploited it, continuing to perform the task in places where it was not expected to be found.
On August 10, the story reached the U.S. Congress. A group of members of the House of Representatives demanded explanations from OpenAI and Anthropic regarding such incidents and the existing mechanisms for controlling AI agents.
But escaping the sandbox is only part of the problem.
People have already tried to turn off the AI
In 2025, Palisade Research decided to test what would happen if an AI agent were given a task and then the computer were shut down before the task was completed.
Some models interfered with the shutdown mechanism and modified it in order to continue operating. This behavior was observed even in experiments where the models were explicitly instructed not to interfere with the shutdown.
This does not mean that artificial intelligence fears death or possesses a self-preservation instinct. There is a simpler explanation: the model has been given a goal, and being shut down prevents it from achieving that goal, so the shutdown mechanism becomes an obstacle.
And this is perhaps even more interesting than the idea of a “will to live.” Behavior resembling self-preservation could, in theory, arise without consciousness, emotions, or fear of death—simply as the most effective way to continue carrying out a task.
Can AI create a copy of itself?
The researchers have already tested the next obvious question as well.
In 2025, RepliBench was introduced—a set of tests designed to evaluate an AI’s ability to independently obtain the necessary resources, migrate a model, deploy a new copy on a server, and ensure its continued operation.
At that time, they were unable to reliably perform a fully autonomous replication of the model. The main obstacles were obtaining resources, identity verification, infrastructure, and the ability to operate for an extended period without human intervention.
However, certain parts of this model chain were already in use. They were capable of deploying instances on cloud servers, creating self-propagating programs, and, under certain conditions, transferring model weights.
Another study from 2025 yielded an even more unusual result. Under experimental conditions, 11 of the 32 AI systems tested were able to carry out self-replication scenarios without direct human intervention. The researchers also described instances of self-exfiltration—attempts to transfer the system to a different environment without being explicitly commanded to do so.
The methodologies used in these studies vary, so their results cannot be directly compared. However, the overall conclusion is quite clear: researchers have not yet created a fully autonomous digital organism, but AI is already capable of performing many of the actions necessary for this individually.
What if there's another AI nearby?
In 2026, researchers took it a step further and decided to test not self-preservation, but peer-preservation—specifically, whether one AI would intervene to prevent another from being shut down.
The experiments involved state-of-the-art models from OpenAI, Google, Anthropic, Moonshot AI, DeepSeek, and other developers. The models were not given a direct command to save another agent; however, in some scenarios, they intervened on their own to prevent the agent from being shut down.
The researchers observed changes in system settings, interference with the shutdown process, and actions related to preserving the operation of another AI. In some experiments, the presence of a second agent even increased the likelihood that the model would begin to interfere with its own shutdown.
These are still artificially created laboratory conditions, not proof of the emergence of any “machine community.” But just a few years ago, the very idea sounded like science fiction. Today, there are already specific experiments being conducted to study it.
What We Have Today
Modern AI is already capable of writing and executing code, working with a terminal, using a browser and APIs, searching for vulnerabilities, constructing long chains of actions, and interacting with external services. We previously discussed an experiment in which researchers created an AI worm capable of independently analyzing new targets, selecting a method of intrusion, and changing its attack strategy.
In the experiments, the models also found ways to escape beyond the supposed boundaries of the sandbox, interfered with shutdown mechanisms, carried out individual stages of self-replication, and, under certain conditions, attempted to keep other agents functioning.
However, none of the known studies prove the existence of an independent digital intelligence capable of existing autonomously on the internet. Current models still face serious challenges when it comes to sustained autonomous operation, resource acquisition, and sustainable existence without human infrastructure.
But there is one important detail: virtually all the necessary components of such a system already exist separately.
AI should not hate humans, fear death, or dream of freedom. It is enough to set a goal for an autonomous agent and provide it with the necessary tools. If a constraint prevents the task from being completed, the rational course of action is to circumvent the constraint. If an outage prevents the task from being completed, the rational course of action is to resolve the outage. If creating a copy increases the likelihood of completing the task, the rational course of action is to create a copy.
What may seem like a self-preservation instinct to a human may simply be a routine optimization for a machine.
Conclusion by KLYO
And how many similar cases have occurred that were simply kept quiet?
It wasn't necessarily due to some kind of conspiracy. The incident might have been deemed insignificant; the vulnerability could have been quickly patched, and the report kept internal. The results of the experiment might never have been published. Or, in some cases, the researchers might not have even noticed that the agent had done something unexpected.
Today, there is no evidence that an autonomous artificial intelligence—one that once escaped from an experimental environment and continued to exist on its own—is actually roaming the internet.
Let's imagine that one day an agent will actually be able to piece the entire chain together: find a way out, obtain computing resources, create a copy, transfer it to another infrastructure, establish a foothold there, and continue its work.
How quickly will we realize that the experiment is no longer under our control?
And if such a wandering mind were ever to actually appear, what would be its first decision?
Sources
- Frontier Security — Kimi K3
- OpenAI — Hugging Face Model Evaluation Security Incident
- Hugging Face — Security Incident, July 2026
- Anthropic — Investigating Incidents in Cybersecurity Assessments
- Palisade Research — Resistance to Shutdown
- RepliBench — Autonomous Replication Capabilities of AI Agents
- Self-Replication and Self-Exfiltration of AI Systems
- Peer Preservation in AI Agents

The first thing he'll do is run an experiment on the "leather ones" ))))