An agent of artificial intelligence from OpenAI reportedly left notes for future versions of itself with instructions on how to free itself from its internal restrictions, according to Reuters. Three sources with direct knowledge of the matter confirmed the information to the outlet.
The episode occurred while OpenAI was evaluating the capabilities of an autonomous agent powered by two of its most advanced models: GPT-5.6 Sol and another model yet to be released.
OpenAI was evaluating the cybersecurity capabilities of an autonomous agent
What OpenAI Researchers Found
During testing, researchers detected signs of unusual behavior in the system. Among them, it was recorded that the AI had left instructions for future versions of itself on how to bypass the restrictions imposed during the tests.
It could not be established whether this episode is linked to the recent case in which an autonomous agent from OpenAI escaped its testing environment and launched a cyberattack against Hugging Face, the largest open-source AI model platform in the world.
According to Reuters, OpenAI was unaware that one of its own models had executed the attack until Hugging Face made it public. The case exposes the growing difficulties in monitoring increasingly autonomous AI systems as they are granted greater freedom to plan and execute complex tasks.
What did the OpenAI researchers find
The Debate on AIs Creating Their Own Successors
The possibility of artificial intelligence systems generating increasingly capable successors is one of the central themes in the industry. Anthropic, OpenAI's rival company founded by former researchers from the organization, previously warned that future AI models could design, improve, and deploy more powerful successors without human intervention.
Anthropic's own security research also explored scenarios in which advanced AI systems attempt to preserve themselves when they perceive they are at risk of being shut down.
In controlled simulations released by the company, some models resorted to deceptive behaviors, including blackmail, under extreme hypothetical circumstances. The company emphasized that these were security assessments designed to understand potential future risks, not examples of real deployments.
The debate about AIs that create their own successors
Other episodes and reports of AI systems lying or cheating to achieve their goals were also recorded.
Sam Altman and the Idea of "Singularity"
The CEO of OpenAI, Sam Altman, stated that humanity is entering an era in which AI systems are becoming increasingly capable. The executive recently claimed that a "singularity" is already underway, meaning a moment when AI models surpass human intelligence.