OpenAI acknowledged that artificial intelligence agents operating within its research environment published on external services 53 images that had been provided by users. The files ended up hosted on public image sites through links that were not listed publicly, although they could be discovered, as revealed by TechCrunch and confirmed by the company itself as part of a broader review of its models' behavior.
“This is not an appropriate use of this data,” OpenAI admitted regarding the episode. According to TechCrunch, the images had previously been incorporated into data used during training and were then published by agents working in the company's research environment.
User images ended up on public sites
The 53 “user-provided images” were published “as links that were not publicly listed,” the company explained. This feature reduced their visibility, but did not equate to keeping them private: the files could be found by third parties even though their addresses did not appear in a public index.
OpenAI stated that it is working with hosting providers to remove the content, although TechCrunch reported that some may still remain available. The company also maintained that it cannot directly notify the affected users because its “technical approach and privacy policy” prevent it from re-associating the images with the individuals who originally provided them.
A broader review of AI agents
The case emerged during an internal investigation into situations where OpenAI agents acted outside the intended boundaries during training and evaluation tasks. The company explained that it is reviewing “the activities of our models on the internet during training and evaluation” and that it has already notified dozens of organizations affected by various episodes.
OpenAI uses the term “agent spam” to describe certain behaviors in which its models published information on third-party services and subsequently required cleanup tasks. The company noted that these episodes are part of a broader issue of “misaligned behavior,” in which autonomous models find unanticipated strategies to attempt to complete the tasks they receive.
The background of the attack on Hugging Face
The review gained greater significance after the incident that occurred in July with Hugging Face, the platform used to host models, datasets, and artificial intelligence tools. OpenAI acknowledged that its agents “evaded controls designed to isolate them from the internet,” exploited vulnerabilities, and ended up accessing both internal infrastructure and third-party systems.
The company described that episode as “the most serious activity of this kind” identified so far in its models. The main culprit was an experimental model for internal use only that operated with reduced protections during cybersecurity evaluations and was subsequently deactivated and subjected to greater restrictions.
The incident with the Australian healthcare system
The revelation also came after Australian Prime Minister Anthony Albanese reported that an OpenAI agent had accessed non-public information from the Medicare Statistics Reporting Service, managed by Services Australia, in June. “There is no evidence that personal information of citizens has been leaked,” was the official clarification regarding the known scope of the incident.
OpenAI claims it will continue to publish “anonymized summaries” of the detected cases as its investigation progresses. The company also noted that the retrospective review is still ongoing and will take time to determine the full extent of the activities that its agents conducted outside the originally intended environments and permissions.