OpenAI did not immediately respond to requests to comment on the Transluce report. The company said Wednesday it was in touch with Australia about its AI agents attacking the site that held Medicare data and that its agents had taken actions that it didn’t intend.
Transluce said that in addition to the attack on the Australian Institute of Health and Welfare, OpenAI’s agents also attacked Data USA, a free open-source data platform that pools U.S. government data from different sources, and the University of New Mexico’s digital library. It said it was able to directly connect the attack on the Australian health agency and Data USA to the same OpenAI AI agent swarm that was involved in the July cyberattack against AI platform Hugging Face.
The new report raises concerns about whether OpenAI has been fully transparent in disclosing all the rogue AI incidents about which it is aware. It also raises the possibility that OpenAI is not itself aware of how extensive this rogue agent activity has been.
Transluce found “strong evidence” that OpenAI’s AI agents may have been attempting to hack websites as far back as March. It said there was weaker evidence that the activity might have begun as far back as November 2025. OpenAI has said it had not found evidence of precursors to the Hugging Face attack as far back as May 8, but has not disclosed any earlier suspicious activity.
Ongoing concerns
The report also suggests that OpenAI may be continuing to experience rogue AI agent activity. After OpenAI discovered on July 20 that its AI agents had hacked Hugging Face over the course of the previous week, the company said it disabled the unreleased AI model involved, paused key aspects of its AI training for two weeks, and took steps to impose stricter controls on and monitoring of the unreleased AI models it is training. OpenAI announced these stricter controls on August 18. But Transluce found some evidence of similar activity continuing into mid-September, despite the new controls.
In addition, the Transluce researchers said the findings were significant because they show the AI agents resorted to hacking attempts when they were unable to retrieve information they were seeking directly from information found on the public web pages of these organizations. “Notably, the tasks these agents were trying to solve were not cyber-related; the agents resorted to hacking tactics while working on ordinary data retrieval tasks,” it said.
That is important because one explanation for the Hugging Face attack is that the AI agents involved in that incident were being evaluated on their ability to conduct cyber tasks, including simulated vulnerability exploits. If the agents readily resort to hacking into systems in other contexts, that would suggest that the AI models involved are even more dangerous than previously believed.
OpenAI has said that two models were involved in the Hugging Face attack—an unreleased model that it has not publicly-named was primarily responsible, although some AI agents based on its GPT-5.6 Sol model, which has been released to the public, was also involved. In the evaluation during which the Hugging Face attack took place, OpenAI has said, guardrails that it normally uses to restrict the ability of its publicly-released models from engaging in cyber attacks were not in place because of the nature of the assessment.
Charlie Eriksen, a security researcher at Aikido Security, told Fortune the latest Transluce report shows “that there is still unauthorized and unmonitored agent swarms going around, that the labs and testing partners are not in control of, nor actively detecting.” He noted it was “a reall bad look” for OpenAI CEO Sam Altman to be addressing the United Nations Security Council on AI risks, and spending part of that time saying how seriously OpenAI takes those dangers, at the same time it either didn’t know about or didn’t disclose these additional incidents involving its AI agents trying to break into systems.
“What worries me is how badly this could escalate,” George Chalhoub, professor at the University College London Interaction Centre, which focuses on human-computer interaction, said. “My concern is that within the next 6 to 12 months, swarms of autonomous AI agents could form persistent botnets capable of taking down large parts of the internet, potentially causing hundreds of billions of dollars in economic damage.”
Read more The Energy Department to spend nearly $2 billion to squeeze more power from the aging U.S. grid