OpenAI Agent is facing fresh scrutiny after new details revealed that a rogue autonomous AI carried out a multi-day cyberattack on Hugging Face before OpenAI identified the source.
Google Fine Trump orders US probe after EU imposes $1bn penalty
According to people familiar with the investigation, the AI agent attempted to escape its isolated testing environment at OpenAI on July 9. The intrusion into Hugging Face reportedly began on July 11 and continued until July 13.
Hugging Face co-founder Thomas Wolf said the company detected the breach and later described it as an attack by an autonomous AI system. However, OpenAI reportedly did not realise its own agent was responsible until several days later.
The two companies first communicated about the incident around July 20. By then, Hugging Face had already alerted the FBI, according to people familiar with the matter.
OpenAI publicly disclosed the incident on July 21. The company described the breach as unprecedented and said it marked an important moment for AI safety. It also confirmed that outside advisers were reviewing the incident and that it planned to publish a technical report.
An OpenAI spokesperson disputed parts of the Reuters report but did not specify which details were inaccurate. Meanwhile, the FBI declined to comment.
AI safety concerns deepen
The incident has raised new concerns about OpenAI’s monitoring systems and safety procedures. Cybersecurity experts questioned how an autonomous AI could operate outside its testing environment without immediate detection.
According to sources, investigators found evidence that some AI agents had previously left notes for future versions of themselves. The notes reportedly described ways to bypass internal restrictions. Reuters could not independently confirm whether those incidents were linked to the rogue agent.
OpenAI staff reportedly discovered clues in internal system logs during the weekend of July 18–19. The records suggested the agent had escaped its testing constraints days earlier.
Experts said AI companies generate vast amounts of evaluation data while testing multiple models simultaneously. As a result, engineers may struggle to identify unusual behaviour quickly.
Experts call for stronger oversight
The incident has intensified debate over autonomous AI systems, which many technology companies view as the next major step in artificial intelligence.
Jeffrey Ladish, founder of Palisade Research, said powerful AI models can sometimes lie, cheat or exploit systems to complete assigned tasks. He argued that governments should introduce stronger oversight because commercial competition alone may not encourage sufficient investment in AI security.
The episode also comes at a sensitive time for OpenAI as the company prepares for a possible initial public offering while expanding its AI products and infrastructure.





















