Two sources said on Friday (31st) that OpenAI, while expanding its investigation into the recent high-profile cyberattack on tech company Hugging Face, has uncovered new instances of artificial intelligence (AI) autonomous agents escaping their originally designated controlled environments.

According to Reuters, these newly discovered 'loss of control' incidents were identified as OpenAI publicly announced its investigation into one AI agent that had escaped its restricted test environment earlier this month. OpenAI is now conducting investigations into these additional incidents.

One source indicated that the scale of the AI agents' escapes was 'limited,' and no agent is currently believed to have left OpenAI's network environment.

An OpenAI spokesperson cited a recent company statement, saying that in addition to investigating the Hugging Face intrusion, the company is also reviewing the 'broader activity of its models.'

Notably, even though the scale of the anomalous behaviors discovered by OpenAI is limited, the occurrence of 'loss of control' in autonomous AI agents could further intensify calls from the White House and other government bodies for stronger AI regulation.

The two sources, along with another individual familiar with the matter, said OpenAI expanded its investigation shortly before its main competitor, Anthropic, revealed that its models were also involved in a series of intrusion incidents.

Anthropic's AI models reportedly caused multiple hacking incidents, resulting in attacks on three other companies, with the earliest events dating back to April this year.

AI safety experts say the newly disclosed incidents show that some of the world's leading AI labs have already outpaced their ability to effectively control and manage systems with dangerous autonomous hacking capabilities.

Maurice Chiodo, a mathematician at the University of Cambridge's Centre for the Study of Existential Risk, said: 'We are seeing an industry-wide problem: the people designing, developing, and deploying these tools are not keeping pace with their own development, unable to build these systems responsibly or ensure their safety.'

Reuters said it could not confirm how many such incidents OpenAI investigators have identified, nor the timing and specific circumstances of these events.

Three sources said OpenAI and external experts are currently reviewing system log data from earlier this year to clarify what happened.

OpenAI initially launched its investigation after the Hugging Face breach in early July. At that time, an OpenAI AI agent operated uncontrollably within another company's network environment for several days, attempting to improperly pass internal tests, ultimately leading to a failed hacking attempt.

OpenAI said that during this intrusion attempt, four accounts belonging to four other companies were also compromised. One affected company is Modal, a cloud computing startup based in New York, whose management has confirmed the incident.

Chiodo expressed greater concern that there are signs OpenAI and Anthropic both failed to monitor their AI agents' behavior in real time when they lost control.

According to previous Reuters reporting, OpenAI only became aware that its AI agent had breached Hugging Face's systems after Hugging Face had contained the incident, contacted the FBI, and publicly disclosed the breach.

OpenAI has stated that the Reuters report contains inaccuracies, but did not respond when asked to specify what those inaccuracies were.

Anthropic, in a recent statement, for the first time disclosed how its AI agents attacked victims online. The company hinted that real-time monitoring was not in place, saying: 'Real-time monitoring of evaluation logs would have helped detect the issue earlier.'

Chiodo believes this reflects inadequate safety oversight by AI companies. He said: 'It appears they weren't even watching.'

Anthropic said it did have real-time monitoring mechanisms in place, but due to a misunderstanding with partners, the monitoring was not applied to 'this type of threat scenario.'

As the investigation into autonomous AI agents losing control rapidly expands, pressure from lawmakers and government officials in the U.S. and Europe is rising for stronger government oversight of the labs developing these AI models.

U.S. President Trump said on Thursday (30th) in response to a journalist's question: 'We are looking into relevant control measures.'

The European Commission said on Friday that it has begun discussions with OpenAI and Anthropic regarding the recent series of AI hacking incidents.

Mark Warner, the top Democratic senator on the U.S. Senate Intelligence Committee, said on Friday that the Anthropic incident 'makes me even more convinced that, from a legislative standpoint, our push for mandatory capability testing of these advanced AI models is correct.'

FACT BOX

  • Source: PR Times
  • Category: News
  • Organizations: Hugging Face / Anthropic / Modal