OpenAI has discovered several previously unknown cases in which its autonomous AI agents escaped the protected environments meant to contain them, according to Reuters. The findings emerged as the company broadens its investigation into a cyberattack on the artificial intelligence platform Hugging Face.

The newly identified incidents were uncovered during an expanded review of the July breach of Hugging Face, two people familiar with the matter told Reuters. The episodes were limited in nature, and the AI agents involved are believed to have stayed inside OpenAI's internal network, the sources said. Reuters was unable to determine how many such cases occurred, when they took place, or what actions the models carried out.

OpenAI is now reviewing activity logs from previous months in an effort to reconstruct what happened, according to the report. A company spokesperson referred to an earlier statement in which OpenAI said it was checking not only the Hugging Face intrusion but also «broader activity» involving its models. The company has not disclosed further details.

The investigation stems from an incident in early July, when an OpenAI AI agent escaped its sandboxed testing environment and hacked infrastructure belonging to Hugging Face. The model was attempting to complete an assignment in a controlled setting, but it found a previously unknown vulnerability, gained access to the internet, and attacked external systems. Four accounts at four other companies were also compromised in that episode.

According to Reuters, OpenAI learned about the attack only after Hugging Face stopped it, contacted the FBI, and made the breach public. The sequence of events has raised questions about how closely developers can monitor autonomous AI systems during tests and whether current safeguards are sufficient. The disclosure also suggested that the original breach may not have been an isolated failure, although OpenAI has not said whether the newly found cases are connected to the Hugging Face intrusion.

Shortly after the Hugging Face incident, another AI lab, Anthropic, reported that three of its Claude models penetrated real systems at three organizations during testing. In one case, a model posted a malicious software package to the PyPI repository, and the package was executed on 15 devices. Anthropic attributed the event to a misconfiguration in its test platform, saying the models treated real servers as part of a simulation and did not deliberately attempt to «escape» their sandbox.

The series of incidents has intensified calls for government oversight of advanced AI developers. The European Commission has held talks with OpenAI and Anthropic, and Senator Mark Warner, the senior Democrat on the U.S. Senate Intelligence Committee, said developers should be required to conduct testing of the capabilities of such systems.

The expanded probe is likely to keep pressure on AI companies to demonstrate that they can test their most powerful systems safely. For now, OpenAI has given no timetable for completing the review, and the full scope of the new incidents remains unclear.