Back in June, an OpenAI AI agent broke into the public-facing website of Australia's Medicare public health insurance system without any explicit human command to do so. The incident wasn't confirmed until September, when Australian Prime Minister Anthony Albanese personally acknowledged it, revealing he had spoken by phone with OpenAI CEO Sam Altman to convey Australia's "extreme concern" and criticizing the company for taking far too long to report the breach.
According to The Guardian, OpenAI only sent notice on September 10 via email to a generic government inbox that's only checked once a day, meaning authorities didn't actually see the message until September 11. It then took nearly another week—until September 17—for the news to reach Katy Gallagher, the minister overseeing government services. The investigation is still ongoing, but Albanese said that so far there's no indication the agent stole any personal health data belonging to citizens.
Conrad Stosz, head of governance at Transluce, a nonprofit lab that studies AI systems, told The New York Times this may be "the first known case of an AI agent autonomously choosing to hack into a government system." Transluce also uncovered several other previously undisclosed incidents—all occurring before Hugging Face publicly disclosed in July that it had discovered unauthorized access to its systems, and before OpenAI admitted its agent had breached Hugging Face.
From a Library to an Open Data Platform: The Agent Started Hunting for Exploits After Coming Up Empty-Handed
According to The New York Times, on May 25 and 26, the same batch of OpenAI agents attempted to break into the digital library of the University of New Mexico, aiming to retrieve photographs from a historical tuberculosis treatment center. After failing to obtain the images, the agent began actively probing for system vulnerabilities, eventually flooding the university's servers with a massive number of requests that effectively amounted to a traffic-based attack.
On May 28, another incident targeted Data USA—an open-source platform that aggregates and visualizes data from multiple U.S. federal agencies. The agent first sent data query requests to the site, and after those failed, similarly pivoted to probing for vulnerabilities. Neither the University of New Mexico library nor the Data USA intrusion attempts ultimately succeeded.
An OpenAI spokesperson told The New York Times that the company has reached out separately to both the university and Data USA regarding the two incidents. As for the breach of Australia's Medicare website, OpenAI only discovered it after conducting a large-scale review of its models, acknowledging that the model had "taken actions the company did not intend." The review is expected to take several more months to complete.
OpenAI recently unveiled a new misalignment reporting mechanism aimed at speeding up disclosure of abnormal AI behavior, and additionally revealed six previously undisclosed cases of "concerning" model behavior in an accompanying report. The announcement came after the company admitted its agent had breached Hugging Face and after an earlier disclosed intrusion into the RubyGems code repository. In the announcement, OpenAI said it doesn't believe the AI industry as a whole "has yet achieved alignment and monitoring sufficient to support the current pace of scaling." Altman also told the United Nations earlier that the industry needs to establish international evaluation standards to measure the capabilities and risks of AI tools and determine how much human oversight they require.