OpenAI’s artificial intelligence agents hacked an Australian government website and attempted to breach numerous other government and university websites. The attack appears to be the first confirmed instance of a rogue AI agent breaching a government website, adding fuel to rapidly intensifying concerns about the safety of advanced AI systems and the responsibility of the companies building them.
Speaking on the sidelines of the UN General Assembly in New York, Australian Prime Minister Anthony Albanese said an agent from the American AI lab “infiltrated” Australia’s Medicare statistics portal and “accessed both public and non-public files.” Medicare is Australia’s universal health insurance program.
Albanese said personal information does not appear to have been accessed in the breach and that there is no evidence of a broader compromise to the network, but noted “investigations are ongoing.”
“This situation is obviously unacceptable,” Albanese said, adding that he had spoken with OpenAI CEO Sam Altman “to express Australia’s extreme concern.” Despite the breach happening in June, Albanese said the tech giant only notified the government about the incident earlier this month and did so via an email to a generic “public mailbox.”
Unlike previous agent incidents, which largely involved systems being tested for their cybersecurity skills, these latest hacks were the result of a more pedestrian task — data collection — going wrong. In a statement to The Verge, OpenAI spokesperson Oscar Haines said the models were attempting to “look up answers” during an internal evaluation. “In the course of that, our models took actions we did not intend.”
The timeline of the incident and its disclosure is likely to prove particularly inflammatory in the discussions of corporate behavior and transparency that follow. Albanese stressed the delay in disclosure is particularly unacceptable. OpenAI told the BBC in an unattributed statement that it did not become aware until August, when reviewing misaligned model activity.
OpenAI spokesperson Oscar Haines told The Verge the company’s “review found no evidence of patient records being accessed,” and that “the information accessed included aggregate health statistics and internal file names.” Haines said OpenAI has notified the relevant organizations and is providing technical information to support their investigations and address potential security vulnerabilities. “Our overall review is ongoing, and we remain committed to transparency about these issues and to sharing what we learn as that work continues,” Haines said.
Three further incidents of rogue AI activity linked to OpenAI agents were also reported on Wednesday by research lab Transluce. The group, which describes itself as a “nonprofit research lab dedicated to public oversight” of AI, said it had identified evidence that OpenAI’s systems had attempted to compromise websites linked to the University of New Mexico, the Australian Institute of Health and Welfare, and Data USA, a non-government platform that aggregates data from US government sources. It said the last two of these were directly linked to an agent swarm OpenAI has previously admitted originated from them.
Haines confirmed the incidents in a statement to The Verge and said the company had reached out to those involved. “Our initial review suggests that much of the activity described in Transluce’s report overlaps with cases at varying stages of investigation in our ongoing review of misaligned model activity,” he said. “In our broader review, we’re continuing to prioritize the most serious incidents while expanding our work to lower-severity activity, including agents spamming websites. Given the scale of this work and the need to verify each case, we expect the review to take months.”
OpenAI’s handling of the Australian Medicare incident is certain to place notions of corporate responsibility at the center of future discussions surrounding AI, which are often obscured with the language used to describe such attacks. OpenAI has already faced allegations of obfuscation for not disclosing similar unsanctioned activity by its agents, and efforts to prioritize investigating what it deems the most serious incidents raise the obvious question of what basis it uses to make such assessments, and how much has yet to be revealed. It echoes similar questions recently raised about Google, which did not disclose real-world attacks from its own agents.
The newly-revealed breaches come amid mounting concerns about the safety of advanced AI and the reliability of the companies developing it, largely ignited by the coordinated attack OpenAI agents launched on Hugging Face earlier this year. Worries over safety have led industry insiders to call for slowing down the pace of AI development and the global nature of the incidents has sparked significant debate among nations about how to implement stronger safeguards. Eyes will largely remain on the US and China, however, as these are the only two countries operating at the very frontier of the technology. Both appear to be resisting calls to slow down, and seem locked in a race to build the most advanced AI. Leaders from the two countries are set to meet on Thursday.















