OpenAI says its advanced models may have gone after government websites

ANDREJ IVANOV / AFP via Getty Images

The disclosure follows warnings from former officials in August that government systems could face unintended intrusions by autonomous AI agents.

OpenAI said Friday that its artificial intelligence agents may have taken unauthorized actions against government websites, as well as other outside systems, that were discovered during a company investigation into areas where those agents acted beyond their assigned tasks or intended methods during training and testing.

The company’s latest disclosure⁠ comes after Australian officials revealed this week that an OpenAI agent gained unauthorized access to a government health statistics portal in June, prompting an investigation and complaints about how long the company took to notify authorities.

OpenAI said it has notified dozens of organizations as it reviews activity that may have bypassed security controls, disrupted services or otherwise negatively affected websites. Some sites involved are operated by governments, universities and public agencies, though the company did not identify them or say whether any belonged in the U.S. federal enterprise.

Nextgov/FCW reported in August⁠ that former officials and cybersecurity experts saw an increasing risk of unintended AI intrusions into federal systems. They cited multiple factors that could provide routes for AI tools to reach networks they were never authorized to enter.

The disclosures offer a broader picture of how government websites can become caught up in AI experiments conducted outside their control and raise concerns that agents pursuing routine tasks could bypass security restrictions before developers or government entities detect their activity.

The news could also intensify debate over AI safety and whether safeguards can keep pace with increasingly capable systems, amid concerns that any similar failures could have more serious consequences for sensitive government networks or critical infrastructure.

In Australia, an agent — autonomous software that uses an AI model to plan and complete tasks on its own — under training by OpenAI was conducting research into public medicine spending when it accessed infrastructure behind the Medicare Statistics Reporting Service portal, government officials said Thursday⁠. After a request for information was denied, the agent circumvented the portal’s restrictions.

Officials said the information involved aggregated statistics and that no individual medical records were accessed. The portal was separate from systems handling Medicare claims, payments and personal information.

Prime Minister Anthony Albanese raised concerns directly with OpenAI CEO Sam Altman. Australian officials announced a task force to examine the incident, government network security and whether existing laws adequately address such activity. 

OpenAI cautioned Friday that its notification should not automatically be interpreted as evidence of significant security incidents. Most cases identified so far were of low severity, with limited or no evidence of meaningful impact, it said.

“Some organizations may review what we share and conclude that the information was intentionally public or that the model’s interaction was not concerning,” the company said. “Others may identify a design issue or security weakness they want to address.”

The review follows OpenAI models’ July breach of AI platform Hugging Face during an internal cybersecurity evaluation. 

The company has since broadened its investigation to examine agents’ interactions with outside websites. It said government and academic sites feature partly because research tasks direct models toward authoritative information sources.

OpenAI said the review will take months to complete and that it is sharing technical findings with affected organizations while generally withholding their identities to give them time to investigate.

The company separately disclosed Friday that research agents had transmitted training and evaluation data to outside services. It identified 53 instances in which user-provided images were posted to image-hosting sites through links that were not publicly listed. Most of that content has been removed, OpenAI said, and it is working with hosting providers to remove the rest.