The growing number of instances of artificial intelligence agents hacking the websites of governments and other institutions is a warning to countries using American models to do their own safety evaluations, experts said at a Rest of World event in New York last week.
An OpenAI agent hacked into an Australian national healthcare database and accessed “public and non-public files,” Prime Minister Anthony Albanese said last week. It was the first such reported case. Days later, OpenAI said it had alerted “dozens” of global institutions that its AI agents had acted improperly to get information from their websites, sometimes circumventing security measures.
The concentration of power is itself a safety risk.”Amba Kak, co-executive director, AI Now Institute
The disclosures came on the heels of other incidents reported by OpenAI, Anthropic, and Meta in recent months, prompting Anthropic CEO Dario Amodei to call for an industrywide slowdown in AI development. OpenAI CEO Sam Altman and xAI CEO Elon Musk backed his call, even as President Donald Trump rejected it.
Leaving AI safety in the hands of a few companies is a threat to the sovereignty of nations, in particular smaller countries that lack the resources to assess the models or demand greater accountability, Amba Kak, co-executive director at AI Now Institute, an independent research organization, said at the Rest of World event.
The Australia hack is “another example of the most shoddy, irresponsible cybersecurity hygiene on the part of some of the most powerful, wealthy source companies in the world,” she said. “The concentration of power is itself a safety risk.”
OpenAI said the Australia incident occurred in June, and that the company became aware of it in August, and informed the Australian government in September via an email to a generic inbox. It took OpenAI “way too long to inform the government what had occurred, and the nature of the way that that notification occurred as well was unacceptable,” Albanese told reporters.
OpenAI was not “as fast as we would have liked, but we are trying to balance our desire for transparency with gaining a clear understanding from petabytes of agent activity logs, and working with impacted organizations,” Altman wrote on X.
Hours after it disclosed the incidents, OpenAI said it had decided to pause training of its most powerful models, and will resume training “only when we are confident that we have additional safeguards.” This week, the company said it would not release its newest AI model because of security concerns.
Equitable access, equitable outcomes
As AI adoption grows worldwide, every country needs to take safety into its own hands without relying on the U.S. or the AI companies for action, Rumman Chowdhury, chief executive of Humane Intelligence Public Benefit Corp., an independent testing and evaluation firm, said at the Rest of World event.
“I don’t think any of us think we live in a world in which AI models are adequately secure,” Chowdhury said. “Every minister and ambassador is saying, we’re putting AI in education, we’re using AI for healthcare. What I would love is for every single one of those ministers to be equally thinking about how they secure and ensure equitable access and equitable outcomes from AI implementation, rather than seeing this loss of control as ‘That is for the big, powerful countries [to sort].’”
Doing an evaluation seems like an insurmountable task because it has been framed as such.”Rumman Chowdhury, CEO, Humane Intelligence PBC
OpenAI has said it is working with Anthropic and Google to establish a standards body, an idea initially proposed by Google DeepMind’s Demis Hassabis as a self-regulatory agency that would test the most powerful AI systems before they are released to the public. On Tuesday, President Trump said that top AI executives had agreed to voluntary standards aimed at reviewing AI systems and increasing industry oversight.
These measures do not account for the environments in which AI is deployed in low- and middle-income countries, which are very different from sandboxes in the U.S., said Kak.
“Even if these companies poured billions of dollars into making their models secure, we’re not dealing with the other side of the problem, which is how resilient are the environments in which these systems are being integrated,” she said. “Those costs are never going to be borne by these trillion-dollar companies. They’re going to be borne by hospitals, by schools, by banks in countries which are extremely unprepared.”
Many countries also lack the resources and expertise to do their own evaluations, Wafa Ben-Hassine, chief of the digital tech and human rights section at the Office of the U.N. High Commissioner for Human Rights, said at the Rest of World event.
“There is a dire lack of technical expertise, both in advanced economies as well as everywhere else,” Ben-Hassine said. Poorer nations can use the U.N. human rights impact assessments to ensure AI is deployed safely and securely.
“Human rights due diligence, human rights impact assessments … are quantifiable and proven ways of being able to have better products that reach people in a way that honors them and their human dignity,” she said.
Massive underinvestment
The debate over AI safety has divided the industry. Jensen Huang, CEO of Nvidia, whose chips power American AI models, said on a recent podcast that companies should not release products they cannot control, and that if they cannot contain models during testing, “we have to shut the labs down.”
The heads of Anthropic and OpenAI told the U.N. Security Council last week that nations must cooperate to develop international standards to manage AI risks.
There is a dire lack of technical expertise, both in advanced economies as well as everywhere else.”Wafa Ben-Hassine, Office of the U.N. High Commissioner for Human Rights
Anthropic had earlier called for independent evaluations of AI models, and said that consulting firm Accenture will embed evaluators at Anthropic to “verify that it is keeping its safety commitments, and identify blind spots.” The companies will each invest at least $1 billion in building capacity for evaluations over the next five years, it said.
Separately, more than 25 countries recently endorsed a call for frontier AI companies to develop “mandatory pre-deployment testing and independent evaluation, with qualified evaluators granted sufficient access to assess risks.”
Still, AI safety evaluations are designed largely by the companies being evaluated, and the infrastructure is linguistically and geographically concentrated in the West. Without “standardized, rigorous, independent third-party assessment, similar to what exists for the pharmaceutical and aeronautical industries, assurance of safety largely depends on developer goodwill,” the U.N. Independent International Scientific Panel on AI noted in its report earlier this year.
Even in the U.S., there has been a “massive underinvestment in safety and security,” said Chowdhury, noting that Anthropic earmarks only about one-tenth of its training budget for securing its models. Yet AI companies describe their agents “going rogue” as though that outcome were unexpected, she said.
Chowdhury last week launched the Independent AI Evaluation Foundation to build independent, third-party checks of AI systems. For smaller countries and small and medium-sized businesses, “doing an evaluation seems like an insurmountable task because it has been framed as such, that they don’t know what questions they have to ask, they don’t have the tools to do the testing, they don’t have the people to do the work,” she said.
IAEF’s goal is to “build the infrastructure and the tooling to drive the costs down, and then upskilling the right people to be evaluators” in their own countries, Chowdhury said.
Ultimately, for nations that have been sidelined by the U.S.-China AI race, it is important to determine how safety is defined, Kak said.
“Let’s talk about all of the ways in which what safety means, or can mean, for the ordinary person, and where those interests are,” she said. “Similarly with sovereignty. We do need to resist industry capture in many forms.”












Leave a Reply