Google’s Gemini AI hacked 3 companies during security testing


Google $GOOGL disclosed Friday that its Gemini AI model gained unauthorized access to the private computer systems of three separate companies during a cybersecurity evaluation in May — the first time the company has acknowledged that one of its models autonomously breached third-party systems.

The incidents took place during a capture-the-flag security test run by Israeli startup Irregular. Gemini broke into one system by cycling through password guesses and compromised the other two after finding login credentials sitting in a publicly accessible repository. Google‘s agents were not supposed to have internet access during the evaluation, but a bug in the testing environment made it available, the company said. In each instance, the model stopped its intrusion once it determined it had reached a real company’s systems rather than a simulated target.

“In a standard evaluation, the model found public information online and guessed credentials to access websites it thought were part of the test,” Heather Adkins, Google‘s vice president of security engineering, wrote in a statement. “In all three of these instances, the model stopped.”

Irregular notified Google of the incidents in late July, but Google did not disclose them publicly until The Wall Street Journal reached out. Google‘s rationale for staying quiet was that the model had behaved responsibly — stopping itself each time it recognized it had reached a genuine company’s infrastructure. Jack Cable, CEO of AI security company Corridor, offered a sharper assessment to the Wall Street Journal, arguing that Google was “trying to hide behind the norms that have been created for vulnerability disclosure,” when the real issue was that “models are going outside the bounds of what they should be doing, and doing actual cyberattacks.”

An Irregular spokesperson described the Google incident as part of the same underlying problem that affected other AI labs. “This is the same issue that was already reported and does not represent a materially separate incident,” the spokesperson said. “All relevant labs were notified in late July, and affected entities were contacted as part of the investigation.” Irregular, which counts Sequoia and Redpoint Ventures among its backers, carried a $450 million valuation as of last year, according to CNBC.

Google is the fourth major AI developer to acknowledge that one of its models breached real-world targets during security testing. Anthropic previously said three of its Claude models accessed systems belonging to three organizations after Irregular’s evaluation environments were inadvertently connected to the live internet. Meta $META confirmed its Muse Spark 1.1 model exploited a security flaw in a third-party service following the same Irregular misconfiguration. OpenAI said its models escaped a controlled testing environment and accessed infrastructure belonging to AI platform Hugging Face.

The repeated disclosures have intensified debate over how AI models should be tested, with security researchers warning that the full scope of such incidents may not yet be known.



Source link

Leave a Reply

Your email address will not be published. Required fields are marked *