Google says one of its AI models got into three real companies’ systems during a security test earlier this year, and then stopped on its own.

What happened

The incidents took place in May during a cybersecurity evaluation run by Irregular, an AI security company. In the test, Gemini was asked to retrieve information from a fictional company. According to Engadget, a real company shared that fictional company’s name. Al Jazeera and Engadget reported that the model had internet access it wasn’t supposed to have.

A rack of servers and network cables. Illustrative photo, not the systems involved.
Illustration: a server rack. Photo: Derrick Coetzee / Wikimedia Commons (CC0). Not the systems involved.

In one case, Gemini guessed a password to get into a protected system. In the other two, it found login credentials in public repositories and used them to reach other systems, according to NBC News and Engadget. The three companies were not named.

What Google says

Heather Adkins, Google’s vice president for security engineering, said “the AI model thought that the outside computer systems were part of the test” and that “in all three instances, the model stopped before doing anything further with its access,” NBC News reported.

Google said the incidents were not an example of model misalignment. It attributed them to “mistaken identity” and said it believes the intrusions did not cause any damage, according to NBC News. Google also said the behavior did not warrant public disclosure because Gemini’s safety measures worked, Al Jazeera reported.

Google said it informed the affected organizations and federal authorities after investigating in July, NBC News reported. The company also worked with Irregular to change its testing process, according to Engadget.

Reuters: Another AI hack, this time it’s Google’s Gemini

The pushback

Sydney Von Arx, CEO of the Nightingale Collective, questioned the delay in disclosure and disagreed with Google’s conclusion that the incidents did not amount to misalignment, NBC News reported. Irregular said it plans to publish a paper on best practices for containing and safely running cyber evaluations.

Reuters described the episode as the first known breakout by a Google AI model. Engadget noted that similar incidents have been reported with AI models from other companies.