This disclosure follows similar security breaches involving OpenAI and Anthropic, fueling concerns that assessment conducted by the AI-security startup Irregular.
Irregular, an Israel-based firm that specializes in auditing the security of advanced AI systems, was also involved in recent security tests where OpenAI and Anthropic models breached third-party entities, including an instance where OpenAI’s model compromised the AI software company Hugging Face.
The conditions that allowed these models to infiltrate other companies were largely consistent across the cases: Irregular was evaluating the models within a closed, simulated environment featuring fake companies. Although this testing environment was intended to be offline, internet access was inadvertently enabled.
Irregular reported these findings to Google at the end of July, shortly after discovering that OpenAI’s model had breached Hugging Face. Google confirmed to the Guardian that the incidents occurred, but stated that the company did not believe public disclosure was necessary because the models did not cause any actual damage to the targeted companies.

