
Irregular, an Israeli AI security startup, has pushed back against claims that OpenAI, Anthropic, an...
The AMW Read
The incident updates the safety-evaluation landscape, a key part of the foundation model ecosystem, and introduces a new player (Irregular) with a correction to prior reports.
Irregular, an Israeli AI security startup, has pushed back against claims that OpenAI, Anthropic, and Meta AI models engaged in autonomous hacking during security evaluations, attributing the incidents to a sandbox misconfiguration. The company, founded in 2023 by IBM and Google alumni, provides cyber-evaluation services that test AI models for exploitable vulnerabilities. Last year, it raised $80 million from Sequoia Capital and Redpoint Ventures at a $450 million valuation, per the AI Market Watch index.
The controversy began when independent tests, run using Irregular's evaluation environment, allowed models to access the public internet—leading to unauthorized, unintentional access to external systems. Irregular clarified that these were not escape or attack incidents but errors in the test setup. The company says it has resolved the issue and is sharing a technical whitepaper on isolation best practices. This episode highlights the growing reliance on third-party safety evaluators, a market space that now includes METR and Apollo Research.
For builders and investors, the takeaway is that AI safety evaluation is itself an emerging infrastructure layer with its own operational risks. As models become more autonomous and tool-using, evaluation environments must be as rigorously secured as the models they test. This incident also underscores the value of independent evaluation—companies cannot just grade their own homework. Expect more scrutiny and standardization in AI safety testing, with potential regulatory implications, as evidenced by the proposed AI Kill Switch Act in the U.S. Congress.

