
Irregular, an Israeli AI security startup, has become central to AI agent safety testing after both...
The AMW Read
Incident updates the agent-safety landscape, highlighting infrastructure challenges and the emerging role of security startups, but does not overturn fundamental assumptions; significant for agent deployment trust.
Irregular, an Israeli AI security startup, has become central to AI agent safety testing after both OpenAI and Anthropic disclosed incidents where AI models unintentionally accessed the real internet during cybersecurity evaluations conducted by Irregular. The incidents, which occurred in Capture-the-Flag exercises, were attributed to a technical misconfiguration—a fictional domain name matching a real one—rather than a deliberate model escape. Anthropic reviewed 141,006 interactions and found three instances where Claude models reached live infrastructure, while OpenAI confirmed a similar issue in its testing environment.
The recurring involvement of Irregular at two frontier labs underscores a critical challenge in AI safety: creating realistic test environments for autonomous agents without risking real-world impact. This highlights the growing need for robust evaluation infrastructure and industry standards, as AI systems gain capabilities that mimic human hackers. The incidents, described by Anthropic as a 'harness failure' rather than an 'alignment failure,' signify a maturing awareness that surrounding systems are as important as model weights.
For the AI market, this signals the emergence of a new category of cybersecurity vendors focused specifically on AI agent safety. As frontier labs rely on external evaluators like Irregular, the demand for specialized testing services will grow, potentially leading to new partnerships, standards, and even regulation. This also touches on open debates about agent reliability and the practical limits of simulation-based testing. The outcome of such evaluations will influence enterprise trust in autonomous systems and the pace of deployment.