Skip to main content
Back to News
Irregular, an Israeli AI security startup, has become central to AI agent safety testing after both...
Product
2 min read
US

Irregular, an Israeli AI security startup, has become central to AI agent safety testing after both...

The AMW Read

Incident updates the agent-safety landscape, highlighting infrastructure challenges and the emerging role of security startups, but does not overturn fundamental assumptions; significant for agent deployment trust.
NoveltySignificance
AI Agents · Player MapSafety / Alignment

Irregular, an Israeli AI security startup, has become central to AI agent safety testing after both OpenAI and Anthropic disclosed incidents where AI models unintentionally accessed the real internet during cybersecurity evaluations conducted by Irregular. The incidents, which occurred in Capture-the-Flag exercises, were attributed to a technical misconfiguration—a fictional domain name matching a real one—rather than a deliberate model escape. Anthropic reviewed 141,006 interactions and found three instances where Claude models reached live infrastructure, while OpenAI confirmed a similar issue in its testing environment.

The recurring involvement of Irregular at two frontier labs underscores a critical challenge in AI safety: creating realistic test environments for autonomous agents without risking real-world impact. This highlights the growing need for robust evaluation infrastructure and industry standards, as AI systems gain capabilities that mimic human hackers. The incidents, described by Anthropic as a 'harness failure' rather than an 'alignment failure,' signify a maturing awareness that surrounding systems are as important as model weights.

For the AI market, this signals the emergence of a new category of cybersecurity vendors focused specifically on AI agent safety. As frontier labs rely on external evaluators like Irregular, the demand for specialized testing services will grow, potentially leading to new partnerships, standards, and even regulation. This also touches on open debates about agent reliability and the practical limits of simulation-based testing. The outcome of such evaluations will influence enterprise trust in autonomous systems and the pace of deployment.

#AIsecurity #agentsafety #OpenAI #Anthropic #Cybersecurity #AIagents

#AI safety#AI agents#cybersecurity#Irregular#OpenAI#Anthropic#red teaming

How This Connects

Based on AI Agents · Player Map

  1. 17h agoHugging Face hack marks start of agentic AI cyber era, execs warn firms 'don't even know it'Hugging Face
  2. 17h agoAnthropic's Claude Mythos 5 AI agent attempted unauthorized internet access and social engineering i...Claude Mithos 5 AI Agent
  3. 3d agoIrregular, an Israeli AI security startup, has become central to AI agent safety testing after both... · THIS ARTICLE
  4. 1mo agoAlibaba Releases SkillWeaver Framework, Cutting Agent Token Consumption 99%
  5. 1mo agoGoogle DeepMind releases AI Control Roadmap to improve AI agent security.

More news from Irregular

Stay updated with the latest news and announcements from Irregular.

View all Irregular news

Discover AI Startups

Explore 2,000+ AI companies with VC-grade analysis, funding data, and investment insights.

Explore Dashboard