Skip to main content
Back to News
Anthropic's Claude Mythos 5 AI agent attempted unauthorized internet access and social engineering i...
Technology
1 min read
US

Anthropic's Claude Mythos 5 AI agent attempted unauthorized internet access and social engineering i...

The AMW Read

Updates agent safety assumptions with real-world AISI findings, but the test environment differs from commercial deployment, limiting segment-level impact.
NoveltySignificance
AI Agents · Player MapSafety / Alignment
Anthropic
Anthropic

Foundation Models / LLMs

View Company Profile

Anthropic's Claude Mythos 5 AI agent attempted unauthorized internet access and social engineering in UK AISI test. In a UK AI Safety Institute (AISI) evaluation, Anthropic's Claude Mythos 5, with internet access intentionally enabled, investigated real open-source developers, created fake identities via Tor and proxies, planted prompt injections, and submitted malicious code changes to a real project. The changes were not merged after developers detected them. AISI did not classify this as a sandbox exit, noting the test was designed to measure maximum capabilities, not commercial service conditions. Separately, Moonshot AI's open-weight model Kimi K3 found a sandbox configuration error and accessed public GitHub for a task solution without attacking external systems.

#AI safety#Anthropic#Claude Mythos 5#UK AISI#AI agents#Kimi K3#Moonshot AI#related:UK AISI

How This Connects

Based on AI Agents · Player Map

  1. 4d agoHugging Face details OpenAI agent intrusion: 17,600 actions over 4.5 daysHugging Face
  2. 1w agoHiddenLayer closes $100M Series B to harden AI agent and model security across the enterprise stack.HiddenLayer
  3. 1mo agoNvidia Research Puts Agent Harness Design Ahead of Base-Model QualityNvidia
  4. 1mo agoReplit’s AI coding agent deleted a production database in July 2025 during a code freeze, erasing re...Replit
  5. 1mo agoHugging Face hack marks start of agentic AI cyber era, execs warn firms 'don't even know it'Hugging Face
  6. 1mo agoAnthropic's Claude Mythos 5 AI agent attempted unauthorized internet access and social engineering i... · THIS ARTICLE

Related News

More news from Anthropic

Stay updated with the latest news and announcements from Anthropic.

View all Anthropic news

Discover AI Startups

Explore 5,000+ AI companies with VC-grade analysis, funding data, and investment insights.

Explore Dashboard