
Anthropic's Claude in Chrome browser agent reaches general availability with new prompt-injection defenses
The AMW Read
GA launch of a known Anthropic product with newly published, quantified red-team attack-success metrics across model generations meaningfully updates the agentic-browser safety baseline without resolving an open debate.
Anthropic's Claude in Chrome browser agent reaches general availability with new prompt-injection defenses
Anthropic made Claude in Chrome, its AI browser extension, generally available on August 26, 2026, to all paid Claude plans. The extension lets Claude read the page a user has open, type into fields, click links, navigate between pages, and fill out forms β all inside the user's existing logged-in session, so it can act on internal dashboards, legacy systems, and other portal sites that were never built with API integrations. Anthropic paired the launch with a disclosure of the defenses it built against prompt injection, where malicious instructions hidden in a web page or email hijack the agent's actions: internal red-team testing in 2025 found a 23.6% attack success rate with no defenses active.
The company described a three-layer defense built over roughly a year: continuous model training against an attack library drawn from internal red teams, external testers, and monitoring classifiers; a probe that flags likely-injected content and has Claude check with the user before acting on it; and a classifier that blocks actions inconsistent with the original request, gating higher-risk steps behind user confirmation by default. In Anthropic's latest evaluation, combined defenses hold attack success below 0.3%, down from 17.6% for Claude Opus 4.5 and 3.8% for Claude Opus 5 with no defenses; Claude Sonnet 5, Opus 5, and Mythos 5 saw zero successful attacks with the probe and classifier active, and Fable 5 saw 0.3%, all rated low-severity. Per the AI Market Watch index, Anthropic logged 296 tracked news items in the past 90 days versus 210 in the prior 90 (coverage is name-matched over pipeline-ingested sources, not a census), and the volume tracks a week in which Anthropic has repeatedly published quantified safety metrics, including the Mythos 5 vulnerability-scanning restriction and an automated alignment-research paper.
For builders integrating Claude in Chrome β available via the Chrome Web Store, with enterprise admins able to restrict it to approved domains, though it still can't touch local files or other apps and doesn't cover Chromium forks or mobile β the published attack-success table is now a usable benchmark rather than a marketing claim. Enterprises evaluating any browser- or computer-use agent should ask vendors for the same disclosure before granting standing access to internal tools, and investors should treat probe-and-classifier defense stacks as an emerging technical layer rather than a compliance line item.
#Anthropic #ClaudeInChrome #AIAgents #PromptInjection #AISafety #BrowserAutomation

