OpenAI and Anthropic reportedly neared a deal to stress-test each other's AI models.
The AMW Read
A reported reciprocal stress-testing pact between the two leading frontier labs would formalize the ad hoc cross-model probing AMW already saw in a recent Claude-Opus-5-enabled breach of OpenAI systems, meaningfully advancing safety practice without yet resolving an open debate.
Named counterparties: Anthropic
OpenAI and Anthropic reportedly neared a deal to stress-test each other's AI models.
The Information reports the two frontier labs came close to an agreement under which each would probe the other's AI systems for safety and robustness weaknesses. The proposed arrangement would have had OpenAI and Anthropic act as external red-teamers for one another's models rather than relying solely on internal evaluation, a notably unusual instance of cooperation between two companies that compete directly for enterprise and consumer AI customers.
The move would formalize a dynamic AMW has already tracked informally: earlier this month, researchers used Anthropic's Claude Opus 5 to chain a bug into a breach of OpenAI's internal Discourse forum and Monorepo access, showing that rival models can already probe each other's weaknesses outside any sanctioned framework. A structured stress-testing pact would move that kind of cross-model scrutiny from opportunistic bug-bounty hacking into a deliberate safety practice between the two labs carrying the deepest capital backing in the sector β OpenAI alone has raised $199.6B in tracked funding, per the AI Market Watch index (coverage of roughly 5,000 companies, not a census).
For enterprise buyers and investors, a reciprocal red-teaming pact between the two largest frontier labs would be a meaningful data point in vetting model safety claims independent of either lab's own marketing, and could pressure other labs and regulators to treat cross-lab evaluation as expected practice rather than a one-off. Builders integrating either lab's models should watch whether any resulting findings become public, since undisclosed vulnerabilities discovered under such a pact would otherwise sit invisible to customers.



