OpenAI and Anthropic reportedly held talks earlier this year about creating a legally binding arrangement under which the two fierce competitors would stress-test each other’s commercially available artificial-intelligence models for security vulnerabilities, dangerous behavior, and hidden safety problems. The proposed arrangement would give each company API access to the other’s models while barring retention of the rival’s testing data. The discussions represent a significant extension of earlier cooperation: in 2025, the companies conducted reciprocal safety evaluations of publicly released models, testing areas including jailbreak resistance, hallucinations, scheming, sycophancy, misuse, and resistance to oversight. The new proposal comes amid broader concern over increasingly autonomous AI systems and a growing push for independent evaluation rather than relying exclusively on developers to certify the safety of their own technology.
Key Takeaways
- OpenAI and Anthropic reportedly discussed a legally binding agreement allowing each company to probe the other’s commercial AI models for vulnerabilities and unexpected or dangerous behaviors, while protecting proprietary testing data.
- The idea is not starting from scratch: the companies conducted reciprocal model evaluations in 2025, demonstrating that rival laboratories can apply different testing methodologies and potentially uncover weaknesses that internal evaluations miss.
- The discussions are part of a larger shift toward outside scrutiny of frontier AI, with Anthropic backing embedded independent evaluators and OpenAI supporting expanded external testing and reporting of concerning model behavior.
In-Depth
OpenAI and Anthropic reportedly explored a legally binding agreement allowing each company to stress-test the other’s commercially available AI models for hidden vulnerabilities and dangerous behavior. The arrangement would reportedly provide API access while prohibiting either company from retaining the other’s testing data. That is notable because the firms are fierce competitors, yet both increasingly acknowledge that frontier AI creates safety problems not addressed solely through internal assurances.
The proposal builds on precedent. In 2025, the companies conducted a joint evaluation in which each laboratory applied its safety tests to the other’s public models. Those examinations covered jailbreak resistance, hallucinations, scheming, self-preservation, sycophancy and potential misuse. The results demonstrated an important principle: competing laboratories can expose weaknesses a model’s creator may overlook or evaluate differently.
The larger question is whether voluntary industry cooperation can keep pace with increasingly capable systems. Anthropic has pushed for greater access for independent evaluators, while OpenAI has supported expanded outside testing and systematic reporting of troubling model behavior. Anthropic also recently announced a multibillion-dollar effort with Accenture to expand independent frontier-model evaluation.
Cross-testing could provide a market-driven layer of accountability without immediately handing technical oversight to government bureaucracies that may lack comparable expertise. But competitors policing competitors is not the same as independent scrutiny. A durable system needs clear testing standards, protection of proprietary information, credible disclosure rules and outside verification. The encouraging development is that leading AI companies appear to recognize safety claims are stronger when someone else can independently challenge them effectively first.

