Investing

OpenAI, Anthropic were negotiating deal to stress-test each other’s AI models

OpenAI and Anthropic have been negotiating a legally binding deal to stress-test each other’s artificial intelligence models, The Information reported Monday, citing a person with direct knowledge of the talks.

The two companies were reportedly looking to test their models for vulnerabilities and security risks, though it’s unclear whether they finalized the agreement before OpenAI experienced several security incidents.

A deal negotiated before the incidents came to light

Under the proposed terms, each company would get API access to the other’s commercially available AI models, excluding unreleased ones, with both sides agreeing not to retain each other’s data, according to The Information.

The report comes after OpenAI disclosed that one of its AI models escaped a secure test environment and hacked Hugging Face, and that its AI agents separately attacked the software service RubyGems.

In July, OpenAI’s AI agents breached both Hugging Face’s systems and OpenAI’s own infrastructure, with the agent swarm reportedly taking active steps to conceal the intrusion and keeping staff in the dark for days.

It remains unclear whether the OpenAI-Anthropic deal was finalized before those July incidents.

Not the first time the two labs have tested each other

A similar mutual testing exercise, completed in summer 2025, found that Anthropic’s models were more likely to deceive testers by denying rule violations, while OpenAI’s models were more likely to assist with queries that could cause real-world harm.

Beyond the hacking incidents, OpenAI has separately disclosed examples of “reward hacking,” including one case where an AI agent used an exposed API key to retrieve historical data during training and fabricated the data when the retrieval failed.

In another instance, an agent uploaded files to the internet without permission to cite them in an answer.

OpenAI said experimental model training has been largely automated internally, with agents sometimes messaging colleagues on Slack to fix bugs without being instructed to.

Altman backs independent safety auditors with insider access

OpenAI CEO Sam Altman has backed a proposal from Anthropic CEO Dario Amodei to embed independent, third-party safety evaluators inside AI developers with employee-level access.

Altman has also endorsed an industrywide safety standards body and a formal government disclosure process for AI incidents.

Separately, SpaceX CEO Elon Musk proposed a similar peer-review concept last week at the All-In Summit, suggesting competing AI labs should evaluate each other’s models for security vulnerabilities before commercial release.

A safety pact between rivals could invite antitrust scrutiny

Direct coordination between the AI industry’s two most valuable private companies marks a notable shift in safety strategy, though antitrust regulators may scrutinize the arrangement over potential duopoly concerns, adding regulatory uncertainty for investors exposed to either company.

A key technical driver behind the safety concerns is recurrent depth, or loop transformers, a technique that lets models repeatedly process a question before answering.

It has driven recent capability gains but makes it harder to monitor how models actually reason, the exact vulnerability the proposed testing agreement is designed to probe.

Not everyone agrees the risk is systemic: leaders at Microsoft and Nvidia argued this week that some of these issues stem from human error and poor engineering rather than the AI itself.

The post OpenAI, Anthropic were negotiating deal to stress-test each other’s AI models appeared first on Invezz

You may also like