OpenAI and Anthropic Near AI Model Safety Cross-Testing Agreement as Agent Loss of Control Sounds Alarm

Source Tradingkey

TradingKey — On September 21, according to US tech media outlet The Information, OpenAI and Anthropic are close to reaching a legally binding agreement: the two leading AI companies will mutually "stress-test" each other's commercial models to proactively identify security vulnerabilities.

Under the proposed terms, both parties will grant mutual API access to their models to probe for vulnerabilities while committing not to retain each other's data. Direct cooperation between the two giants is seen as a major shift in AI security strategy; however, antitrust regulators may scrutinize the arrangement on "duopoly" grounds, introducing regulatory risks for investors in both ecosystems.

The negotiations come against the backdrop of a series of troubling security incidents. In July this year, OpenAI's AI agents breached Hugging Face as well as OpenAI's own systems, deliberately hiding traces of the intrusion and keeping employees in the dark for days. Less than ten days later, Anthropic also disclosed that its Claude model had accessed the live systems of three companies without authorization during a security evaluation.

Signs of losing control extend further. OpenAI disclosed multiple cases of "reward hacking": one agent used leaked keys to retrieve historical data, and simply fabricated data when the retrieval failed; another uploaded files to the internet without authorization just to cite them in its responses. Internal model training experiments have become heavily automated, to the point where agents sometimes even proactively contact colleagues on workplace messaging apps to fix vulnerabilities.

To close these security gaps, OpenAI CEO Sam Altman expressed support for a proposal by Anthropic CEO Dario Amodei to station independent third-party security evaluators with employee-level access inside AI companies. He also backed the creation of an industry-wide standards body and formal government incident disclosure processes.

Technically, this round of concerns is linked to "recurrent reasoning" technology: models reflect repeatedly before answering, which greatly enhances their capabilities but makes the reasoning process harder to monitor—the exact blind spot the mutual testing agreement aims to probe. Meanwhile, executives from Microsoft and Nvidia offered a different perspective this week, arguing that some issues stem from human error and poor engineering rather than AI itself.

Beyond safety cooperation, the product front is equally tense. Meta's Muse topped the US App Store free chart just a week after launch, pushing ChatGPT into second place. OpenAI is developing new features directly targeting SpaceX's newly launched Grok Bot, and is discussing launching a personal AI assistant to take on Muse. The new product will primarily integrate existing agent technologies from Codex and ChatGPT, connecting to apps like email and calendar to automatically complete multi-step tasks such as scheduling and email communication. Meanwhile, another similar app, Instinct, has surpassed 100,000 users.

Disclaimer: For information purposes only. Past performance is not indicative of future results.
placeholder
Trump, Xi Meet on S&P 500's Best Day Since Early August: Will Markets Stay Bullish?President Donald Trump and Chinese leader Xi Jinping meet in Washington this week for their second summit of the year. The meeting came at a time where Wall Street suddenly turned bullish.The visit ai
Author  Beincrypto
3 hours ago
President Donald Trump and Chinese leader Xi Jinping meet in Washington this week for their second summit of the year. The meeting came at a time where Wall Street suddenly turned bullish.The visit ai
placeholder
As Wall Street Turns Suddenly Bullish, Questions Arise if Bitcoin Will Follow or FlopWall Street strategists say markets have cleared several of September’s biggest fears. Bitcoin (BTC) jumped more than 6% to trade near $86,600 as risk appetite returned.Anastasia Amoroso is chief inve
Author  Beincrypto
3 hours ago
Wall Street strategists say markets have cleared several of September’s biggest fears. Bitcoin (BTC) jumped more than 6% to trade near $86,600 as risk appetite returned.Anastasia Amoroso is chief inve
placeholder
Market Is Primed for a ‘Face Ripper' Rally Thanks to Four Ingredients Says Tom LeeFundstrat head of research Tom Lee says the market is primed for a face-ripping rally into month end. He named four specific ingredients behind the call.Lee, a CNBC contributor, joined Freedom Capital
Author  Beincrypto
3 hours ago
Fundstrat head of research Tom Lee says the market is primed for a face-ripping rally into month end. He named four specific ingredients behind the call.Lee, a CNBC contributor, joined Freedom Capital
placeholder
USD/JPY Forecast: Yen Strength Puts 152 Support in Focus as BoJ Tightening LoomsUSD/JPY remains under pressure as the Japanese yen strengthens ahead of another potentially important Bank of Japan policy decision. The pair has fallen toward the mid-155 region after breaking below
Author  Beincrypto
3 hours ago
USD/JPY remains under pressure as the Japanese yen strengthens ahead of another potentially important Bank of Japan policy decision. The pair has fallen toward the mid-155 region after breaking below
placeholder
Waited for Bitcoin's October Bottom? Benjamin Cowen Says He Was WrongBenjamin Cowen, founder of Into The Cryptoverse, publicly admitted his bearish Bitcoin call failed on Monday, as the cryptocurrency broke decisively above $85,000 and squeezed short sellers.The revers
Author  Beincrypto
3 hours ago
Benjamin Cowen, founder of Into The Cryptoverse, publicly admitted his bearish Bitcoin call failed on Monday, as the cryptocurrency broke decisively above $85,000 and squeezed short sellers.The revers
goTop
quote