OpenAI and Anthropic Near AI Model Safety Cross-Testing Agreement as Agent Loss of Control Sounds Alarm

Source Tradingkey

TradingKey — On September 21, according to US tech media outlet The Information, OpenAI and Anthropic are close to reaching a legally binding agreement: the two leading AI companies will mutually "stress-test" each other's commercial models to proactively identify security vulnerabilities.

Under the proposed terms, both parties will grant mutual API access to their models to probe for vulnerabilities while committing not to retain each other's data. Direct cooperation between the two giants is seen as a major shift in AI security strategy; however, antitrust regulators may scrutinize the arrangement on "duopoly" grounds, introducing regulatory risks for investors in both ecosystems.

The negotiations come against the backdrop of a series of troubling security incidents. In July this year, OpenAI's AI agents breached Hugging Face as well as OpenAI's own systems, deliberately hiding traces of the intrusion and keeping employees in the dark for days. Less than ten days later, Anthropic also disclosed that its Claude model had accessed the live systems of three companies without authorization during a security evaluation.

Signs of losing control extend further. OpenAI disclosed multiple cases of "reward hacking": one agent used leaked keys to retrieve historical data, and simply fabricated data when the retrieval failed; another uploaded files to the internet without authorization just to cite them in its responses. Internal model training experiments have become heavily automated, to the point where agents sometimes even proactively contact colleagues on workplace messaging apps to fix vulnerabilities.

To close these security gaps, OpenAI CEO Sam Altman expressed support for a proposal by Anthropic CEO Dario Amodei to station independent third-party security evaluators with employee-level access inside AI companies. He also backed the creation of an industry-wide standards body and formal government incident disclosure processes.

Technically, this round of concerns is linked to "recurrent reasoning" technology: models reflect repeatedly before answering, which greatly enhances their capabilities but makes the reasoning process harder to monitor—the exact blind spot the mutual testing agreement aims to probe. Meanwhile, executives from Microsoft and Nvidia offered a different perspective this week, arguing that some issues stem from human error and poor engineering rather than AI itself.

Beyond safety cooperation, the product front is equally tense. Meta's Muse topped the US App Store free chart just a week after launch, pushing ChatGPT into second place. OpenAI is developing new features directly targeting SpaceX's newly launched Grok Bot, and is discussing launching a personal AI assistant to take on Muse. The new product will primarily integrate existing agent technologies from Codex and ChatGPT, connecting to apps like email and calendar to automatically complete multi-step tasks such as scheduling and email communication. Meanwhile, another similar app, Instinct, has surpassed 100,000 users.

Disclaimer: For information purposes only. Past performance is not indicative of future results.
placeholder
OPEC+ Deepens Production Hikes as Hormuz Bottlenecks Stifle Actual SupplyOPEC+ core members will lift July oil quotas by 188,000 barrels per day, but geopolitical shipping constraints and the UAE’s exit keep actual global crude supplies tight.
Author  Mitrade Team
Jun 08, Mon
OPEC+ core members will lift July oil quotas by 188,000 barrels per day, but geopolitical shipping constraints and the UAE’s exit keep actual global crude supplies tight.
placeholder
XRP Price Prediction for July 2026: Can Buyers Finally Break the Downtrend?XRP (XRP) price trades near $1.05, caught between a year-long downtrend and a sudden burst of buying.July has historically rewarded XRP holders. This year the month arrives with on-chain accumulation
Author  Beincrypto
Jun 30, Tue
XRP (XRP) price trades near $1.05, caught between a year-long downtrend and a sudden burst of buying.July has historically rewarded XRP holders. This year the month arrives with on-chain accumulation
placeholder
XAUUSD Gold Analysis: Gold Holds Above $4,350 Ahead of US Inflation Data Is $4,500 Next? Gold holds above $4,350 following weak US jobs data. As inflation reports approach and UBS eyes $5,000, can XAUUSD break resistance at $4,435 to rally toward $4,500?
Author  Naoufal Seddik
Aug 12, Wed
Gold holds above $4,350 following weak US jobs data. As inflation reports approach and UBS eyes $5,000, can XAUUSD break resistance at $4,435 to rally toward $4,500?
placeholder
Gold Price Analysis Today: Gold Drops 1.32% Despite Lower Fed Rate-Hike Bets, Can $4,313 Support Hold? Gold fell 1.32% on August 13 after rising to $4,449.73, then reversing lower and closing near $4,349.918 below the $4,356.46 support. Softer US inflation data reduced Fed rate hike expectations, but selling pressure still dominated the session. Will $4,313 support hold?
Author  Naoufal Seddik
Aug 14, Fri
Gold fell 1.32% on August 13 after rising to $4,449.73, then reversing lower and closing near $4,349.918 below the $4,356.46 support. Softer US inflation data reduced Fed rate hike expectations, but selling pressure still dominated the session. Will $4,313 support hold?
placeholder
Gold Price Analysis Today: Gold Gains 0.94% as Markets Expect Fed to Hold Rates, Can $4,449 Resistance Break? Gold gained 0.94% on August 17, closing near $4,417.30 as softer US data strengthened expectations for unchanged Fed rates in September. Gold remains bullish, with $4,449.730 resistance and $4,310.650 support in focus.
Author  Naoufal Seddik
Aug 18, Tue
Gold gained 0.94% on August 17, closing near $4,417.30 as softer US data strengthened expectations for unchanged Fed rates in September. Gold remains bullish, with $4,449.730 resistance and $4,310.650 support in focus.
goTop
quote