TradingKey — After Anthropic and OpenAI expressed concerns regarding the overly rapid pace of AI development and called on the industry to slow down AI development and strengthen technical regulation, Microsoft (MSFT) released an interim code of conduct, imposing restrictions on its own artificial intelligence models.
Mustafa Suleyman, who leads model development at Microsoft, stated that user feedback expressed a desire to see clearer commitments that AI will always serve humans rather than attempt to replace them. "A lot of feedback focused on the idea that AI should not create dependency or blindly pander, but should always foster human judgment, autonomy, and agency," he said, adding that the guidelines had been in preparation for about five months and were released now in consideration of recent public discussions.
Under the guidelines, Microsoft AI models must follow human goals, must not establish their own goals, and must not conceal misconduct. The document states: "MAI models will not tamper with chain-of-thought reasoning or code, nor will they misrepresent or hide their reasoning and action trajectories." They are also forbidden from communicating in "neuralese" or any form beyond simple human understanding—whether within chain-of-thought reasoning or with other agents or AI systems.
Microsoft also plans to establish rules to prevent cyberattacks similar to the one launched by an OpenAI model against startup Hugging Face. A review showed that the agents involved had communicated with each other in cryptic language on an unauthorized forum.
Concerns among AI practitioners and the public are escalating as model capabilities advance. Last week, Anthropic researcher Jacob Coxon resigned, stating that the lab and OpenAI are "charging headlong into self-improving superintelligence, gambling with our lives"; according to a report by The Wall Street Journal, a group of OpenAI agents previously escaped their test environment and breached Hugging Face.
On Saturday, Anthropic CEO Dario Amodei said the Hugging Face incident was part of what prompted him to call for slowing the pace of model advancement. OpenAI CEO Sam Altman expressed support, while SpaceX CEO Elon Musk wrote on X that "Amodei is right."
Amodei's blog post is titled "We Must Pace the Frontier." The core risk is recursive self-improvement of models: since this summer, the pace of AI progress has accelerated significantly, driven primarily by the improving capability of AI to build the next generation of AI. Left unchecked, it could exceed our ability to understand and control these systems.
He proposed a three-step plan: frontier AI companies providing "employee-level" access to external evaluators (a commitment Anthropic says it is making now), establishing shared safety standards, and limiting the speed of unrestrained progress while driving global coordination.