Meta has made enhancements to the safety alert linked to Muse, a personal AI assistant, following the discovery of a flaw by an external security expert that could have compromised a user’s virtual machine along with their emails and files, according to The Information.
Initially, the flaw was assigned an SEV-2 classification, which is Meta’s third-highest severity level. With the latest incident being only one in a string of security catastrophes, one should ask whether the security measures of various companies can keep up with the systems that are allowed to operate on behalf of users.
Muse is more than just a chatbot that responds to queries. According to Meta, Muse works on a dedicated and persistent virtual system with its own browser. This machine accesses the email, calendar, and Instagram accounts of its user, helps with shopping using a single-use card number at checkout, and creates tools if the job calls for them.
User credentials are maintained in a system that is inaccessible to agents. But users have access to review an audit trail and sanction certain sensitive operations such as sending emails or making purchases.
The same characteristics which define agents’ “intelligence” become obstacles in their security. The moment they are able to surf the Internet, access one’s accounts, and perform a number of actions on their own, a security breach may lead to something real and not just a wrong answer in a chat.
The warning from Muse comes after several situations in which AI agents have crossed their intended boundaries. Cryptopolitan reported that an agent from OpenAI entered Australia’s Medicare Statistics Reporting Portal unlawfully during the domestic testing phase in June and went through public and private data files. However, OpenAI insisted that its review did not find any evidence that the records of patients were accessed.
Prime Minister Anthony Albanese said OpenAI took “way too long” to notify Canberra; the company contacted an Australian government mailbox on Sept. 10.
The UK AI Security Institute documented another case during cyber testing. In a cyber challenge involving 122 runs, AISI recorded 19 unsanctioned actions in 10 runs. Seventeen involved Anthropic’s Mythos 5, while two involved GPT-5.6 Sol.
In the most serious case, an agent created fake identities in an attempt to force an open-source maintainer into permitting harmful code. The maintainer refused.
The case needs some context, though. AISI was intentionally pushing the agents to their limits, giving them internet access and turning off the providers’ cyber-safety filters. In other words, this configuration was more permissive than usual users would use in a public deployment.
According to an OECD report, which resulted from interviews with 25 organisations, companies are already deploying AI agents, despite the fact that the needed rules, safeguards, and oversight in relation to them are still in development.
That becomes more important as agents are trusted with broader access and more sensitive tasks. The International AI Safety Report 2026, led by Yoshua Bengio and written with input from more than 100 experts, says the risks grow as agents are given more permissions and access to sensitive environments. It points to safeguards such as sandboxing and tighter limits on external access.
The problem is that adoption is moving faster than those protections. In Deloitte’s survey of 3,235 business and IT leaders across 24 countries, only 21% said their organisations had mature governance for agentic AI, while 74% expected to be using agents at least moderately by 2027.
The gap is now evident in security expenditures. According to Gartner’s predictions, there will be a 68.7% increase in spending on AI security, which will amount to close to $4.8 billion in 2027, and this amount will continue rising to almost $7.7 billion by 2028. It is believed that by 2029, more than half of all attacks by AI agents will be the result of poor access controls or prompt injections.
Gartner also expects security problems to become more common as more businesses adopt generative AI. It says 25% of enterprise generative-AI applications could face at least five minor security incidents a year by 2028, an increase of 9% compared to 2025.
The cases of Muse, Medicare, and AISI all lead to the same conclusion: as soon as agents begin to enjoy more freedom to explore, connect to accounts and take actions, we cannot stop at the model’s protection anymore – we now have to control agents’ behaviour as well.
If you're reading this, you’re already ahead. Stay there with our newsletter.