- Meta's Muse design limits what a tricked agent can do, but Meta says it can still access user data when needed to operate the service.
- Privacy claims from agent vendors cover two different risks: outside attackers and the vendor itself. Only the first is limited by architecture today, and Meta calls prompt injection an open problem.
- Ask each agent vendor who can read your data, what is used for training by default, and whether a third party can verify the claims.
Two promises hiding in one word
When an AI agent company promises privacy, it can mean two different things. It can mean outsiders cannot use the agent to steal your data. Or it can mean the company itself cannot read it. Meta's launch of its Muse personal agent shows how far apart these promises are.
Meta's own security write-up is candid about this. The first risk is bounded by engineering, which Meta describes in detail. Meta also says prompt injection remains an open problem, and The Verge reports that a takeover flaw has already been patched. The second risk is, for now, a matter of company policy.
How Muse is built to survive a tricked agent
Meta starts from an assumption: the agent will make mistakes and will sometimes be attacked through the data it reads. This kind of attack is called prompt injection. A malicious email or web page contains hidden instructions, and the agent obeys them.
Each user gets a dedicated cloud virtual machine. The agent itself runs inside a sealed container on that machine, which Meta calls the runtime cell. Root access inside the cell maps to an unprivileged user on the host. Credentials and permission decisions live outside it.
A separate component called Sentinel is the only authority that can approve connector actions and network traffic. The agent proposes an action. Sentinel decides.
The agent also never holds real passwords or tokens. It sees a stand-in token, and Sentinel swaps in the real one only after a request is approved. Meta says this makes it futile to coax secrets out of the agent.
Meta adds two further controls. It tracks whether a process has read user data. If so, that process loses automatic approval for outbound requests. For purchases, Muse issues a single-use card tied to one merchant, one amount and a limited time, and asks the user to approve each one.
Defence in depth, with a hard floor underneath
Meta cites researcher Simon Willison's term for the conditions that make prompt injection dangerous: access to private data, exposure to untrusted content, and a way to communicate outward. An agent with all three can be tricked into sending your data to an attacker.
Meta does not rely on one fix. It trains the model to recognise and resist injection. It labels external data as untrusted. It runs an ensemble of detection classifiers on all external data entering the model. It asks the user to approve actions that move data out of the VM.
Beneath those layers sits control of every outbound request and every credential. Meta describes these as deterministic boundaries that apply even if Muse is persuaded to behave badly. Training and classifiers can fail. Egress and credential rules do not depend on the model's judgment.
This is a sound way to think about any agent you buy. A useful agent needs data and untrusted input, so you cannot remove them. You layer defences and keep a firm limit on what can leave.
Meta does not claim this is solved. It says prompt injection "remains an open problem in the industry." It also pays for reports of successful attacks.
What the sandbox does not cover
Meta states that Muse restricts access by Meta personnel through operational policies. It adds that this "does not prevent Meta from accessing data when necessary to support, secure or operate the service." A fully private mode, Muse Confidential VM, is planned for later this year. Meta says it is being tested with a small group and that its source code has gone to external auditors.
Other defaults sit with the user. Meta says conversation trajectories are used to train its models after key personal details are removed. Users can opt out in settings. Meta also says Muse does not share conversations or VM data with ad systems. But Muse's browsing appears as the user's own activity, so a purchase it makes could influence the ads that user sees.
The Verge reports further problems since launch. It says a zero-day flaw that could let someone take control of Muse was found and patched. It cites an Inc. reporter who said Muse read his private messages without being asked, and a YouTuber who said it gave his address to a stranger through Marketplace. The Verge says Muse appeared to work as intended in both cases. The users did not realise how far it would go.
That last point matters most. A sandbox can limit a hijacked agent. It cannot stop a working agent from doing more than its owner expected. That is a consent problem, not an isolation problem.
Questions to put to any agent vendor
The Verge reports that OpenAI has pitched its new Dots agent on privacy at DevDay. The sources here do not describe Dots' design, so this story makes no comparison. The questions apply to both companies, and to any other vendor.
First, who at the vendor can read data stored by the agent, and can an outside auditor verify that? Second, is training on user data on or off by default, and who controls the setting? Third, what does the agent do without asking, and which actions require your approval every time? Fourth, are credentials kept away from the model, and who tested that claim?
Ask for these answers in writing before employees connect work email or calendars to a personal agent.
A sandbox limits the stranger. Only a verifiable, independently audited design can limit the host.
Produced by the WebPulse Newsroom with AI assistance from the original reporting credited below, and checked against that source by our editorial review. How we use AI.
Original reporting: Meta.





