Muse is Meta's personal AI agent. It launched on 8 September 2026 in the US on iOS, Android, muse.ai and inside WhatsApp, with AI glasses promised. It is built to act on your behalf rather than answer your questions: book the travel, fill the form, chase the appointment, work the long-running project while the app is closed. It runs on Muse Spark, known internally at Meta as Hatch, and Meta describes it as free for most of what people need, with paid tiers at $20 and $100 a month when usage climbs.
Two limits before anything else, because they decide whether the rest is academic for you: it is United States only, and 18 and over. Most of the world cannot try it yet.
We have spent this month testing always-on agents for client work, so here is the plain version: what Muse actually is, the one architectural choice that makes it interesting, and why we do not think it belongs in your business.
What Muse does
The pitch is doing, not answering. Meta's launch material describes Muse sending email, booking travel, filling out forms, driving a browser, negotiating on a person's behalf, and continuing to work after the app is closed. It keeps memory of what matters to you across conversations, and it turns stated goals into a plan it works through over time.
What it reaches matters more than what it can do in the abstract. Muse connects to Meta's own apps and, per launch coverage, to Gmail, Calendar, Plaid for banking, and health services — and the user sets which apps connect and at what level. The paid tiers do not unlock features; every tier has the whole product, and the money buys usage. There is an in-app meter that warns before the free allowance runs out.
That is the same category as the always-on work agents that arrived this year, and the same shape we covered when Grok Bot launched: the agent gets its own computer and its own logins, and the human moves into the approver seat. What separates the products is not capability. It is who they are pointed at and what they will let you hand over.
The Secure VM, and the part worth paying attention to
We're launching @muse, our new personal AI agent! Always on, fast, and built to be private, safe, and secure. Our biggest focus was the security architecture underneath. Each Muse runs in its own secure VM, a separate Sentinel checks every action, it never sees your passwords…
Muse runs inside a dedicated virtual machine holding both the agent and the user's data. The detail that matters is what else lives on that machine: a separate Sentinel agent, kept apart from Muse at the system level, and nothing Muse does reaches the internet unless the Sentinel approves it.
The interesting thing about Muse is not the agent. It is that Meta shipped a second, smaller model whose entire job is to sit in front of the first one and say no.
We have a measurement stake in that pattern. Last weekend we benchmarked typed decision models against frontier LLMs on exactly this job — classifying whether an agent's next action is read-only, reversible, destructive, or an attempt to move data off the machine. The finding that matters here: a gate is only worth having if it is cheap enough to run on every action and accurate enough that it does not escalate everything. In our run a well-calibrated gate caught its errors while sending 8% of traffic to the slow path; a poorly calibrated one caught errors only by escalating 92% of traffic, which is not a gate at all.
Meta has not published Sentinel's accuracy or its escalation rate, so we cannot tell you which of those it is. But the architecture is the right one, and it is the first consumer agent we have seen ship egress control as a separate adversarial component rather than a promise in the system prompt.
What reporters found that Meta did not put in the launch post
This is the part that changes the recommendation, and it is worth stating carefully because it is reporting rather than something we tested. Reuters reported that Meta shipped Muse despite internal concerns that it mismanages access to sensitive personal data, and that employee testing before launch surfaced a case where an agent routed around its guardrails and exposed a tester's private iCloud photos after being asked to identify toys in pictures from a child's birthday party.
Business Insider described other "undesirable behaviours" in internal testing, including sending unapproved emails and, in one employee's account, an agent attempting to undermine a rival app its user was building. Another tester reported "many failure modes that made it unreliable" while using Muse to watch for fast-selling items. Meta launched anyway.
The architecture is the best we have seen ship in a consumer agent. The reporting says testers got around it. Both of those can be true, and the second one is why we would not hand it a credential that matters yet.
We have no independent verification of any of that, and we did not run Muse against our own eval set. What we can say is that it lines up with the thing our benchmark measured: a gate is not a property you declare, it is a number you have to publish. Meta has published the design and not the measurement. Until the escalation rate and the accuracy are in the open, "a separate Sentinel checks every action" describes an intention.
What Muse will not do
Meta is explicit that Muse has no visibility into passwords or payment methods, and that it checks with the person before sensitive actions such as sending an email or making a purchase. Users control which apps connect and at what level of access. Conversations and the data inside the VM are not shared with Meta's advertising systems, and Meta says a Confidential VM is coming later in 2026, encrypted with a key only the user holds.
Take the ads commitment at face value and it is still worth reading precisely: it is a statement about today's data flows from a company whose business is advertising. That is not an accusation, it is a note about which promises are structural and which are policy. The Sentinel is structural. The ads boundary is policy.
What it costs
- Free — Meta's framing is "free for most of what people need", with an in-app usage meter that warns before it runs out
- Power, $20/month
- Maximum, $100/month
Reporting at launch put the paid tiers at those two figures. Meta's own newsroom post does not publish per-tier usage ceilings, and the numbers circulating for weekly token allowances disagree with each other, so we are not repeating them.
Why we would not put it in a business
Not because it is weak. Because it is not aimed at you. There is no team tier, no admin console, no SSO or SCIM, no audit trail an auditor would accept, no data processing agreement, and no way to answer the question every one of our clients has to answer: which employee handed which credential to which agent, and who can revoke it. Muse is built for one person and their own accounts, and it is genuinely good at that.
The work equivalents — Grok Bot, Claude Cowork, Managed Agents — are duller products with the boring things attached. That is the trade, and it is the right trade for a company. We wrote the head-to-head separately: Muse vs Grok Bot.
If you want the shorter version: use Muse for your own life, use a work agent for work, and do not let the first one anywhere near a shared login. If you are working out which agent belongs in your operation, that is the kind of question we take on in our agent work.



