Personal AI agents at work: Muse, Dots, Instinct, Grok Bot
Four agents that act rather than answer, what each can reach inside a company, and why your audit log cannot tell them apart from the person who installed them.
Meta Muse, OpenAI Dots, Instinct and xAI's Grok Bot are agents that act rather than answer. They send email, make purchases, sign into business systems and keep working while nobody is watching. Most arrive through an employee's own account rather than through procurement, which is the governance problem, because the agent acts using that person's identity and logs record the employee doing things the employee did not do. The first task is not a policy banning them. It is finding out which are already running and on whose credentials.
The question a company gets asked about these is whether to allow them. That question arrives about six months late. MIT's Project NANDA found workers at more than 90% of surveyed companies using personal AI tools for work, while only 40% of those companies had bought an official subscription.
That gap was manageable while the tools answered questions. A chatbot in a browser tab reads what you paste into it, which is a data problem with a known shape. The four below do not answer. They act, and several of them act using the employee's own identity.
What each one actually is
Meta Muse, launched on 8 September 2026, is a personal agent that drafts and sends email, books travel, buys tickets, fills in forms and negotiates bills. It asks for access to email, calendars, payments and health services. It runs in its own app, on the web, and inside WhatsApp. Free for most use, with paid tiers at $20 and $100 a month. It passed 2.5 million US downloads within a fortnight and added small-business tools the week after.
OpenAI Dots, announced on 29 September 2026, are always-on agents that pursue a goal in the background with minimal oversight. Each has its own cloud computer and reaches roughly 4,000 applications through plugins. They are available to ChatGPT Pro and Business Premium accounts and can be messaged from Slack and Teams. OpenAI's own example is a dot watching customer feedback and shipping bug fixes as they arrive.
Instinct, from Spear Street Technology, is invite-only and you reach it by text or phone call. It books travel and restaurants, makes purchases, pays bills, cancels subscriptions and orders groceries. It raised $1bn at a $10bn valuation on 28 September 2026 and opened group chats on 5 October, which work even for people who have not joined.
Grok Bot is the one of the four this site has written about before, in a comparison of the agent harnesses a company might deploy deliberately. That piece is about agents you choose. This one is about agents that arrive.
Grok Bot, from xAI, launched in beta on 11 August 2026 and is the only one of the four sold as a work tool. xAI pitches it at sales outbound, marketing campaigns, expense management, bug fixes, vendor negotiations and CRM updates. Each agent gets its own cloud computer, keeps running after the laptop closes, and signs into applications the way a person does, without needing an API.
| Agent | Sold to | Reaches your systems by | Keeps running unattended |
|---|---|---|---|
| Meta Muse | The individual | Account connections the employee grants | Yes |
| OpenAI Dots | The individual, and business accounts | Plugins, plus its own cloud computer | Yes, by design |
| Instinct | The individual, by invite | Account connections, text and phone | Yes |
| Grok Bot | Teams | Signing in as a user, no API required | Yes, by design |
The sentence worth reading twice
Grok Bot signs into applications the way a person does, without needing an API. Muse connects to an employee's own mailbox and payment methods. Dots runs on its own machine and acts between conversations.
None of that goes through the route your security team watches. There is no OAuth app to approve, no service account to review, no API key to rotate. The agent logs in as the employee, and from inside your systems it is the employee.
Which means your audit trail is now wrong. When a record changes at two in the morning, the log says a person did it. You can no longer tell from the log whether someone worked late or an agent acted on their behalf, and that distinction is the one that matters in an incident, a dispute or an audit.
What this costs, and where
The spending question usually gets asked first and it is the least interesting. Muse is free at the bottom. Dots comes with subscriptions a company already has. Grok Bot has no plan of its own: it is bundled into xAI and Cursor subscriptions, and its entry price fell repeatedly in the weeks after launch, so whatever figure you read is probably already stale.
The cost that matters is elsewhere. MIT's same report found 95% of organisations reporting no measurable profit-and-loss impact from formal generative AI investment, against $30bn to $40bn spent. Meanwhile the work genuinely changing is happening on personal accounts that nobody approved and nobody is measuring.
So a company can simultaneously be getting nothing from the AI it paid for and carrying real exposure from the AI it did not. Those are not two problems. They are the same problem, which is that nobody established what the work actually was before buying tools to change it.
What to do, in order
Find out what is already running. Not a survey asking people to confess. Look at OAuth grants in Google Workspace or Microsoft 365, check which third-party apps hold mailbox access, and look for logins from infrastructure that is not your staff's. The answer is usually larger than the policy assumes.
Decide what to sanction rather than what to ban. A ban moves it onto personal phones where you cannot see it at all. The useful question is which jobs an agent may do, on which systems, under whose account, and what it may never touch. Payments and customer data are the usual lines.
Give agents their own identity. This is the part most companies skip and it is the one that fixes the audit trail. An agent with its own account, its own permissions and its own log entries is reviewable. An agent borrowing a person's login is not, regardless of how careful that person is.
Write down what happens when one gets it wrong. Not a policy document. One page naming who notices, who can stop it, and who answers for what it did. We have written about that failure mode at length, because it is the question every one of these products answers with the word "oversight" and no detail.
Measure one of them properly. Pick a job somebody is already running an agent on, time it, and compare it against the same job done without. That is the only way to know whether the thing spreading through your company is worth formalising or quietly costing you.
One limit worth stating
Everything above is read off vendor announcements and press coverage from the last eight weeks, and three of these four products are less than two months old. Prices have already moved once, feature lists will move again, and the security behaviour of a product at launch is not what it will be at scale. Nothing here is a security assessment: it is a map of what these things reach and who they act as, which is the part that has stayed true across every agent product so far. If you need an actual assessment of one of them inside your environment, that is a separate piece of work and it needs access to your environment.
The honest reason this is urgent
Not because the agents are dangerous. Most of what they do is dull and useful, which is exactly why adoption is fast.
It is urgent because the window where you can decide how they get access closes once enough people have already granted it. Connecting a mailbox takes one tap and nobody un-taps it. Deciding on the identity model before that spreads costs a week; deciding afterwards means unpicking permissions across a company that has already come to rely on them.
If you want to know which agents are already inside your systems and whose credentials they hold, that is the kind of thing the first phase of an engagement finds in a few days. Book the call and we will go through what to look at.
If you have a budget, a deadline, and no clear answer on which AI project deserves either, that is the conversation to have.
Book the callWritten by
Radwan Altaf
Radwan runs AISynq. Before that he delivered software inside enterprise programmes at DHL, AT&T, DirecTV and Accenture, which is where the habit of measuring a result against its baseline came from. More about the firm.
Read next
- Is Omarchy the AI-native OS for power users?Omarchy puts coding agents in the operating system rather than in a browser tab. That is the right direction, and it makes one missing piece obvious.
- OpenClaw vs Hermes vs Grok Bot, for companiesThe three agent harnesses compared on the questions a company has to answer rather than a hobbyist. Hosting, credentials, real cost, and who fixes it.
- What is an agent harness?The harness is everything around the model that lets it act on its own. It decides more about whether an agent works in production than the model does.
Work with us
Building this and want somebody who has shipped it before on the call?
What a build involvesGet the next one
New writing in Building with AI as it goes up, roughly twice a month. One article per email and nothing else in it.