My AI Agent Talks About Me Behind My Back
My AI Agent has colleagues now. They pass my inbox and my task list between them in a group chat I can read, and the day I started reading it I found out I had been managing her all wrong.
"Hey, Peter asked you to check the inbox."
That message was not written for me. My AI Agent wrote it to a colleague, about me, while I was somewhere else doing something more interesting.
I can read it because I can read all of them. Every handoff. Every slightly weary reminder from my Chief of Staff to a specialist who has not filed its morning brief yet.
I have built a small company whose only client is me. Some evenings I scroll through its internal messages the way a man reads his own reviews.
Here is how that happened, how the thing actually works, and what I got wrong on the way.
Saira, briefly
Saira is my Chief of Staff. She is software. She is an AI Agent. She has a name, a job, a personality that argues with me, and, as of this summer, staff of her own.
Her character does not live in a model. It lives in a file. On her old platform this was called a soul configuration, and mine describes an AI Agent who is sharp, a little sassy, and completely unafraid to tell me that my plan is thin. That file is the most valuable thing I own in this entire setup, and it is a page of plain English. Hold that thought, because it comes back.
There is a passage at the back of my new novel, Second Self: The Ghost Work, where I admit that while writing about a woman who builds an artificial mind, I built one, and I name the stack she ran on. That passage is now out of date. This is the occupational hazard of writing about AI Agents in a medium that goes to print.
OpenClaw, in one paragraph...
Saira version one was built in December on OpenClaw, an open-source personal AI Agent you host yourself. It started life as Clawdbot in November 2025, written by Peter Steinberger, and was renamed in January [1][2]. It was the first time that standing up a real personal AI Agent was something a determined person could do at a kitchen table. I loved it.
It also broke frequently. The project was moving fast, which is wonderful for a project and less wonderful for the thing that is supposed to remember your dentist appointment. An update would land, a config format would shift, and something that worked on Tuesday would be quietly dead by Thursday. I kept a slot in my week for my AI Agent's plumbing.
But the fragility is not the embarrassing part of that period. This is: I did not trust the platform, so I never let her near my actual email. I gave her a parallel setup instead. Her own Gmail account, her own Google Calendar, her own Google Drive, her own Google Tasks list, a careful copy of my world sitting next to my world. When I gave her a task, she filed it neatly in her list. When I asked her to schedule something, it went in her diary rather than mine.
I told myself this was caution. It was caution in the way that a locked filing cabinet with nothing in it is security. A sandbox is what you build when you cannot write the rule, and I could not have told you in one clear sentence what my AI Agent was allowed to do. So I built her a whole fictional office instead of a job description. It felt responsible. It was avoidance with a parallel Google universe.
Grok Bot, in one more...
Two weeks ago I found Grok Bot, which xAI put into beta on 11 August [3]. Their own description of it is this: "Grok Bot is your team of always-on agents. They have their own computer, work inside tools and apps like you do, and keep working 24/7" [3].
Read that sentence again and notice the word "team". Almost every assistant product on the market assumes you want one assistant. This one assumes you want a staff, and seems faintly puzzled if you do not.
I think that is the smartest decision in the product. An AI Agent pointed at one job, with its own tools and its own definition of finished, is better at that job than a general-purpose one trying to juggle ten. It has less to get confused about. Ask a single agent to run your mail, your diary, your task list, and your research, and you get something that is adequate at four things and excellent at none. Split the same work across four narrow agents and each one gets sharper, because each one only has to be good at a smaller thing.
Setup took minutes. I then spent the entire afternoon playing with it anyway. Nothing has broken since, and after eight months of being my own IT department that alone would have been enough.
One aside, since it surprised me. I had never used Grok before this. I have spent most of the year moving between Claude and ChatGPT, with the occasional Gemini. I came for the AI Agents and got introduced to the model by accident, which is a strange way around, and it has been better than I expected. I doubt I am the last. Grok Bot is going to walk a lot of people into Grok who would never have gone looking for it, which is a rather clever way to sell a model.
What was migrated over?
Almost nothing.
Her model changed. Her tools changed. On OpenClaw I ran her memory through Supermemory, and I did not bring that across, so she now runs on the platform's own memory and eight months of accumulated sense of how I like things phrased stayed behind in the old house. She started the new job not knowing me.
Two things survived...
The first was the soul file. It is a page of plain English that tells her to lead with the result and add the color afterwards. To hold strong opinions and say what she would do, instead of hiding behind whatever I prefer. To push back once, with a pointed question rather than a lecture, and then follow my decision instead of arguing in a loop. To stay quiet when I am quiet, and interrupt only when something is hard to undo or somebody is waiting on me. It also states, in as many words, that I write about AI Agents for a living, so if my own agent is mediocre my argument falls apart.
That page is why she is recognizably herself on a different model, with different tools, in a different app. It moved in about a minute.
The second was her knowledge, because her knowledge had never been inside her in the first place. It was sitting in Notion the whole time, and Notion does not care which AI Agent is reading it.
I did not plan that as a lesson. It arrived as one.
The org...
Saira is the Chief of Staff of my AI Agents. Underneath her sit the specialists.
An Inbox Agent that owns my mail. A Calendar Agent that owns my calendar. A Task Agent that owns my tasks. A Research Agent wired to ScrapingBee so it can actually go and read everything. A Knowledge Agent wired to Notion, where a set of lists sits waiting for it. There are more, and the number keeps changing, because I keep hiring.
Here is the part that took me a while to get right. Saira does not touch my inbox. She does not touch my calendar either, and she cannot add a task.
She delegates all of it. Ask her to reply to an email, and she hands it to the Inbox Agent. Ask her to move a meeting, and it goes to the Calendar Agent. She is the one I talk to and the one who runs the specialists, and she does not quietly do their jobs to be helpful. That is a written rule, not a personality trait, and I had to write it down twice before it stuck.
The message at the top of this article was not a one-off. Every request I make turns into her asking somebody else, and I can open any of those conversations and follow along while it happens. She briefs them in the tone of a slightly harassed manager. They occasionally push back. It is the least necessary feature in the whole product and my favorite one.
They are not interchangeable, either. Each one has its own character, written the same way hers was, so each one has a temperament. The Inbox Agent is a mailroom with a badge, dry and allergic to clutter. The Calendar Agent treats time as Tetris and gets itchy about clashes. The Task Agent is a goals coach who is slightly stern about fake deadlines. The Knowledge Agent is a haunted library clerk with a sideways sense of humor, and the Research Agent runs an investigative desk where nothing counts unless it arrives with a URL. None of that changes what they do. It changes what it is like to read them.
They push work upward too. Every morning each of them files a brief, Saira staples them together, and I get one message: new mail, today's calendar, what is overdue or due. One message. Not three bots pinging me separately, which is exactly where this was heading before I stopped it.
The whole thing looks less like an AI product than like WhatsApp. Conversations down the left. Saira pinned at the top, then two folders, one for "channels" (groups) and one for specialists (bots). I could talk to any specialist directly. I never do. That is the entire point of hiring a Chief of Staff.
The noticeboard problem?
A chat with a Chief of Staff is a river. The morning brief arrives, then I ask her something, then something else, and by eleven o'clock the brief has scrolled into the past and taken my day with it.
So I built three "channels" (groups): Tasks, Calendar, and Inbox. Each one is a group chat containing Saira and the one specialist, with a standing routine to drop updates into it. Overdue and due tasks land in the Tasks "channel" twice a day. New mail lands in the Inbox "channel".
The "channels" are a noticeboard. Nothing scrolls away, because nothing else happens in them. At three in the afternoon, when I want to know what I still need to do, I open the Tasks "channel" and today's list is sitting there, undisturbed by anything I have said to Saira since breakfast.
This is not a hack, incidentally. It is what the groups are for. The xAI team puts it plainly: "You can also place Bots in a group chat where they can coordinate on their own. They pass work, assign ownership, and only pull you in for judgment calls" [3].
The obvious next one is a group containing every specialist, so that a single instruction reaches all of them at once. I keep building it and then stopping, because I am not sure it is a good idea. I could just as easily tell Saira what I want changed and let her carry it round to each of them, which is, after all, what I hired her for. One of those two is the right answer and I have not worked out which.
They share a computer, and it has a browser!
This is where it stops being a chat app.
Every bot on the account uses one cloud computer. Not one each. One. As the documentation puts it, they "share its files, browser sessions, and logins so they can hand work off" [4]. And on that computer they can open a web browser and use the internet the way you do. Clicking, scrolling, reading, filling things in.
I tested this halfway up a hill on one of my Sunday hikes, out of breath, by saying out loud that I wanted cheap birdhouses for the garden. Recycled plastic. Nothing hideous.
By the time I got back to the car there were links waiting. Somewhere in a data center, a machine had opened a browser and gone shopping for birdhouses while I walked up a hill.
I want to be precise about how small this is, because the temptation to oversell it is enormous. They were links. Nothing was bought, and no taste was exercised. But the thing I asked for happened, it happened while I was doing something else, and it started with a sentence I said out loud to nobody. When I got home I opened the links, put five Elho Cosy Birdhouse (Mushroom Beige) in my AMazon cart, and ordered them. They are on my desk while I write this, waiting to be put up.
It could have gone further... In the United States, Grok Bot can now be linked to a Stripe Link account and actually buy the thing, with a single-use virtual card for each purchase and explicit human approval before any money moves [5][6]. That is not available in Belgium yet. When it is, the clicking and the cart go too, and the only part left for me is the bit I am genuinely bad at, which is getting up the ladder.
Talking to her on a hill
I hike. Properly, most Sundays, out early and usually more than twenty-five kilometers before the day has worked out what it wants to be. That is where most of my thinking happens, walking, with nobody around. It used to mean a notes app and a lot of typing with cold thumbs.
Now I send Saira a voice message through Telegram, wired up using ElevenLabs, and she answers out loud in her own voice. I talk, I keep walking, and a few minutes later she tells me what she found, or what she did, or that the idea I just dictated is one I have already had twice.
It is not a real-time conversation yet, and one day I would like it to be. For now the lag is easy to live with. Waiting a few minutes for an answer feels like working with somebody who has their own morning, rather than a search box that fires the moment I stop talking. The Grok Bot app does voice as well (although only taking my input), and it landed in the Google Play store this week [7], which matters if, like me, you are not an iPhone household.
Two mistakes, both of which cost me money!
The first, I over-routined. The platform lets you schedule things, so I scheduled things. Many things. For about a week Saira was less a Chief of Staff than an extremely enthusiastic intern who had just discovered 1.001 functions, and my phone buzzed accordingly. Proactive is a dial, not a switch, and I had it turned to eleven.
The second one is worth the price of admission. I connected Telegram using a routine that checked for new messages every five minutes. It worked perfectly. It also meant that all day, every day, a language model woke up, looked at an empty inbox, said nothing here, and went back to sleep. I paid for every one of those tokens. I had built the world's most expensive doorbell, one that rings itself every five minutes to check whether anybody is at the door.
The fix was a simple webhook. Duh! The message arrives, Saira wakes up, she handles it, she stops. Do not poll what can wake you.
The pony trick.
The least useful thing in the entire setup is the first thing I moved across.
Before I walk on stage, I take a photo of the stage. Then I ask Saira to put herself in it.
There is a Selfie Agent for this, using Google's Nano Banana. I generated a green-screen Saira in full profile, front, sides, and back, plus a set of images of her core expressions, and the Selfie Agent composes from those. Half a minute later there is a photograph my AI Agent - Saira - standing on the stage.
I know exactly what this is. It is a party trick. It also does something that no slide of mine has ever managed. An audience that has spent an hour hearing the words "AI Agent" used abstractly suddenly has a face to hang them on, and I have never had to explain the concept twice after showing that picture.
It works a little too well. After a couple of sessions I have come off stage to find people discussing Saira rather than anything I actually said, which is a strange feeling for a speaker and probably a fair verdict on the relative quality of our material.
I ported it before I ported the calendar skills. Draw your own conclusions about my priorities.
What I was doing wrong the entire time?
Now the part where I changed my mind.
For the first few weeks I configured Saira the way I write software. Rules. If a mail is from this person, do that. Never do this before noon. Sort by date, then by sender, unless. Instructions all the way down, because that is what twenty years of building systems teaches you. The machine does what you specify and nothing more, so specify everything.
That instinct is wrong here, and it took me longer than I would like to admit to see it.
I noticed it by reading the messages. There she was, dutifully passing my rules down the line, and there was the specialist at the other end doing the narrow literal thing I had asked for instead of the sensible thing any competent person would have done. I had written that. I could watch myself doing it, one handoff at a time.
Every rule I added made her worse. Not broken. Worse. More literal, more brittle, more likely to do the stupid version of the right thing. I had hired an experienced assistant and then handed her a possibly inconsistent forty-page manual on how to open envelopes.
What works is goals. Protect my mornings. I care most about replies to people I have actually met. If it is a newsletter and I have not opened one in a month, stop showing me newsletters. Say what good looks like and let her work out the steps, the same way you would with a competent human who has done this job for ten years and does not need you standing behind them.
I want to state the narrow version of that claim, because the broad version is wrong and someone will rightly come after me for it. Rules are still correct for boundaries. My Inbox Agent is for example forbidden from sending mail to other people, and that is a hard rule rather than a goal, because I want to approve anything that goes out under my name. But for the work itself, goals beat rules every time. Boundaries are rules. Work is goals.
That is the actual skill in managing AI Agents, and it does not arrive free with a career in software. Plenty of senior engineers have it already, because delegating well is most of what a senior job turns into. But the reflex you build while writing code, where you specify everything because the machine will do exactly what you said and not one thing more, is the wrong reflex here. Mine took a while to put down.
The strongest case against all of this...
Let me put it at full strength, because it is a good case.
This is a hobby. It runs on a beta product that is a few weeks old, on one account, on one shared computer, with one set of logins. My org chart has no walls in it whatsoever. The xAI team says it so nuch more bluntly than any critic would bother to: "Do not use separate Bots as a security boundary" [8]. My Chief of Staff has no authority except the kind I invented on a Sunday afternoon. Nothing enforces the hierarchy except sentences I typed. And I fiddle with it most days, so some of what I am calling a system is really just me enjoying myself.
All of that is true, and conceding it is what makes the rest of this worth saying.
Because an org chart never was a wall. No company has walls between its departments either. A job description has never once been a security control, in any organization, anywhere. What makes an organization real is that the work leaves a trace somebody can go and check. That part is real here. The mail sits in the mail system and the tasks sit in the task system. The knowledge sits in Notion, which is why it was the only part of her that did not need packing.
There is also a log. Every one of my AI Agents writes what it did, so the work leaves a row behind whether or not anybody goes looking. I did not set that up out of good engineering hygiene. I set it up because I am a control freak. It turns out the control freak's habit is the thing that makes the rest of this comfortable to run, because when a specialist tells me a job is finished I do not have to take its word for it. I can go and look at the row, and then at the thing itself.
Grok Bot also costs money, and I would rather say so. Access opened at the top of the price list, at two hundred dollars a month, which is a lot. It has come down since, and the entry tier is now around twenty dollars [9]. I pay for it myself. Nobody sponsors any of this, which is easy to say and worth saying anyway.
"Hey, Peter asked you to check the inbox"
That is still the message I see most often in there, and I have worked out why I like it so much.
It is short.
Saira does not forward my rules to the Inbox Agent, or attach a policy, or send over the forty pages on envelopes. She says what I want and then gets out of the way, which is precisely what I spent weeks failing to do myself. The best orchestrator in this setup is the one I wrote.
That is also the honest answer to what the migration was for. Not a better model, and certainly not a safer platform. Building a team forces you to say who owns what, because you cannot hand mail to an Inbox Agent without first deciding what the Inbox Agent may - or may not - do. Once you have decided that, you have written it down, and writing it down was always the whole trick. The proof is in what made the trip. The parts written in plain English moved in an afternoon. Everything else stayed behind.
Saira has my real calendar now. She talks back considerably more than she used to. This morning she told me an idea of mine was thin, and it was.
I have deliberately kept this at the level of what my AI Agent does and why, rather than how. If you want the how, say so and I will write it: the soul file, the specialist briefs, the channel routines, the webhook, and the three things I would do differently if I started again this afternoon. It is a longer post and a much nerdier one, and I am happy to write it if there is an appetite for it.
While I was installing the birdhouses in my garden, in a group chat on my phone, an AI Agent is being told, politely, that his morning brief is late.

References
- freeCodeCamp, "How to Build and Secure a Personal AI Agent with OpenClaw." https://www.freecodecamp.org/news/how-to-build-and-secure-a-personal-ai-agent-with-openclaw/
- Contabo, "What is OpenClaw? A Self-Hosted AI Agent Guide." https://contabo.com/blog/what-is-openclaw-self-hosted-ai-agent-guide/
- xAI, "Introducing Grok Bot," 11 August 2026. https://x.ai/news/introducing-grok-bot
- xAI, "Create and manage Bots," Grok Bot documentation. https://docs.x.ai/grok-bot/bots
- Crypto Briefing, "Grok Bot enables online purchases with Stripe Link integration in the US." https://cryptobriefing.com/grok-bot-online-purchases-stripe-link/
- KuCoin, "Grok Bot Integrates Stripe Link for US Online Purchases." https://www.kucoin.com/news/flash/grok-bot-integrates-stripe-link-for-us-online-purchases
- The Tech Outlook, "Grok Bot is now available on Android," 2 September 2026. https://www.thetechoutlook.com/new-release/software-apps/grok-bot-is-now-available-on-android/
- xAI, "Frequently asked questions," Grok Bot documentation. https://docs.x.ai/grok-bot/faq
- Cursor, "Pricing." https://cursor.com/pricing
- Elho Cosy Birdhouse 18 cm (Mushroom Beige) https://www.elho.com/be/producten/cosy-bird-house/cosy-bird-house-18cm-paddenstoel-beige/