The Most Impressive Thing Grok Bot Did Was Stop
I asked it to compare the price of a bottle of Coca-Cola across every major Belgian supermarket. It hit a wall at Carrefour, slid the browser window over to me, and asked me to click the box that proves I am not a robot. That moment, not the price list, is the reason I am writing this.
I ignored xAI's Grok for two years.
Not on principle. I had ChatGPT for one kind of work, Claude for another, Gemini when I wanted a third opinion, and no free evening to add a fourth thing to the rotation. Grok was the one I never got around to. Nothing about it looked like it would change my week.
Last week I finally tried it, through their new product called Grok Bot. What changed my mind was not something a model said to me. It was a browser window sliding onto my screen with a checkbox on it. Verify you are human.
A bot had hit a wall it could not climb. So it stopped working, handed me the screen, stepped back, and waited. I clicked the box. It took the screen again and carried on with the job.
That reads like a failure. I think it is the most important thing in the product.
Eight months as my own sysadmin
To explain why, I have to go back to December.
That was when I set up Saira, my personal agent, on OpenClaw. OpenClaw started life in November 2025 as a project called Clawdbot, written by the Austrian developer Peter Steinberger, and was renamed in January [1]. It is open source and you host it yourself. You point it at a model, wire it into the messaging apps you already use, and it goes and does things.
For me it was the first time standing up a personal AI agent was something a determined consumer could do rather than something that needed a platform team. I loved it. Saira read things for me, chased things down, and sat there waiting, which no tool of mine had ever done before.
I also spent a lot of evenings on it.
Getting her running took tinkering. Keeping her running took more. Automatic updates would land, something in the chain would shift, and a thing that worked on Tuesday would not work on Thursday. I started leaving a slot in my week for it, the way you leave a slot for the car.
Somewhere in that stretch I noticed what I had actually built. Not an assistant. A pet with a maintenance schedule. I was not directing an agent. I was operating one, and the operating was the part that ate my time.
What Grok Bot actually is
xAI put Grok Bot into beta on 11 August 2026 [2]. Their own documentation describes the unit of the product this way: "A Bot is a durable AI teammate with a name, a job, its own conversation, and working context that develops over time" [3].
Two words in that sentence are doing the work: durable, and job.
Underneath, the architecture is simpler than the marketing suggests, and better. Your account gets one computer in the cloud. It has a browser, a filesystem, and a terminal. Every bot you create works on that same machine, with the same files and the same logins. You can have up to fifty bots and group chats on an account, combined [3]. Bots can hand work to each other, and each one can be running its own task at the same time.
The picture most people have of this is fifty little robots, each with its own laptop. The real picture is one desk. One computer on the desk. One keyboard. Fifty colleagues who take turns sitting down at it.
That single image explains almost everything else about the product, including the parts I do not like, and I will come back to it.
The work also does not depend on me. The documentation is blunt about it: closing the app, the laptop, or the phone does not stop a job that is already running [3]. It runs on their machine, not mine. Nothing to keep awake, nothing to patch.
If you have read my piece on what an AI agent actually is, all of this sits in the harness rather than the model. The intelligence was already there last year. What changed is the scaffolding around it.
I created a bot whose only job is to look after my other bots, naming conventions included. That is the sort of thing you do at eleven at night when a setup is fun rather than fragile.
Setup took me about as long as signing into a new laptop. No tinkering. Plug and play, and not plug and pray.
A bottle of Coca-Cola
The run that made me sit up went like this.
I asked it, in plain language, to compare the price of a bottle of Coca-Cola across the major Belgian supermarkets, tell me the cheapest and the most expensive, and flag any promotions running that day. I did not tell it which shops, which sizes, or where to look.
The first thing it said back was about method:
I'm on the official store sites now so I don't guess from comparison blogs. I'll send the cheapest and most expensive once those pages are in.
That sentence is the whole reason the exercise worked. There are dozens of price comparison sites for Belgian groceries. Reading one would have been faster and would have produced a confident, plausible, stale answer. Instead it went to the shops.
A few minutes later it came back with a table, and with something better than a table.
It reported Colruyt and Delhaize both at 2,38 euro for a 1,5 liter bottle of regular Coca-Cola, Albert Heijn a cent higher at 2,39. It noted that Aldi does not sell a 1,5 liter at all, and that Lidl's webshop carries no branded Coca-Cola bottle, only its own cola. It found the running promotions, five plus one free on one-liter bottles at Colruyt, a six-pack deal at Delhaize, and worked out that both land at about 1,78 euro per liter.
Then it told me where it had failed. Carrefour had blocked the browser. It gave me the pack price it could see from elsewhere, said plainly that it could not open the store pages, and asked: want me to retry Carrefour if you click the human check?
So I clicked the human check.
You're through Cloudflare. Cookie banner is up. I'll accept that and grab the Coca-Cola prices.
Carrefour opened. Two liters at 2,95 euro, which worked out to 1,48 per liter and quietly became the best price per liter of the whole exercise, beating the Aldi two-liter it had ranked first a few minutes earlier. It revised its own answer without being asked to.
Two things in that run matter more to me than the prices.
When a shop was closed to it, it said so, instead of inventing a number or quietly falling back on a comparison blog. And when it could not get through at all, it did not spin and it did not pretend. It asked for ten seconds of my hands.
I should be honest about the ceiling here. This is a bot reading retail websites on one afternoon, and a bot reading a website can obviously still misread it. I'm not sure I would publish those prices as fact for a new business. What I would trust is the shape of the answer, including the holes in it, because it showed me where the holes were.
The seam is the product
xAI has a name for the thing that happened at Carrefour. Their documentation puts it in one line: "Passwords, two-factor codes, CAPTCHAs, and similar human-only steps use a computer takeover" [4].
Read what that sentence gives up.
The agent is not trying to be a person. It is built to fail at being a person, deliberately, at exactly the points where being a person is the entire point: passwords, second factors, and the box that asks whether you are human.
Compare that to how automation has worked for the last twenty years. You wanted a script to log into something, so you gave the script the password. You put it in a config file, or a vault if you were careful, and then you hoped. The credential lived wherever the automation lived, which meant the blast radius of any mistake was the size of your login.
Grok Bot inverts it. The bot keeps the session. I keep the secret. It never sees my password because the design never asks it to.
Call that a usability detail if you like. It is the difference between an agent you can point at your actual life and one you can only point at a sandbox. Everything I have wanted a personal agent to do for three years has been on the far side of a login screen. Retailers. Municipal portals. Airlines. Banks. That is where a normal person's admin lives, and until now the honest answer was that you either handed over your credentials or you did it yourself.
I spent a book arguing that the scarce skill in all of this is orchestration, deciding who does what and where a human stays in the loop. I did not expect part of the answer to arrive as a button.
The strongest argument against everything I just wrote
Now the other side, and it is a serious one.
A team at the University of Washington tested seven agentic browsers this year and found that four of them let an attacker break the same origin policy, the rule from 1995 that stops one website from reading another one's data. Hidden instructions in a page were enough to turn the agent into the leak [5]. Their co-senior author, David Kohlbrenner, put it about as directly as an academic can:
Browser agents aren't ready for the public. Even if you're a relatively savvy user, if these agents have access to a browser that contains your credentials, your email, your bank account, whatever it is, you should not trust that these systems are ready to truly protect your information.
This is not a hypothetical. In March, Unit 42, the threat intelligence team at Palo Alto Networks, reported the first real-world case of a web page carrying instructions written for the machine reading it rather than the human. One page used two dozen separate techniques to hide them, zero-size fonts, suppressed CSS, script-generated text [6].
And xAI hands the objection more ammunition than any critic has. Their own documentation says that shared computer files and sign-ins are not isolated between bots, and follows it with a warning I have not seen a vendor write so plainly: "Do not use separate Bots as a security boundary" [4].
Go back to the desk. One desk, one computer, one drawer with every login in it, and fifty colleagues taking turns. The reason the takeover feels so smooth is the same reason there is nothing between one bot and another. There are no walls in that room.
So I will concede the point. The computer takeover solves the credential handover. It does nothing whatsoever about a poisoned web page telling my bot to do something I never asked for. Those are two different problems, and this product has honestly fixed one of them.
What I will actually defend
The narrower claim I think survives contact with all of that is this.
The industry has spent years asking how to safely give an agent your keys, and the answer that is now shipping is to not give it the keys at all. The human steps into the machine's session for one action and steps back out. That is a real design advance, it is not specific to xAI, and I expect every serious agent platform to end up with some version of it, because the alternative is a credential vault that an attacker only has to reach once.
What has not been solved is what happens between those moments, when an agent with your logged-in life on its screen reads a page written by someone who wants something from it. Nobody has shipped a good answer to that. I have written before about how easily an agent overreaches when a legitimate route fails, and this is the same shape of problem wearing different clothes.
I should also be straight about the price. Grok Bot has no standalone subscription. It rides along with a top tier, and mine runs two hundred euros a month [2]. For a product that is days old, with no published security architecture and no independent measurement of how reliably it finishes real work, that is a lot of money to spend on being early. I am spending it with my eyes open, and I would not tell you to.
The moment it stops
I ignored Grok for two years because I was judging the category on the wrong axis.
I was comparing models. Which one writes better, which one makes fewer things up. On that axis the four big assistants have been close enough for a while that switching is mostly a matter of taste, which is exactly why I never bothered.
The interesting differences moved somewhere else while I was not looking. They moved into the harness: whether the thing survives an update without my help, whether it runs when my laptop is shut, and above all what it does when it hits a wall.
This is also not one vendor having a good week. Most of the serious agent platforms shipped this year were aimed at companies, at teams and seats and admin consoles. The lane nobody was really serving was mine: one person, at home, with a login and an errand to run.
Eight months ago the state of the consumer art was an open source project I babysat on my own cloud hardware, and I was delighted to have it. Today it is a product that hands me the wheel for ten seconds and takes it back. That gap, in eight months, is the story. The vendor is almost incidental. Something better than Grok Bot will exist by the time you have finished arguing about whether Grok Bot is good.
So here is the test I would use on whichever one you are offered next. Do not ask it to do something impressive. Ask it to do something it cannot possibly finish on its own, and watch what happens when it gets to the wall.
If it invents an answer, close the tab.
If it stops and asks for your hands, pay attention. That is the part that is new.

References
- On OpenClaw, its origin as Clawdbot in November 2025, the January 2026 rename, and self-hosting: freeCodeCamp, "How to Build and Secure a Personal AI Agent with OpenClaw," https://www.freecodecamp.org/news/how-to-build-and-secure-a-personal-ai-agent-with-openclaw/ and Contabo, "What is OpenClaw: Self-Hosted AI Agent Guide," https://contabo.com/blog/what-is-openclaw-self-hosted-ai-agent-guide/
- On the 11 August 2026 beta launch and on Grok Bot having no standalone subscription: Unite.AI, "xAI Launches Grok Bot, Always-On AI Teammates With Their Own Cloud Computers," https://www.unite.ai/xai-launches-grok-bot-always-on-ai-teammates-with-their-own-cloud-computers/ and Reworked, "xAI Wants In on the Enterprise With Grok Bot," https://www.reworked.co/collaboration-productivity/xai-launches-grok-bot-ai-agents-in-beta/
- xAI, "Create and manage Bots," Grok Bot documentation. https://docs.x.ai/grok-bot/bots
- xAI, "Frequently asked questions," Grok Bot documentation. https://docs.x.ai/grok-bot/faq
- On the University of Washington study of seven agentic browsers, presented at the Agents in the Wild Workshop in July 2026: TechXplore, "Some agentic AI browsers may come with major cybersecurity risks," https://techxplore.com/news/2026-06-agentic-ai-browsers-major-cybersecurity.html
- Unit 42, Palo Alto Networks, "Fooling AI Agents: Web-Based Indirect Prompt Injection Observed in the Wild," 3 March 2026. https://unit42.paloaltonetworks.com/ai-agent-prompt-injection/
- My own article, "What an AI Agent Actually Is (For Now)," 20 June 2026. https://www.petervanhees.com/what-an-ai-agent-actually-is/
- Anotherone of my own articles, "Your Agent Has No Sense of Proportion," 18 August 2026. https://www.petervanhees.com/your-agent-has-no-sense-of-proportion/