On a Friday afternoon last week, one of my agents stopped in the middle of a job it was perfectly able to finish.
It was publishing a new version of SEO Monster, our open source SEO tool, to the official registry that AI assistants read when they look for tools. The publish failed. The agent worked out why in about a minute: the registry only lets members of our GitHub organization publish under its name, and my membership in that organization was set to private. The fix was one command. The agent had the command. It also had my GitHub login sitting on the machine.
It stopped and asked me instead, because making my membership public is a change to my account, visible to anyone who looks, and it is not the agent's to make. I read the message, agreed, ran the command myself, and the publish went through a few minutes later.
Nothing about that moment was dramatic. That is the point. The interesting question with agents is not whether they can do something. Mine can do almost everything I can do at a keyboard. The question is which of those things they should do without me, and how I decide.
The rule, as I actually said it
I have tried to write this rule down several times and the tidy versions never survived contact with a real week. The version that has survived is the one I said out loud on 28 September, when an agent brought me a design fix for our ERPClaw website and asked if it could ship: is the plan fully ready, has it been simulated, are the bugs fixed? If all three are yes, implement it.
It sounds almost too simple. In practice it does a lot of work, because each of those three words has a specific meaning in how we operate.
A plan is ready when it names the exact files and lines it will change. Not a direction, a diff.
Simulated means the change has run somewhere that is not production and been checked the way a user would see it. For that design fix, that meant loading every one of the site's 180 pages at phone width, in light and dark mode, and confirming nothing spilled off the screen.
Bugs fixed means whatever the simulation found got fixed and the simulation ran again. It found four that time, which is exactly why the step exists.
What my agents do alone
A surprising amount, and more every month.
They read anything: search data, analytics, server logs, every page of every site. They run audits. They draft pages, emails and plans. They build sites locally and run every check we have. They commit and push work to our private repositories, which used to need my approval. It moved to "just do it" once two checks became automatic and impossible to skip: a scan for secrets and personal data in anything outgoing, and a check that the destination repository is private.
The scan is not decorative. Today it refused to let an agent push a note because the note contained an email address. The agent removed the address and pushed again. I found out afterwards, which is how it should be.
That is the shape of everything on the "alone" list. The action is reversible, the scope is narrow, and a gate checks it that the agent cannot argue with.
What still needs me
Anything sent under my name. My agents draft outreach and leave it in my drafts folder; I read it and press send. This week that was five emails to people who write lists of accounting and ERP software, with five more drafted and waiting for me. When one of them bounced, the agent found the right address, drafted the resend, and left that for me too.
Anything on this site. It is my voice, so I read every piece before it goes live, including this one.
Changes to accounts, like the GitHub membership. Legal and policy wording. Spending money.
And deploys. My agents prepare a release completely: the build, the checks, a snapshot of the live site to roll back to, a dry run that lists every file that would change or be deleted. Then they wait for my yes, and once it is live they send me the list of pages to check. My part is small. It matters anyway, because a person who knows what the release is for decides that it goes, and looks at the result.
When the line moves, and when it moves back
The line moves when a gate earns trust. Private pushes moved to "alone" because the secret scan and the privacy check made the bad outcome close to impossible, not because I got tired of approving them.
It also moves back, and the clearest example is recent. For a few months we ran a separate system that scheduled agents to work on their own: blog proposals every weekday, outreach drafts every week, audits every month. It was busy. When I finally looked at the board, 35 topic proposals were sitting there waiting for a review nobody had given, and every one of those runs had used real capacity. Autonomy without a reader is not leverage. It is noise with a bill attached. I shut the scheduler down on 27 September and moved the work back into sessions where I am in the loop.
What the gates cost, and what they catch
Waiting costs something. A fix that could ship at noon sometimes ships in the evening because I was in meetings. Once, my own sequencing got ahead of the gate: I sent emails saying a new version was live before the agent had actually published it. The agent noticed the mismatch and flagged it, and the fix was quick, but it was a reminder that the gates protect me from me as well.
They also catch things. On the same design release, an agent reported that a different site had zero pages wider than a phone screen. A second, independent check found six, including the home page. They had been broken for a while and nobody had noticed. Six pages is a small number. The lesson is not small: an agent's report is a claim, and a claim gets checked before it gets trusted.
Where this leaves me
I do not think the right question is how much to trust an agent. Trust is a feeling, and feelings do not scale across a dozen products and a team of fifteen, especially when we build and run the AI we recommend ourselves. The better question is what an action needs before it happens: a plan specific enough to check, a rehearsal somewhere safe, and a gate that does not care how confident anyone sounds. When those exist, I get out of the way. When they do not, I stay in it, and the agent waits. It is the same idea we built into ERPClaw, where the model decides what you meant but never what the books say.
If you want the executive version of this, with a framework rather than a founder's week, we have written two pieces on the AvanSaber blog: when to trust AI with the decision, and what AI agent governance checklists miss.