My AI Agent Does Not Need to Be Autonomous All the Time
The more I build AI agents, the less interested I am in making them fully autonomous. Control, visibility, and good handoffs often matter more.
A lot of agent demos aim for full autonomy.
Give the agent a goal. Let it figure out the steps. Watch it use 20 tools. Then wait and see if it finishes.
That can look good in a demo. I am not convinced it is what I want in a product.
People want the result, not autonomy
If an agent can finish a task safely without asking me anything, that is fine.
If there is a real decision in the middle, I usually want to see it.
"Send this email to 5,000 customers" should not be: the model thinks, the model decides, the model sends.
A better flow is: the model prepares, the user reviews, the user approves, the system sends.
The model still did most of the work. The person kept control of the expensive decision.
I care less about how much the agent does
The best agent I have built is not the one that does the most.
I care about three things.
Can I understand what it is doing? Can I stop it? Can I recover when something goes wrong?
These questions sound dull next to autonomy. They matter once an agent has real tools.
An agent that can create, delete, publish, send, or spend money needs tighter limits than an agent that only answers questions.
I keep autonomy inside limits
My current model is simple.
Let the model handle fuzzy decisions. Keep important constraints outside the model.
The model can decide which content idea is relevant. The application decides whether the user has enough credits.
The model can draft a post. The application checks length and required fields.
The model can choose a tool. The application decides whether that tool is allowed.
That split is easier to reason about.
Asking for approval is not a failure
I used to think asking the user for approval meant the agent was not good enough.
I do not think that now. Sometimes the approval is the product.
A designer does not want a model to publish a brand announcement with no review.
A developer may want an agent to prepare a pull request and still review the diff.
A finance team may want an agent to prepare a payment and require approval before it runs.
The agent can remove 90% of the work without owning the final decision. That is still useful.
Autonomy should be earned
The more reliable an agent becomes, the more I am willing to let it do.
I start with suggest, then review, then execute.
Then I move toward suggest, execute low-risk actions, and ask before high-risk actions.
Some workflows can become fully automatic later.
I would rather give an agent more control after I have evidence than assume it on day one.
That is the less exciting approach. It is the one I trust.