I Spent Too Much Time Making My AI Agent Smarter
A lot of AI engineering is not about finding a better model. Sometimes the real problem is that the system around the model is badly designed.
I used to think the answer was simple.
If an AI feature produced bad output, I needed a better model.
So I changed models. I rewrote prompts. I added more context. I added more examples.
The output got better. Six weeks later I had a system that was harder to understand and still failed in annoying ways.
That is when I started looking at the problem differently.
The model was not always the problem
A model only sees what I give it.
If retrieval is bad, the model gets bad information. If the context is too large, the important parts are harder to find. If the instructions conflict, the model has to guess which one matters. If nothing checks the output, a confident mistake reaches the user.
I stopped asking which model I should use. I started asking where the system loses quality.
That second question is more useful.
An AI feature is a sequence of decisions
A production AI feature is usually a sequence.
Input, then retrieval, then context, then the model, then tools, then validation, then output.
The model is one step in that sequence.
This sounds obvious. It changes how I debug.
When something goes wrong, I look for the first bad state instead of staring at the final answer.
Was the retrieved document wrong?
Was the right document retrieved but ranked badly?
Did the prompt drop an important constraint?
Did the model call the wrong tool?
Did the result fail validation?
That gives me something concrete to fix.
A longer prompt is not a substitute for code
Prompting matters.
I do not want a 2,000-line prompt that tries to hold every rule. That is another piece of software, and it is harder to test.
I prefer smaller prompts with explicit inputs, one clear job, and checks around them.
If a rule must not be broken, I enforce it in code when I can.
Models are useful for judgment. Code is better at hard constraints.
What I added
I started putting three things around AI features.
- Structured inputs and outputs.
- A check after generation.
- Logs that show the actual path of decisions.
The third one has mattered more than I expected.
Without those logs, debugging an AI feature is guessing. With them, I can see that the model was fine and retrieval sent the wrong five documents.
That is a different engineering problem.
The useful changes have been ordinary
The most useful AI improvements I have made lately did not come from a clever prompt.
They came from better retrieval, smaller context, clearer contracts between steps, evaluation sets, retries, validation, and logs.
None of that makes a good demo. It makes a system I can trust a little more.
That is what I care about now.
I still try new models. I do not assume a model upgrade is the answer.