What I Actually Look at When an AI Feature Feels Bad
When an AI feature feels inconsistent, I no longer immediately rewrite the prompt. I break the problem into measurable parts.
When an AI feature feels bad, the usual sequence is short.
The output is bad. I change the prompt. The output is different. I change the prompt again. After a while I cannot tell whether anything improved.
I have done this. I do not start there anymore. I try to make the problem measurable first.
1. What did the model actually receive?
This is the first thing I check.
The prompt in the source code is not always the prompt the model received.
Retrieval may have added the wrong context. A previous message may be missing. A template may have injected something I did not expect. The context may have grown too large.
Before I change the model, I want to see the actual input.
2. Did the model understand the task?
A bad result does not always mean a bad model. Sometimes the task is ambiguous.
"Write a good post" means very little.
"Write a 120-word LinkedIn post for a technical founder, explain X using one concrete example, avoid generic motivational language, and end with a practical takeaway" is much easier to judge.
Specific inputs make better tests.
3. Did it follow the important constraints?
This is where structured output helps.
If I need JSON, I want JSON. If I need five suggestions, I want five. If a field is required, I validate it.
I do not want to find these problems after the result reaches the user.
4. Is the problem consistent?
This one matters a lot.
If the same input fails every time, I probably have a system problem. If it fails one time out of ten, I have a reliability problem.
Those need different fixes. One needs better logic. The other may need retries, a better model, better context, or a different workflow.
5. Can I measure improvement?
This is the part I ignored early on.
I would change a prompt and decide it felt better. That is not a test.
Now I prefer a small evaluation set. I run the old version. I run the new version. I compare them on the same cases.
The evaluation does not need to be perfect. It needs to be consistent enough to catch obvious regressions.
What changed for me
I stopped treating the output as something I could not inspect.
It is still probabilistic. The system around it does not have to be.
I can make inputs explicit. I can validate outputs. I can log decisions. I can build evaluation sets. I can measure latency and cost. I can retry failures. I can put checks around risky actions.
Once I do that, this work looks like the rest of engineering. That makes it easier to change without guessing.