all writing
Sep 26, 2026#ai#llm#engineering#evaluation#agents

What I Actually Look at When an AI Feature Feels Bad

When an AI feature feels inconsistent, I no longer immediately rewrite the prompt. I break the problem into measurable parts.

When an AI feature feels bad, the usual sequence is short.

The output is bad. I change the prompt. The output is different. I change the prompt again. After a while I cannot tell whether anything improved.

I have done this. I do not start there anymore. I try to make the problem measurable first.

1. What did the model actually receive?

This is the first thing I check.

The prompt in the source code is not always the prompt the model received.

Retrieval may have added the wrong context. A previous message may be missing. A template may have injected something I did not expect. The context may have grown too large.

Before I change the model, I want to see the actual input.

2. Did the model understand the task?

A bad result does not always mean a bad model. Sometimes the task is ambiguous.

"Write a good post" means very little.

"Write a 120-word LinkedIn post for a technical founder, explain X using one concrete example, avoid generic motivational language, and end with a practical takeaway" is much easier to judge.

Specific inputs make better tests.

3. Did it follow the important constraints?

This is where structured output helps.

If I need JSON, I want JSON. If I need five suggestions, I want five. If a field is required, I validate it.

I do not want to find these problems after the result reaches the user.

4. Is the problem consistent?

This one matters a lot.

If the same input fails every time, I probably have a system problem. If it fails one time out of ten, I have a reliability problem.

Those need different fixes. One needs better logic. The other may need retries, a better model, better context, or a different workflow.

5. Can I measure improvement?

This is the part I ignored early on.

I would change a prompt and decide it felt better. That is not a test.

Now I prefer a small evaluation set. I run the old version. I run the new version. I compare them on the same cases.

The evaluation does not need to be perfect. It needs to be consistent enough to catch obvious regressions.

What changed for me

I stopped treating the output as something I could not inspect.

It is still probabilistic. The system around it does not have to be.

I can make inputs explicit. I can validate outputs. I can log decisions. I can build evaluation sets. I can measure latency and cost. I can retry failures. I can put checks around risky actions.

Once I do that, this work looks like the rest of engineering. That makes it easier to change without guessing.