all writing
Oct 3, 2026#ai#engineering#evaluation#agents#production

What I check before an AI feature is production

Before I call an AI feature production, I want a one-sentence task, a cost for a bad answer, a retrieval path I can inspect, ten frozen examples, and a cap on spend.

People ask me what to check before an AI feature goes to real users. I use the same short list.

The model is not the first item.

1. Can I name the task in one sentence?

If I cannot, the feature is not ready. "Help the user" is not a task. "Turn this founder note into a LinkedIn draft and wait for approval" is a task.

A vague task produces a vague prompt, and then a vague eval. I fix the task before I touch the prompt.

2. What does a bad answer cost?

I sort actions into three buckets.

  • Read only. A wrong answer is annoying. The user can ignore it.
  • A draft a person reviews. A wrong answer wastes a minute.
  • A write that spends money, sends a message, or changes a record. A wrong answer is a real incident.

The third bucket does not run without a person. I would rather ship a suggest, review, execute path than an agent that "just does it."

3. Where does the context come from?

If the feature needs facts, I want a retrieval path I can inspect. The query, the chunks, and the reason those chunks won.

If I cannot see that, I am debugging a feeling. Rewriting the prompt will not fix a missing document or a stale one.

4. Do I have a frozen set of examples?

Ten real inputs with an expected outcome is enough to start. I run them when I change the prompt, the model, or a tool.

Without that set, "it feels better" is the only test, and it lies.

5. What are the caps?

Steps, tokens, time, and money per run. One bad prompt should not be able to spend the budget. I also want a way to cancel a run when the user leaves.

What I do next

I write the task, the cost of a bad answer, and the ten examples before I call the feature production. If any of those three is missing, it is still a demo.