all writing
Jun 11, 2026#ai#agents#engineering#production#llm

How to build production-ready AI agents: Beyond the deterministic loop

Most agent tutorials stop at a while loop that calls a model and some tools. That is a demo. I care about retrieval, a runtime that can resume and stop, evals, and limits on what a run can do.

Most tutorials on building an AI agent show the same thing: a while loop, a model call, a tool registry, and a parser. That gets me a demo. It does not get me something I would put in front of users.

The gap is the systems work after the loop. This is the part I actually have to build.

1. The loop is the easy 10%

Every agent I have seen comes down to this:

while (!done) {
  const step = await model.generate({ messages, tools });
  if (step.toolCalls) {
    const results = await Promise.all(step.toolCalls.map(runTool));
    messages.push(step, ...results);
  } else {
    done = true;
  }
}

If I stop here, I have a chatbot that sometimes calls a function. The other 90%, the part that decides whether the agent can ship, is everything around that loop.

2. Retrieval is a pipeline, not one vector search

"Embed the docs and do cosine similarity" is the most common reason I have seen an agent look bad in production.

A real retrieval pipeline has at least the pieces below.

Chunking that respects structure

Split on real boundaries, such as headings, function definitions, and paragraphs. Do not split only on a fixed token window.

Keep the path in metadata, from the document to the section to the subsection, so the model knows where a chunk came from.

Hybrid retrieval

Combine dense search (embeddings) with sparse search (BM25 or Postgres full text).

Dense search catches a paraphrase. Sparse search catches exact identifiers and rare tokens.

Take the union, then rerank.

Reranking

A cross-encoder reranker over the top 50 candidates beats a bigger embedding model over the top 5. It is the change that helps most, and it is the one most teams skip.

Query rewriting

The user's question is rarely the right query. I have the agent rewrite it before retrieval: expand acronyms, add synonyms, and split a question that is really several questions.

Freshness and authority

Boost recent docs. Lower the rank of deprecated ones. If the corpus has versions, filter to the right version before similarity.

The agent loop calls this pipeline like any other tool. The pipeline is what decides whether the answer is any good.

3. The runtime matters more than the framework

Frameworks like LangChain or the AI SDK give primitives. The runtime, the part that owns execution, is still something I have to build.

Step persistence

Write every model call, tool call, and tool result to durable storage before the next step. Crashes happen. Resuming from step 7 of 12 is the difference between an annoyance and a broken run.

Idempotency keys on tool calls

If a step retries, the payment must not charge twice and the email must not send twice. Hash the run id, step id, tool name, and arguments, and reject a duplicate at the tool boundary.

Cancellation

Users close tabs. A long-running agent has to see an abort signal at every await, pass it into tool calls, and close streams.

Budget limits

I want hard caps on steps, tokens, wall-clock time, and dollars per run.

Warn before the cap. Without a cap, one bad prompt can spend the account's budget.

Structured logging

Log every step as a row, not a blob:

{
  run_id,
  step_index,
  model,
  latency_ms,
  input_tokens,
  output_tokens,
  tool_name,
  error_kind
}

I expect to query this often.

4. Treat tools like an API

The fastest way I know to make an agent unreliable is to give it 40 tools with vague descriptions. I treat the tool list like a public API.

Narrow input schemas

Use Zod or the same kind of schema. Require the fields that matter. Prefer enums over free strings. Do not use any.

The schema does half the work of the prompt.

Small, structured results

Return the smallest JSON the next step needs. Trim long arrays, paginate, and summarize.

Every token I return is a token the model has to read again on later steps.

Explicit failures

{ ok: false, reason: "not_found" }

That is better than throwing. The model can recover from a structured error. It cannot recover from a stack trace.

Approval before writes

Anything that spends money, sends a message, or writes to a system of record needs a person to approve it. At minimum, the agent should be able to call a dry run first.

5. Evals are how I know it works

I cannot tell whether an agent works by clicking around. I need an eval harness.

A frozen test set

Real user queries, with an expected outcome. That can be a final answer, the tools that should be called, or a rubric.

Replay that does not drift

Same inputs, same seed, same tool stubs, same trace. Without that, I cannot tell whether a regression is my code or the model.

A model as judge, checked by a person

For open-ended output I use a second model as a judge, and I check that judge against human labels on a sample.

Run evals on every change

I run them on every prompt change, every model swap, and every tool edit.

I tie the score to CI. A 4% drop on the gold set should block the deploy.

6. I want the trace, not only the log line

When an agent fails, the question that matters is: what did the model see at step 6?

I want trace-first logging, whether I build it or adopt it.

One trace per run

One span per step, with the inputs, outputs, tool calls, and token counts attached.

Searchable by the fields I actually filter on

user id, tool name, error kind, model, and latency bucket.

Sampling

Sample to control cost. Always keep errors, and always keep runs a user flagged.

Dashboards of totals come second. The trace view is where I spend the time.

7. What I want in place before real users

  • Step-level persistence, with a way to resume
  • Idempotent tool execution
  • A cap per run on steps, tokens, dollars, and wall-clock time
  • Streaming responses, with cancellation
  • A retrieval pipeline that is hybrid and then reranked
  • An eval harness wired into CI
  • Traces I can open per run
  • Rate limits per user and per tool
  • Approval before a destructive tool runs
  • A fallback model when a provider is down
  • Redaction of personal data on inputs and on logged traces

If any of these are missing, I still have a prototype.

What I take from this

Building an agent loop is easy. Building one that keeps working for real users, real traffic, and real failures is a systems problem. The model is one component.

The work I care about is retrieval, the runtime, the tools, evals, traces, and the infrastructure around them.

That is the work. It is ordinary. It is also what separates a demo from something I can leave running.

Further reading