all writing
Oct 9, 2026#ai#agents#engineering#production#reliability

How I handle retries in an AI agent without doing things twice

Retries are where agents cause real damage. I persist every step, put an idempotency key on every write, and only retry what is safe to retry.

Retries are where an AI agent does real damage. A model call that fails is cheap to try again. A tool call that sent an email, charged a card, or published a post is not.

This is how I set up retries in an agent runtime so a failure costs time, not a duplicate action.

1. Split steps into safe and unsafe

Every step falls into one of two groups.

  • Safe to repeat. A model call, a search, a read from the database. Running it twice wastes tokens and nothing else.
  • Not safe to repeat. Anything that writes outside the run. Sending, paying, publishing, updating a record.

I retry the first group freely, with backoff. The second group never retries without a check.

2. Save each step before the next one starts

Before the agent moves on, I write the step to storage: the input, the output, the tool name, and the status.

If the process dies at step 7, the run resumes at step 7. It does not start over and repeat steps 1 to 6. Without this, "retry the run" means "repeat every write in the run."

3. Put an idempotency key on every write

For every unsafe tool call, I build a key from the run id, the step number, the tool name, and the arguments.

The tool checks that key before it acts. If it has seen the key, it returns the saved result and does nothing new.

That check lives at the tool, not in the prompt. A model cannot be trusted to remember it already sent something.

4. Know the difference between failed and unknown

A timeout is the dangerous case. The request may have worked. The response just never came back.

I treat that as unknown, not failed. Before retrying, the tool asks the outside service whether the action happened. If it cannot check, the step stops and waits for a person.

5. Cap the retries

Each step gets a small retry limit. Each run gets caps on steps, tokens, time, and money.

An agent stuck in a retry loop is still spending. A cap turns a bug into an error I can read.

6. Log it so I can see it

Every retry gets its own row: the attempt number, the error type, and what decided the next move. When a user asks why something happened twice, or did not happen at all, I want the answer in one query.

What I do next

Before an agent gets a new write tool, I ask one question: what happens if this call runs twice? If the answer is not "nothing," the tool needs a key and a check before it ships.