When a step fails, replan — don't restart. The difference between an agent that finishes and one that burns tokens in circles is what survives a mistake.

Two agent patterns look similar and behave nothing alike under failure. In the simpler loop, the agent thinks, acts, observes — and when the observation says the attempt failed, it begins again from the top. Everything already established is re-derived, every prior step repeated, and the cost of a single wrong turn is the whole task so far.
The alternative splits the roles. A planner produces a sequence of steps; an executor runs them; when a step fails, only the plan goes back for revision. The work already completed stays completed. This sounds like a minor architectural preference and is actually the difference between an agent that converges on long tasks and one that thrashes — because on any task of real length, failure is not exceptional, it is the normal case that has to be absorbed rather than avoided.
The generalisation is worth stating: recovery cost, not reasoning quality, is what makes long-horizon agent work viable. A slightly weaker model that loses nothing when it stumbles will outperform a stronger one that restarts, because the number of stumbles grows with task length while the value of raw capability does not. Designing for the failure path is designing for the common path.
Which is why the surrounding controls matter more than they look. State that persists across cycles, checkpoints you can roll back to, a retry cap so a loop cannot spiral, and an explicit stop rule. None of it is intelligence; all of it determines whether intelligence gets to finish. The model generates — the loop governs, and governance here means preserving what was already earned.