Bounded Retries
Examine how an agent should respond when an operation keeps failing and recovery may or may not be safe.
Original scenario-based practice challenge for Agentic Architect Lab. This is not an official certification question.
Scenario
An operation is failing repeatedly. How should the agent decide between retrying, stopping, and escalating?
A deployment agent provisions a search index as part of a release. Some failures are transient, such as rate limits or short-lived network interruptions. Other failures are permanent, such as an invalid schema or missing required configuration. The platform team wants the agent to recover automatically when it is reasonable to do so, but to stop causing noise and escalate once further retries are unlikely to help.
Which design best handles retries under those constraints?