Skip to main content

retry

The step-level retry field automatically re-runs a failed step using a configurable backoff strategy, making workflows resilient to transient failures.

steps:
- command: terraform apply vpc -auto-approve
retry:
max_attempts: 3
backoff_strategy: exponential
initial_delay: 5s
max_delay: 60s

Fields​

max_attempts
Maximum number of attempts. Omit for unlimited attempts.
max_elapsed_time
Total time budget for all attempts and backoff delays combined.
backoff_strategy
Delay strategy: constant, linear, or exponential. Defaults to constant.
initial_delay
Delay before the first retry.
max_delay
Maximum delay between retries.
multiplier
Multiplier for linear or exponential backoff.
random_jitter
Jitter factor from 0.0 to 1.0 added to delays.
conditions
List of regular expressions identifying which failures are retryable. For the http step type, patterns are matched against "<status-code> <response-body>", letting you retry responses that would otherwise fail fast (e.g., a specific 4xx). Omit to use each step type's default retry classification.

Omitting a parameter disables that limit or uses the default behavior. Explicit zero values are invalid where a positive duration or count is required.

Retries and timeout​

For shell, atmos, script, and tflint steps, the step's timeout bounds the whole step, not each attempt. One deadline covers every attempt and every backoff wait. When it elapses, Atmos cancels the running attempt, starts no further attempts, and fails the step with step timed out. Use max_elapsed_time to cap retries on their own. The http step type is different: its timeout is a per-attempt request limit.