# Retries and Backoff: Recover Without Amplifying Failure (/docs/distributed-systems/retries-and-backoff)



# Retries and Backoff: Recover Without Amplifying Failure [#retries-and-backoff-recover-without-amplifying-failure]

## TL;DR [#tldr]

Picture a congested highway exit during the evening rush hour. A minor fender bender slows traffic to a crawl. Instead of yielding and merging smoothly, every driver behind the jam floors the accelerator, swerves around, and rams back into the merge lane simultaneously every two seconds on the dot. The result is catastrophic gridlock: a brief, manageable slowdown instantly detonates into an impassable, multi-mile pileup. In distributed systems, this exact failure mode is known as a **retry storm**—where well-intentioned automated retries weaponize client traffic and DDoS an already struggling downstream service straight into an unrecoverable outage.

> 💡 &#x2A;*Rule of thumb:** Retry only transient failures on idempotent operations within the remaining deadline; always apply exponential backoff with full jitter and designate a single retry owner.

* **Classify before retrying:** Retries are only effective against **transient** transport blips, rate-limiting throttle responses (`429`), or temporary server overload (`503`). Deterministic errors (malformed input, invalid credentials, business logic violations) will never self-heal and must never be retried.
* **Idempotency is non-negotiable for state mutations:** A network retry is a brand new attempt, not an undo action. Retrying an ambiguous write without an **idempotency key** or reconciliation path will create duplicate charges, double-booked seats, and ghost inventory entries.
* **Exponential backoff with full jitter:** Immediate retries batter degraded services. Space successive attempts with **exponential backoff**, and always inject randomized **jitter** to break resonance so an entire fleet of clients never retries in locked phase.
* **Single retry owner:** Establish clear **retry ownership** at a single architectural boundary. If the client, API gateway, and intermediate microservice each independently configure 3 retries, a single failed query multiplies into `3 × 3 × 3 = 27` destructive attempts on the deepest database layer.
* **Fatal pitfall:** &#x2A;*Uncoordinated multi-layer retries with fixed intervals.** Nesting un-jittered retry loops across upstream tiers creates an exponential feedback loop that transforms a 200 ms downstream latency fluctuation into a total cascading outage.

<Mermaid
  chart="flowchart LR
  F[Failure] --> C{Transient + safe?}
  C -->|No| S[Stop / reconcile]
  C -->|Yes| B{Budget remains?}
  B -->|No| S
  B -->|Yes| R[Backoff + jitter -> retry]"
/>

<TermBox term="Retryability">
  A failure is **retryable** when a later attempt can change the outcome. Temporary transport, throttling, or availability failures can qualify; invalid input, bad credentials, and stable business rejection are usually deterministic.
</TermBox>

## Classify before retrying [#classify-before-retrying]

| Signal                                 | Default decision                      |
| -------------------------------------- | ------------------------------------- |
| transient transport error              | retry if safety and budget hold       |
| `429` / `503`                          | back off; honor valid `Retry-After`   |
| auth / validation / business rejection | do not retry unchanged input          |
| ambiguous write timeout                | require idempotency or reconciliation |

<TermBox term="Retry Budget">
  A **retry budget** bounds extra work by attempt count, elapsed time, remaining deadline, or shared quota.
</TermBox>

## Make writes safe first [#make-writes-safe-first]

If attempt one may have committed, a new logical request can duplicate a payment or reservation. Reuse the same idempotency identity or reconcile state first. Retry does not create idempotency.

## Back off, jitter, stay inside the deadline [#back-off-jitter-stay-inside-the-deadline]

```text
100 ms -> 200 ms -> 400 ms -> capped delay
```

**Exponential backoff** spaces attempts farther apart. **Jitter** randomizes waits so a fleet does not wake together. Bound attempts and delay, and never sleep past the remaining deadline.

<Mermaid
  chart="sequenceDiagram
  participant A as Client A
  participant B as Client B
  participant S as Service
  S--xA: failure
  S--xB: failure
  A->>A: jitter 120 ms
  B->>B: jitter 310 ms
  A->>S: retry
  B->>S: retry"
/>

## Choose one retry owner [#choose-one-retry-owner]

If Client, Service A, and Service B each allow three total attempts, one original request can drive `3 × 3 × 3 = 27` attempts at the deepest dependency. Pick the layer with context to judge value, safety, and remaining deadline. Existing SDK retries count too.

<TermBox term="Retry Storm">
  A **retry storm** is overload amplified by retries: recovery traffic becomes part of the incident.
</TermBox>

## Protect the fleet [#protect-the-fleet]

During broad failure, a shared **retry budget** or **token bucket** can suppress retries. AWS SDK standard mode uses a token-backed retry quota. Observe original requests separately from attempts, retry reason, exhausted budget, and saturation.

## Production scenario [#production-scenario]

A gateway calls Checkout → Inventory → database. During a latency spike every layer retries twice, uses deterministic backoff, and reservation writes lack stable idempotency identity.

**Impact:** one user request can drive up to 27 database attempts; synchronized retries create bursts, saturation worsens, and ambiguous writes can duplicate reservations.

**Root cause:** retry ownership was spread across layers, failures were not classified, writes were unsafe to repeat, and timing synchronized the fleet.

**Correct pattern:** choose one retry owner, retry only transient failures, preserve one logical request identity, bound attempts inside the remaining deadline, use exponential backoff with jitter, honor valid `Retry-After`, and stop when the shared retry budget is depleted. Measure original requests separately from attempts.

## Self-check [#self-check]

Your SDK and service wrapper already retry. A dependency starts returning `503`. Should the gateway add another retry loop?

<details>
  <summary>
    Show the reasoning
  </summary>

  No. Recover existing retry layers, choose one owner, and coordinate the total attempt/deadline budget before changing policy.
</details>

## Retry operating checklist [#retry-operating-checklist]

* [ ] Retry transient failures, not deterministic ones.
* [ ] Make ambiguous writes idempotent or reconcile first.
* [ ] Fit retries inside the remaining deadline.
* [ ] Assign one owner and include SDK retries.
* [ ] Cap backoff, add jitter, respect `Retry-After`.
* [ ] Bound fleet retries and observe request vs attempt volume.

## Agent rule [#agent-rule]

Before adding retries, recover failure class, operation safety, deadline, existing retries, ownership, bounds, backoff/jitter, server hints, fleet budget, and telemetry. Reject uncoordinated retry loops.

## Sources [#sources]

Primary references verified on **2026-09-15**: [RFC 9110](https://www.rfc-editor.org/rfc/rfc9110.html), [AWS retry guidance](https://docs.aws.amazon.com/wellarchitected/latest/reliability-pillar/rel_mitigate_interaction_failure_limit_retries.html), and [AWS SDK retry behavior](https://docs.aws.amazon.com/sdkref/latest/guide/feature-retry-behavior.html).
