Engineering

Retry Storms and Idempotency: Integration Failure Patterns

A retry added to mask an intermittent failure is how a small problem becomes an outage. The mechanics of retry amplification, and the idempotency guarantees that make retries safe.

Retry Storms and Idempotency: Integration Failure Patterns

A retry storm occurs when retry logic amplifies load against a system that is already degraded, converting partial failure into total failure. Retries are only safe when paired with exponential backoff, jitter, a circuit breaker, and idempotency keys that make duplicate delivery harmless.

Key takeaways
  • Naive retries are positive feedback on a degrading system.
  • Backoff without jitter synchronises clients and makes the problem worse.
  • Without idempotency keys, a successful retry of a request that actually succeeded is a duplicate, not a recovery.

How a retry becomes an outage

A dependency slows down. Requests start timing out. Your client retries each one three times. You have now tripled the load on a system that was already struggling, which increases timeouts, which triggers more retries. The dependency goes from degraded to unavailable, and your retry logic is the reason.

This is the amplification problem, and it is worse in integrations than in internal services because the dependency is shared across all of that provider's customers, all of whom are retrying simultaneously.

The four requirements for safe retries

Exponential backoff

Fixed-interval retries do not reduce pressure. Each attempt should wait substantially longer than the last, with a hard ceiling on total attempts and total elapsed time.

Jitter

The requirement most often skipped. Without randomisation, every client that failed at the same moment retries at the same moment, producing synchronised waves. Randomising the delay across the backoff interval spreads the load and is a one-line change.

A circuit breaker

Backoff limits a single caller's pressure; it does not stop you calling a dependency that is comprehensively down. After a threshold of consecutive failures, stop calling entirely for a cooling period, then probe with a single request before resuming. This converts a slow cascading failure into a fast, contained one.

Idempotency

The correctness requirement, and the one that makes the other three safe. If a request times out, you do not know whether it was processed. Retrying without an idempotency key risks a duplicate: a second charge, a second record, a second notification.

  • Where the provider supports idempotency keys, generate one per logical operation and reuse it across every retry of that operation. Do not generate a fresh key per attempt, which defeats the mechanism entirely.
  • Where it does not, make the operation naturally idempotent: upsert on a stable business key rather than insert, and check for existence before creating.
The mask problem

The most damaging retry is the one added deliberately to make an intermittent failure go away. It works, the error rate falls to zero, and the underlying fault continues: now consuming three times the request budget and invisible to monitoring. Months later the fault resurfaces with distorted symptoms. Record every retry added as a mitigation on the integration's timeline, with an explicit note that it is still in place.

What not to retry

Retrying the wrong error class wastes quota and delays the real fix:

  • 4xx client errors other than 429, retrying a malformed request produces the same malformed request.
  • 401 and 403: refresh credentials, do not retry blindly. See token expiry.
  • 429: retry, but honour Retry-After rather than applying your own backoff. Ignoring it is how rate limiting escalates to suspension.
  • Non-idempotent operations without a key: do not retry at all; surface the ambiguity instead.

Budgets, not just limits

A mature pattern is a retry budget: cap retries as a proportion of total requests over a rolling window, a few percent is typical, rather than per request. Under broad failure the budget is exhausted quickly and retries stop automatically, which is exactly the behaviour you want and exactly what per-request limits fail to deliver.

Choosing the parameters

Reasonable defaults, which matter more than precision here:

  • Base delay a little above the dependency's typical recovery time. Too short and you are adding load during the window it needs to recover.
  • Multiplier of two. Higher converges on the ceiling too fast to be useful; lower barely reduces pressure.
  • Full jitter. Randomise across the whole interval rather than adding a small random component. Partial jitter leaves clients substantially synchronised.
  • Attempt ceiling of three to five for interactive paths, higher only for background work where latency is not visible.
  • Total elapsed cap as well as an attempt cap. Five attempts with exponential backoff can span minutes, which is not acceptable behind a user request.
  • Retry budget of a few percent of total requests over a rolling window, which is the control that actually stops a storm.

Honour Retry-After where the provider sends it, in preference to all of the above. Ignoring it is the most reliable way to turn rate limiting into suspension.

Frequently asked questions

What is a retry storm?

A failure mode where retry logic amplifies load against an already-degraded dependency, increasing timeouts and triggering further retries until partial failure becomes total failure. It is amplified in integrations because every customer of that provider retries simultaneously.

Why does retry backoff need jitter?

Without randomisation, all clients that failed at the same moment retry at the same moment, producing synchronised load waves that keep the dependency saturated. Randomising delay across the backoff interval spreads the load and is a trivial change.

When is it unsafe to retry a request?

When the operation is not idempotent and the provider offers no idempotency key, because a timed-out request may have succeeded and retrying creates a duplicate. Also for 4xx client errors other than 429, where the same request will fail identically.

Stop rediscovering the same integration failure

Traxivo correlates the signals your tools already produce into one incident timeline, recognises a recurrence as a recurrence, and drafts the follow-up with the evidence attached. Nothing is sent without a named approver.

See how Traxivo works Browse use cases

Related reading