← blog
JavaScriptOctober 7, 2026 · 6 min read

Stop Retrying Immediately. Exponential Backoff Fixes It.

Retrying a failed fetch() right away feels responsible — until the server that failed was already overloaded, and five instant retries per client turn a blip into an outage. Exponential backoff with jitter is the fix.

Parsa Jiravand · Frontend engineer · building bestpractic
Stop Retrying Immediately. Exponential Backoff Fixes It.

The API call fails. Your retry logic kicks in — of course it does, you wrote it for exactly this. Five attempts, back to back, as fast as the event loop allows. The first one failed in 40ms. By 200ms you've fired all five.

That's the whole bug. Not that you retried — that you retried immediately, which is the one thing a failing server can least afford.

Here's the version almost everyone writes first:

JavaScript
1
2
3
4
5
6
7
8
9
10
11
async function fetchWithRetry(url, retries = 5) { for (let i = 0; i < retries; i++) { try { const res = await fetch(url); if (res.ok) return res; } catch { // network error — fall through and try again } } throw new Error(`Failed after ${retries} attempts: ${url}`); }

It reads like good engineering. A flaky request gets a second chance, then a third, up to five, before you give up and surface an error. Test it against a healthy endpoint that you fail on purpose for one call, and it works perfectly — the second attempt succeeds, nobody notices.

The test that never runs is the one where the server is the problem. A deploy goes out with a bad config, a dependency times out, a traffic spike pushes response times past your fetch timeout. Now every client hitting that endpoint gets a failure — and every client's retry loop fires its five attempts in the same few hundred milliseconds. The server that was struggling under normal load now gets 5x the request volume, instantly, from every caller that was already in flight. That's not recovery. That's a second, self-inflicted spike landing directly on top of the first one.

This has a name — a retry storm — and it's a well-documented failure mode in distributed systems, not a hypothetical. The fix isn't "retry less." It's "retry differently."

Three things, and none of them are exotic:

  • A growing delay between attempts, so a struggling server gets breathing room instead of an immediate second hit.
  • Randomness in that delay, so thousands of clients that all failed at the same moment don't all retry at the same moment too.
  • A reason not to retry everything — a 404 will fail the same way on attempt five as attempt one. Retrying it five times doesn't help; it just makes you look five times slower at reporting the real error.

Put those together and you get exponential backoff with jitter: each retry waits roughly twice as long as the last, capped at some maximum, and the exact wait is randomized within that window instead of fixed.

JavaScript
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
async function fetchWithRetry(url, { retries = 4, baseDelay = 300, maxDelay = 8000 } = {}) { for (let attempt = 0; ; attempt++) { let res; try { res = await fetch(url, { signal: AbortSignal.timeout(5000) }); } catch (err) { if (attempt === retries) throw err; await backoff(attempt, baseDelay, maxDelay); continue; } if (res.ok) return res; if (res.status < 500 && res.status !== 429) return res; // a client error won't fix itself if (attempt === retries) return res; await backoff(attempt, baseDelay, maxDelay, res.headers.get("Retry-After")); } } function backoff(attempt, base, cap, retryAfter) { if (retryAfter) { const seconds = Number(retryAfter); const ms = Number.isFinite(seconds) ? seconds * 1000 : new Date(retryAfter).getTime() - Date.now(); return sleep(Math.max(0, ms)); } const windowMs = Math.min(cap, base * 2 ** attempt); return sleep(Math.random() * windowMs); // "full jitter" — pick anywhere in the window, not the edge } const sleep = (ms) => new Promise((r) => setTimeout(r, ms));

Walk through what changed, because each line is answering one of the three gaps above:

  • AbortSignal.timeout(5000) caps how long a single attempt can hang before it counts as a failure and moves on to the backoff — without this, one slow attempt can eat your whole retry budget just sitting in fetch's pending state.
  • res.status < 500 && res.status !== 429 stops the loop from wasting attempts on a 400 or 404 — those are the server telling you the request itself is wrong, and no amount of waiting changes that. 429 (rate limited) and 5xx (server trouble) are the ones worth retrying.
  • Retry-After is a real HTTP response header — servers send it on 429 and 503 to tell clients how long to wait before trying again (it also shows up on some 3xx redirects, for a different reason) — as a number of seconds, or an HTTP date. When it's there, use it instead of guessing; the server is telling you its own recovery estimate.
  • Math.random() * windowMs is "full jitter": instead of every client waiting exactly 300ms, then exactly 600ms, then exactly 1200ms — which just re-synchronizes everyone's retries into new, smaller storms — each client picks a random point inside that window. Spread the same total number of retries over time instead of in lockstep, and the server sees a trickle instead of a wave.

Runs right in your browser — poke at it and watch the concept react live.

One more thing the naive loop glossed over: it retried everything, including requests that aren't safe to repeat. A GET is idempotent — running it five times has the same effect as running it once. A POST that creates an order is not. If that request actually reached the server and created the order, but the response got lost on the way back, retrying blindly creates a second order.

The honest fix is smaller in scope than it sounds: only wrap genuinely idempotent requests (GET, HEAD, a PUT that fully replaces a resource) in automatic retry by default. For a POST you need retried — payment confirmation, for instance — the real fix is an idempotency key: a unique ID you generate once and send with every attempt, so the server can recognize "I've already done this one" and return the original result instead of doing it twice. That's a server-side contract, not something fetchWithRetry can fake on its own — which is exactly why it's worth calling out instead of quietly auto-retrying a POST and hoping.

None of this makes requests succeed that were always going to fail — a backoff loop retrying against a server that's down for an hour will still, eventually, give up. What it changes is the shape of the load your retries put back on a server that's recovering: spread out instead of synchronized, bounded instead of infinite, and aimed only at the failures that stand a chance of succeeding on a second try.

The three-line fix really is three lines — a growing delay, a random window inside it, and a check for which status codes deserve a second attempt. Getting there meant admitting the first version wasn't "retrying too little" or "retrying too much." It was retrying at exactly the wrong tempo.

What's your current retry count, and have you ever actually watched what it does to a server that's already struggling?

Instant feedback, a hint on every question, and an explanation for each answer — right or wrong.


🚀 Want more like this? Every guide, playground, and quiz lives on bestpractic.org — open it and sign up free so the next one finds you.

Thanks for reading! Let's stay connected:

Keep reading

One post a day, in your inbox

Each one with a runnable playground and a quiz. No pitch, no digest, unsubscribe in one click.

0 comments

Sign in to join the discussion, like comments, and save articles for later.