← ~/blog

Why Your Retry Logic Is a DDoS Cannon Pointed at Yourself

 /  systems  /  301 words

The service went down for two minutes. The outage lasted forty. The difference between those two numbers was our own retry logic.

Here is the mechanism. A downstream dependency blips. Every caller times out and retries, immediately, three times. Traffic to the struggling service just tripled at the exact moment it could least afford it. It falls over harder. More timeouts, more retries. By the time the original blip resolved, we had built a standing wave of our own traffic that kept the service pinned. Recovery required manually shedding load, which is a fancy way of saying we turned things off until the screaming stopped.

Retries are not free. Every retry is a bet that the failure was transient and cheap, and when the failure is actually overload, the bet doubles down on the exact wrong thing.

The fixes are old and known and still skipped in most codebases I read. Exponential backoff so retry two waits longer than retry one. Jitter, meaning randomness added to the wait, because without it every client that failed together retries together, and you get synchronized thundering herds arriving in neat waves. Cap the total retries. And most important, a retry budget: if more than say ten percent of recent requests are retries, stop retrying entirely and fail fast, because at that point retries are the attack.

One more subtle one. Retry at one layer only. We had retries in the HTTP client, in the service mesh, and in the application logic. Three layers of three retries is 27 attempts per user click. Nobody designed that. It emerged, like mold.

Your retry policy is load bearing infrastructure. Write it down, test it under simulated failure, and assume the service you are retrying against is one bad minute away from being a service you are attacking.