Circuit Breakers: Fail Fast, Recover Faster, Stop Paging Me
When a dependency gets slow, the naive caller gets slow with it. Every request dutifully waits out its full timeout, threads pile up holding open connections, and a problem in one service becomes a resource exhaustion problem in yours. Slow is contagious in a way that down is not, because down fails fast and slow fails at maximum expense.
A circuit breaker is the machine version of a very tired SRE saying "stop calling them, they're clearly having a moment." It watches your calls to a dependency, and when failures cross a threshold, it trips open and starts rejecting calls immediately, without attempting them. No waiting out timeouts, no thread pileup. Your service stays healthy and serves whatever degraded answer it can, cached data, a default, an honest error.

That log shows the full lifecycle. Five failures in the window trips it open. For 30 seconds everything gets rejected instantly and we serve slightly stale cached inventory, which customers cannot distinguish from fresh. Then half open, the polite knock on the door: three probe requests go through. All succeed, breaker closes, normal service resumes. Total human involvement: zero. That last part is the point. Before breakers, this exact scenario paged someone.
Practical notes from running these. Tune the threshold per dependency, five failures in ten seconds is right for a chatty service and wrong for one you call once a minute. Always define the fallback, a breaker with no fallback just converts slow errors into fast errors, which helps you but not the user. And put a metric on state transitions, because a breaker flapping open and closed all day is telling you something the individual errors were hiding.
The philosophy underneath: your availability should not be the minimum of every dependency's availability. Breakers are how you decouple those numbers.