Blog
Backoff and jitter aren’t enough when overload control meets the database
Do backoff and jitter really stop overload? See how circuit breakers and predictable storage latency close the gap.
Blog
Do backoff and jitter really stop overload? See how circuit breakers and predictable storage latency close the gap.
When a distributed system starts to struggle under load, retries can quickly turn a difficult situation into a much larger set of problems. A request times out; the client retries; other requests hit the same issue; and the system ends up spending real capacity on work that only exists because earlier work failed.
Exponential backoff and jitter are the standard first-line defenses. Backoff makes clients wait progressively longer between attempts; jitter randomizes that wait so thousands of clients don't retry in lockstep. Together, they stop a struggling system from being hit by a synchronized wall of retries. However, they don’t determine how much work the system can safely handle. Backoff and jitter can make retries less frequent and less synchronized, but they don’t eliminate the underlying demand. If a service can sustainably process 10,000 operations per second while clients collectively generate 15,000 operations per second, spreading those operations more evenly doesn’t change the fundamental imbalance. The clients may be perfectly coordinated, but the system is still receiving more work than it can process.
Overload control relies on mechanisms such as admission control, load shedding, and the feedback loops that connect them. Rather than asking only, “When should a client retry?” the system also needs to ask, “Given what I’m seeing right now, how much work can I safely accept?” A retry policy follows rules defined in advance, while a control loop continuously adjusts to the system’s changing conditions.
The distinction applies across distributed systems, regardless of the technology stack. What matters is how well the control loop responds as conditions change; if feedback arrives too late or the system is responding to the wrong signal, demand can continue to outpace available capacity even when the control mechanism itself is functioning as designed.
An admission controller is only as effective as the signal it uses to make decisions. Whether that signal is queue depth, tail latency, rejection rate, resource utilization, or something else, the control loop depends on it accurately reflecting the system’s current condition. When the signal indicates that the system is approaching its limit, the controller needs to reduce or reject incoming work; as capacity becomes available, it can begin admitting more work. The goal is to keep demand aligned with the amount of work the system can safely handle at any given moment.
This means that the signal must stay legible precisely when it matters most: under load, as the system approaches its limit. A signal that only tells the truth under comfortable conditions and degrades exactly when pressure rises is adding to the overload rather than measuring it, as the controller is making admission decisions off noise it cannot distinguish from a real signal.
When the admission control signal relies on conditions deep in the stack, the database can significantly influence what that signal reports. Database latency contributes to the time a request spends in the system, queueing can indicate growing pressure, and resource contention can make individual requests increasingly expensive. As those conditions change, the admission controller is making decisions based on signals that may originate well below the layer where those decisions are made.
Database behavior can change significantly as load increases, making the signals an admission controller relies on less reliable precisely when they matter most. A database whose latency is largely driven by a warm cache may appear predictable until the cache-miss rate changes and a portion of requests suddenly takes a much slower path. That shift can distort the tail latency signal the admission controller is watching, making it harder to distinguish a system approaching saturation from one simply experiencing a change in workload.
The same dynamic can occur when internal database mechanisms add work as pressure increases. Unbounded replication queues or cascading internal retries can cause the database to generate additional load on top of the load it is already handling. At this point, the system is no longer simply reporting congestion to the layer above it; its own behavior is contributing to the congestion, creating another feedback loop that can accelerate the path toward overload. And rather than surfacing as a bug, this shows up as an admission controller that looks well-designed on paper, yet still lets a peak event through because the database quietly changed the meaning of "current condition" underneath it.
In a benchmark conducted by McKnight Consulting Group comparing Redis and Aerospike, both systems were held to fixed throughput targets: one comfortably within capacity, one near it, one slightly beyond it. At the same time, two throughput metrics were tracked: the dependable rate (the rate the system actually holds, not the average) and rate jitter (how much that rate moves around, run to run).
At a 2.4 million operations-per-second target, Redis's dependable rate fell to 2.10 million with a rate jitter of 4.4%, and its p99.99 latency reached 9.19 ms against a 0.66 ms median, more than six milliseconds of hidden tail that only shows up once load is pushed toward the edge. Run in memory on the same hardware and workload, Aerospike maintained a dependable rate of 2.39 million at that same target, with rate jitter never rising above 0.2%; run in its durable, disk-backed configuration, it maintained 2.33 million with jitter at or below 1.3%.
At the throughput level where Redis’ signal begins to degrade, the difference in behavior becomes more apparent. Redis shows increasing jitter and tail latency, roughly six milliseconds above its median, while Aerospike maintains a tighter relationship between the two. An admission controller built on top of Redis is therefore making decisions based on a signal that has already become less consistent as the system approaches its limits. With Aerospike, the signal remains more predictable as throughput increases.
This difference is rooted in the write path rather than in tuning alone. Aerospike applies backpressure to keep replication queues short and limit the conditions that can trigger retry cascades, while placing explicit bounds on the I/O and network operations that can otherwise become sources of tail latency outliers. The result is overload control embedded in the data path itself, operating one layer below where admission control is typically applied.
The same gap becomes especially visible during peak demand. Teams don’t necessarily lose confidence in their admission-control logic; they lose confidence in the layers underneath it when those systems behave differently under sustained pressure. A database that performs predictably at 60% utilization but degrades sharply at 90% can undermine every control mechanism built on top of it, precisely when that control matters most.
This is why peak-readiness testing needs to go beyond merely validating whether an overload mechanism correctly rejects or admits work. It also needs to establish that the underlying system maintains predictable behavior as utilization rises, workload mix changes, access patterns shift, and individual nodes fail.
Despite these limitations, admission control, load shedding, and backoff and jitter each address an important part of the overload problem and remain essential to a resilient system. However, the database should not be treated as a passive dependency that overload control simply protects from the outside. If the layer generating the control loop’s signal begins to degrade under the same conditions the loop is designed to manage, the control system itself becomes less reliable. You may have an overload-control mechanism in place, but its decisions are only as sound as the signal informing them.
The question, then, is not only where the system rejects work when it becomes overloaded. It is whether the systems underneath the admission controller continue to provide reliable signals as load approaches its limits. If those signals become less predictable under pressure, the control loop can end up making its most important decisions based on its least reliable information.
For a deeper understanding and more insights, explore these additional resources.
See moreBlog
Blog
Blog
Blog