Backpressure: Controlling Flow in Distributed Systems · Daniel Tinizaray

Backpressure: Controlling Flow in Distributed Systems · Daniel Tinizaray

Every system with queues eventually reaches the point where production outpaces consumption capacity. The difference between a mature system and a fragile one is not whether that moment arrives, but what happens when it does. Without backpressure, overload becomes unbounded memory, cascading latency, and a restart that wipes the problem (and the data). With backpressure, the system says "no" in a controlled way before collapsing.

The problem: the queue that hides overload

A queue is useful for absorbing spikes, but it is also the best mechanism for hiding a capacity problem. If a producer emits 1,000 events per second and the consumer processes 800, the queue grows by 200 events per second indefinitely. While nobody watches the queue depth, everything seems fine: tests pass, dashboards are green, and alerts sleep. Three hours later, memory runs out, messages expire, or processing latency is so high that the result is no longer useful.

The core lesson is simple: an unbounded queue is not a buffer, it is a deferral of failure. Backpressure exists to turn that deferral into an explicit, manageable signal.

What backpressure really is

Backpressure is the propagation of the "I am saturated" signal from the consumer upstream to the producer, so the producer can slow down. It is not a single technique, but a family of combined decisions:

  • Overload detection: queue depth metrics, processing latency, resource usage, or canaries that signal stress.
  • A decision about the excess: reject, degrade, sample, or reroute the surplus work.
  • Upstream propagation: the producer must find out and adjust its rate, not keep pushing.

Practical implementation patterns

1. Explicit queue limits

Every queue should have a defined maximum. Upon reaching it, behavior must be deliberate: reject the publication, overwrite the oldest data, or switch to a degraded mode. A 10,000-message limit with an explicit rejection policy beats an "infinite" queue that only postpones the incident.

2. Load shedding at the edge

Before accepting work the system cannot process in time, it is better to reject it early with a clear response. In HTTP APIs, this means returning an overload status when saturation is detectable. A millisecond rejection is cheaper for everyone than a 30-second timeout.

3. Rate limiting at the source

Backpressure works best when complemented with rate control at the origin. Algorithms like token bucket or leaky bucket let you smooth spikes or tolerate bursts depending on the case. A generic configuration example:

rate_limit: { capacity: 500, refill_rate: 100/s } // 500 burst, 100 req/s sustained

4. Downstream timeouts and cancellation

A slow consumer should be able to cancel work no longer worth finishing. Per-stage timeouts, with latency budgets propagated in request headers, prevent a slow component from consuming resources across the entire call tree.

5. Graceful degradation

Under pressure, the system can reduce the fidelity of its work: sample events, disable non-essential enrichment, or switch to batch mode. The goal is to keep the essential service running instead of trying to keep everything and losing it all.

Common mistakes

  • Unbounded retries: when a rejected producer retries immediately, backpressure becomes a self-inflicted attack. Always use exponential backoff with jitter.
  • Alerting on CPU instead of queues: queue depth and oldest-message age are usually earlier signals than resource consumption.
  • Buffering at every layer: "just in case" buffers in every component multiply latency and hide the real bottleneck.
  • Backpressure as a purely technical matter: if the product team does not know that service quality degrades beyond a certain volume, the decision will be made mid-incident instead of at design time.
Key idea: backpressure is not an optimization, it is a contract. Define in advance which work is accepted, which is rejected, and which is degraded, and make that visible in both metrics and service documentation. A system that fails fast and predictably is more reliable than one that promises never to fail.

Design checklist

  • Does every queue have a maximum limit and a defined overflow policy?
  • Does the producer reduce its rate upon rejection, or only retry?
  • Are latency budgets propagated between services?
  • Are queue depth and oldest-message age monitored?
  • Is the degraded mode tested, not just documented?

Backpressure is one of those invisible investments: nobody notices when it works, but it defines the difference between a traffic spike and an incident page. Design the "no" before you need it.


Enjoyed this article?

If you're dealing with these challenges in your company, let's talk. No obligation. 30 minutes to understand your situation.

Book a free call →