API Rate Limiting and Throttling: Production Strategies · Daniel Tinizaray

API Rate Limiting and Throttling: Production Strategies · Daniel Tinizaray

Every public or multi-tenant API eventually needs rate limits. Without them, a single badly-behaved client (or a bug in an integration script) can exhaust database connections, blow up your compute bill, and degrade the experience for everyone else. Rate limiting is not an optional feature: it is a design property of the backend.

Limiting is not punishment: it is capacity protection

A well-designed rate limiter serves three goals: it protects infrastructure from abusive spikes, guarantees fairness across clients (avoiding the noisy neighbor problem), and turns overload into an explicit, predictable signal instead of mysterious timeouts. The difference between good and bad rate limiting usually lies in communication: a 429 with clear headers is useful; a generic 503 is not.

Algorithms: which one for which case

  • Fixed window: a simple counter per time window (e.g. 100 req/min). Cheap and easy, but suffers the boundary problem: a client can send 2x the limit right across the edge of two windows.
  • Token bucket: tokens regenerate at a constant rate and each request consumes one. Allows controlled bursts, ideal for public APIs with per-plan quotas.
  • Sliding window: a continuous window that eliminates the boundary problem. Higher precision at the cost of more state; recommended for sensitive endpoints like payments or expensive searches.
  • Leaky bucket: smooths output to a constant rate, useful for protecting fragile downstream services (e.g. webhook delivery pipelines).
Rule of thumb: token bucket for developer APIs, sliding window for strict per-tenant limits, leaky bucket for protecting consumers. There is no universal algorithm — there are per-endpoint trade-offs.

The distributed problem

In a single process, an in-memory counter is enough. But in production there are almost always multiple replicas behind a load balancer, each with its own view of the counter. The standard solution is to externalize state to a shared store such as Redis, where every instance atomically increments the counter. With Lua scripts or commands like INCR + EXPIRE, you get an atomic, fast operation:

// Distributed counter in Redis (fixed window)
const key = `rl:${clientId}:${Math.floor(Date.now() / 60000)}`;
const count = await redis.incr(key);
if (count === 1) await redis.expire(key, 60);
if (count > limit) return res.status(429).json({ error: 'rate_limited' });

For high-traffic scenarios, the logarithmic sliding window in Redis (sorted sets plus ZREMRANGEBYSCORE) offers precision with bounded memory. If shared-store latency is a concern, a hybrid pattern works well: an approximate local counter per node plus periodic synchronization, accepting a small margin of error in exchange for fewer round-trips.

Communication: headers and graceful degradation

  • Use the standard headers RateLimit-Limit, RateLimit-Remaining and RateLimit-Reset (and Retry-After on every 429).
  • Document limits per plan and per endpoint: an invisible limit feels like a bug.
  • Apply tiered limits: per API key, per tenant, plus stricter limits on expensive endpoints (uploads, search, inference).
  • Consider grace periods or slightly padded limits for historically well-behaved clients; hard throttling should be the exception.

Common mistakes

Limiting only at the gateway and not in the services (one internal bypass defeats everything), using the IP address as the sole key behind proxies (it clusters legitimate users), not persisting counters across redeploys, and failing to monitor the 429 rate as a product signal: a spike in rejections means a client needs a higher plan or has a broken integration. Well-instrumented rate limiting is also a source of business data.

Bottom line: start simple (token bucket at the gateway + Redis), measure, and refine per endpoint. Complexity in rate limiting should be earned with data, not assumed upfront.


Enjoyed this article?

If you're dealing with these challenges in your company, let's talk. No obligation. 30 minutes to understand your situation.

Book a free call →