Enter the system: distributed rate limiting, explored in 3D
Loading the system…
Distributed rate limiting, without the 3D.
A simplified teaching model: one client, a limit of 5 requests per 3 seconds. Not the production implementation.
- 01 · Ingress
Roughly 500,000 OTP requests. One bot attack.
Every request starts the same way: a signal crossing the internet toward an application. The backend behind it can only absorb so much. Something has to say no.
- 02 · The problem
Three instances. Three private counters.
Each instance counts only what it sees. Traffic is split between them, so every instance believes the client is behaving. Together they let far more through than the limit allows.
- 03 · Shared state
One window. One answer.
Every instance asks Redis the same question: how many requests has this client made in the last three seconds? Expired timestamps are removed, the rest are counted, and the decision is global.
- 04 · Explore
Your turn.
Send a burst. Compare both designs on identical traffic. Select an instance, or inspect Redis.
- 05 · The world
One system. There are more.
Rate limiting is one system in a larger world: how it works, what gets built from it, and how it is explained.
How each request is decided
Both designs use a sliding window of timestamps. With local counters, each instance keeps its own list. With shared state, all instances use one list in Redis.
- Remove timestamps older than the window.
- Count the timestamps that remain.
- If the count is below the limit, record this request's timestamp and allow it.
- Otherwise reject it. Rejected requests are not recorded.
With a limit of 5 and three instances, local counters can allow up to 15 requests in one window, because each instance applies the limit to its own share. Shared state allows 5.