- API design
- API security
- developer guide
- web engineering
API Rate Limiting: Protect Capacity Without Surprising Clients
, by System Admin

Rate limits protect APIs from overload, accidental request loops and concentrated bursts of traffic. A useful limit does more than reject requests: it gives clients a predictable way to slow down and try again.
Choose what you are limiting
Start by deciding which resource needs protection and how requests should be counted. A service may apply limits per authenticated account, API key, route, resource, or across the whole system. For unauthenticated traffic, an IP address can be one signal, but shared networks and changing addresses make it an imperfect identity on its own.
Different operations may have different costs. A lightweight read and an expensive report-generation request need not consume the same allowance. If work varies substantially, use weighted units or separate limits for costly endpoints instead of treating every request as equivalent.
Select a limit algorithm deliberately
- Fixed window: counts requests in set intervals. It is simple, but traffic can bunch at the boundary between two windows.
- Sliding window: tracks recent activity more smoothly, at the cost of additional state or computation.
- Token bucket: replenishes a pool of tokens over time. It can allow short bursts while bounding the sustained rate.
- Leaky bucket: processes work at a steadier pace, which can help smooth bursts before a constrained downstream service.
There is no universally best algorithm. Match the policy to the resource and the traffic pattern you can safely serve. For a distributed service, make sure counters or token state are shared or consistently partitioned; independent per-instance counters can unintentionally multiply the allowance.
Make throttling useful to clients
When a client exceeds its allowance, HTTP 429 Too Many Requests is the standard response. Include a concise explanation and, when you can calculate it, a Retry-After header indicating when a retry is appropriate. The header can express a delay in seconds or an HTTP date.
Clients should respect that guidance. For retries without a server-provided delay, exponential backoff with jitter helps avoid every client retrying at the same instant. Retrying immediately in a tight loop turns a temporary limit into sustained load.
Keep the policy understandable
Document the scope of each limit, the relevant window or refill behaviour, and which operations consume units. If you publish rate-limit response headers, keep their meaning and reset timing consistent. Avoid promising a particular quota unless the service actually enforces it.
Consider legitimate bursts, pagination, batch operations, and long-running tasks. A queue or asynchronous job endpoint may be better than allowing clients to repeatedly submit expensive work synchronously. Separate overload protection from abuse controls: a request may be valid but temporarily over capacity, or it may need to be denied for a different security reason.
Test failure paths, not only the happy path
- Verify that requests below the threshold succeed and excess requests receive the documented response.
- Check that counters behave as intended across window boundaries and across service instances.
- Confirm that retries follow
Retry-Afterand that client backoff includes jitter. - Load-test the rate limiter itself; a limiter that becomes the bottleneck is not protecting the service.
- Monitor throttling by route and identity category without logging credentials or sensitive request bodies.
Rate limits work best as an explicit contract between a service and its clients: protect scarce capacity, explain the constraint, and make recovery safe and predictable.