REST API design: the decisions that make an API pleasant to use
An API is a product with developers as users. Most of what makes one good is consistency, not cleverness.
Rate limiting caps how many requests a client can make in a window. Token bucket is the best general-purpose algorithm because it allows short bursts while enforcing a sustained average; implement it in Redis so limits are shared across instances, key it by user or API key rather than IP where possible, and return 429 with Retry-After and rate limit headers.
| Algorithm | Burst behaviour | Accuracy | Cost |
|---|---|---|---|
| Fixed window | Allows 2× at boundaries | Poor | Cheapest |
| Sliding window log | Exact | Best | Memory per request |
| Sliding window counter | Smooth | Good | Low |
| Token bucket | Configurable burst | Good | Low |
| Leaky bucket | Smooths output | Good | Low |
Token bucket is the default we reach for. It has two intuitive parameters — bucket size (how big a burst you tolerate) and refill rate (the sustained limit) — which map directly to how product owners think about limits.
The most specific stable identity available: API key, then user id, then session, and IP address only as a last resort.
-- Atomic token bucket in Redis (called via EVAL)
local tokens = tonumber(redis.call("HGET", KEYS[1], "tokens") or ARGV[1])
local last = tonumber(redis.call("HGET", KEYS[1], "ts") or ARGV[3])
local delta = math.max(0, ARGV[3] - last)
tokens = math.min(ARGV[1], tokens + delta * ARGV[2])
if tokens < 1 then return {0, tokens} end
redis.call("HSET", KEYS[1], "tokens", tokens - 1, "ts", ARGV[3])
redis.call("EXPIRE", KEYS[1], 3600)
return {1, tokens - 1}The key requirement is atomicity. A naive GET-then-SET from multiple instances races, and under exactly the load you are trying to limit. Use a script, a Redis module built for this, or your platform's managed rate limiter.
HTTP/1.1 429 Too Many Requests
RateLimit-Limit: 100
RateLimit-Remaining: 0
RateLimit-Reset: 37
Retry-After: 37
Content-Type: application/problem+json
{ "title": "Rate limit exceeded", "status": 429,
"detail": "You have exceeded 100 requests per minute. Retry in 37 seconds." }Both. Edge limits absorb volumetric abuse cheaply before it reaches your infrastructure; application limits understand user identity and per-endpoint cost. They solve different problems.
Measure your legitimate p99 client behaviour and set the limit several times above it. Start in log-only mode to see who would have been blocked before you enforce.
Key by IP plus the submitted identifier and check a Redis counter at the top of the handler before doing any work. Unauthenticated forms are the most abused endpoints on most sites.
No — give them higher limits instead. Compromised accounts and buggy integrations both generate exactly the traffic you need protection from.
ROVQIX Engineering
Engineering team, ROVQIX
The ROVQIX engineering team builds and maintains web platforms, APIs and infrastructure for clients across SaaS, ecommerce and enterprise. These notes come out of real production work — deploys, incidents, migrations and audits.
ROVQIXdesigns and builds production web platforms — Next.js front ends, Node.js APIs and the infrastructure behind them. Tell us what you're building and we'll scope it with you.
An API is a product with developers as users. Most of what makes one good is consistency, not cleverness.
Caching is easy until the cache expires. Everything interesting about Redis in production happens in the seconds after a popular key disappears.
Not a compliance document. This is the list we actually work through before a client site handles its first real user.
No spam. Just the occasional case study and craft breakdown.