A FastAPI-based distributed rate limiting service backed by Redis. The project supports multiple caller identity strategies, pluggable rate limiting algorithms, runtime rule management, and atomic Redis Lua scripts for correctness under concurrent traffic.
The internal boundaries are production-shaped: identity resolution, rule lookup, limiter selection, counter storage, and API routing are isolated so each piece can evolve without rewriting the request path.
- Sliding Window and Token Bucket algorithms
- Atomic Redis Lua scripts
- Multiple identity resolution strategies
- Runtime rule management
- Prometheus metrics endpoint
- FastAPI + OpenAPI support
- Distributed rate limiting with Redis
- Quick Start
- Docker
- Configuration
- API Usage
- Architecture
- Request Flow
- Rule Model
- Algorithms
- Redis Data Model
- Design Decisions
- Project Layout
- Development Checks
- Known Gaps
Copy the example environment file before starting the service:
cp src/app/.env.example src/app/.envEdit src/app/.env if you need different limits, headers, logging, or upstream
proxy settings.
Use Compose for local development:
docker compose up --buildCompose builds the FastAPI app image from Dockerfile and starts a Redis
container from redis:8-alpine on the same machine. This is convenient for
development because the app and Redis come up together.
The API will be available at:
http://127.0.0.1:8000/healthhttp://127.0.0.1:8000/docshttp://127.0.0.1:8000/openapi.json
If you prefer running the app directly on your host, create a virtual environment, install dependencies, and start Redis yourself:
python3 -m venv limiter
source limiter/bin/activate
pip install -r requirements.txt
docker run --rm -p 6379:6379 redis:8-alpine
PYTHONPATH=src python -m uvicorn app.main:app --reloadcompose.yaml is intended for development. It starts both services needed by
the app:
app: builds this project withDockerfileand exposes port8000.redis: runsredis:8-alpine, exposes port6379, and includes a health check used by the app startup dependency.
Create src/app/.env from the example file before using Compose:
cp src/app/.env.example src/app/.env
docker compose up --buildInside Compose, the app uses REDIS_URL=redis://redis:6379 so it can reach the
Redis service by its Compose service name.
The Dockerfile builds only the rate limiter application. It does not include
or start Redis. Use this image in production when Redis is provided separately
by your server, platform, or managed Redis service.
Example:
docker build -t distributed-rate-limiter .
docker run --rm -p 8000:8000 --env-file src/app/.env distributed-rate-limiterFor production, set REDIS_URL in the environment to the external Redis
instance, for example:
REDIS_URL=redis://redis.example.internal:6379/0IDENTITY_RESOLUTION_ORDER controls how a request is mapped to a caller. The
first resolver that finds a value wins.
UPSTREAM_BASE_URL controls gateway mode. Leave it empty for demo mode, where
accepted backend requests return a local explanation response. Set it to your
protected backend, for example https://api.example.com, to forward accepted
non-admin traffic to that server.
Settings are defined in src/app/core/config.py using pydantic-settings.
Values are loaded from environment variables and .env.
| Setting | Purpose |
|---|---|
REDIS_URL |
Redis connection URL used by the shared Redis client. |
DEFAULT_LIMIT |
Fallback sliding-window request limit. |
DEFAULT_WINDOW_SECONDS |
Fallback sliding-window duration. |
IDENTITY_RESOLUTION_ORDER |
Ordered list of identity types to attempt. |
API_KEY_HEADER |
Header used by ApiKeyResolver. |
USER_ID_HEADER |
Header used by UserIdResolver. |
PHONE_HEADER |
Header used by PhoneResolver. |
OAUTH_ACCOUNT_HEADER |
Header used by OAuthAccountResolver. |
RULE_CACHE_TTL_SECONDS |
In-process rule cache TTL. |
UPSTREAM_BASE_URL |
Optional protected backend base URL for accepted non-admin traffic. |
LOG_LEVEL |
Loguru logging level. |
Supported identity types:
api_keyip_addressuser_idjwt_subcurrently skipped because JWT verification is not configuredphoneoauth_accountcustommodel enum exists, resolver not yet implemented
curl http://127.0.0.1:8000/healthResponse:
{
"status": "ok",
"redis": "ok"
}curl http://127.0.0.1:8000/metricsReturns Prometheus plaintext if prometheus_client is installed. Otherwise it
returns an empty Prometheus-compatible plaintext response.
curl -X POST http://127.0.0.1:8000/rules \
-H "Content-Type: application/json" \
-d '{
"name": "ip-100-per-minute",
"identity_type": "ip_address",
"priority": 10,
"enabled": true,
"config": {
"algorithm": "sliding_window",
"limit": 100,
"window_seconds": 60
}
}'curl -X POST http://127.0.0.1:8000/rules \
-H "Content-Type: application/json" \
-d '{
"name": "api-key-bursty",
"identity_type": "api_key",
"priority": 20,
"enabled": true,
"config": {
"algorithm": "token_bucket",
"capacity": 50,
"refill_rate": 5
}
}'refill_rate is tokens per second. The example allows a burst of 50 requests
and refills at 5 requests per second.
curl http://127.0.0.1:8000/rulescurl http://127.0.0.1:8000/rules/{rule_id}identity_type and config are intentionally immutable after creation because
they define the uniqueness key in Redis.
curl -X PATCH http://127.0.0.1:8000/rules/{rule_id} \
-H "Content-Type: application/json" \
-d '{
"name": "ip-200-per-minute",
"priority": 15,
"enabled": true
}'curl -X DELETE http://127.0.0.1:8000/rules/{rule_id}The middleware hashes identity values before building Redis keys. The usage endpoint accepts the raw identity value and constructs the same internal key.
curl "http://127.0.0.1:8000/usage/client-123?identity_type=api_key"You can also target a specific rule:
curl "http://127.0.0.1:8000/usage/client-123?rule_id={rule_id}"curl -X POST "http://127.0.0.1:8000/usage/client-123/reset?identity_type=api_key"The service has two path categories:
- Local/admin paths are handled by this FastAPI app:
/health,/metrics,/docs,/redoc,/openapi.json,/rules, and/usage/.... - Backend paths are everything else, such as
/orders,/users/42, or/checkout. These requests are rate limited first. If allowed, they are either forwarded toUPSTREAM_BASE_URLor answered locally in demo mode.
For backend paths, the middleware attaches standard rate-limit headers:
X-RateLimit-Limit: 100
X-RateLimit-Remaining: 99
X-RateLimit-Reset: 1753881072
Retry-After: 1When UPSTREAM_BASE_URL is empty, accepted backend requests return a demo
response instead of being forwarded:
curl http://127.0.0.1:8000/orders \
-H "X-API-Key: client-123"{
"detail": "Rate limit check passed",
"mode": "demo",
"message": "No UPSTREAM_BASE_URL is configured, so this request was not forwarded. Set UPSTREAM_BASE_URL to proxy accepted traffic to your protected backend.",
"would_forward": {
"method": "GET",
"path": "/orders",
"query": ""
}
}When UPSTREAM_BASE_URL is set, accepted backend requests are proxied to the
protected server using the same method, path, query string, body, and forwarded
headers:
UPSTREAM_BASE_URL=https://api.example.comGET http://127.0.0.1:8000/orders?status=open
-> rate limit check
-> GET https://api.example.com/orders?status=open
When a request is blocked:
HTTP/1.1 429 Too Many Requests{
"detail": "Rate limit exceeded",
"limit": 100,
"remaining": 0,
"reset_at": 1753881072.25,
"retry_after": 1.0
}Blocked requests are not forwarded upstream.
flowchart LR
Client[Client Request]
MW[RateLimitMiddleware]
Resolver[CompositeIdentityResolver]
Cache[RuleCache]
Repo[RedisRuleRepository]
Factory[LimiterFactory]
Limiter[RateLimiter]
Counter[DistributedCounter]
Lua[Redis Lua Script]
Redis[(Redis)]
Route[Local Admin Route]
Upstream[Protected Backend]
Demo[Demo Response]
Client --> MW
MW -->|local/admin path| Route
Route --> MW
MW -->|backend path| Resolver
Resolver --> MW
MW --> Cache
Cache -->|cache miss| Repo
Repo --> Redis
Cache --> MW
MW --> Factory
Factory --> Limiter
Limiter --> Counter
Counter --> Lua
Lua --> Redis
Redis --> Lua
Lua --> Counter
Counter --> Limiter
Limiter --> MW
MW -->|allowed + UPSTREAM_BASE_URL set| Upstream
Upstream --> MW
MW -->|allowed + no upstream| Demo
MW -->|blocked| Blocked[429 Response]
| Module | Responsibility |
|---|---|
app.main |
FastAPI app factory, lifespan, middleware, exception handlers. |
app.api.middleware |
Hot-path rate limiting, local route detection, and upstream/demo forwarding. |
app.api.routes |
Admin/debug HTTP API for rules and usage. |
app.api.dependencies |
Dependency wiring for Redis, repository, cache, factory, resolvers. |
app.utils.identity |
Identity extraction from request headers/client IP. |
app.rules.models |
Pydantic rule schemas and discriminated algorithm configs. |
app.rules.repository |
Redis-backed rule persistence and uniqueness constraints. |
app.rules.cache |
In-process cache-through layer for rule lookups. |
app.limiter.* |
Algorithm implementations and result model. |
app.storage.counter |
Redis script loading and atomic counter operations. |
app.storage.scripts |
Lua scripts executed atomically inside Redis. |
app.core.* |
Settings, logging, and domain exceptions. |
sequenceDiagram
participant C as Client
participant M as RateLimitMiddleware
participant I as IdentityResolver
participant RC as RuleCache
participant RR as RuleRepository
participant LF as LimiterFactory
participant L as RateLimiter
participant R as Redis
participant A as Local Admin Route
participant U as Protected Backend
C->>M: HTTP request
alt local/admin path
M->>A: handle locally
A-->>M: response
M-->>C: response
else backend path
M->>M: Resolve identity and rule
M->>I: resolve(request)
I-->>M: Identity(type, value)
M->>RC: get_for_identity(identity.type)
alt cache miss or expired
RC->>RR: get_for_identity(identity.type)
RR->>R: HGETALL rules:index
R-->>RR: rule JSON values
RR-->>RC: highest priority enabled rule
end
RC-->>M: RateLimitRule
M->>LF: create(rule.config.algorithm)
LF-->>M: limiter
M->>L: check(rate_limit_key, rule)
L->>R: EVALSHA Lua script
R-->>L: allowed, remaining, reset_at
L-->>M: RateLimitResult
alt allowed and UPSTREAM_BASE_URL set
M->>U: forward original request
U-->>M: upstream response
M-->>C: response + X-RateLimit headers
else allowed and UPSTREAM_BASE_URL empty
M-->>C: demo response + X-RateLimit headers
else blocked
M-->>C: 429 + X-RateLimit headers
end
end
Rules are Pydantic models using a discriminated union on config.algorithm.
That means algorithm-specific fields live inside config, and FastAPI/OpenAPI
can validate the correct shape automatically.
classDiagram
class RateLimitRule {
string id
string name
IdentityType identity_type
int priority
bool enabled
TokenBucketConfig | SlidingWindowConfig config
}
class TokenBucketConfig {
algorithm = token_bucket
int capacity
float refill_rate
}
class SlidingWindowConfig {
algorithm = sliding_window
int limit
int window_seconds
}
RateLimitRule --> TokenBucketConfig
RateLimitRule --> SlidingWindowConfig
Rule selection today is by identity type, not identity value:
- Resolve caller identity, for example
api_key:sk_live_.... - Select enabled rules where
rule.identity_type == identity.type. - Choose the rule with the highest
priority. - If no rule matches, synthesize the default sliding-window rule from settings.
- Use the identity value only for counter keying, so each caller gets an independent bucket/window.
This keeps rule count low: one API-key rule can apply to all API keys, while the counters remain per API key.
Implementation:
- Python:
src/app/limiter/sliding_window.py - Lua:
src/app/storage/scripts/sliding_window.lua - Redis structure: sorted set
Behavior:
- Remove entries older than
now - window_seconds. - Count current entries in the sorted set.
- If count is below
limit, add the current request timestamp. - Compute
reset_atfrom the oldest request still in the window. - Set key TTL to approximately
window_seconds.
Why this algorithm:
- Accurate rolling window, not fixed-window boundary approximation.
- Simple Redis representation.
- Good fit for low-to-medium per-identity request volumes.
Tradeoff:
- Stores one sorted-set member per accepted request inside the window.
- For extremely high-cardinality/high-throughput identities, memory can grow more than a token bucket.
Implementation:
- Python:
src/app/limiter/token_bucket.py - Lua:
src/app/storage/scripts/token_bucket.lua - Redis structure: hash
Behavior:
- Load current
tokensandupdated_at. - Refill tokens by elapsed time times
refill_rate. - Cap tokens at
capacity. - Consume one token if available.
- Store updated state and compute
reset_at. - Set TTL based on time to refill the bucket.
Why this algorithm:
- Allows bursts up to
capacity. - Smooths traffic using a continuous refill rate.
- Stores constant-size state per identity/rule pair.
Tradeoff:
- It does not enforce an exact rolling count over a window.
- It is better described as sustained rate plus burst allowance.
Rules are stored in a Redis hash:
rules:index
{rule_id} -> json(RateLimitRule)
The repository also maintains a uniqueness index:
rules:by-identity-algo
{identity_type}:{algorithm} -> {rule_id}
This prevents two rules with the same immutable (identity_type, algorithm)
pair. The create path uses Redis WATCH/transaction semantics to avoid races.
Counter keys are built by middleware:
rate-limit:{algorithm}:{rule_id}:{identity_type}:{sha256(identity_value)}
Examples:
rate-limit:sliding_window:rule-1:ip_address:4f9c...
rate-limit:token_bucket:rule-2:api_key:a12b...
Why hash identity values:
- Avoid leaking API keys, phone numbers, OAuth account IDs, or user IDs into Redis key names.
- Keep Redis keys predictable in length.
- Preserve deterministic lookup for usage inspection/reset.
ZSET rate-limit:sliding_window:{rule_id}:{identity_type}:{digest}
score = request timestamp
member = "{timestamp}:{count}"
HASH rate-limit:token_bucket:{rule_id}:{identity_type}:{digest}
tokens -> float
updated_at -> unix timestamp
The middleware does not know how an identity is extracted. It only asks an
IdentityResolver for an opaque Identity.
Why:
- New identity types can be added without touching rate limit algorithms.
- Auth concerns stay separate from bucketing concerns.
- Rule selection and Redis keys consume a stable internal object.
CompositeIdentityResolver tries resolvers in IDENTITY_RESOLUTION_ORDER.
Why:
- Deployments can choose their preferred identity hierarchy.
- API-key traffic can take precedence over IP fallback.
- Internal services can use user IDs while public traffic falls back to IP.
RuleCache caches rules by IdentityType for RULE_CACHE_TTL_SECONDS (30 seconds by default). Rule updates may therefore take up to one cache TTL to become visible across workers
Why:
- Rule lookup is on every request.
- Rules change rarely compared to traffic volume.
- A small TTL reduces Redis reads while keeping admin changes reasonably fresh.
Admin rule mutations call cache.invalidate(...) in the current process. In a
multi-process or multi-node deployment, use a Redis pub/sub invalidation channel
or keep TTL low.
Both algorithms execute as Redis Lua scripts.
Why:
- Check-and-consume must be atomic.
- Multiple app instances can safely share Redis.
- The app avoids multi-command race conditions under concurrent traffic.
LimiterFactory maps AlgorithmType to a concrete limiter.
Why:
- Middleware does not branch on algorithm internals.
- Adding a new limiter is a registry change.
- Route/rule schemas remain the source of algorithm selection.
If no explicit rule matches an identity type, the repository returns a default sliding-window rule from settings.
Why:
- There is always a safety net.
- Startup does not depend on seed data.
- Configuration can define baseline protection before custom rules exist.
src/app/
├── api/
│ ├── dependencies.py # FastAPI dependency graph
│ ├── middleware.py # request-time rate limiting pipeline
│ └── routes.py # health, metrics, rules, usage endpoints
├── core/
│ ├── config.py # pydantic-settings configuration
│ ├── exceptions.py # domain exceptions
│ └── logging.py # structured loguru setup
├── limiter/
│ ├── base.py # RateLimiter interface
│ ├── factory.py # algorithm registry/factory
│ ├── result.py # RateLimitResult
│ ├── sliding_window.py # sliding window implementation
│ └── token_bucket.py # token bucket implementation
├── rules/
│ ├── cache.py # TTL cache-through rule lookup
│ ├── loader.py # YAML/JSON rule loader
│ ├── models.py # rule/config Pydantic models
│ └── repository.py # Redis rule repository
├── storage/
│ ├── counter.py # Lua script registration/invocation
│ ├── redis_client.py # shared Redis client lifecycle
│ └── scripts/
│ ├── sliding_window.lua
│ └── token_bucket.lua
├── utils/
│ └── identity.py # identity resolver implementations
└── main.py # FastAPI app factory and lifespan
Compile the package:
PYTHONPATH=src python -m compileall -q src/appCheck app import and OpenAPI generation:
PYTHONPATH=src python - <<'PY'
from fastapi import FastAPI
from app.api.routes import router
from app.api.middleware import RateLimitMiddleware
app = FastAPI()
app.include_router(router)
app.add_middleware(RateLimitMiddleware)
print(sorted(app.openapi()["paths"]))
PYRun the server:
PYTHONPATH=src python -m uvicorn app.main:app --reloadAll mutable rate limit state lives in Redis, so multiple FastAPI workers or
service instances can share the same quota state. The in-process RuleCache is
per worker, so cross-worker rule invalidation is eventually consistent via TTL
unless a broadcast invalidation mechanism is added.
Startup pings Redis in main.lifespan(). If Redis is unavailable, startup fails
instead of accepting traffic without enforcement.
The token bucket Lua script uses Redis TIME, which gives a central clock for
all application instances. The sliding window code currently passes time.time()
from Python into Lua. For stricter cross-node consistency, move sliding-window
time sourcing into Redis Lua as well.
Current identity resolvers extract caller-asserted headers and direct peer IP. They do not authenticate API keys, verify JWT signatures, normalize phone numbers, or trust proxy forwarding headers. In production, put authentication before this service or implement verified resolvers.
The middleware returns rate limit headers. Logging is configured with Loguru and a correlation-id context variable. Metrics endpoint plumbing exists, but the current code does not yet record per-request counters/histograms in middleware.
JWTResolveris intentionally not implemented until JWT library, issuer, audience, key source, and algorithms are configured.CUSTOMidentity type exists in the enum but has no resolver.- No built-in cross-process rule cache invalidation yet.
- No repository-level rule matching by specific identity value yet.
- Add or reuse an
IdentityType. - Implement
IdentityResolverinsrc/app/utils/identity.py. - Register it in
get_identity_resolver()insrc/app/api/dependencies.py. - Add the identity type to
IDENTITY_RESOLUTION_ORDER.
- Add a config model in
src/app/rules/models.py. - Add the algorithm to
AlgorithmType. - Implement
RateLimiter. - Add any required Redis script/counter operation.
- Register the limiter in
LimiterFactory.
src/app/rules/loader.py can load YAML or JSON rule lists and seed them
idempotently through RuleRepository.seed_defaults(). Wire it into
main.lifespan() when a RULES_FILE setting is added.