Skip to content

About

Learn to scale API's

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

6 Commits

Folders and files

Repository files navigation

scalable-api

Live: web / docs / playground · perf-api · secure-api

Two Spring Boot services, each built to prove a different engineering discipline:

Service Proves Stack
perf-api Raw per-core throughput WebFlux + Netty, zero I/O hot path
secure-api Airtight auth + rate limiting Spring Security + JWT, Bucket4j + Redis

Both ship with Swagger/OpenAPI docs, Actuator + Prometheus metrics, structured JSON logging with per-request correlation IDs, and a Dockerfile ready for Railway. A third piece, web, is a Next.js + shadcn/ui site with the architecture diagrams, API reference, and a live playground that calls the running services directly from your browser.

Why two services, two stacks

A single "does everything" service can't credibly demonstrate either claim: a service fast enough to chase huge throughput numbers has no business doing password hashing and database writes on the hot path, and a service doing real auth work has no business skipping the safety checks that cost CPU. So perf-api has no auth, no database, no blocking I/O — the ceiling is the JVM and the network stack, nothing else. secure-api has no throughput ambitions — it spends its CPU budget on bcrypt, JWT verification, and Redis round-trips instead, because that's the point.

perf-api — throughput

  • WebFlux on Netty, not Servlet/Tomcat — event-loop, non-blocking, no thread-per-request ceiling.
  • Three endpoints: GET /api/v1/ping (the benchmark target — no I/O, no allocation beyond the response record), GET /api/v1/echo/{value} (proves it's really serializing, not returning a constant), POST /api/v1/hash (real SHA-256 CPU work, so the benchmark isn't measuring a pure no-op).
  • jackson-module-afterburner for faster (de)serialization, compression explicitly off (it's a CPU tax we'd rather spend on requests), RFC 7807 error bodies via spring.webflux.problemdetails.enabled=true (no hand-written exception handler needed).
  • No auth, no persistence, no Redis — nothing on the hot path that isn't the request itself.

secure-api — auth & rate limiting

  • JWT access tokens (15 min TTL) + rotating opaque refresh tokens (7 day TTL). Refresh tokens are never JWTs — they're random 256-bit values, stored only as a SHA-256 hash, so a database read alone can't be replayed.
  • Reuse detection: every refresh rotates the token (old one is revoked, new one issued). If a revoked token is ever presented again, that's a stolen-token replay — every session for that user is revoked immediately, not just the one request.
  • bcrypt (strength 12) password hashing + account lockout after 5 failed logins (15 min).
  • RBAC via Spring Security (ROLE_USER / ROLE_ADMIN), enforced in the filter chain, not scattered through controller code.
  • Distributed rate limiting (Bucket4j + Redis, token-bucket): every instance behind a load balancer consumes from the same bucket, not its own in-memory count. Two tiers, chosen automatically by whether the request carries a valid JWT at the point the filter runs — no path-based special-casing needed:
    • Anonymous / per-IP: 20 req/min (covers /auth/**, i.e. brute-force/credential-stuffing attempts on login).
    • Authenticated / per-user: 100 req/min.
    • Every response carries X-RateLimit-Limit / X-RateLimit-Remaining; a rejected request gets 429 + Retry-After + an RFC 7807 body.
  • Virtual threads enabled (spring.threads.virtual.enabled=true) — this service does blocking JPA/Redis I/O, so Project Loom removes the platform-thread-pool as a ceiling.
  • Postgres (via Flyway migrations) for users/refresh tokens, Redis for rate-limit buckets.

Every request is traceable

Both services tag every request with an id (reused from an inbound X-Request-Id header if the caller sent one, generated otherwise), echo it back in the response, and log one structured JSON line per request (Elastic Common Schema, via Spring Boot's built-in logging.structured.format.console — no custom Logback config). On secure-api the id is also pushed into SLF4J MDC before the request reaches the security filter chain, so every log line for that request — an auth failure, a lockout, a rate-limit rejection — carries the same id, not just the access-log line. Given an X-Request-Id from a bug report, grep finds everything that happened for that request in one pass.

API docs

Both services expose Swagger UI:

  • perf-api: http://localhost:8081/swagger-ui.html
  • secure-api: http://localhost:8082/swagger-ui.html (click Authorize with an access token from POST /api/v1/auth/login to try protected endpoints)

For a nicer read (and a live playground that fires real requests from your browser), see web — npm run dev there, or run the whole stack via docker compose up below.

Running locally

Requires Docker.

docker compose up --build

This starts everything: perf-api (:8081), secure-api (:8082, with Postgres + Redis), Prometheus (:9090), Grafana (:3000, anonymous viewer access — no login needed), and the web site (:3100) with the docs/architecture/playground pages wired up against the two local services.

To run a service directly against mvn instead (e.g. for faster iteration), start just the infra:

docker compose up postgres redis
cd secure-api && mvn spring-boot:run   # or: cd perf-api && mvn spring-boot:run

For the frontend alone: cd web && npm install && npm run dev (defaults to perf-api/secure-api on localhost:8081/8082 — override with NEXT_PUBLIC_PERF_API_URL / NEXT_PUBLIC_SECURE_API_URL).

Evidence

Load-test scripts live in benchmarks/k6; raw results (JSON) are committed under benchmarks/results so the numbers below are reproducible, not asserted.

brew install k6   # once
k6 run --summary-export=benchmarks/results/perf-api-throughput.json \
  benchmarks/k6/perf-api-throughput.js

perf-api throughput

Local run, docker compose (shared CPU with the k6 client itself — see caveat below), ramping perf-api-throughput.js up to a 30,000 req/s target against GET /api/v1/ping. Full raw output: perf-api-throughput.txt / .json.

Metric Value
Peak sustained rate 30,000 req/s (target fully met, no throttling)
Total requests 1,342,136 over the full ramp
Failures 0.00%
Latency (avg / p90 / p95) 8.2ms / 27.5ms / 42.1ms

Methodology note, found the hard way: pushing the k6 target past ~30k/s on this single laptop — where the load generator and the containers under test share the same CPU cores — makes k6 itself the bottleneck: measured throughput actually drops (to ~17k/s at a 45k/s target) and latency climbs, because the client runs out of room to generate load, not because perf-api is struggling. 30k/s is the highest target that stays clean end-to-end on this hardware. This is exactly why "1M req/s" is described here as an architectural property (non-blocking Netty event loop, zero allocation on the hot path, Netty native transport, no shared mutable state — so it scales horizontally with zero coordination) rather than a number a single laptop can produce: doing so credibly needs a distributed load generator, not a bigger k6 run.

The same benchmark against the actual Railway deployment

Run from a home machine over the public internet against the live single-container deployment — real network latency, real shared-CPU PaaS hardware, no localhost shortcuts. This is the number that matters for "what do I actually get," and it's reported here even though it's ~75x smaller than the local one, on purpose:

Target Result Evidence
30,000 req/s (same script, unmodified) 770 req/s achieved, 14.2% failed, p95 606ms — the container saturates and starts timing out perf-api-railway-throughput-overload.txt
400 req/s (right-sized for one container) 400 req/s, 0.00% failures, p95 217ms perf-api-railway-throughput.txt

That's the actual honest headline: one Railway hobby/pro container cleanly serves ~400 req/s of this endpoint over the open internet. Getting closer to the local 30k number on Railway is a horizontal-scaling exercise (more replicas behind Railway's load balancer — the architecture already supports it, nothing in perf-api holds per-instance state), not a code change.

Proving the scaling lever, not just claiming it

Bumped perf-api to 3 replicas (numReplicas: 3 via Railway's API — confirmed by watching instance in the /ping response rotate across different container IDs), re-ran the same benchmark methodology to find the new clean ceiling:

Replicas Clean sustained rate Failures p95
1 400 req/s 0.00% 217ms
3 550 req/s 0.87% 288ms

Real evidence: perf-api-railway-scaled-3x.txt.

Honestly, not a 3x multiplier — +37%, not +200%. Tripling container count didn't triple throughput because the containers were never the only bottleneck in this path: Railway's shared edge/proxy layer, per-connection TLS/DNS overhead, and the single k6 client generating load from one machine all sit outside the replica count and cap the improvement. That's a genuinely useful finding, not a disappointing one — it's the difference between a real measurement and a marketing number. Scaled back to 1 replica after capturing this (to avoid leaving extra billed compute running); bump numReplicas on the perf-api service in Railway to restore.

secure-api rate limiting

secure-api-anonymous-rate-limit.js fires 30 requests at GET /api/v1/public/ping from one IP (configured limit: 20/min):

allowed (200) limited (429)
20 10

Exact match to the configured app.rate-limit.anonymous-capacity: 20. Every 429 carried a Retry-After header and X-RateLimit-Remaining: 0. Full output: secure-api-anonymous-rate-limit.txt.

secure-api-authenticated-rate-limit.js registers+logs in a throwaway user, then fires 110 requests at GET /api/v1/users/me with its access token (configured limit: 100/min):

allowed (200) limited (429)
100 10

Exact match to app.rate-limit.authenticated-capacity: 100 — and proof the two tiers are actually different, not just configured to look different. Full output: secure-api-authenticated-rate-limit.txt.

Against the live deployment, the anonymous test reproduces the same exact 20/10 split — same Redis token-bucket, same result, network latency doesn't change the math there since 30 sequential requests still complete in a couple of seconds. The authenticated test is the more interesting result: over the public internet each request takes ~85–250ms instead of ~1–3ms locally, so 110 sequential requests take ~14s wall-clock instead of well under a second — long enough for Bucket4j's continuous greedy refill (not a fixed-window reset) to add tokens back mid-burst, so more than 100 succeed. That's not a bug in the limiter; it's the token-bucket algorithm behaving exactly as designed, and it only becomes visible once requests are slow enough to spread across real time — a detail the fast local network hides entirely.

Deploying to Railway

Already live at the links above — three Railway services in one project, each pointed at this repo, secure-api wired to Railway's Postgres and Redis plugins via variable references. What follows is how to reproduce it.

One real gotcha hit along the way, worth knowing before you try this yourself: Railway's remote BuildKit builder rejects anonymous --mount=type=cache (works fine with local Docker Desktop, fails cloud builds with a cryptic "missing an id argument" / cache-key-prefix error whose exact required format isn't documented) — the Dockerfiles here don't use build cache mounts at all as a result, trading some rebuild speed for not fighting an undocumented requirement.

Each service is a separate Railway service pointed at the same repo, same root directory, with a per-service config path:

  1. New Railway project → Deploy from GitHub repo → select this repo.
  2. First service: Settings → set Config-as-code file path to perf-api/railway.json. No database/Redis needed.
  3. Second service (New → GitHub Repo, same repo again): Settings → set Config-as-code file path to secure-api/railway.json.
  4. On the secure-api service: Add → Database → PostgreSQL and Add → Database → Redis (Railway auto-injects PGHOST/PGPORT/PGDATABASE/PGUSER/PGPASSWORD and REDISHOST/REDISPORT/REDISPASSWORD — application.yml already reads those directly, no glue code).
  5. On secure-api, set one required variable: JWT_SECRET — a unique Base64 256-bit+ value (openssl rand -base64 32). Without it, the service falls back to a dev-only secret and logs a loud warning on startup; it will still run, just not securely.
  6. Both services already have a Dockerfile-based build (railway.json → DOCKERFILE builder) and a health check at /actuator/health.

Optional: CORS_ALLOWED_ORIGINS on secure-api and perf-api (comma-separated origins; defaults to * — safe here since auth uses Authorization: Bearer, not cookies, so there's no credentialed-CORS risk).

Third service — the web site:

  1. New → GitHub Repo (same repo again) → Settings → set Root Directory to web (this one does not need the repo-root trick the two Java services use — it's a standalone Node app, and web/railway.json already points at web/Dockerfile).
  2. Set two build-time variables so the playground defaults point at your deployed services instead of localhost — NEXT_PUBLIC_PERF_API_URL and NEXT_PUBLIC_SECURE_API_URL (the public URLs Railway gives the other two services). These are baked into the client bundle at build time, not read at runtime, so a redeploy is required after changing them — the playground also lets you override the base URL per-request without redeploying anything.

Project layout

perf-api/     WebFlux throughput service
secure-api/   Auth + rate-limiting service
web/          Next.js + shadcn/ui site — docs, architecture diagrams, live playground
benchmarks/   k6 scripts + committed results (the evidence)
monitoring/   docker-compose Prometheus/Grafana stack (local only, not deployed)

About

Learn to scale API's

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages