Live: web / docs / playground · perf-api · secure-api
Two Spring Boot services, each built to prove a different engineering discipline:
| Service | Proves | Stack |
|---|---|---|
perf-api |
Raw per-core throughput | WebFlux + Netty, zero I/O hot path |
secure-api |
Airtight auth + rate limiting | Spring Security + JWT, Bucket4j + Redis |
Both ship with Swagger/OpenAPI docs, Actuator + Prometheus metrics, structured JSON logging with
per-request correlation IDs, and a Dockerfile ready for Railway. A third piece, web, is a
Next.js + shadcn/ui site with the architecture diagrams, API reference, and a live playground that
calls the running services directly from your browser.
A single "does everything" service can't credibly demonstrate either claim: a service fast enough
to chase huge throughput numbers has no business doing password hashing and database writes on the
hot path, and a service doing real auth work has no business skipping the safety checks that cost
CPU. So perf-api has no auth, no database, no blocking I/O — the ceiling is the JVM and the
network stack, nothing else. secure-api has no throughput ambitions — it spends its CPU
budget on bcrypt, JWT verification, and Redis round-trips instead, because that's the point.
- WebFlux on Netty, not Servlet/Tomcat — event-loop, non-blocking, no thread-per-request ceiling.
- Three endpoints:
GET /api/v1/ping(the benchmark target — no I/O, no allocation beyond the response record),GET /api/v1/echo/{value}(proves it's really serializing, not returning a constant),POST /api/v1/hash(real SHA-256 CPU work, so the benchmark isn't measuring a pure no-op). jackson-module-afterburnerfor faster (de)serialization, compression explicitly off (it's a CPU tax we'd rather spend on requests), RFC 7807 error bodies viaspring.webflux.problemdetails.enabled=true(no hand-written exception handler needed).- No auth, no persistence, no Redis — nothing on the hot path that isn't the request itself.
- JWT access tokens (15 min TTL) + rotating opaque refresh tokens (7 day TTL). Refresh tokens are never JWTs — they're random 256-bit values, stored only as a SHA-256 hash, so a database read alone can't be replayed.
- Reuse detection: every refresh rotates the token (old one is revoked, new one issued). If a revoked token is ever presented again, that's a stolen-token replay — every session for that user is revoked immediately, not just the one request.
- bcrypt (strength 12) password hashing + account lockout after 5 failed logins (15 min).
- RBAC via Spring Security (
ROLE_USER/ROLE_ADMIN), enforced in the filter chain, not scattered through controller code. - Distributed rate limiting (Bucket4j + Redis, token-bucket): every instance behind a load
balancer consumes from the same bucket, not its own in-memory count. Two tiers, chosen
automatically by whether the request carries a valid JWT at the point the filter runs — no
path-based special-casing needed:
- Anonymous / per-IP: 20 req/min (covers
/auth/**, i.e. brute-force/credential-stuffing attempts on login). - Authenticated / per-user: 100 req/min.
- Every response carries
X-RateLimit-Limit/X-RateLimit-Remaining; a rejected request gets429+Retry-After+ an RFC 7807 body.
- Anonymous / per-IP: 20 req/min (covers
- Virtual threads enabled (
spring.threads.virtual.enabled=true) — this service does blocking JPA/Redis I/O, so Project Loom removes the platform-thread-pool as a ceiling. - Postgres (via Flyway migrations) for users/refresh tokens, Redis for rate-limit buckets.
Both services tag every request with an id (reused from an inbound X-Request-Id header if the
caller sent one, generated otherwise), echo it back in the response, and log one structured JSON
line per request (Elastic Common Schema, via Spring Boot's built-in
logging.structured.format.console — no custom Logback config). On secure-api the id is also
pushed into SLF4J MDC before the request reaches the security filter chain, so every log line for
that request — an auth failure, a lockout, a rate-limit rejection — carries the same id, not just
the access-log line. Given an X-Request-Id from a bug report, grep finds everything that
happened for that request in one pass.
Both services expose Swagger UI:
perf-api:http://localhost:8081/swagger-ui.htmlsecure-api:http://localhost:8082/swagger-ui.html(click Authorize with an access token fromPOST /api/v1/auth/loginto try protected endpoints)
For a nicer read (and a live playground that fires real requests from your browser), see
web — npm run dev there, or run the whole stack via docker compose up below.
Requires Docker.
docker compose up --buildThis starts everything: perf-api (:8081), secure-api (:8082, with Postgres + Redis),
Prometheus (:9090), Grafana (:3000, anonymous viewer access — no login needed), and the
web site (:3100) with the docs/architecture/playground pages wired up against the two
local services.
To run a service directly against mvn instead (e.g. for faster iteration), start just the infra:
docker compose up postgres redis
cd secure-api && mvn spring-boot:run # or: cd perf-api && mvn spring-boot:runFor the frontend alone: cd web && npm install && npm run dev (defaults to perf-api/secure-api
on localhost:8081/8082 — override with NEXT_PUBLIC_PERF_API_URL / NEXT_PUBLIC_SECURE_API_URL).
Load-test scripts live in benchmarks/k6; raw results (JSON) are committed under
benchmarks/results so the numbers below are reproducible, not asserted.
brew install k6 # once
k6 run --summary-export=benchmarks/results/perf-api-throughput.json \
benchmarks/k6/perf-api-throughput.jsLocal run, docker compose (shared CPU with the k6 client itself — see caveat below), ramping
perf-api-throughput.js up to a 30,000 req/s target against
GET /api/v1/ping. Full raw output: perf-api-throughput.txt / .json.
| Metric | Value |
|---|---|
| Peak sustained rate | 30,000 req/s (target fully met, no throttling) |
| Total requests | 1,342,136 over the full ramp |
| Failures | 0.00% |
| Latency (avg / p90 / p95) | 8.2ms / 27.5ms / 42.1ms |
Methodology note, found the hard way: pushing the k6 target past ~30k/s on this single
laptop — where the load generator and the containers under test share the same CPU cores — makes
k6 itself the bottleneck: measured throughput actually drops (to ~17k/s at a 45k/s target)
and latency climbs, because the client runs out of room to generate load, not because perf-api is
struggling. 30k/s is the highest target that stays clean end-to-end on this hardware. This is
exactly why "1M req/s" is described here as an architectural property (non-blocking Netty event
loop, zero allocation on the hot path, Netty native transport, no shared mutable state — so it
scales horizontally with zero coordination) rather than a number a single laptop can produce: doing
so credibly needs a distributed load generator, not a bigger k6 run.
Run from a home machine over the public internet against the live single-container deployment — real network latency, real shared-CPU PaaS hardware, no localhost shortcuts. This is the number that matters for "what do I actually get," and it's reported here even though it's ~75x smaller than the local one, on purpose:
| Target | Result | Evidence |
|---|---|---|
| 30,000 req/s (same script, unmodified) | 770 req/s achieved, 14.2% failed, p95 606ms — the container saturates and starts timing out | perf-api-railway-throughput-overload.txt |
| 400 req/s (right-sized for one container) | 400 req/s, 0.00% failures, p95 217ms | perf-api-railway-throughput.txt |
That's the actual honest headline: one Railway hobby/pro container cleanly serves ~400 req/s of this endpoint over the open internet. Getting closer to the local 30k number on Railway is a horizontal-scaling exercise (more replicas behind Railway's load balancer — the architecture already supports it, nothing in perf-api holds per-instance state), not a code change.
Bumped perf-api to 3 replicas (numReplicas: 3 via Railway's API — confirmed by watching
instance in the /ping response rotate across different container IDs), re-ran the same
benchmark methodology to find the new clean ceiling:
| Replicas | Clean sustained rate | Failures | p95 |
|---|---|---|---|
| 1 | 400 req/s | 0.00% | 217ms |
| 3 | 550 req/s | 0.87% | 288ms |
Real evidence: perf-api-railway-scaled-3x.txt.
Honestly, not a 3x multiplier — +37%, not +200%. Tripling container count didn't triple
throughput because the containers were never the only bottleneck in this path: Railway's shared
edge/proxy layer, per-connection TLS/DNS overhead, and the single k6 client generating load from
one machine all sit outside the replica count and cap the improvement. That's a genuinely useful
finding, not a disappointing one — it's the difference between a real measurement and a marketing
number. Scaled back to 1 replica after capturing this (to avoid leaving extra billed compute
running); bump numReplicas on the perf-api service in Railway to restore.
secure-api-anonymous-rate-limit.js fires 30
requests at GET /api/v1/public/ping from one IP (configured limit: 20/min):
| allowed (200) | limited (429) |
|---|---|
| 20 | 10 |
Exact match to the configured app.rate-limit.anonymous-capacity: 20. Every 429 carried a
Retry-After header and X-RateLimit-Remaining: 0. Full output:
secure-api-anonymous-rate-limit.txt.
secure-api-authenticated-rate-limit.js
registers+logs in a throwaway user, then fires 110 requests at GET /api/v1/users/me with its
access token (configured limit: 100/min):
| allowed (200) | limited (429) |
|---|---|
| 100 | 10 |
Exact match to app.rate-limit.authenticated-capacity: 100 — and proof the two tiers are actually
different, not just configured to look different. Full output:
secure-api-authenticated-rate-limit.txt.
Against the live deployment, the anonymous test reproduces the same exact 20/10 split — same Redis token-bucket, same result, network latency doesn't change the math there since 30 sequential requests still complete in a couple of seconds. The authenticated test is the more interesting result: over the public internet each request takes ~85–250ms instead of ~1–3ms locally, so 110 sequential requests take ~14s wall-clock instead of well under a second — long enough for Bucket4j's continuous greedy refill (not a fixed-window reset) to add tokens back mid-burst, so more than 100 succeed. That's not a bug in the limiter; it's the token-bucket algorithm behaving exactly as designed, and it only becomes visible once requests are slow enough to spread across real time — a detail the fast local network hides entirely.
Already live at the links above — three Railway services in one project, each
pointed at this repo, secure-api wired to Railway's Postgres and Redis plugins via variable
references. What follows is how to reproduce it.
One real gotcha hit along the way, worth knowing before you try this yourself: Railway's remote
BuildKit builder rejects anonymous --mount=type=cache (works fine with local Docker Desktop,
fails cloud builds with a cryptic "missing an id argument" / cache-key-prefix error whose exact
required format isn't documented) — the Dockerfiles here don't use build cache mounts at all as a
result, trading some rebuild speed for not fighting an undocumented requirement.
Each service is a separate Railway service pointed at the same repo, same root directory, with a per-service config path:
- New Railway project → Deploy from GitHub repo → select this repo.
- First service: Settings → set Config-as-code file path to
perf-api/railway.json. No database/Redis needed. - Second service (New → GitHub Repo, same repo again): Settings → set Config-as-code file
path to
secure-api/railway.json. - On the
secure-apiservice: Add → Database → PostgreSQL and Add → Database → Redis (Railway auto-injectsPGHOST/PGPORT/PGDATABASE/PGUSER/PGPASSWORDandREDISHOST/REDISPORT/REDISPASSWORD—application.ymlalready reads those directly, no glue code). - On
secure-api, set one required variable:JWT_SECRET— a unique Base64 256-bit+ value (openssl rand -base64 32). Without it, the service falls back to a dev-only secret and logs a loud warning on startup; it will still run, just not securely. - Both services already have a Dockerfile-based build (
railway.json→DOCKERFILEbuilder) and a health check at/actuator/health.
Optional: CORS_ALLOWED_ORIGINS on secure-api and perf-api (comma-separated origins;
defaults to * — safe here since auth uses Authorization: Bearer, not cookies, so there's no
credentialed-CORS risk).
Third service — the web site:
- New → GitHub Repo (same repo again) → Settings → set Root Directory to
web(this one does not need the repo-root trick the two Java services use — it's a standalone Node app, andweb/railway.jsonalready points atweb/Dockerfile). - Set two build-time variables so the playground defaults point at your deployed services
instead of localhost —
NEXT_PUBLIC_PERF_API_URLandNEXT_PUBLIC_SECURE_API_URL(the public URLs Railway gives the other two services). These are baked into the client bundle at build time, not read at runtime, so a redeploy is required after changing them — the playground also lets you override the base URL per-request without redeploying anything.
perf-api/ WebFlux throughput service
secure-api/ Auth + rate-limiting service
web/ Next.js + shadcn/ui site — docs, architecture diagrams, live playground
benchmarks/ k6 scripts + committed results (the evidence)
monitoring/ docker-compose Prometheus/Grafana stack (local only, not deployed)