A self-hosted API gateway for llama.cpp (or any OpenAI-compatible backend). Adds bearer token auth, concurrency limiting with queue tracking, and per-user token metrics — then exposes it publicly via Cloudflare Tunnel with no open router ports.
bash <(curl -fsSL https://raw.githubusercontent.com/hcastc00/llm-gateway/main/install.sh)The installer will:
- Install Docker if needed (Linux only; macOS requires Docker Desktop)
- Prompt for your upstream LLM URL, concurrency limit, API keys, and Cloudflare Tunnel token
- Write
.envandusers.json - Build and start the containers
Cloudflare Edge → cloudflared → gateway :8000 → llama.cpp :8001
↑
also on LAN at host-ip:8000
- No open ports —
cloudflaredconnects outbound to Cloudflare. Configure your public hostname in the Zero Trust dashboard:Service: http://gateway:8000 - LAN access —
http://<server-ip>:8000works directly on your local network - Bearer auth — every
/v1/*request requires a valid API key fromusers.json - Concurrency queue —
MAX_CONCURRENTlimits simultaneous upstream requests; streaming clients receive their queue position immediately via SSE
Copy .env.example to .env and fill in your values:
UPSTREAM_URL=http://192.168.1.46:8001 # your llama.cpp server
MAX_CONCURRENT=1 # max parallel requests
TUNNEL_TOKEN=your_cloudflare_token_here # from Zero Trust dashboardusers.json maps API keys to usernames:
{"sk-abc123": "alice", "sk-xyz789": "bob"}./users.sh list # show all users and keys
./users.sh add # add interactively
./users.sh add sk-newkey username # add inline
./users.sh remove username # remove all keys for a userChanges take effect immediately — no restart needed.
| Endpoint | Auth | Description |
|---|---|---|
GET /health |
No | Liveness check |
GET /metrics |
No | Queue depth, TPS, latency, per-user token totals |
/v1/* |
Bearer | Proxied to upstream |
# Start with Docker Compose
docker compose up --build
# Run locally for development
pip install -r gateway/requirements.txt
UPSTREAM_URL=http://localhost:8001 USERS_FILE=./users.json uvicorn gateway.main:app --reload --port 8000