Optimize ollama endpoint to support more than one-request at a time, reducing latency
- Python 3.11+
- Ollama running locally (or reachable over network)
Using uv (recommended):
uv syncUsing pip:
pip install -e .uv run optollamaBy default the server binds to 0.0.0.0:8000.
Configured via environment variables (powered by pydantic-settings).
OLLAMA_BASE_URL(default:http://127.0.0.1:11434)
Example:
export OLLAMA_BASE_URL="http://localhost:11434"
uv run optollamaPOST /jobs (202 Accepted)
Request body:
{
"model": "llama3",
"prompt": "Write a haiku about autumn.",
"options": { "temperature": 0.7 }
}Curl example:
curl -X POST http://localhost:8000/jobs \
-H "Content-Type: application/json" \
-d '{
"model": "llama3",
"prompt": "Write a haiku about autumn.",
"options": { "temperature": 0.7 }
}'Response (example):
{
"id": "8b9b1f2e-2d1c-4b4b-9c68-1a2b3c4d5e6f",
"status": "queued",
"created_at": "2025-10-30T12:34:56.000000+00:00",
"updated_at": "2025-10-30T12:34:56.000000+00:00",
"model": "llama3",
"prompt": "Write a haiku about autumn.",
"options": { "temperature": 0.7 },
"result": null,
"error": null
}GET /jobs/{job_id}
curl http://localhost:8000/jobs/<job_uuid>When finished, status becomes completed and result contains the model output. If a failure occurs, status is failed and error describes the issue.
GET /jobs
curl http://localhost:8000/jobsResponse shape:
{
"items": [ { /* Job */ }, ... ],
"total": 1
}- This server uses an in-memory queue and store (no database). Data resets on restart.
- The worker processes one job at a time and calls Ollama's
/api/generateendpoint withstream=false. - Update the Ollama base URL via
OLLAMA_BASE_URLif your Ollama instance is remote.
- Entrypoint is
optollama.server:app(FastAPI). You can run with uvicorn directly:
uv run uvicorn optollama.server:app --host 0.0.0.0 --port 8000