Skip to content

Latest commit

 

History

2 Commits

Folders and files

Repository files navigation

optimized-ollama

Optimize ollama endpoint to support more than one-request at a time, reducing latency

Getting started

Requirements

  • Python 3.11+
  • Ollama running locally (or reachable over network)

Install

Using uv (recommended):

uv sync

Using pip:

pip install -e .

Run the API server

uv run optollama

By default the server binds to 0.0.0.0:8000.

Configuration

Configured via environment variables (powered by pydantic-settings).

  • OLLAMA_BASE_URL (default: http://127.0.0.1:11434)

Example:

export OLLAMA_BASE_URL="http://localhost:11434"
uv run optollama

API

Schedule a job

POST /jobs (202 Accepted)

Request body:

{
  "model": "llama3",
  "prompt": "Write a haiku about autumn.",
  "options": { "temperature": 0.7 }
}

Curl example:

curl -X POST http://localhost:8000/jobs \
  -H "Content-Type: application/json" \
  -d '{
    "model": "llama3",
    "prompt": "Write a haiku about autumn.",
    "options": { "temperature": 0.7 }
  }'

Response (example):

{
  "id": "8b9b1f2e-2d1c-4b4b-9c68-1a2b3c4d5e6f",
  "status": "queued",
  "created_at": "2025-10-30T12:34:56.000000+00:00",
  "updated_at": "2025-10-30T12:34:56.000000+00:00",
  "model": "llama3",
  "prompt": "Write a haiku about autumn.",
  "options": { "temperature": 0.7 },
  "result": null,
  "error": null
}

Get job status/result

GET /jobs/{job_id}

curl http://localhost:8000/jobs/<job_uuid>

When finished, status becomes completed and result contains the model output. If a failure occurs, status is failed and error describes the issue.

List all jobs

GET /jobs

curl http://localhost:8000/jobs

Response shape:

{
  "items": [ { /* Job */ }, ... ],
  "total": 1
}

Notes

  • This server uses an in-memory queue and store (no database). Data resets on restart.
  • The worker processes one job at a time and calls Ollama's /api/generate endpoint with stream=false.
  • Update the Ollama base URL via OLLAMA_BASE_URL if your Ollama instance is remote.

Development

  • Entrypoint is optollama.server:app (FastAPI). You can run with uvicorn directly:
uv run uvicorn optollama.server:app --host 0.0.0.0 --port 8000

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages