AgentForge is a self-hosted multi-agent orchestration project for demonstrating agent planning, worker collaboration, bounded context, skill feedback, and prompt versioning.
The project is intentionally runnable without a dataset, an API key, or internet
access. By default it uses deterministic mock model output; setting
LLM_PROVIDER=deepseek enables live DeepSeek calls through an OpenAI-compatible API.
- Planner + Worker + Reviewer orchestration with LangGraph typed state.
- Context engineering using structured state, sliding-window selection, periodic summaries, and tool-result pruning.
- Skill Registry that tracks calls, success rate, last execution time, prompt versions, and feedback-triggered prompt optimization.
- Collaboration traceability through JSONL events and a Markdown table.
- Docker Compose isolation for one-command startup and multi-instance demos.
The bundled scenario is a job-facing project coach:
Given an Agent/LLM job description and the current AgentForge project idea, the system plans the project, extracts job-facing signals, drafts resume bullets, reviews truthfulness, and returns a collaboration log.
It does not need any external dataset. The scenario is embedded in
src/agentforge/scenarios.py and can be replaced with any user task.
python3 -m venv .venv
source .venv/bin/activate
python -m pip install -e ".[dev]"
uvicorn agentforge.api:app --reloadCopy .env.example to .env for local configuration:
cp .env.example .envKeep real credentials only in .env; it is ignored by Git.
curl -X POST http://127.0.0.1:8000/demoOr:
curl -X POST http://127.0.0.1:8000/tasks \
-H "Content-Type: application/json" \
-d '{"scenario_id":"job_agentforge","use_live_llm":false}'Useful endpoints:
GET /healthGET /scenariosPOST /demoPOST /tasksGET /context/{task_id}POST /feedback
DeepSeek is optional. Set local environment variables:
LLM_PROVIDER=deepseek
LLM_MODEL=deepseek-v4-flash
DEEPSEEK_BASE_URL=https://api.deepseek.com
DEEPSEEK_API_KEY=your-local-keyThen run a live task:
curl -X POST http://127.0.0.1:8000/tasks \
-H "Content-Type: application/json" \
-d '{"scenario_id":"job_agentforge","use_live_llm":true}'AgentForge uses a hybrid policy:
- Write: keep structured task state outside the raw prompt.
- Select: keep only recent high-value turns in the active prompt.
- Compress: periodically summarize older messages into JSON.
- Isolate: prune verbose tool-call details while preserving decisions.
This is more effective for a small self-hosted agent than pushing the entire history into a long-context window, because it lowers cost, keeps tool noise out of prompts, and still persists the decision trail.
Context artifacts are stored under data/context/.
data/skills.json stores local skill metadata:
- system prompt
- call count
- success/failure count
- failure rate
- last execution time
- prompt history versions
When a skill failure rate exceeds SKILL_FAILURE_THRESHOLD after
SKILL_MIN_ATTEMPTS, AgentForge appends a prompt revision and preserves the old
version for rollback.
Feedback example:
curl -X POST http://127.0.0.1:8000/feedback \
-H "Content-Type: application/json" \
-d '{"skill":"planning","success":false,"feedback":"missing assumptions"}'Each Compose project gets its own network and named volumes. Start one instance:
docker compose -p agentforge-a --env-file .env.a up --build -dStart a second isolated instance with a different port and instance identifier:
docker compose -p agentforge-b --env-file .env.b up --build -dExample .env.a:
INSTANCE_ID=agentforge-a
APP_PORT=8001Example .env.b:
INSTANCE_ID=agentforge-b
APP_PORT=8002The -p value namespaces the Compose network and volumes, while APP_PORT
prevents host-port conflicts. Visit /health on each port to verify isolation.
pytest
ruff check .