Skip to content

Latest commit

 

History

28 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Harbor Adapters

Docs Cookbook DOI

Release blog: https://harbor-index.org/

Paper: https://arxiv.org/abs/2609.04298


This repository is the standalone home for Harbor adapters — the code that converts external benchmarks (SWE-Bench, Aider Polyglot, GPQA, AIME, and 80+ more) into Harbor's task format so they can be run, scored, and shared through the Harbor harness. It was split out of the main harbor-framework/harbor monorepo so adapters can be versioned and contributed independently.

Harbor itself — the CLI and evaluation framework — lives in harbor-framework/harbor and is an external dependency of this repo (install with uv tool install harbor). See the Harbor docs and the Harbor Cookbook for the framework itself.

What is an adapter?

An adapter translates a benchmark into Harbor tasks. Each generated task is a directory with:

  • task.toml — configuration and metadata
  • instruction.md — the natural-language task given to the agent
  • environment/ — Dockerfile / environment definition
  • tests/ — verification scripts (test.sh writes the reward to /logs/verifier/reward.txt)
  • solution/ (optional) — the oracle / reference solution

An adapter's job is to parse an upstream benchmark and emit those task directories reproducibly.

Repository structure

adapters/
├── src/                       # one directory per adapter — the heart of this repo (85 adapters)
│   └── <adapter-name>/        # a self-contained, independent uv project
│       ├── pyproject.toml            # package name = "harbor-<adapter-name>-adapter"
│       ├── uv.lock                   # each adapter locks its own dependencies
│       ├── README.md                 # adapter docs (parity results, reproduction, notes)
│       ├── parity_experiment.json    # parity results vs. the original benchmark
│       ├── adapter_metadata.json     # adapter metadata
│       ├── <adapter-name>.yaml       # reference config for running the adapter
│       └── src/<adapter_name>/       # the Python package (dashes → underscores)
│           ├── adapter.py            # task-generation logic (<AdapterName>Adapter class)
│           ├── main.py               # CLI entry point (--output-dir/--limit/--overwrite/--task-ids)
│           └── task-template/        # task.toml, instruction.md, environment/, solution/, tests/
├── docs/                      # adapter authoring guides (see Documentation below)
├── skills/                    # Agent skills for adapter work
├── scripts/                   # Helper scripts (validation, parity summary)
└── .github/workflows/         # CI: adapter review + parity summary

Every adapter under src/ is an independent uv project with its own pyproject.toml and uv.lock. There is no repo-wide virtualenv or lockfile — you work inside a single adapter directory. This mirrors the layout described in docs/adapters.mdx.

Installation

Install the Harbor CLI (the harness that runs the tasks these adapters generate):

uv tool install harbor
# or: pip install harbor

Using an adapter

Each adapter generates task directories you can then run with Harbor:

cd src/<adapter-name>
uv sync                                              # install this adapter's deps
uv run <adapter-name> --output-dir /path/to/output   # generate task directories

uv run <adapter-name> maps to <adapter_name>.main:main via the adapter's [project.scripts]. Once the tasks exist, run them with the Harbor harness, e.g.:

harbor run -p /path/to/output -a claude-code -m "anthropic/claude-opus-5"

See a specific adapter's own README.md for its exact commands, parity results, and any special setup.

Creating a new adapter

Use the create-adapter skill in skills/create-adapter/ — it scaffolds the adapter with harbor adapter init and hands off to the authoritative spec, docs/adapters.mdx. New adapters must live at src/<adapter-name>/.

Key conventions (enforced by scripts/validate_adapter.py):

  • pyproject.toml name = harbor-<folder>-adapter.
  • [project.scripts] has <folder> = "<adapter_name>.main:main".
  • Adapter code lives at src/<adapter_name>/ (dashes → underscores); adapter.py defines an <AdapterName>Adapter class whose run(self) writes tasks under self.output_dir.
  • Task names must be stable across runs and unique / registry-safe.

Citation

If you use Harbor adapters in academic work, please cite the paper:

@misc{shi2026harbor,
  title         = {Harbor Adapters and Harbor-Index: Infrastructure and a Curated Meta-Dataset for Large-Scale Agentic Evaluation},
  author        = {Shi, Lin and Lin, Haowei and Zhu, Zixuan and Zhou, Xiaoyue and Li, Xiang and Lin, Xiangning and Deng, Yaxuan and Xu, Han and Li, Yuangang and Li, Shanda and Chen, Zizhao and Xing, Hanwen and Raj, Harsh and Chen, Bo and Shi, Quan and Dillmann, Steven and Gao, Yipeng and Khanna, Puneesh and Lu, Ruofan and Zhou, Chao Beyond and Yang, Michael and Zhang, Robert and Chai, Siyuan and Chang, Jiayu and Chen, Yizhao and Chen, Xiaokun and Dai, Yiwei and Yang, Wenting and Liu, Hange and Liu, Minghao and Wang, Zihan and others},
  year          = {2026},
  eprint        = {2609.04298},
  archivePrefix = {arXiv},
  primaryClass  = {cs.LG},
  url           = {https://arxiv.org/abs/2609.04298}
}

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages