From f3e242c5ff19305cc6d596a4f3633d9d3a71bfc1 Mon Sep 17 00:00:00 2001 From: Claude Date: Tue, 18 Aug 2026 22:45:38 +0000 Subject: [PATCH 1/5] docs: refresh site around case studies and patterns The site described a Goose-first, single-guide world that no longer matches what these repositories do. Rebuild it around the repositories themselves. New: - docs/case-studies/: one page per repository running agents on real curation. Ontologies (GO, Uberon, Mondo, Cell Ontology, EFO) and the mechs (DisMech, AI Gene Review, CommunityMech, HabitatMech and the CultureBot family). - docs/patterns/: nine practices that appear in more than one repository, each linking to the repositories that use it. - docs/evidence.md: where the skill, subagent and workflow counts come from. - docs/tutorials/first-curation-session.md: a cloud-first first session, replacing the Goose, OWL-MCP and Protege walkthrough. - docs/reference/harnesses.md: replaces six client stub pages. Changed: - Goose is now documented as historic throughout. It stays described because it is still deployed, but it is no longer recommended for new setups. - Canonical site URL is now ai4curation.io/aidocs, matching the CNAME. - CLAUDE.md is a pointer to AGENTS.md, so this repository follows the one-source-of-instructions pattern it documents. - set-up-github-actions.md leads with Claude Code Action; the dragon-ai-agent and Goose route moves to a historic section. - author-skills.md no longer ends in a TODO, and its skill counts are current. Removed: - Six client comparison stubs, examples.md, and the Goose tutorial. Their content is covered by the case studies and the harnesses page. Infrastructure: - mkdocs build is now --strict in deploy, and a new docs-checks workflow runs the strict build plus a link check on pull requests and weekly. Outbound links to other repositories are the main thing that rots here. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_017tKFksxqHZJNLoH1zJ5WXm --- .github/workflows/deploy_documentation.yml | 2 +- .github/workflows/docs-checks.yml | 46 +++ AGENTS.md | 25 ++ CLAUDE.md | 95 +----- README.md | 4 +- docs/case-studies/ai-gene-review.md | 71 +++++ docs/case-studies/cell-ontology.md | 54 ++++ docs/case-studies/communitymech.md | 89 ++++++ docs/case-studies/dismech.md | 113 +++++++ docs/case-studies/efo.md | 65 ++++ docs/case-studies/go-ontology.md | 51 +++ docs/case-studies/habitatmech.md | 73 +++++ docs/case-studies/index.md | 58 ++++ docs/case-studies/mondo.md | 53 ++++ docs/case-studies/uberon.md | 50 +++ docs/evidence.md | 86 +++++ docs/examples.md | 85 ----- docs/faq.md | 46 ++- docs/glossary.md | 40 ++- docs/how-tos/author-skills.md | 30 +- docs/how-tos/build-agentic-harness.md | 8 +- .../create-agentic-curation-pipeline.md | 20 +- docs/how-tos/instruct-github-agent.md | 8 +- docs/how-tos/integrate-ai-into-your-kb.md | 35 +-- docs/how-tos/set-up-github-actions.md | 297 +++++++----------- docs/index.md | 54 +++- docs/patterns/fast-and-slow-validation.md | 66 ++++ docs/patterns/ground-identifiers.md | 69 ++++ docs/patterns/guard-untrusted-input.md | 49 +++ docs/patterns/human-regulating-the-loop.md | 54 ++++ docs/patterns/index.md | 17 + docs/patterns/one-source-of-instructions.md | 56 ++++ docs/patterns/review-checklists.md | 48 +++ docs/patterns/scanners.md | 50 +++ docs/patterns/shared-control.md | 48 +++ docs/patterns/skills-before-automation.md | 66 ++++ docs/reference/agentic-tools.md | 12 +- docs/reference/claude-skills.md | 2 +- docs/reference/client-apps.md | 21 -- docs/reference/clients/claude-code.md | 19 -- docs/reference/clients/claude-desktop.md | 17 - docs/reference/clients/codex-cli.md | 14 - docs/reference/clients/gemini-cli.md | 13 - docs/reference/clients/goose.md | 15 - .../reference/{clients => }/github-copilot.md | 4 +- docs/reference/github-integrations.md | 4 +- docs/reference/harnesses.md | 69 ++++ docs/tutorials/first-curation-session.md | 139 ++++++++ docs/tutorials/ontology-editing-with-ai.md | 159 ---------- docs/tutorials/tutorials-for-curators.md | 21 +- mkdocs.yml | 42 ++- 51 files changed, 1933 insertions(+), 699 deletions(-) create mode 100644 .github/workflows/docs-checks.yml create mode 100644 docs/case-studies/ai-gene-review.md create mode 100644 docs/case-studies/cell-ontology.md create mode 100644 docs/case-studies/communitymech.md create mode 100644 docs/case-studies/dismech.md create mode 100644 docs/case-studies/efo.md create mode 100644 docs/case-studies/go-ontology.md create mode 100644 docs/case-studies/habitatmech.md create mode 100644 docs/case-studies/index.md create mode 100644 docs/case-studies/mondo.md create mode 100644 docs/case-studies/uberon.md create mode 100644 docs/evidence.md delete mode 100644 docs/examples.md create mode 100644 docs/patterns/fast-and-slow-validation.md create mode 100644 docs/patterns/ground-identifiers.md create mode 100644 docs/patterns/guard-untrusted-input.md create mode 100644 docs/patterns/human-regulating-the-loop.md create mode 100644 docs/patterns/index.md create mode 100644 docs/patterns/one-source-of-instructions.md create mode 100644 docs/patterns/review-checklists.md create mode 100644 docs/patterns/scanners.md create mode 100644 docs/patterns/shared-control.md create mode 100644 docs/patterns/skills-before-automation.md delete mode 100644 docs/reference/client-apps.md delete mode 100644 docs/reference/clients/claude-code.md delete mode 100644 docs/reference/clients/claude-desktop.md delete mode 100644 docs/reference/clients/codex-cli.md delete mode 100644 docs/reference/clients/gemini-cli.md delete mode 100644 docs/reference/clients/goose.md rename docs/reference/{clients => }/github-copilot.md (97%) create mode 100644 docs/reference/harnesses.md create mode 100644 docs/tutorials/first-curation-session.md delete mode 100644 docs/tutorials/ontology-editing-with-ai.md diff --git a/.github/workflows/deploy_documentation.yml b/.github/workflows/deploy_documentation.yml index 023a2bb..fbb7ff0 100644 --- a/.github/workflows/deploy_documentation.yml +++ b/.github/workflows/deploy_documentation.yml @@ -35,7 +35,7 @@ jobs: run: uv sync --all-extras - name: Build with MkDocs - run: uv run mkdocs build + run: uv run mkdocs build --strict - name: Upload artifact uses: actions/upload-pages-artifact@v3 diff --git a/.github/workflows/docs-checks.yml b/.github/workflows/docs-checks.yml new file mode 100644 index 0000000..908b780 --- /dev/null +++ b/.github/workflows/docs-checks.yml @@ -0,0 +1,46 @@ +name: Docs checks + +on: + pull_request: + push: + branches: + - main + schedule: + # Weekly, to catch links that rot in the repos this site points at. + - cron: "0 6 * * 1" + workflow_dispatch: + +permissions: + contents: read + +jobs: + build: + name: Build site + runs-on: ubuntu-latest + steps: + - uses: actions/checkout@v4 + - uses: astral-sh/setup-uv@v5 + - uses: actions/setup-python@v5 + with: + python-version: "3.12" + - run: uv sync --all-extras + - name: Build with MkDocs + run: uv run mkdocs build --strict + + links: + name: Check links + runs-on: ubuntu-latest + steps: + - uses: actions/checkout@v4 + - name: Check links in Markdown + uses: lycheeverse/lychee-action@v2 + with: + args: >- + --no-progress + --max-retries 2 + --accept 200,206,429 + --exclude-path docs/overrides + --exclude 'github\.com/ai4curation/agent-watcher' + 'docs/**/*.md' + 'README.md' + fail: true diff --git a/AGENTS.md b/AGENTS.md index a10e2fa..9913726 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -68,3 +68,28 @@ curation workflows. - Follow existing documentation structure instead of creating new sections unnecessarily. - Emphasize how AI improves existing workflows rather than replacing curators. + +## Site structure + +The site has two sections that carry most of its value: + +- `docs/case-studies/`: one page per repository that runs agents on real + curation. Each page follows the same shape: what to look at, what works, what + to copy first, and gaps. +- `docs/patterns/`: practices that appear in more than one repository. Each page + names at least two repositories that use it. + +Do not centralize material that belongs in another repository. Link to the file +or folder instead. If a page disagrees with the repository it describes, the +repository is right. + +Counts of skills, subagents, and workflows come from +[agent-watcher](https://github.com/ai4curation/agent-watcher). Cite the dated +report you took them from, as `docs/evidence.md` does. + +## Checks + +Run `uv run mkdocs build --strict` before committing. The `docs-checks.yml` +workflow runs the same build and a link check on every pull request, and the +link check also runs weekly to catch links that rot in the repositories this +site points at. diff --git a/CLAUDE.md b/CLAUDE.md index 79343a9..1bc0d89 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -1,90 +1,9 @@ -# CLAUDE.md for aidocs +# CLAUDE.md -This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository. +Read [AGENTS.md](AGENTS.md). It is the authoritative guidance for all agents +working in this repository, including Claude Code. -## Project Overview - -**aidocs** is a documentation repository for **AI4Curators**, providing practical guides for curators and maintainers of knowledge bases to integrate AI into their workflows. The project focuses on immediate, actionable integration strategies rather than theoretical discussions. - -### Core Mission -- Help curators integrate AI agents into existing GitHub-based workflows -- Provide plugins and tools for existing chat UIs -- Support ontology editing and curation workflows with AI assistance - -## Repository Structure - -``` -aidocs/ -├── docs/ # MkDocs documentation source -│ ├── how-tos/ # Practical how-to guides -│ ├── tutorials/ # Step-by-step tutorials -│ ├── reference/ # Technical reference materials -│ │ └── clients/ # Documentation for various AI clients -│ └── overrides/ # MkDocs theme customizations -├── src/aidocs/ # Python package source -├── tools/ # Utility scripts and tools -├── mkdocs.yml # MkDocs configuration -└── pyproject.toml # Python project configuration -``` - -## Technology Stack - -- **Documentation**: MkDocs with Material theme -- **Python**: >=3.9 (for tooling and scripts) -- **Deployment**: GitHub Pages -- **Build System**: Hatchling - -## Key Files and Their Purpose - -- `mkdocs.yml`: Configures the documentation site structure, theme, and plugins -- `pyproject.toml`: Python project metadata and dependencies -- `docs/index.md`: Main landing page content -- `docs/glossary.md`: Terminology definitions for the domain -- `docs/how-tos/`: Practical implementation guides -- `docs/reference/clients/`: Documentation for various AI client applications - -## Development Guidelines - -### Documentation Standards -- Focus on **practical, immediately actionable** content -- Provide step-by-step guides over theoretical explanations -- Include real-world examples and use cases -- Maintain consistency with the existing Material theme styling - -### Content Categories -1. **How-tos**: Task-oriented guides for specific implementation scenarios -2. **Tutorials**: Comprehensive learning-oriented walkthroughs -3. **Reference**: Technical specifications and API documentation -4. **Glossary**: Domain-specific terminology and definitions - -### Building and Testing -- Use `mkdocs serve` for local development -- Documentation is automatically deployed to GitHub Pages -- Test all links and examples before committing - -## Target Audience - -- **Primary**: Curators and maintainers of knowledge bases and ontologies -- **Secondary**: Developers integrating AI into existing curation workflows -- **Focus**: Practitioners who need immediate, working solutions over academic discussions - -## AI Agent Integration - -This repository serves as both documentation and a practical example of AI agent integration: -- GitHub agents can directly contribute to documentation -- Examples demonstrate real-world AI-assisted curation workflows -- Reference materials guide implementation of similar systems - -## Contributing Guidelines - -When working on this repository: -1. **Prioritize practical value** - every guide should be immediately applicable -2. **Test all examples** - ensure code snippets and procedures work as documented -3. **Maintain consistency** - follow existing patterns in documentation structure -4. **Focus on integration** - emphasize how AI enhances existing workflows rather than replacing them - -# important-instruction-reminders -Do what has been asked; nothing more, nothing less. -NEVER create files unless they're absolutely necessary for achieving your goal. -ALWAYS prefer editing an existing file to creating a new one. -NEVER proactively create documentation files (*.md) or README files. Only create documentation files if explicitly requested by the User. +Do not add instructions here. Keeping one file avoids the drift that happens +when `CLAUDE.md`, `AGENTS.md`, and `.github/copilot-instructions.md` are +maintained separately. See +[One source of instructions](docs/patterns/one-source-of-instructions.md). diff --git a/README.md b/README.md index 47f78c2..3994120 100644 --- a/README.md +++ b/README.md @@ -4,9 +4,9 @@ This repository contains documentation about using AI to assist with curation, p ## Documentation -📖 **[Visit the full documentation](https://ai4curation.github.io/aidocs/)** +📖 **[Visit the full documentation](https://ai4curation.io/aidocs/)** -The complete guides, tutorials, and reference materials are available on our [GitHub Pages site](https://ai4curation.github.io/aidocs/). +The complete guides, tutorials, and reference materials are available on our [GitHub Pages site](https://ai4curation.io/aidocs/). ## About diff --git a/docs/case-studies/ai-gene-review.md b/docs/case-studies/ai-gene-review.md new file mode 100644 index 0000000..5805183 --- /dev/null +++ b/docs/case-studies/ai-gene-review.md @@ -0,0 +1,71 @@ +# AI Gene Review + +A review of existing Gene Ontology annotations, one file per gene, with a public +voting interface for feedback. + +**Repository**: [ai4curation/ai-gene-review](https://github.com/ai4curation/ai-gene-review) +**App**: [Browse gene reviews](https://ai4curation.io/ai-gene-review/app/index.html) +**Product**: `genes///-ai-review.yaml` +**Agent surface**: local sessions, GitHub mention, pull request review + +## What to look at + +| Path | What it is | +| --- | --- | +| [`.claude/skills/`](https://github.com/ai4curation/ai-gene-review/tree/main/.claude/skills) | Fifteen skills, including `annotation-reviewer`, `core-function-synthesizer`, `gocam-curation`, and `aigr-pr-review` | +| [`.claude/hooks/`](https://github.com/ai4curation/ai-gene-review/tree/main/.claude/hooks) | Hooks that run around agent actions | +| `genes/` | One directory per gene, holding the review and its cached references | +| `src/ai_gene_review/schema/` | The LinkML schema | + +## What works + +### Every annotation gets a verdict and a reason + +A review does not rewrite an annotation. It records an action against the +existing one, such as `ACCEPT`, `MODIFY`, or `REMOVE`, together with the reason: + +```yaml +existing_annotations: + - term: + id: GO:0005515 + label: protein binding + action: MODIFY + reason: "While evidence is strong, 'protein binding' is uninformative..." +``` + +This is reviewable in a way that a rewritten file is not. A curator reads the +verdict and the reason, and does not have to reconstruct what changed. + +### Identifiers carry their labels + +Every term appears as an identifier and a label together. A wrong pair fails +validation, because an agent would have to fabricate both consistently to get +past the check. See +[Make identifiers hard to fake](../patterns/ground-identifiers.md). + +### Non-curators can give feedback + +The generated pages carry thumbs-up and thumbs-down controls, and there is a +longer [evaluation form](https://go.lbl.gov/gene-eval) for detailed review. A +domain expert can disagree with an agent without learning Git. + +This is the part most projects skip. Validation catches fabricated evidence. +Only a human who knows the gene catches a claim that is well-cited and wrong. + +### Skills split reviewing from synthesizing + +`annotation-reviewer` judges what exists. `core-function-synthesizer` writes what +the gene does. `aigr-pr-review` reviews the pull request. Three jobs, three +skills, three sets of instructions that do not interfere. + +## What to copy first + +Copy the action-and-reason record shape. If your agents change existing curated +statements, record the verdict and the reason next to the original instead of +replacing it. Review gets much cheaper. + +## Gaps + +The public voting data is feedback, not yet a measurement. Turning votes into a +number you can track over releases is still open work, and it is the obvious +next step for anyone copying this pattern. diff --git a/docs/case-studies/cell-ontology.md b/docs/case-studies/cell-ontology.md new file mode 100644 index 0000000..c0e5c82 --- /dev/null +++ b/docs/case-studies/cell-ontology.md @@ -0,0 +1,54 @@ +# Cell Ontology + +Cell Ontology runs two agent workflows and has no skills, subagents, or +commands. It is the most useful negative example on this site. + +**Repository**: [obophenotype/cell-ontology](https://github.com/obophenotype/cell-ontology) +**Product**: `src/ontology/cl-edit.owl` +**Agent surface**: GitHub mention, pull request review + +## What to look at + +| Path | What it is | +| --- | --- | +| [`CLAUDE.md`](https://github.com/obophenotype/cell-ontology/blob/master/CLAUDE.md) | The whole agent configuration | +| [`.github/workflows/ai-agent.yml`](https://github.com/obophenotype/cell-ontology/blob/master/.github/workflows/ai-agent.yml) | Runs an agent on a GitHub mention | +| [`.github/workflows/clara-review.yml`](https://github.com/obophenotype/cell-ontology/blob/master/.github/workflows/clara-review.yml) | Reviews pull requests | + +## What works + +There is one instructions file. Most repositories in +[Block A](index.md#block-a-ontologies) keep the same text in both `CLAUDE.md` +and `.github/copilot-instructions.md` and have to edit both. Cell Ontology does +not have that problem. + +`CLAUDE.md` is specific about search. It tells the agent that `cl-edit.owl` has +one axiom per line, and gives exact `grep` commands for finding a term by +identifier and by label. Concrete commands beat general advice. + +It is also specific about identifiers. New term requests use the `CL_99xxxxx` +range, defined in `src/ontology/cl-idranges.owl`, and the file says never to +guess a term identifier or a PubMed identifier. + +## What to copy first + +Copy the search section. Three worked `grep` examples against your own edit file +will save more agent time than a page of prose. + +## Gaps + +Two workflows send real work to the agent, and nothing underneath breaks that +work into parts. Every run starts from the same general instructions. + +`CLAUDE.md` contains this line: + +> DO NOT bother doing your own greps over the file, or looking for other files, +> unless otherwise asked, you will just waste time. + +That is a fix for something that went wrong once, written as a prohibition. A +search skill would do the same job and would also tell the agent what to do +instead of what to avoid. Prohibitions accumulate; skills compose. + +Of the repositories we track, this one would gain the most from a small set of +skills. [Uberon's](uberon.md) `identifier-validator` and `metadata-checker` are +a reasonable starting pair. diff --git a/docs/case-studies/communitymech.md b/docs/case-studies/communitymech.md new file mode 100644 index 0000000..0de6e79 --- /dev/null +++ b/docs/case-studies/communitymech.md @@ -0,0 +1,89 @@ +# CommunityMech and the CultureBot mechs + +A knowledge base of microbial communities, and the clearest evidence that the +[DisMech](dismech.md) pattern transfers to a different domain and a different +group. + +**Repository**: [CultureBotAI/CommunityMech](https://github.com/CultureBotAI/CommunityMech) +**Product**: `kb/communities/*.yaml`, one file per community + +CommunityMech says so directly in its README: *"Adapted from Monarch +Initiative's dismech"*. It is one of six knowledge bases in the +[CultureBotAI](https://github.com/CultureBotAI) organization built on the same +shape. + +## What to look at + +| Path | What it is | +| --- | --- | +| `src/communitymech/schema/communitymech.yaml` | The LinkML schema | +| `kb/communities/` | One YAML file per community | +| `src/communitymech/export/kgx_export.py` | Koza transform to knowledge graph edges | +| `src/communitymech/export/browser_export.py` | Faceted browser export | +| `justfile` | Validation and generation commands | + +The commands and layout match DisMech closely: `just validate`, +`just validate-terms`, `just validate-references`. If you have read one +repository you can navigate the other. + +## The family + +| Repository | What it curates | Scale | +| --- | --- | --- | +| [CultureMech](https://github.com/CultureBotAI/CultureMech) | Culture media recipes | 10,657 recipes from 10 sources | +| [MediaIngredientMech](https://github.com/CultureBotAI/MediaIngredientMech) | Media ingredients and their ontology mappings | 995 mapped, 136 unmapped | +| [CommunityMech](https://github.com/CultureBotAI/CommunityMech) | Microbial communities and their interactions | Tens of communities | +| [HabitatMech](habitatmech.md) | Microbial habitats | About 3,200 habitat records | +| [TraitMech](https://github.com/CultureBotAI/TraitMech) | Microbial traits, seeded from METPO | 354 records, 233 reviewed | +| [proteintraitsmech](https://github.com/CultureBotAI/proteintraitsmech) | Protein traits | Early | + +## What works + +### One shape, six knowledge bases + +Each repository curates a different kind of thing, but all of them use one YAML +file per record, a LinkML schema, ontology-grounded terms, evidence with +citations, and a `justfile`. A curator who learns one can work in any of them. +So can an agent. + +This is the strongest argument for the mech pattern. It was not designed to be +reusable, and it turned out to be. + +### The repositories feed each other + +CultureMech aggregates 10,657 media recipes. MediaIngredientMech takes the +ingredients out of those recipes and curates their ontology mappings, then +exports the validated mappings back. The output of one curation project is the +input to the next, and both sides are files under version control. + +### Curation events are recorded, including AI assistance + +MediaIngredientMech's schema has a `CurationEvent` class that records who made a +change, when, and whether an LLM helped. Provenance is part of the data model +rather than something reconstructed from Git history afterwards. + +### Cross-repository work has its own agent layer + +[culturebotai-claw](https://github.com/CultureBotAI/culturebotai-claw) +coordinates agents across the repositories for jobs no single repository owns: +schema changes that have to propagate, synchronized releases, and integration +testing. It also assigns a model tier per job, so documentation runs on a cheap +model and integrity auditing runs on an expensive one. + +This layer is early. Several of its agents are not built yet. Read it for the +idea rather than as a finished tool. + +## What to copy first + +If you are starting a new knowledge base, copy the repository layout from +CommunityMech or [HabitatMech](habitatmech.md) rather than from DisMech. +DisMech carries years of accumulated automation. These are smaller and easier to +read as a starting point. + +## Gaps + +Some `justfile` targets in the CommunityMech README are planned rather than +built. Check the `justfile` before you rely on a command. + +Several READMEs in the family contain absolute paths from a developer's laptop +in their setup instructions. Use `git clone` and ignore those lines. diff --git a/docs/case-studies/dismech.md b/docs/case-studies/dismech.md new file mode 100644 index 0000000..8d5c054 --- /dev/null +++ b/docs/case-studies/dismech.md @@ -0,0 +1,113 @@ +# DisMech + +A knowledge base of disease mechanisms. Most edits come from agents. Humans +review the process rather than every entry. + +**Repository**: [monarch-initiative/dismech](https://github.com/monarch-initiative/dismech) +**Site**: [dismech.monarchinitiative.org](https://dismech.monarchinitiative.org/app/) +**Product**: `kb/disorders/*.yaml`, one file per disorder +**Agent surface**: local sessions, Claude Code on the web, GitHub mention, and +about 28 workflows + +DisMech is the most developed agentic curation setup we know of. Read its +[`CONTRIBUTING.md`](https://github.com/monarch-initiative/dismech/blob/main/CONTRIBUTING.md) +first. It is a guide to running a curation project with agents, not only a guide +to this repository. + +## What to look at + +| Path | What it is | +| --- | --- | +| [`CONTRIBUTING.md`](https://github.com/monarch-initiative/dismech/blob/main/CONTRIBUTING.md) | How the project works, written for humans | +| [`.claude/skills/`](https://github.com/monarch-initiative/dismech/tree/main/.claude/skills) | Seventeen skills, including `curate-next`, `dismech-terms`, `dismech-references`, and `dismech-pr-review` | +| [`.github/workflows/`](https://github.com/monarch-initiative/dismech/tree/main/.github/workflows) | Review, scanners, guards, and site builds | +| `.github/agent-config.yaml` | Which model each automated job uses | +| `.github/cron-profiles.yaml` | How often each scanner runs | +| `src/dismech/schema/dismech.yaml` | The LinkML schema | +| `docs/explanation/design-decisions.md` | Why the project is built this way | +| `justfile` | Every validation command | + +## What works + +### The philosophy is written down + +`CONTRIBUTING.md` states the rules of the project in plain language: + +* Assume that issues and comments are AI-generated. If you write something + yourself and want that known, mark it `[human authored]`. +* Contributors are not expected to check agent output themselves. Review is a + process, not a personal duty. +* Use top-tier models. Weaker models produce work that review has to send back, + which costs everyone time. +* Be bold. Every team member is encouraged to do work that ends in a pull + request. + +Most projects leave these questions unanswered and each contributor guesses. + +### Humans regulate the loop + +The project asks curators to fix patterns rather than entries: + +> look for *patterns* where results are suboptimal; curate examples and +> counter-examples; work with agent to integrate this into the process + +It also asks curators to review the reviews, and to check whether each automated +process is too eager or not eager enough. See +[Regulate the loop](../patterns/human-regulating-the-loop.md). + +### Human documentation is treated as perishable + +`CONTRIBUTING.md` warns that it may be out of date, and tells you to ask the +agent instead: + +> "I want to contribute. How?" +> "Explain what this repo is" +> "I noticed a problem on one of the pages -- what should I do?" + +The repository is the current answer. The document is a snapshot. + +### Validation has a fast loop and a slow loop + +`just count-verified-snippets` checks evidence quotes against a local cache in +seconds, so you can run it while you curate. `just validate-disorders` runs the +full schema, term, and reference sweep, and is meant to run once before you open +the pull request. See +[Fast and slow validation](../patterns/fast-and-slow-validation.md). + +### Scanners find work + +Separate workflows scan the literature for new papers, look for incomplete +entries, and move stalled issues and pull requests forward. Low-risk scanners +may use cheaper models. A `low_effort` label lets a human assign a task to a +cheaper model by hand. See [Scanners](../patterns/scanners.md). + +### The untrusted surface is guarded + +Pull requests from forks are closed, because GitHub does not give fork workflows +the secrets that automated review needs. A separate workflow guards untrusted +comments, and only registered controllers can summon the agent. + +## What to copy first + +Copy `CONTRIBUTING.md` and edit it for your project. Deciding who is accountable +for agent output, and saying so in writing, costs nothing and prevents the most +common argument. + +If you are starting a new knowledge base rather than adding agents to an +existing one, copy the repository shape instead: one YAML file per record, a +LinkML schema, and a `justfile` with fast and slow validation targets. See +[Create an agentic curation pipeline](../how-tos/create-agentic-curation-pipeline.md). + +## Gaps + +The project describes itself as alpha and experimental, and says so on the +front page along with a clear statement that it is not medical advice. +Validation proves that citations exist, that quoted text is exact, and that +ontology terms are real. It does not prove that a claim is scientifically +correct. That distinction is stated in the README, and every project running +this pattern should state it too. + +Twenty-eight workflows is a lot to hold in your head. The repository handles +this by telling you to ask an agent for the current picture rather than reading +a list. That works, but it means a newcomer cannot audit the automation without +running an agent. diff --git a/docs/case-studies/efo.md b/docs/case-studies/efo.md new file mode 100644 index 0000000..a8fdf94 --- /dev/null +++ b/docs/case-studies/efo.md @@ -0,0 +1,65 @@ +# EFO + +The Experimental Factor Ontology. EFO treats the main agent as an orchestrator +that routes work to specialists, and keeps one review checklist for three +different agent runtimes. + +**Repository**: [EBISPOT/efo](https://github.com/EBISPOT/efo) +**Product**: `src/ontology/efo-edit.owl` +**Agent surface**: local sessions, started with a command + +## What to look at + +| Path | What it is | +| --- | --- | +| [`CLAUDE.md`](https://github.com/EBISPOT/efo/blob/master/CLAUDE.md) | The orchestrator role and its routing table | +| [`AGENTS.md`](https://github.com/EBISPOT/efo/blob/master/AGENTS.md) | A short pointer to the authoritative files | +| [`.claude/agents/`](https://github.com/EBISPOT/efo/tree/master/.claude/agents) | Specialist subagents | +| [`.github/agents/`](https://github.com/EBISPOT/efo/tree/master/.github/agents) | The same specialists for GitHub Copilot | +| [`.claude/commands/`](https://github.com/EBISPOT/efo/tree/master/.claude/commands) | Commands, including `/efo-ticket` | +| [`.mcp.json`](https://github.com/EBISPOT/efo/blob/master/.mcp.json) | Two MCP servers: OLS and artl-mcp | +| `docs/agents-documentation/efo-pr-review-checklist.md` | One review checklist, shared by three reviewers | + +## What works + +**`CLAUDE.md` gives the agent a role.** It says the agent is the orchestrator, +then gives a table of which specialist handles which kind of request. Most +repositories describe the project; EFO describes the job. + +**`AGENTS.md` is a pointer, not a copy.** It is about twenty lines and it names +`.github/copilot-instructions.md` and `CLAUDE.md` as the authoritative sources. +Nothing has to be kept in sync by hand. Every other Block A repository +duplicates its instructions across two files. See +[One source of instructions](../patterns/one-source-of-instructions.md). + +**One checklist, three reviewers.** The Claude, Copilot, and Codex reviewers all +apply `docs/agents-documentation/efo-pr-review-checklist.md`. The Codex reviewer +has its own procedure file that points at the same checklist. When the review +standard changes, one file changes. + +**`/efo-ticket ` is a front door.** A curator does not have to +describe the workflow. The command holds it. + +**MCP servers are declared in the repository.** `.mcp.json` gives every session +the same two tools: the EBI Ontology Lookup Service over HTTP, and `artl-mcp` +for literature. Curators do not configure these themselves. + +## What to copy first + +Copy `AGENTS.md`. It is the cheapest fix on this site: replace your duplicated +instructions file with a pointer, and the drift problem disappears today. + +After that, copy the shared review checklist. One file that all your reviewers +read is worth more than three reviewers with three standards. + +## Gaps + +EFO has no workflow that runs an agent when someone mentions it on an issue. The +orchestrator and its specialists are reachable from a local session only. GO, +Uberon, and Cell Ontology all have that surface. This looks like a deliberate +choice, but it means EFO issues do not get agent responses. + +The specialists exist twice, once under `.claude/agents/` and once under +`.github/agents/`, with names that differ only in case. The two sets can drift. +The review checklist shows the fix EFO already knows: one authoritative file, +referenced from both places. diff --git a/docs/case-studies/go-ontology.md b/docs/case-studies/go-ontology.md new file mode 100644 index 0000000..7e09c06 --- /dev/null +++ b/docs/case-studies/go-ontology.md @@ -0,0 +1,51 @@ +# GO ontology + +The Gene Ontology edit file, with the largest library of agent skills of any +ontology repository we track. + +**Repository**: [geneontology/go-ontology](https://github.com/geneontology/go-ontology) +**Product**: `src/ontology/go-edit.obo` +**Agent surface**: GitHub mention, pull request review, local sessions + +## What to look at + +| Path | What it is | +| --- | --- | +| [`.claude/skills/`](https://github.com/geneontology/go-ontology/tree/master/.claude/skills) | Ten skills, one per recurring job | +| [`.github/workflows/ai-agent.yml`](https://github.com/geneontology/go-ontology/blob/master/.github/workflows/ai-agent.yml) | Runs an agent when someone mentions it in an issue | +| [`.github/workflows/claude-code-review.yml`](https://github.com/geneontology/go-ontology/blob/master/.github/workflows/claude-code-review.yml) | Reviews pull requests | +| [`CLAUDE.md`](https://github.com/geneontology/go-ontology/blob/master/CLAUDE.md) | Repository instructions | + +The ten skills are `chemical-entity`, `design-pattern`, `external-term-lookup`, +`mapping`, `odk-make`, `pr-review`, `reaction`, `research`, +`taxon-constraint`, and `term-obsoletion`. + +## What works + +The skill names map onto jobs a GO editor already recognizes. An editor who has +handled a taxon constraint ticket knows what `taxon-constraint` is for. This +makes the setup readable to curators, not only to developers. + +Skills also keep instructions out of the main context. `CLAUDE.md` stays short +because the detail lives in the skill that needs it. See +[Break work into skills](../patterns/skills-before-automation.md). + +`odk-make` is worth a look on its own. It wraps the Ontology Development Kit +commands so the agent runs the same build steps an editor runs, instead of +inventing its own. + +## What to copy first + +Copy the idea, not the files. List the five tickets your editors handle most +often. Write one skill for each. Name each skill after the ticket type. + +## Gaps + +There is no command or subagent that sequences the skills. Each run plans its +own path from issue to pull request. EFO solves this with an +[orchestrator and a `/efo-ticket` command](efo.md). GO has the workflow wiring +to make the same approach useful. + +`CLAUDE.md` and `.github/copilot-instructions.md` hold the same text in two +files. Both need editing by hand when guidance changes. See +[One source of instructions](../patterns/one-source-of-instructions.md). diff --git a/docs/case-studies/habitatmech.md b/docs/case-studies/habitatmech.md new file mode 100644 index 0000000..f566743 --- /dev/null +++ b/docs/case-studies/habitatmech.md @@ -0,0 +1,73 @@ +# HabitatMech + +A knowledge base of microbial habitats that merges four source vocabularies into +one record per habitat. It is the best worked example of using agents to +harmonize sources that disagree. + +**Repository**: [CultureBotAI/HabitatMech](https://github.com/CultureBotAI/HabitatMech) +**Site**: [Browse the corpus](https://culturebotai.github.io/HabitatMech/) +**Product**: `data/habitats//.yaml` + +HabitatMech follows the [DisMech](dismech.md) pattern and is part of the +[CultureBot family](communitymech.md). + +## The problem it solves + +The same habitat has a different name in every source: + +| Source | How it names marine sediment | +| --- | --- | +| JGI GOLD | `Environmental > Aquatic > Marine > Sediment` | +| BacDive | `Marine-sediment` | +| PREGO | `ENVO:00002113` | +| Madin et al. | `ENVO:00002113` | + +## What to look at + +| Path | What it is | +| --- | --- | +| `data/habitats/` | One YAML file per habitat, about 3,200 records | +| [Term requests page](https://culturebotai.github.io/HabitatMech/term-requests.html) | Gaps this project is asking ENVO to fill | +| `src/habitatmech/schema/` | The LinkML schema | + +## What works + +### The merge is the product + +Each source name becomes a source concept. Every source concept resolves to an +identifier. Source concepts that resolve to the same identifier merge into one +record that keeps all of their attestations. + +So `data/habitats/terrestrial/soil.yaml` is one record, grounded in +`ENVO:00001998`, that knows GOLD saw 26,399 organisms there and that PREGO and +Madin's literature curation associate 8,715 and 2,934 taxa with it +independently. The disagreement between sources is kept, not flattened. + +### Minted identifiers are honest about being minted + +Where an ontology term is defensible, the record uses the ontology CURIE. Where +none is, the project mints a content-hashed `habitatmech:` CURIE instead of +forcing a bad ontology match. You can tell the two apart by looking. + +This matters for agents. Forcing a term match is exactly the kind of plausible +error an agent makes, and a project that permits a local identifier removes the +pressure to guess. + +### Gaps are published as term requests + +The project renders the terms it needs and cannot find as a public +[term requests page](https://culturebotai.github.io/HabitatMech/term-requests.html) +for the ontology community. Curation that finds a gap produces a request rather +than a workaround. + +## What to copy first + +Copy the term request page. If your curation depends on an ontology you do not +control, publish what you needed and could not find. It turns a private +annoyance into a contribution the ontology maintainers can act on. + +## Gaps + +The corpus was seeded in a batch from kg-microbe. As with any seeded knowledge +base, coverage reflects what the sources contained, not what the domain +contains. Check the README for the seed date before treating counts as current. diff --git a/docs/case-studies/index.md b/docs/case-studies/index.md new file mode 100644 index 0000000..f1c8091 --- /dev/null +++ b/docs/case-studies/index.md @@ -0,0 +1,58 @@ +# Case studies + +Every page in this section describes one repository that runs agents on real +curation work. Each page points at the files and folders you can open and copy. + +We do not centralize the material here. The repositories are the source of +truth. If a page disagrees with the repository, trust the repository and +[open an issue](https://github.com/ai4curation/aidocs/issues). + +## Block A: ontologies + +These are established OBO ontologies that added agents to an existing editor +workflow. The ontology file stays the product. Agents work through issues and +pull requests. + +| Repository | Best known for | Read this first | +| --- | --- | --- | +| [GO](go-ontology.md) | The largest skill library of any ontology repository | `.claude/skills/` | +| [Uberon](uberon.md) | Subagents for specialist jobs | `.claude/agents/` | +| [Mondo](mondo.md) | A Goose-based agent alongside Claude Code subagents | `.github/workflows/ai-agent.yml` | +| [Cell Ontology](cell-ontology.md) | Live agent traffic with almost no supporting structure | `CLAUDE.md` | +| [EFO](efo.md) | An orchestrator that routes work to specialists | `CLAUDE.md` | + +## Block B: the mechs + +These are newer knowledge bases built for agents from the start. The record is +a YAML file. Validation is strict. Most edits come from agents. + +| Repository | What it curates | Read this first | +| --- | --- | --- | +| [DisMech](dismech.md) | Disease mechanisms | `CONTRIBUTING.md` | +| [AI Gene Review](ai-gene-review.md) | Gene Ontology annotations | `.claude/skills/` | +| [CommunityMech](communitymech.md) | Microbial communities | `src/communitymech/schema/` | +| [HabitatMech](habitatmech.md) | Microbial habitats | `data/habitats/` | + +Four more repositories in the CultureBot family follow the same shape: +[CultureMech](https://github.com/CultureBotAI/CultureMech) (culture media), +[MediaIngredientMech](https://github.com/CultureBotAI/MediaIngredientMech) +(media ingredients), +[TraitMech](https://github.com/CultureBotAI/TraitMech) (microbial traits), and +[proteintraitsmech](https://github.com/CultureBotAI/proteintraitsmech) (protein +traits). See [CommunityMech](communitymech.md) for how the family fits together. + +## How to read a case study + +Each page has the same four sections: + +* **What to look at** lists paths in the repository. +* **What works** describes practices worth copying. +* **What to copy first** is the shortest useful step. +* **Gaps** records what the repository has not solved. These sections are + honest, not critical. Every setup here has gaps. + +## Where the observations come from + +Counts of skills, subagents, and workflows come from agent-watcher, which scans +these repositories on a schedule and publishes dated reports. See +[Evidence](../evidence.md). diff --git a/docs/case-studies/mondo.md b/docs/case-studies/mondo.md new file mode 100644 index 0000000..553ed94 --- /dev/null +++ b/docs/case-studies/mondo.md @@ -0,0 +1,53 @@ +# Mondo + +The Mondo Disease Ontology. Mondo is the clearest example of a repository whose +automated agent and whose subagent library are on different tracks. + +**Repository**: [monarch-initiative/mondo](https://github.com/monarch-initiative/mondo) +**Product**: `src/ontology/mondo-edit.obo` +**Agent surface**: GitHub mention through dragon-ai-agent, local Claude Code sessions + +## What to look at + +| Path | What it is | +| --- | --- | +| [`.github/workflows/ai-agent.yml`](https://github.com/monarch-initiative/mondo/blob/master/.github/workflows/ai-agent.yml) | The automated agent, built on dragon-ai-agent and Goose | +| [`.claude/agents/`](https://github.com/monarch-initiative/mondo/tree/master/.claude/agents) | Six subagents for Claude Code | +| [`.config/goose/`](https://github.com/monarch-initiative/mondo/tree/master/.config/goose) | Goose configuration used by the workflow | +| [`CLAUDE.md`](https://github.com/monarch-initiative/mondo/blob/master/CLAUDE.md) | Repository instructions | + +The subagents are `deep-research-specialist`, `design-pattern-advisor`, +`identifier-validator`, `metadata-checker`, `ontology-reasoner`, and +`task-coordinator`. Four of these names also appear in +[Uberon](uberon.md). + +## What works + +Mondo has run agent-assisted curation longer than most of the repositories on +this site. Its issue tracker is a good place to read real agent pull requests +from end to end, including the ones that went wrong: +[issues involving dragon-ai-agent](https://github.com/monarch-initiative/mondo/issues?q=involves%3Adragon-ai-agent). + +The subagents are a working library. If you want a starting set for an OBO +ontology, read Mondo's and Uberon's together and take the four they share. + +## What to copy first + +Read the merged agent pull requests before you copy any configuration. Mondo has +enough history to show you what agent output looks like after review, which is +more useful than any template. + +## Gaps + +The only automated agent workflow runs dragon-ai-agent and Goose. The six +subagents under `.claude/agents/` are Claude Code features. An agent triggered +by `ai-agent.yml` therefore cannot reach them. They are available in local +Claude Code sessions only. + +This is not necessarily wrong, but it is undocumented. If you have both, say in +`CLAUDE.md` which surface each one serves. If you want the subagents reachable +from GitHub, add a workflow that runs Claude Code the way +[GO](go-ontology.md) and [Uberon](uberon.md) do. + +Goose is no longer a tool we recommend for new setups. See +[Harnesses](../reference/harnesses.md). diff --git a/docs/case-studies/uberon.md b/docs/case-studies/uberon.md new file mode 100644 index 0000000..1a9bdea --- /dev/null +++ b/docs/case-studies/uberon.md @@ -0,0 +1,50 @@ +# Uberon + +The multi-species anatomy ontology. Uberon splits agent work across eight +subagents instead of skills. + +**Repository**: [obophenotype/uberon](https://github.com/obophenotype/uberon) +**Product**: `src/ontology/uberon-edit.obo` +**Agent surface**: GitHub mention, pull request review, local sessions + +## What to look at + +| Path | What it is | +| --- | --- | +| [`.claude/agents/`](https://github.com/obophenotype/uberon/tree/master/.claude/agents) | Eight subagents | +| [`.github/workflows/ai-agent.yml`](https://github.com/obophenotype/uberon/blob/master/.github/workflows/ai-agent.yml) | Runs an agent on a GitHub mention | +| [`CLAUDE.md`](https://github.com/obophenotype/uberon/blob/master/CLAUDE.md) | Repository instructions | + +The subagents are `deep-research-specialist`, `design-pattern-advisor`, +`identifier-validator`, `metadata-checker`, `ntr-term-researcher`, +`ontology-reasoner`, `ontology-term-lookup`, and `task-coordinator`. + +## What works + +A subagent runs in its own context and returns a result. This suits jobs that +read a lot and report a little. `deep-research-specialist` can read twenty +papers and hand back three sentences, and the main session never carries the +twenty papers. + +`identifier-validator` and `metadata-checker` are checkers. They run after the +edit and look for problems. Splitting the checker from the editor is useful: +the checker has no stake in the edit being correct. + +`ntr-term-researcher` handles new term requests, the most common ticket type in +an anatomy ontology. Uberon named a subagent after its highest-volume job. + +## What to copy first + +Take `identifier-validator` and `metadata-checker`. They are the two subagents +that transfer to any OBO ontology with the least change. Mondo already runs +both. + +## Gaps + +Uberon has one skill and eight subagents. Mondo has six subagents with four of +the same names. Neither repository states whether the shared subagents are kept +in sync or were copied once and left to drift. If you copy them, record where +you copied them from. + +`CLAUDE.md` and `.github/copilot-instructions.md` hold the same text. See +[One source of instructions](../patterns/one-source-of-instructions.md). diff --git a/docs/evidence.md b/docs/evidence.md new file mode 100644 index 0000000..8e01edd --- /dev/null +++ b/docs/evidence.md @@ -0,0 +1,86 @@ +# Evidence + +We measure what these repositories actually do, rather than what they say they +do. This page describes where the observations on this site come from. + +## agent-watcher + +`ai4curation/agent-watcher` scans repositories where agents are deployed and +publishes dated reports. + +!!! note + + The agent-watcher repository is currently private, so the links in this + section need access. The counts below are reproduced here so you can read + them without it. + +It runs two jobs: + +**Per-repository activity reports.** For each watched repository, it collects +recently updated issues and pull requests, marks the ones that involve an agent, +and asks a model to write a plain-language assessment of what is working and +what is not. Reports appear as dated issues, one per repository per report date. + +**A weekly cross-repository setup review.** This one compares the ontology +repositories against each other: instruction files, workflow conventions, and +whether repeated work is broken into skills, commands, or subagents. + +The setup reviews are the source of the skill, subagent, and workflow counts in +our [case studies](case-studies/index.md). This is the snapshot from the setup review of 10 August 2026: + +| Repository | Instructions | Subagents | Skills and commands | Agent workflows | +| --- | ---: | ---: | ---: | ---: | +| [EFO](case-studies/efo.md) | 3 | 8 | 2 | 1 | +| [GO](case-studies/go-ontology.md) | 2 | 0 | 10 | 3 | +| [Mondo](case-studies/mondo.md) | 2 | 6 | 4 | 2 | +| [Cell Ontology](case-studies/cell-ontology.md) | 1 | 0 | 0 | 2 | +| [Uberon](case-studies/uberon.md) | 2 | 8 | 1 | 3 | + +Counts change. Check the latest report before you rely on a number here. + +### Why this design is worth copying + +Two parts of agent-watcher transfer to other monitoring jobs. + +**The collector does not judge.** It produces a neutral document of counts and +timelines. The model reads that document and writes the assessment. Keeping the +measurement separate from the judgment means you can check both. + +**Reading and writing are separated.** The watcher reads the repositories it +watches with a read-only token, and publishes its reports into a different +repository. It cannot modify anything it observes. See +[Guard the untrusted surface](patterns/guard-untrusted-input.md). + +## Execution traces + +agent-watcher also mines and publishes agent execution traces under +`public-traces/`. Each trace covers one agent pull request: the pull request, +the workflow runs behind it, and the log. + +Traces are the only way to answer questions like which tool the agent called +before it produced a wrong answer, or how many attempts a task took. Reports +tell you what happened. Traces tell you why. + +## Provenance in the repository + +[ai-blame](https://github.com/ai4curation/ai-blame) extracts provenance from +agent execution traces and gives you line-level attribution: which lines an +agent wrote, when, and in what session. + +Some projects record provenance in the data itself instead. The CultureBot +mechs have a `CurationEvent` class in their schema that records who made each +change and whether a model assisted. See +[CommunityMech](case-studies/communitymech.md). + +## What we do not measure yet + +* **Quality over time.** We can count agent pull requests. We cannot yet say + whether the entries they produce are getting better. +* **What review catches.** Automated reviewers request changes. Nobody has + measured which categories of error they reliably catch and which they miss. +* **Community feedback as a signal.** + [AI Gene Review](case-studies/ai-gene-review.md) collects votes from domain + experts. Turning those into a number you can track across releases is open + work. + +These are the three most useful things anyone reading this site could build. diff --git a/docs/examples.md b/docs/examples.md deleted file mode 100644 index efb5d97..0000000 --- a/docs/examples.md +++ /dev/null @@ -1,85 +0,0 @@ -# Example Repositories Using AI4Curation Setup - -## Ontology Repositories - -### Mondo Disease Ontology -**Repository**: [monarch-initiative/mondo](https://github.com/monarch-initiative/mondo) - -**Key Files to Explore**: -- [ai-agent.yml](https://github.com/monarch-initiative/mondo/blob/master/.github/workflows/ai-agent.yml) - GitHub Actions workflow configuration -- [CLAUDE.md](https://github.com/monarch-initiative/mondo/blob/master/CLAUDE.md) - Agent instructions -- [.config/goose/config.yaml](https://github.com/monarch-initiative/mondo/tree/master/.config/goose) - Goose configuration - -**Agent Platform**: Goose (via dragon-ai-agent) - -**Agent Activity**: [Search for @dragon-ai-agent mentions and PRs](https://github.com/monarch-initiative/mondo/issues?q=involves%3Adragon-ai-agent) - -### Cell Ontology (CL) -**Repository**: [obophenotype/cell-ontology](https://github.com/obophenotype/cell-ontology) - -**Key Files to Explore**: -- CLAUDE.md - Agent instructions - -**Agent Platform**: Goose (via dragon-ai-agent) - -**Agent Activity**: [Search for @dragon-ai-agent mentions and PRs](https://github.com/obophenotype/cell-ontology/issues?q=involves%3Adragon-ai-agent) - -### Uberon Anatomy Ontology -**Repository**: [obophenotype/uberon](https://github.com/obophenotype/uberon) - -**Key Files to Explore**: -- [PR #3580](https://github.com/obophenotype/uberon/pull/3580) - GitHub Copilot integration - -**Agent Platform**: GitHub Copilot, Goose - -**Agent Activity**: [Search for @dragon-ai-agent mentions and PRs](https://github.com/obophenotype/uberon/issues?q=involves%3Adragon-ai-agent) - -### EFO (Experimental Factor Ontology) -**Repository**: [EBISPOT/efo](https://github.com/EBISPOT/efo) - -**Agent Platform**: GitHub Copilot - -**Agent Activity**: [Search for Copilot activity](https://github.com/EBISPOT/efo/issues?q=author%3Aapp%2Fcopilot) - -## Agentic Infrastructure - -### ai-blame -**Repository**: [ai4curation/ai-blame](https://github.com/ai4curation/ai-blame) - -Provenance extraction from agent execution traces. Demonstrates line-level attribution for AI-assisted edits. - -- [Docs](https://ai4curation.github.io/ai-blame) | [PyPI](https://pypi.org/project/ai-blame/) - -### curation-skills -**Repository**: [ai4curation/curation-skills](https://github.com/ai4curation/curation-skills) - -Reusable skill packs for ontology and biocuration tasks. Example of structuring agent behavior for consistency. - -- [Skills Article](https://anthropic.com/engineering/equipping-agents-for-the-real-world-with-agent-skills) - -### noctua-mcp -**Repository**: [geneontology/noctua-mcp](https://github.com/geneontology/noctua-mcp) - -MCP server for GO-CAM editing — an example of wrapping a domain-specific API for agent access. - -- [PyPI](https://pypi.org/project/noctua-mcp/) - -### ICBO 2025 AI Tutorial -**Repository**: [ai4curation/icbo-ai-tutorial](https://github.com/ai4curation/icbo-ai-tutorial) - -Tutorial materials for agent-assisted ontology curation workflows, with exercises and social coding patterns. - -- [Tutorial Page](https://go.lbl.gov/icbo-2025-ai) | [Recording](https://youtu.be/_9Re39yB7EE) | [Zenodo](https://zenodo.org/records/18653147) - -## Documentation Repositories - -### AI4Curators Documentation -**Repository**: [ai4curation/aidocs](https://github.com/ai4curation/aidocs) (this repository) - -**Key Files to Explore**: -- `.github/workflows/` - GitHub Actions for AI agent integration -- `CLAUDE.md` - Repository-specific instructions - -**Agent Platform**: Claude Code, Goose - -**Agent Activity**: [Search for @dragon-ai-agent mentions and PRs](https://github.com/ai4curation/aidocs/issues?q=involves%3Adragon-ai-agent) diff --git a/docs/faq.md b/docs/faq.md index fe05a49..c9aaafe 100644 --- a/docs/faq.md +++ b/docs/faq.md @@ -4,14 +4,14 @@ ### What is AI4Curators? -AI4Curators is a project focused on providing practical guides for curators and maintainers of knowledge bases to integrate AI into their workflows. Rather than theoretical discussions, we emphasize immediate, actionable integration strategies that work with existing GitHub-based workflows. +AI4Curators provides practical guides for curators and maintainers of knowledge +bases who want to use AI agents in the workflows they already have. -Our core mission includes: -- Helping curators integrate AI agents into existing GitHub-based workflows -- Providing plugins and tools for existing chat UIs -- Supporting ontology editing and curation workflows with AI assistance +We do not centralize guidance. We point at repositories that run agents on real +curation work and explain what they do, so you can copy a working setup. See +[Case studies](case-studies/index.md). -The project serves both as documentation and a practical example of AI agent integration, where GitHub agents can directly contribute to documentation and examples demonstrate real-world AI-assisted curation workflows. +This repository is itself run this way. Agents contribute to it directly. ## GitHub Copilot Integration @@ -35,7 +35,7 @@ GitHub Copilot's coding agent includes a firewall that restricts internet access - ✅ Works: `https://raw.githubusercontent.com/oborel/obo-relations/refs/heads/master/ro-base.owl` - ❌ Blocked: `http://purl.obolibrary.org/obo/ro/ro-base.owl` -For more details, see the [GitHub Copilot](reference/clients/github-copilot.md) documentation and GitHub's guide on [customizing the agent firewall](https://docs.github.com/en/copilot/how-tos/use-copilot-agents/coding-agent/customize-the-agent-firewall). +For more details, see the [GitHub Copilot](reference/github-copilot.md) documentation and GitHub's guide on [customizing the agent firewall](https://docs.github.com/en/copilot/how-tos/use-copilot-agents/coding-agent/customize-the-agent-firewall). ## Agent Harness and Infrastructure @@ -66,4 +66,34 @@ Use **[ai-blame](https://github.com/ai4curation/ai-blame)**, which extracts prov - **[oak-mcp](https://github.com/monarch-initiative/oak-mcp)** — ontology search, traversal, and operations via OAK - **[owl-mcp](https://github.com/monarch-initiative/owl-mcp)** — general OWL ontology operations -These give agents structured access to domain-specific operations instead of raw file manipulation. \ No newline at end of file +These give agents structured access to domain-specific operations instead of raw file manipulation. + +### Which harness should I use? + +Claude Code or Codex. Both work. Your repository setup matters more than the +choice between them. See [Harnesses](reference/harnesses.md). + +We previously recommended Goose for curators who were not comfortable at the +command line. We no longer do for new setups. + +### Do I have to install anything? + +No. [Claude Code on the web](https://code.claude.com/docs/en/claude-code-on-the-web) +runs the session in a cloud container, clones the repository, and opens the pull +request for you. See [Your first curation session](tutorials/first-curation-session.md). + +### Who is responsible for checking what my agent wrote? + +Your project has to answer this, and write the answer down. + +[DisMech](case-studies/dismech.md) states in its `CONTRIBUTING.md` that +contributors are not assumed to have verified everything their agent produced, +and that review is the process that catches problems. Other projects hold the +contributor responsible. Either works. Leaving it unstated does not. See +[Regulate the loop](patterns/human-regulating-the-loop.md). + +### Does validation mean the content is correct? + +No. Validation proves that a cited paper exists, that a quoted sentence is +exact, and that an ontology term is real. It does not prove that a claim is +scientifically correct. Say so where your readers can see it. diff --git a/docs/glossary.md b/docs/glossary.md index b68fb1e..dafd177 100644 --- a/docs/glossary.md +++ b/docs/glossary.md @@ -49,12 +49,15 @@ An automation platform provided by GitHub that allows users to automate software ## Goose -An AI application and host environment that that allows you to interact with an AI model while also accessing local files, running local code, or using [MCPs](#model-context-protocol-mcp) +An AI application and host environment that lets you interact with a model while +also accessing local files, running local code, or using +[MCPs](#model-context-protocol-mcp). It comes as a desktop app and a command +line app. -Goose is actually two separate programs - -* A Desktop app (recommended for non-technical tasks) -* A Terminal / Command Line app, similar to Claude Code. +Goose is historic on this site. Earlier guidance recommended it for curators who +were not comfortable at the command line. We no longer recommend it for new +setups, though it is still deployed in some repositories. See +[Harnesses](reference/harnesses.md). ## OBO Format A standardized file format used for creating and exchanging ontologies, particularly prevalent in the biomedical and life sciences domains. For more details, see the [OBO Flat File Format Specification](http://owlcollab.github.io/oboformat/doc/GO.format.obo-1_4.html). @@ -123,3 +126,30 @@ Here the community can be conceived of as containing both humans and AI agents. ## Tool In the context of AI agents, a tool refers to a specific function or capability that an AI model can access to perform actions beyond text generation, such as reading files, executing code, making API calls, or interacting with external systems. + +## Mech + +An informal name for a knowledge base built on the pattern established by +[DisMech](case-studies/dismech.md): one YAML file per record, a LinkML schema, +ontology-grounded terms, evidence with citations that are checked +automatically, and curation carried out mostly by agents. Used by +[several knowledge bases](case-studies/communitymech.md) beyond the original. + +## Scanner + +A scheduled job that looks for curation work and creates it, rather than waiting +for a human to ask. Examples are literature scans, gap scans, and sweeps for +stalled issues. See [Scanners find the work](patterns/scanners.md). + +## Skill + +A folder holding a `SKILL.md` file that describes one job for an agent. The +agent loads it when that job comes up and ignores it otherwise. See +[Break work into skills](patterns/skills-before-automation.md). + +## Subagent + +A separate agent run started by the main session, with its own context. Useful +for jobs that read a great deal and report a little, because the reading does +not stay in the main session. See +[Break work into skills](patterns/skills-before-automation.md). diff --git a/docs/how-tos/author-skills.md b/docs/how-tos/author-skills.md index b39d5fb..949cd8f 100644 --- a/docs/how-tos/author-skills.md +++ b/docs/how-tos/author-skills.md @@ -185,7 +185,7 @@ The principles above are easier to absorb from real, working skill sets. Both re ### Monarch dismech -[`monarch-initiative/dismech`](https://github.com/monarch-initiative/dismech) is a knowledge base of disease mechanisms, stored as LinkML-schema YAML files with bindings to ontologies like MONDO, HPO, and GO. It ships ~14 skills covering the full curation lifecycle, and is a good example of using skills as **standard operating procedures** rather than just capability shims. +[`monarch-initiative/dismech`](https://github.com/monarch-initiative/dismech) is a knowledge base of disease mechanisms, stored as LinkML-schema YAML files with bindings to ontologies like MONDO, HPO, and GO. It ships seventeen skills covering the full curation lifecycle, and is a good example of using skills as **standard operating procedures** rather than just capability shims. Patterns worth copying: @@ -195,7 +195,7 @@ Patterns worth copying: ### GO ontology -[`geneontology/go-ontology`](https://github.com/geneontology/go-ontology) is the development repo for the Gene Ontology. Its eight skills (`design-pattern`, `taxon-constraint`, `term-obsoletion`, `external-term-lookup`, `mapping`, `reaction`, `chemical-entity`, `research`) are a clean illustration of the "[consider not writing a skill](#consider-not-writing-a-skill)" point inverted: the agent already knows ontology basics, so these skills exist to encode **latent, project-specific knowledge that overrides its defaults**. +[`geneontology/go-ontology`](https://github.com/geneontology/go-ontology) is the development repo for the Gene Ontology. Its ten skills (`design-pattern`, `taxon-constraint`, `term-obsoletion`, `external-term-lookup`, `mapping`, `reaction`, `chemical-entity`, `research`, `odk-make`, `pr-review`) are a clean illustration of the "[consider not writing a skill](#consider-not-writing-a-skill)" point inverted: the agent already knows ontology basics, so these skills exist to encode **latent, project-specific knowledge that overrides its defaults**. Patterns worth copying: @@ -203,6 +203,28 @@ Patterns worth copying: - **Domain SOPs with explicit judgement calls.** [`taxon-constraint`](https://github.com/geneontology/go-ontology/tree/master/.claude/skills/taxon-constraint) encodes when *not* to act: prefer parsimonious (broader) constraints, skip constraints already inherited from CL/UBERON, and remove them when a term is obsoleted. - **Cross-referencing skills for composition.** The `design-pattern` skill defers chemical terms to the `chemical-entity` skill rather than duplicating that knowledge — modular skills that point at each other instead of repeating content. -## Optional: Create a skills marketplace +For more on both repositories, see the +[DisMech](../case-studies/dismech.md) and [GO](../case-studies/go-ontology.md) +case studies. -TODO +## Optional: create a skills marketplace + +Once several repositories in your organization need the same skill, stop copying +it. A marketplace is a repository of skills that others install from, so a fix +lands in one place. + +[curation-skills](https://github.com/ai4curation/curation-skills) holds the +shared ontology and biocuration skills for this organization. The CultureBot +mechs keep a shared `culturebot-skills` repository and use it the same way +across their six knowledge bases. + +Use a marketplace when a skill is genuinely general, such as term lookup or +reference validation. Keep skills that encode one project's conventions in that +project. A shared skill that has to describe five projects' conventions helps +none of them. + +Whichever you choose, record where a skill came from. The main risk with copied +skills is silent divergence: [Uberon](../case-studies/uberon.md) and +[Mondo](../case-studies/mondo.md) share four subagent names, and neither +repository says whether the shared ones are kept in sync. See +[One source of instructions](../patterns/one-source-of-instructions.md). diff --git a/docs/how-tos/build-agentic-harness.md b/docs/how-tos/build-agentic-harness.md index a17b96a..c356ae0 100644 --- a/docs/how-tos/build-agentic-harness.md +++ b/docs/how-tos/build-agentic-harness.md @@ -26,7 +26,7 @@ The ai4curation ecosystem already provides the building blocks. Here's how they ### 1. Prompt preset management — System instructions -Every repository should have a `CLAUDE.md` (for Claude Code / GitHub agents) and/or `.goosehints` (for Goose) checked into the root. These files tell the agent what the repository is, what conventions to follow, and what tools to use. +Every repository should have a `CLAUDE.md` checked into the root. It tells the agent what the repository is, what conventions to follow, and what tools to use. If you also need `AGENTS.md` for Codex or `.github/copilot-instructions.md` for GitHub Copilot, make them pointers rather than copies. See [One source of instructions](../patterns/one-source-of-instructions.md). **Examples:** @@ -84,7 +84,7 @@ GitHub branch protection rules ensure agents can't merge directly to main. Every │ │ │ ┌──────────────┐ ┌──────────────┐ │ │ │ CLAUDE.md │ │ curation- │ System │ -│ │ .goosehints │ │ skills │ Instructions │ +│ │ AGENTS.md │ │ skills │ Instructions │ │ └──────────────┘ └──────────────┘ │ │ │ │ ┌──────────────┐ ┌──────────────┐ │ @@ -108,7 +108,7 @@ GitHub branch protection rules ensure agents can't merge directly to main. Every │ └─────────────────────────────────┘ │ │ │ │ ┌─────────────────────────────────┐ │ -│ │ AI Agent (Claude, Goose) │ The Agent │ +│ │ AI Agent (Claude, Codex) │ The Agent │ │ └─────────────────────────────────┘ │ └─────────────────────────────────────────────────────┘ ``` @@ -131,5 +131,5 @@ As you mature your setup, add: ## Further reading - [Agentic tools reference](../reference/agentic-tools.md) — detailed documentation for each tool -- [Example repositories](../examples.md) — see harnesses in action +- [Example repositories](../case-studies/index.md) — see harnesses in action - [Agentic tooling on GitHub](https://github.com/topics/ai4curation) diff --git a/docs/how-tos/create-agentic-curation-pipeline.md b/docs/how-tos/create-agentic-curation-pipeline.md index b91aacd..84e1b0e 100644 --- a/docs/how-tos/create-agentic-curation-pipeline.md +++ b/docs/how-tos/create-agentic-curation-pipeline.md @@ -5,11 +5,15 @@ repository where the canonical knowledge base is a set of YAML files, validated by LinkML and reviewed through GitHub pull requests. The main example is -[DisMech](https://github.com/monarch-initiative/dismech), a disorder mechanisms -knowledge base. The same pattern is also used in derived mechanism repositories -such as `community-mech`, and in -[ai-gene-review](https://github.com/ai4curation/ai-gene-review), where each YAML -file is an AI-assisted gene review rather than a disease mechanism record. +[DisMech](../case-studies/dismech.md), a disorder mechanisms knowledge base. The +same pattern is used in +[AI Gene Review](../case-studies/ai-gene-review.md), where each YAML file is a +gene review, and in the six +[CultureBot mechs](../case-studies/communitymech.md), which cover microbial +communities, habitats, traits, culture media, and media ingredients. + +Read the case studies first if you want to see the finished result before you +build one. This page assumes you already have the basic GitHub agent wiring from [How to create an AI agent for GitHub actions](set-up-github-actions.md). For @@ -404,8 +408,8 @@ Before opening a PR: ``` Do not put every detailed workflow in the top-level instructions. Use -[skills](../reference/claude-skills.md) for task-specific procedures that agents -can load when relevant. DisMech has skills for +[skills](../patterns/skills-before-automation.md) for task-specific procedures +that agents can load when relevant. DisMech has skills for [ontology term work](https://github.com/monarch-initiative/dismech/blob/main/.claude/skills/dismech-terms/SKILL.md), [reference validation and repair](https://github.com/monarch-initiative/dismech/blob/main/.claude/skills/dismech-references/SKILL.md), [compliance improvement](https://github.com/monarch-initiative/dismech/blob/main/.claude/skills/dismech-compliance/SKILL.md), @@ -555,7 +559,7 @@ pattern is: - records remain fully explicit - validation and review check whether a claimed conformance is plausible -This is a good way to derive repositories such as `community-mech` from the +This is a good way to derive repositories such as [CommunityMech](../case-studies/communitymech.md) from the DisMech pattern while changing the domain model and content. ## 13. Touching Ontology Workflows diff --git a/docs/how-tos/instruct-github-agent.md b/docs/how-tos/instruct-github-agent.md index a9c6bea..295b588 100644 --- a/docs/how-tos/instruct-github-agent.md +++ b/docs/how-tos/instruct-github-agent.md @@ -36,8 +36,10 @@ but check with your repo maintainer first for local procedures. ### AI system instructions See the file `~/CLAUDE.md` in the top level of the repo. Other AI -applications may use different files -- for example, [goose](../glossary.md#goose) uses -`~/.goosehints`, but these will typically be symlinked. +applications may use different files. Codex reads `AGENTS.md` and GitHub +Copilot reads `.github/copilot-instructions.md`. Keep one file authoritative and +make the others pointers to it, rather than maintaining copies. See +[One source of instructions](../patterns/one-source-of-instructions.md). The instructions are in natural language and should be equally readable by humans or AI. @@ -50,7 +52,7 @@ Depending on what AI host runner is used, the configuration file that controls which agents are available may be in different places. - For claude code, this will be in `.claude/settings.json` -- For goose, this will be in `.config/goose/config.yaml` +- For Goose, which is [historic](../reference/harnesses.md), this is in `.config/goose/config.yaml` ## Prompting guide diff --git a/docs/how-tos/integrate-ai-into-your-kb.md b/docs/how-tos/integrate-ai-into-your-kb.md index 410b313..edf2c2d 100644 --- a/docs/how-tos/integrate-ai-into-your-kb.md +++ b/docs/how-tos/integrate-ai-into-your-kb.md @@ -21,29 +21,28 @@ Examples: - [CLAUDE.md in Uberon repo](https://github.com/obophenotype/uberon/blob/master/CLAUDE.md) -## Tip 3: Train curators to use simple tool-enabled AI applications (e.g [Goose](../glossary.md#goose)) +## Tip 3: Let curators run agents in the cloud -Many AI hosts such as Claude Code or various VS code plugins are suboptimal for non-technical users. AI applications such as Claude Desktop may be better, but currently it's hard to configure. +The two things that stop curators are installing software and getting access to +a model. Both go away if you point them at a browser. -At the time of writing, we recommend Goose as an AI app/host, due to these features: +[Claude Code on the web](https://code.claude.com/docs/en/claude-code-on-the-web) +clones the repository, runs the session in a cloud container, and opens the pull +request. A curator needs a paid Claude plan and write access to your repository. -- Ease of configuring MCPs -- Choice of either Desktop version (for non-devs) or Command Line (for devs) -- Ability to use multiple models including proxies. +For a whole team at once, a shared hosted environment removes the account +problem as well. The Gene Ontology consortium runs one at +[geneontology/go-jupyter](https://github.com/geneontology/go-jupyter), where each +participant gets a workspace with an agent already configured and no key of +their own. -See the [Installation Guide](https://block.github.io/goose/docs/getting-started/installation/) +Whichever route you choose, write down the setup steps in your repository. +DisMech's [`CONTRIBUTING.md`](https://github.com/monarch-initiative/dismech/blob/main/CONTRIBUTING.md) +is a good model: it covers the environment configuration, the network access +setting, and the environment variables its tools need. -As an example, this video shows how to configure: - - +We previously recommended Goose here. We no longer do for new setups. See +[Harnesses](../reference/harnesses.md). ## Tip 4: Validate agent outputs automatically diff --git a/docs/how-tos/set-up-github-actions.md b/docs/how-tos/set-up-github-actions.md index d6966dc..019e08a 100644 --- a/docs/how-tos/set-up-github-actions.md +++ b/docs/how-tos/set-up-github-actions.md @@ -1,213 +1,140 @@ -# How to create an AI agent for GitHub actions +# Set up GitHub Actions -This assumes that you are already using GitHub as your source of truth for -content, O3-guidelines style. +This guide adds an agent to a repository you already manage on GitHub. It +assumes your content is in the repository and that you have basic quality +control actions running. -It also assumes you have some familiarity with GitHub actions, and -have basic QC actions set up. If you are managing an ODK-compliant -repo this is certainly the case. +By the end you will have three workflows, which is the setup +[GO](../case-studies/go-ontology.md) and [Uberon](../case-studies/uberon.md) +both run. -## IMPORTANT - GitHub repo configuration -When using AI agents, ensure that your `main` repository branch has [GitHub branch protection rules](https://docs.github.com/en/repositories/configuring-branches-and-merges-in-your-repository/managing-protected-branches/about-protected-branches) enabled. Specifically, the `main` branch should have at least these settings configured: -- Require pull request reviews before merging -- Require at least one PR reviewer to approve the PR -- Do not allow bypassing the above settings +## Protect the main branch first -## Quick setup with Claude Code +Do this before anything else. Once an agent can open pull requests, your branch +protection is what stops unreviewed content reaching your product. -If you have [Claude Code](../reference/clients/claude-code.md) installed, you can use the `install-github-app` command for a streamlined setup process. This command will authenticate you and create a pull request with GitHub Actions configuration: +Set these rules on `main` or `master`: + +* Require a pull request before merging. +* Require at least one approving review. +* Do not allow anyone to bypass these settings. + +## Add the workflows + +The quickest route is to let Claude Code do it: ```bash claude install-github-app ``` -This approach automatically handles authentication and creates the necessary GitHub Actions workflow for you. For more details, see the [Claude Code GitHub Actions documentation](https://docs.anthropic.com/en/docs/claude-code/github-actions). - -If you prefer manual setup or need more customization, continue with the manual configuration steps below. - -## Set up `ai.yml` - -This might look something like this: - -https://github.com/monarch-initiative/mondo/blob/master/.github/workflows/ai-agent.yml - -```yaml -name: Dragon AI Agent GitHub Mentions - -on: - issues: - types: [opened, edited] - issue_comment: - types: [created, edited] - pull_request: - types: [opened, edited] - pull_request_review_comment: - types: [created, edited] - -jobs: - check-mention: - runs-on: ubuntu-latest - outputs: - qualified-mention: ${{ steps.detect.outputs.qualified-mention }} - prompt: ${{ steps.detect.outputs.prompt }} - user: ${{ steps.detect.outputs.user }} - item-type: ${{ steps.detect.outputs.item-type }} - item-number: ${{ steps.detect.outputs.item-number }} - controllers: ${{ steps.detect.outputs.controllers }} - steps: - - name: Checkout repository - uses: actions/checkout@v4 - - - name: Detect AI mention - id: detect - uses: dragon-ai-agent/github-mention-detector@v1.0.0 - with: - github-token: ${{ secrets.PAT_FOR_PR }} - fallback-controllers: 'cmungall' - - respond-to-mention: - needs: check-mention - if: needs.check-mention.outputs.qualified-mention == 'true' - permissions: - contents: write - pull-requests: write - issues: write - runs-on: ubuntu-latest - steps: - - name: Checkout repository - uses: actions/checkout@v4 - with: - fetch-depth: 0 - token: ${{ secrets.PAT_FOR_PR }} - - - name: Respond with AI Agent - uses: dragon-ai-agent/run-goose-obo@v1.0.4 - with: - anthropic-api-key: ${{ secrets.ANTHROPIC_API_KEY }} - openai-api-key: ${{ secrets.CBORG_API_KEY }} - github-token: ${{ secrets.PAT_FOR_PR }} - prompt: ${{ needs.check-mention.outputs.prompt }} - user: ${{ needs.check-mention.outputs.user }} - item-type: ${{ needs.check-mention.outputs.item-type }} - item-number: ${{ needs.check-mention.outputs.item-number }} - controllers: ${{ needs.check-mention.outputs.controllers }} - agent-name: 'Dragon-AI Agent' - branch-prefix: 'dragon_ai_agent' - robot-version: 'v1.9.7' -``` +This authenticates you and opens a pull request that adds the workflow files. +See the +[Claude Code GitHub Actions documentation](https://docs.claude.com/en/docs/claude-code/github-actions) +for what it creates. -This assumes using [goose-ai-obo-action](https://github.com/ai4curation/goose-ai-obo-action/) - -## Set up repository secrets - -The URL should be something like `https://github.com/monarch-initiative/mondo/settings/secrets/actions` - -A typical setup might include: - -* `ANTHROPIC_API_KEY` -* `CBORG_API_KEY` -* `PAT_FOR_PR` - -The key names will correspond to what you have in the action.yml above - -## Configure the agent, including default MCPs - -Create a folder `.config/[goose](../glossary.md#goose)` with a file `config.yaml`. - -Examples: - -- [.config/goose for Mondo](https://github.com/monarch-initiative/mondo/tree/master/.config/goose) - -Here is an example: - -```yaml -OPENAI_HOST: https://api.cborg.lbl.gov -OPENAI_BASE_PATH: v1/chat/completions -GOOSE_MODEL: anthropic/[claude-sonnet](../glossary.md#claude) -GOOSE_PROVIDER: openai -extensions: - developer: - bundled: true - display_name: Developer - enabled: true - name: developer - timeout: 300 - type: builtin - git: - args: - - mcp-server-git - bundled: null - cmd: uvx - description: Git version control system integration - enabled: false - env_keys: [] - envs: {} - name: git - timeout: 300 - type: stdio - memory: - bundled: true - display_name: Memory - enabled: true - name: memory - timeout: 300 - type: builtin - owlmcp: - args: - - owl-mcp - bundled: null - cmd: uvx - description: '' - enabled: false - env_keys: [] - envs: {} - name: owlmcp - timeout: 300 - type: stdio - pdfreader: - args: - - mcp-read-pdf - bundled: null - cmd: uvx - description: Read large and complex PDF documents - enabled: false - env_keys: [] - envs: {} - name: pdfreader - timeout: 300 - type: stdio -``` +To do it by hand, copy from a repository that already works. All three of these +run [`anthropics/claude-code-action`](https://github.com/anthropics/claude-code-action): + +| Workflow | What it does | Copy from | +| --- | --- | --- | +| `ai-agent.yml` | Runs an agent when someone mentions it in an issue or comment | [GO](https://github.com/geneontology/go-ontology/blob/master/.github/workflows/ai-agent.yml) | +| `claude-code-review.yml` | Reviews every pull request | [GO](https://github.com/geneontology/go-ontology/blob/master/.github/workflows/claude-code-review.yml) | +| `copilot-setup-steps.yml` | Prepares the environment for the GitHub Copilot coding agent | [Uberon](https://github.com/obophenotype/uberon/blob/master/.github/workflows/copilot-setup-steps.yml) | + +## Add the secrets + +Set these at `https://github.com/OWNER/REPO/settings/secrets/actions`. + +| Secret | What it is for | +| --- | --- | +| `ANTHROPIC_API_KEY` or `CLAUDE_CODE_OAUTH_TOKEN` | Lets the action call the model. You need one of the two. | +| `PAT_FOR_PR` | A token that lets the agent open pull requests, if the default token is not enough. | + +Agent runs cost money. Set a spending limit on the account that owns the key +before you turn on anything that runs on a schedule. + +## Write the instructions file + +Create `CLAUDE.md` in the repository root. Tell the agent: + +* Which file is the editable product, and which files are generated. +* How to search that file, with worked commands. +* Your identifier rules, including the range for new terms. +* The validation command to run before opening a pull request. + +[Cell Ontology's `CLAUDE.md`](https://github.com/obophenotype/cell-ontology/blob/master/CLAUDE.md) +is a good short example of the search and identifier sections. + +Keep one authoritative file. If you also need `AGENTS.md` or +`.github/copilot-instructions.md`, make them pointers rather than copies. See +[One source of instructions](../patterns/one-source-of-instructions.md). -The `extensions` section defines the default MCP plugins. +## Declare your MCP servers in the repository -This setup is configured to use Anthropic claude-sonnet via a LiteLLM proxy (CBORG): +Add `.mcp.json` so every session gets the same tools without the curator +configuring anything. [EFO](../case-studies/efo.md) declares two: -```yaml -OPENAI_HOST: https://api.cborg.lbl.gov -OPENAI_BASE_PATH: v1/chat/completions -GOOSE_MODEL: anthropic/claude-sonnet -GOOSE_PROVIDER: openai +```json +{ + "mcpServers": { + "OLS-MCP": { + "type": "http", + "url": "http://www.ebi.ac.uk/ols4/api/mcp" + }, + "artl-mcp": { + "command": "uvx", + "args": ["artl-mcp"] + } + } +} ``` -You can use a similar setup for your own LiteLLM proxy if you have one. Note: if you find this complex, please upvote [this issue](https://github.com/block/goose/issues/2507). +Ontology term lookup is the one to add first. It is the most common source of +fabricated identifiers, and a lookup tool removes the need to guess. See +[Make identifiers hard to fake](../patterns/ground-identifiers.md). -Things are easier if you just want to talk straight to a provider like Anthropic, no proxy, all you need is `GOOSE_MODEL`. But note that this likely means you are using a personal API key. Be aware that agentic AI usage can be costly. +## Control who can start a run -## Set up .goosehints +A mention in an issue comment starts an agent run, and on a public repository +anyone can write that comment. Restrict who can summon the agent before you turn +the workflow on. See +[Guard the untrusted surface](../patterns/guard-untrusted-input.md). -For many of my repos, I have a `CLAUDE.md` and I symlink `.goosehints` to that, because I am too lazy to write different instructions for different agents. In practice it might be better to tune instructions. +Note also that GitHub does not give repository secrets to workflows triggered +from a fork. Pull requests from forks get no automated review. Decide now +whether you will grant contributors branch access on the origin repository, as +[DisMech](../case-studies/dismech.md) does, or accept that fork contributions +are reviewed by humans only. -What you put in there depends on your own use case. Do careful evaluations if you can't, but otherwise do vibe tests and iterate. +## Enable the GitHub Copilot coding agent -Some examples here: +Copilot works from issues and produces pull requests inside GitHub, without a +local session. Uberon's +[pull request 3580](https://github.com/obophenotype/uberon/pull/3580) shows the +two files you need to change. -- [CLAUDE.md on Mondo](https://github.com/monarch-initiative/mondo/blob/master/CLAUDE.md) +Copilot's agent runs behind a firewall that blocks most external hosts by +default, including `purl.obolibrary.org`. See +[GitHub Copilot](../reference/github-copilot.md) for how to allowlist the hosts +your workflows need. +## Then add validation -## NEW: enable GitHub copilot to do reviews and PRs +Workflows that run agents are the easy half. The half that decides whether this +works is the checking that happens afterwards. -Change two files, as documented here: +Add term and reference validation to continuous integration before you let +agents open many pull requests. See +[Make identifiers hard to fake](../patterns/ground-identifiers.md) and +[Fast and slow validation](../patterns/fast-and-slow-validation.md). -https://github.com/obophenotype/uberon/pull/3580 +## Historic: dragon-ai-agent and Goose +Several OBO repositories run an older setup, in which `ai-agent.yml` calls +`dragon-ai-agent` and the agent itself is Goose, configured in +`.config/goose/config.yaml` with a `.goosehints` file for instructions. +[Mondo](../case-studies/mondo.md) still runs this way. +Keep it working where it is deployed. Do not start there. One consequence is +visible in Mondo: subagents defined under `.claude/agents/` are Claude Code +features, so an agent started by a Goose-based workflow cannot reach them. diff --git a/docs/index.md b/docs/index.md index e8db23e..e1bf0e8 100644 --- a/docs/index.md +++ b/docs/index.md @@ -1,12 +1,48 @@ -# AI Guide +# AI4Curators -This site will grow into a collection of how-tos and reference guides -for curators and maintainers of [knowledge bases](glossary.md#knowledge-base-kb) to integrate AI into their -workflows. +Practical guides for curators and maintainers of +[knowledge bases](glossary.md#knowledge-base-kb) who want to use AI agents in +the workflows they already have. -The aim here is to provide practical guides that can be *immediately integrated* -into existing curation workflows. This means things like: +Agents are curating real ontologies and knowledge bases today. This site points +you at those repositories and explains what they do, so you can copy working +setups instead of designing your own from scratch. -- agents that plug into the existing [GitHub](https://github.com) repos that curators and ontology editors use, - empowered to make changes to curated artefacts -- plugins for existing chat UIs \ No newline at end of file +## Start here + +**If you curate**, and you want to try an agent on a real task: + +1. Read [Your first curation session](tutorials/first-curation-session.md). +2. Read [Instruct the GitHub agent](how-tos/instruct-github-agent.md) for how to + ask an agent for work through an issue. +3. Look at [DisMech](case-studies/dismech.md) to see what a project run this way + looks like from the inside. + +**If you maintain a repository**, and you want agents working in it: + +1. Read the [case studies](case-studies/index.md) for repositories like yours. +2. Read the [patterns](patterns/index.md) they have in common. +3. Follow [Set up GitHub Actions](how-tos/set-up-github-actions.md) to add an + agent to your repository. +4. Add validation before you add automation. See + [Make identifiers hard to fake](patterns/ground-identifiers.md). + +## What is on this site + +| Section | What it holds | +| --- | --- | +| [Case studies](case-studies/index.md) | One page per repository that runs agents on real curation | +| [Patterns](patterns/index.md) | Practices that appear in more than one of them | +| [How-tos](how-tos/instruct-github-agent.md) | Step-by-step tasks | +| [Reference](reference/harnesses.md) | Harnesses, tools, and GitHub integrations | +| [Evidence](evidence.md) | What we measure across these repositories | +| [Glossary](glossary.md) | Terms used on this site | + +## How this site works + +We do not centralize guidance here. Repositories are the source of truth, and +they change faster than documentation does. Each page points at files and +folders you can open. + +If a page disagrees with the repository it describes, trust the repository and +[tell us](https://github.com/ai4curation/aidocs/issues). diff --git a/docs/patterns/fast-and-slow-validation.md b/docs/patterns/fast-and-slow-validation.md new file mode 100644 index 0000000..0b7367a --- /dev/null +++ b/docs/patterns/fast-and-slow-validation.md @@ -0,0 +1,66 @@ +# Fast and slow validation + +**Use it when** full validation is too slow to run while curating. + +Checking that quoted evidence appears in the cited papers is slow. It may fetch +full text. Ontology term checking can pull down large databases. If the only +validation command you have takes several minutes, agents stop running it, and +errors reach the pull request. + +## The pattern + +Provide two commands. + +**A fast check for the curation loop.** It runs in seconds, uses the local +cache, and takes any number of files. Agents and curators run it constantly. + +**A slow sweep before the pull request.** It runs everything, and it is the same +thing continuous integration runs. You run it once, at the end. + +[DisMech](../case-studies/dismech.md) does exactly this: + +```bash +# Fast: check evidence snippets against the local reference cache. +just count-verified-snippets kb/disorders/YourFile.yaml + +# Slow: the batched schema, terms, and references sweep that CI runs. +just validate-disorders kb/disorders/YourFile.yaml +``` + +Its `CLAUDE.md` labels them, so the agent knows which one belongs in the loop +and which one belongs at the end. + +## Cache the expensive things + +The fast check only works if the reference text is already local. DisMech keeps +a reference cache and gives you one command to fill it: + +```bash +just fetch-reference PMID:12345678 +``` + +The instructions then say never to create a cache file by hand. That rule +matters: a hand-written cache file would let fabricated evidence pass the check +that exists to catch fabricated evidence. + +## Watch out for large downloads + +Ontology term checking backed by OAK downloads SQLite databases on demand. Some +are large: `ncbitaxon` is around 13.5 GB unpacked. In a cloud session or on a +metered connection, an interrupted download can kill a validation run. + +Two things help: + +* Make single-file validation trust the committed caches, and only query the + ontology for terms that are not cached yet. +* Provide a command that fetches the databases deliberately, with resume and + retry, so you can do it before you need it. + +DisMech has `just fetch-ontology-dbs` for the second. + +## Who does this + +* [DisMech](../case-studies/dismech.md): `count-verified-snippets` and + `validate-disorders`. +* [CommunityMech](../case-studies/communitymech.md) and the other CultureBot + mechs use the same split, with `just validate` and `just validate-references`. diff --git a/docs/patterns/ground-identifiers.md b/docs/patterns/ground-identifiers.md new file mode 100644 index 0000000..2fbb229 --- /dev/null +++ b/docs/patterns/ground-identifiers.md @@ -0,0 +1,69 @@ +# Make identifiers hard to fake + +**Use it when** agents write ontology terms, citations, or quotes. + +An agent can produce `GO:0006954` when it means something else, or cite a paper +that does not contain the sentence it quotes. Both look right. Neither is +caught by a schema check that only tests the shape of the data. + +## The pattern + +Never record an identifier on its own. Record it with something that must match +it. + +Instead of this: + +```yaml +term: GO:0006954 +``` + +Do this: + +```yaml +term: + id: GO:0006954 + label: inflammatory response +``` + +To pass a check, the agent now has to get both parts right and consistent with +the ontology. Guessing one is easy; guessing a matching pair is not. + +Apply the same rule to evidence. A citation on its own is weak. A citation plus +the exact sentence it supports can be checked against the paper: + +```yaml +evidence: + - reference: PMID:12345678 + supports: SUPPORT + snippet: "Exact quote from the paper" + explanation: "Why this supports the claim" +``` + +## The tools + +* [linkml-term-validator](https://github.com/linkml/linkml-term-validator) + checks that a term exists and that the label matches. +* [linkml-reference-validator](https://github.com/linkml/linkml-reference-validator) + checks that the quoted text appears in the cited reference. + +Both run in continuous integration, so a pull request that fabricates a term or +a quote fails before a human reads it. + +## Who does this + +* [DisMech](../case-studies/dismech.md) runs both validators, plus a fast + snippet check you can run while you curate. +* [AI Gene Review](../case-studies/ai-gene-review.md) pairs every term with its + label throughout the schema. +* [HabitatMech](../case-studies/habitatmech.md) goes further: where no ontology + term fits, it mints a local identifier instead of forcing a match. + +## What this does not do + +Validation proves that a citation exists, that a quote is exact, and that a +term is real. It does not prove that the claim is correct. Say so where your +readers can see it. DisMech puts this on its front page. + +## Read more + +* [Make identifiers hallucination-resistant](../how-tos/make-ids-hallucination-resistant.md) diff --git a/docs/patterns/guard-untrusted-input.md b/docs/patterns/guard-untrusted-input.md new file mode 100644 index 0000000..dc85911 --- /dev/null +++ b/docs/patterns/guard-untrusted-input.md @@ -0,0 +1,49 @@ +# Guard the untrusted surface + +**Use it when** anyone on the internet can comment on your issues. + +If a GitHub mention starts an agent run, then the text of that comment reaches +your agent. On a public repository, anyone can write that text. Issue bodies, +comments, review text, and the content of papers your agent fetches are all +input you do not control. + +## The pattern + +Control who can start a run, and check what reaches the agent. + +**Keep a list of who can summon the agent.** [DisMech](../case-studies/dismech.md) +requires you to be a registered controller in a JSON file in the repository. +A mention from anyone else does nothing. + +**Guard comment content.** DisMech runs a separate workflow, +`untrusted-comment-guard.yml`, over comment text before it drives an agent. + +**Close pull requests from forks.** These get no automated review anyway, +because GitHub withholds secrets from fork workflows. DisMech closes them +automatically and asks contributors to push branches to the origin repository +instead. + +**Do not let documentation trigger the agent.** DisMech's mention keyword is +ignored inside inline code spans and fenced code blocks. That way a page +documenting the keyword does not summon the agent every time someone reads it. + +**Scope credentials tightly.** Give the agent tokens for a test server, not +production. Give the watcher read access, not write. Instructions are guidance; +credentials are the control. + +## Separate reading from writing + +Agents that only read are far less risky than agents that write. If your agent +summarizes, triages, or reports, do not give it write access at all. Add write +access one surface at a time, and record what each one can reach. + +[agent-watcher](../evidence.md) is built this way. It reads the repositories it +watches with a read-only token and publishes its reports to a different +repository, so it cannot modify anything it observes. + +## Treat fetched content as data + +An agent that reads a paper, a web page, or a comment is reading text that +someone else wrote. Text that looks like an instruction is still just text in a +document. This applies to your validation caches too, which is why DisMech +forbids hand-written cache files and requires them to be fetched by a command. diff --git a/docs/patterns/human-regulating-the-loop.md b/docs/patterns/human-regulating-the-loop.md new file mode 100644 index 0000000..a851b34 --- /dev/null +++ b/docs/patterns/human-regulating-the-loop.md @@ -0,0 +1,54 @@ +# Regulate the loop + +**Use it when** you cannot review every agent edit yourself. + +"Human in the loop" means a person checks each item. That works until agents +produce more work than your curators can read. Then the person becomes the +bottleneck, and the usual result is that review turns into rubber-stamping. + +[DisMech](../case-studies/dismech.md) uses a different phrase: *human +regulating the loop*. Curators work on the process that produces entries, not +on each entry. + +## The pattern + +Spend curator time on these three jobs. + +**Find patterns, not errors.** One wrong entry is a fix. The same wrong entry +five times is a process problem. Look for the repeated kind, then curate +examples and counter-examples into the instructions or the skill that produced +them. + +**Review the reviews.** Your automated reviewer has its own failure modes. It +misses some things and it fixates on others. Read a sample of its reviews and +ask what it did not catch, and what it complained about that did not matter. +Then change the checklist. + +**Tune how eager each process is.** Scanners and automated reviewers each have a +cadence and a threshold. Some fire too often and create noise. Some never fire +and miss work. This is a dial, and someone has to be responsible for setting it. + +## Say who is accountable + +The hardest part is not technical. Projects break down when nobody has said +whether a contributor is answerable for what their agent wrote. + +DisMech states its answer in `CONTRIBUTING.md`: contributors are not assumed to +have verified everything their agent produced, and the default assumption is +that issues and comments are AI-generated. If you write something yourself and +want that known, you mark it `[human authored]`. + +You do not have to make the same choice. You do have to make one, and write it +down. + +## Who does this + +* [DisMech](../case-studies/dismech.md) states the policy in `CONTRIBUTING.md` + and asks contributors to curate the process. +* [AI Gene Review](../case-studies/ai-gene-review.md) collects feedback from + domain experts through voting and an evaluation form, so disagreement reaches + the project without requiring Git. + +## Read more + +* [Pull request reviews for agent improvement](../how-tos/pr-reviews-for-agent-improvement.md) diff --git a/docs/patterns/index.md b/docs/patterns/index.md new file mode 100644 index 0000000..c71a127 --- /dev/null +++ b/docs/patterns/index.md @@ -0,0 +1,17 @@ +# Patterns + +Practices that appear in more than one repository. Each page is short, and +each one names the repositories that use it so you can go and read the real +version. + +| Pattern | Use it when | +| --- | --- | +| [Make identifiers hard to fake](ground-identifiers.md) | Agents write ontology terms, citations, or quotes | +| [Regulate the loop](human-regulating-the-loop.md) | You cannot review every agent edit yourself | +| [Break work into skills](skills-before-automation.md) | The same instructions keep getting repeated | +| [One source of instructions](one-source-of-instructions.md) | You have more than one agent instructions file | +| [Fast and slow validation](fast-and-slow-validation.md) | Full validation is too slow to run while curating | +| [Scanners find the work](scanners.md) | Agents only act when a human remembers to ask | +| [One review checklist](review-checklists.md) | Different reviewers apply different standards | +| [Shared control](shared-control.md) | Curation happens in a web tool, not in files | +| [Guard the untrusted surface](guard-untrusted-input.md) | Anyone on the internet can comment on your issues | diff --git a/docs/patterns/one-source-of-instructions.md b/docs/patterns/one-source-of-instructions.md new file mode 100644 index 0000000..7fba09b --- /dev/null +++ b/docs/patterns/one-source-of-instructions.md @@ -0,0 +1,56 @@ +# One source of instructions + +**Use it when** you have more than one agent instructions file. + +Different agents read different files. Claude Code reads `CLAUDE.md`. Codex and +several others read `AGENTS.md`. GitHub Copilot reads +`.github/copilot-instructions.md`. The obvious response is to put the same text +in all of them, and then keep them in sync by hand. + +That does not hold. In practice one file gets updated and the others do not, and +your agents start behaving differently for reasons nobody can see. + +## The pattern + +Pick one authoritative file. Make the others short pointers to it. + +[EFO](../case-studies/efo.md) does this. Its `AGENTS.md` is about twenty lines, +and it says where the real instructions are: + +> The authoritative curation workflow, routing, and domain rules live in +> **`.github/copilot-instructions.md`** (full guide) and **`CLAUDE.md`** +> (orchestration). Read the relevant one before making ontology changes. + +The pointer file can still carry anything genuinely specific to that agent. EFO +uses `AGENTS.md` to tell Codex where its own read-only review procedure lives. + +This site follows the same pattern: `AGENTS.md` holds the guidance and +`CLAUDE.md` points at it. + +## What this replaces + +As of August 2026, [GO](../case-studies/go-ontology.md), +[Mondo](../case-studies/mondo.md), and [Uberon](../case-studies/uberon.md) each +keep `CLAUDE.md` and `.github/copilot-instructions.md` with the same text in +both. Nothing cross-references anything. Every guidance change needs two edits. + +[Cell Ontology](../case-studies/cell-ontology.md) avoids the problem by having +only one file. + +## The same rule applies to subagents + +EFO defines its specialists twice, once under `.claude/agents/` and once under +`.github/agents/`, with names that differ only in case. The two sets can drift +in the same way. + +EFO has already solved this for review: the Claude, Copilot, and Codex reviewers +all read one checklist file. Apply that fix to specialists too. Generate the +copies from one source, or make one a pointer. + +## How to do it today + +1. Choose the file that has the best content. Make it authoritative. +2. Replace the others with a pointer of a few lines. +3. Keep only agent-specific notes in the pointer files. + +This takes about ten minutes and removes a whole class of drift. diff --git a/docs/patterns/review-checklists.md b/docs/patterns/review-checklists.md new file mode 100644 index 0000000..620768c --- /dev/null +++ b/docs/patterns/review-checklists.md @@ -0,0 +1,48 @@ +# One review checklist + +**Use it when** different reviewers apply different standards. + +Once you have an automated reviewer, the checklist it applies becomes the real +quality standard of your project. It is worth treating as a curated artifact. + +## The pattern + +Keep the checklist in one file. Point every reviewer at it. + +[EFO](../case-studies/efo.md) has three reviewers: Claude, GitHub Copilot, and +Codex. All three apply `docs/agents-documentation/efo-pr-review-checklist.md`. +The Codex reviewer has its own procedure file, and that file points at the same +checklist. When the standard changes, one file changes. + +[DisMech](../case-studies/dismech.md) keeps its rubric in a skill, +`dismech-pr-review`, and configures which model runs it in +`.github/agent-config.yaml`. [GO](../case-studies/go-ontology.md) and +[AI Gene Review](../case-studies/ai-gene-review.md) both have review skills too, +called `pr-review` and `aigr-pr-review`. + +## Separate the review from the model + +Two things change at different rates. The checklist changes when your standards +change. The model changes when a better one appears. Keep them in different +files so neither change disturbs the other. + +## The review decides something + +An automated review that leaves a comment is advice. An automated review that +marks a pull request "changes requested" or "ready to merge" is a gate. DisMech +does the second. This is what makes it possible for contributors not to check +their own agent's work. + +Make it a gate only when you trust the checklist. That trust has to be earned by +reading its reviews for a while, which is the point of +[reviewing the reviews](human-regulating-the-loop.md). + +## Automated review does not work on forks + +GitHub does not give repository secrets to workflows triggered from a fork, so a +fork-based pull request gets no automated review. + +Projects that depend on automated review have to handle this. DisMech asks +contributors to push branches to the origin repository instead of forking, and +closes fork pull requests automatically. That means granting branch access to +contributors, which is a governance decision, not a technical one. diff --git a/docs/patterns/scanners.md b/docs/patterns/scanners.md new file mode 100644 index 0000000..6ce6fe3 --- /dev/null +++ b/docs/patterns/scanners.md @@ -0,0 +1,50 @@ +# Scanners find the work + +**Use it when** agents only act if a human remembers to ask. + +Most agent setups are reactive. Someone opens an issue, mentions the agent, and +the agent responds. The work that nobody files never happens. + +A scanner is a scheduled job that looks for work and creates it. + +## The pattern + +Run each scanner on its own schedule. Each one answers a different question. + +| Scanner | Question it answers | +| --- | --- | +| Literature scan | What new papers affect entries we already have? | +| Knowledge gap scan | Which entries are missing required fields? | +| Stalled work scan | Which issues and pull requests have gone quiet? | +| Compliance scan | Which entries fall below our quality bar? | + +[DisMech](../case-studies/dismech.md) runs all four. Its cadences live in +`.github/cron-profiles.yaml`, separate from the workflow files, so you can +change how often something runs without editing the job. + +## Match the model to the risk + +Not every scanner needs your best model. DisMech runs its pull request reviewer +on a top-tier model always, and lets low-risk scanners use cheaper ones. The +per-job model choice lives in `.github/agent-config.yaml`. + +It also has a `low_effort` label, so a human can send one specific task to a +cheaper model by hand. + +[culturebotai-claw](../case-studies/communitymech.md) assigns a tier per agent +in the same way: documentation on a cheap model, integrity auditing on an +expensive one. + +## Start with one + +A literature scanner is usually the best first one. It produces issues a curator +can judge quickly, and it does not change any curated content on its own. + +Set the cadence low to begin with. A scanner that files thirty issues a day gets +ignored within a week, and an ignored scanner is worse than no scanner. + +## Keep tuning it + +Scanner cadence is one of the dials that +[regulating the loop](human-regulating-the-loop.md) is about. Ask regularly +whether each scanner is too eager or not eager enough. diff --git a/docs/patterns/shared-control.md b/docs/patterns/shared-control.md new file mode 100644 index 0000000..abc5513 --- /dev/null +++ b/docs/patterns/shared-control.md @@ -0,0 +1,48 @@ +# Shared control + +**Use it when** curation happens in a web tool, not in files. + +Much curation does not live in a Git repository. It lives in a curation +application backed by a database: Noctua for GO-CAM models, and similar tools +elsewhere. You cannot give an agent a pull request workflow for these. + +## The pattern + +Give the agent an MCP server that wraps the same API the curation interface +already uses. The agent and the curator then write to the same database through +the same path. + +The curator keeps the interface they know. When the agent makes a change, it +appears in that interface. The curator can edit alongside the agent, in the same +session, and see the result of each agent action as it happens. + +This is different from an agent that edits files and hands you a diff. Nothing +is staged. You watch it happen and you can intervene. + +## Why the existing API matters + +Wrapping the API the interface already uses means the agent is subject to the +same rules the interface is: the same validation, the same permissions, the same +data model. An agent given direct database access has none of that. + +## Point it at a test server first + +An agent with write access to a production curation database can do real damage. +Restrict the credentials to a development or test server while you are learning +what the agent does, and say so in your agent instructions as well as enforcing +it in the token. + +Credentials are the control that works. Instructions alone are not enough. + +## Who does this + +* [noctua-mcp](https://github.com/geneontology/noctua-mcp) wraps the + Noctua and Barista API for GO-CAM editing. +* [oak-mcp](https://github.com/monarch-initiative/oak-mcp) gives agents ontology + search and traversal through the Ontology Access Kit. +* [EFO](../case-studies/efo.md) declares the EBI Ontology Lookup Service MCP + server in `.mcp.json`, so every session gets the same term lookup tool. + +## Read more + +* [Agentic tools](../reference/agentic-tools.md) diff --git a/docs/patterns/skills-before-automation.md b/docs/patterns/skills-before-automation.md new file mode 100644 index 0000000..4735ca6 --- /dev/null +++ b/docs/patterns/skills-before-automation.md @@ -0,0 +1,66 @@ +# Break work into skills + +**Use it when** the same instructions keep getting repeated. + +A skill is a folder with a `SKILL.md` file that describes one job. The agent +loads it when the job comes up, and ignores it otherwise. This keeps the main +instructions file short and gives each recurring task a place to live. + +## The pattern + +Name skills after the work, not after the technology. A curator should +recognize the name. + +[GO](../case-studies/go-ontology.md) has ten skills, and the names are ticket +types: `taxon-constraint`, `term-obsoletion`, `design-pattern`, +`external-term-lookup`. An editor who has handled one of those tickets knows +what the skill does. + +Write a skill when you notice yourself explaining the same thing twice. + +## Automation without decomposition does not work well + +[Cell Ontology](../case-studies/cell-ontology.md) is the example to learn from. +It has two workflows sending real work to an agent and no skills, subagents, or +commands. Every run starts from the same general instructions and plans its own +approach. + +The symptom shows up in the instructions file. Cell Ontology's `CLAUDE.md` +contains this: + +> DO NOT bother doing your own greps over the file, or looking for other files, +> unless otherwise asked, you will just waste time. + +That is a fix for something that went wrong, written as a prohibition. +Prohibitions accumulate and start to conflict. A skill tells the agent what to +do instead, and composes with the others. + +## Skills or subagents? + +Both split work up. They are not the same. + +| | Skill | Subagent | +| --- | --- | --- | +| Runs in | The current session | Its own context | +| Good for | A job that needs specific instructions | A job that reads a lot and reports a little | +| Cost | Instructions loaded on demand | A separate run | + +[Uberon](../case-studies/uberon.md) uses subagents for research and checking. +`deep-research-specialist` can read twenty papers and return three sentences, +and the main session never carries the twenty papers. + +Use a subagent when you want the reading kept out of the main context. Use a +skill for everything else. + +## Who does this + +* [GO](../case-studies/go-ontology.md): ten skills, no subagents. +* [Uberon](../case-studies/uberon.md): eight subagents, one skill. +* [DisMech](../case-studies/dismech.md): seventeen skills, including one that + holds the pull request review rubric. +* [AI Gene Review](../case-studies/ai-gene-review.md): fifteen skills, split + between reviewing and synthesizing. + +## Read more + +* [Create curation skills](../how-tos/author-skills.md) diff --git a/docs/reference/agentic-tools.md b/docs/reference/agentic-tools.md index d741ab1..2099df6 100644 --- a/docs/reference/agentic-tools.md +++ b/docs/reference/agentic-tools.md @@ -60,9 +60,15 @@ MCP server for ontology operations via the Ontology Access Kit ([OAK](../glossar ## System Instructions -### CLAUDE.md / .goosehints +### CLAUDE.md and AGENTS.md -Configuration files checked into the root of your repository that provide system-level instructions to AI agents. These serve as prompt preset management — different repositories can have different instructions tailored to their domain and workflows. +Files checked into the root of your repository that tell agents what the +repository is, which files are editable, and what conventions to follow. +Different agents read different files: Claude Code reads `CLAUDE.md`, Codex +reads `AGENTS.md`, and GitHub Copilot reads `.github/copilot-instructions.md`. + +Keep one of them authoritative and make the others pointers to it. See +[One source of instructions](../patterns/one-source-of-instructions.md). - **When to use**: Always. Every repository that uses AI agents should have system instructions. - See [Instruct the GitHub agent](../how-tos/instruct-github-agent.md) @@ -70,6 +76,8 @@ Configuration files checked into the root of your repository that provide system ## Further Reading +- [Case studies](../case-studies/index.md) — these tools in the repositories that use them +- [Patterns](../patterns/index.md) — the practices they support - [Build your agentic harness](../how-tos/build-agentic-harness.md) — how these tools compose into a harness - [Browse agentic tooling on GitHub](https://github.com/topics/ai4curation) - [ai4curation GitHub org](https://github.com/ai4curation) diff --git a/docs/reference/claude-skills.md b/docs/reference/claude-skills.md index 5c4975b..bad1ef5 100644 --- a/docs/reference/claude-skills.md +++ b/docs/reference/claude-skills.md @@ -66,4 +66,4 @@ These skills are particularly valuable for: - **Full Documentation**: [https://github.com/ai4curation/curation-skills](https://github.com/ai4curation/curation-skills) - **Related Guides**: [Integrate AI into your KB](../how-tos/integrate-ai-into-your-kb.md) -- **Client Setup**: [Claude Code](clients/claude-code.md) +- **Client Setup**: [Claude Code](harnesses.md) diff --git a/docs/reference/client-apps.md b/docs/reference/client-apps.md deleted file mode 100644 index 9431c49..0000000 --- a/docs/reference/client-apps.md +++ /dev/null @@ -1,21 +0,0 @@ -# Client Apps - -AI client applications that can be used for curation workflows. Each has different strengths depending on your role and technical comfort level. - -| Client | Type | Best For | MCP Support | -|--------|------|----------|-------------| -| [Claude Code](clients/claude-code.md) | Terminal CLI | Developers, power users | Yes | -| [Claude Desktop](clients/claude-desktop.md) | Desktop app | General use with local files | Yes | -| [Codex CLI](clients/codex-cli.md) | Terminal CLI | OpenAI model users | Limited | -| [Gemini CLI](clients/gemini-cli.md) | Terminal CLI | Google model users | Limited | -| [GitHub Copilot](clients/github-copilot.md) | GitHub-integrated | Automated PR workflows | No (uses GitHub tools) | -| [Goose](clients/goose.md) | Desktop + CLI | Non-technical curators, easy MCP setup | Yes | - -## Choosing a client - -- **Non-technical curators**: Start with [Goose](clients/goose.md) — it has the easiest MCP configuration and a desktop UI -- **Developers**: [Claude Code](clients/claude-code.md) for terminal-based workflows -- **GitHub-first workflows**: [GitHub Copilot](clients/github-copilot.md) for automated issue-to-PR workflows -- **General desktop use**: [Claude Desktop](clients/claude-desktop.md) for local file interaction with MCP support - -For details on each client, see the individual pages linked above. diff --git a/docs/reference/clients/claude-code.md b/docs/reference/clients/claude-code.md deleted file mode 100644 index add882e..0000000 --- a/docs/reference/clients/claude-code.md +++ /dev/null @@ -1,19 +0,0 @@ ---- -title: Claude Code -url: https://www.anthropic.com/claude-code -install: https://block.github.io/goose/docs/getting-started/installation -logo: https://upload.wikimedia.org/wikipedia/commons/thumb/8/8a/Claude_AI_logo.svg/1200px-Claude_AI_logo.svg.png -models: "Claude only" -ui_type: CLI -ease_of_use_for_non_technical: 2 -hints_file: CLAUDE.md ---- - - - -Claude Code Ident - -Claude Code is a command line AI client, used primarily by software developers. However, it can be used for any kind of content, including ontologies. The low ease of use score refers to the fact that this a command line interface which is inherently suboptimal for many non-technical users. However, for technical users the experience is minimal yet pleasant. - -Claude Code is a UI counterpart to [Claude Desktop](claude-desktop.md). We include a separate entry for each, because they are quite different in how the interoperate or are configured. - diff --git a/docs/reference/clients/claude-desktop.md b/docs/reference/clients/claude-desktop.md deleted file mode 100644 index 69baa3d..0000000 --- a/docs/reference/clients/claude-desktop.md +++ /dev/null @@ -1,17 +0,0 @@ ---- -title: Claude Desktop -url: "https://claude.ai/download" -logo: https://upload.wikimedia.org/wikipedia/commons/thumb/8/8a/Claude_AI_logo.svg/1200px-Claude_AI_logo.svg.png -models: "Claude only" -ui_type: Desktop -ease_of_use_for_non_technical: 4 ---- - -Claude Desktop Screenshot - -Claude Desktop is a local Desktop application -that provides many of the same functionalities as the web app, but allows you to configure your own MCPs. - -Claude Desktop is a UI counterpart to [Claude Code](claude-code.md). We include a separate entry for each, because they are quite different in how the interoperate or are configured. - -At the time of writing, it is less easy to configure MCPs than it is for Goose. \ No newline at end of file diff --git a/docs/reference/clients/codex-cli.md b/docs/reference/clients/codex-cli.md deleted file mode 100644 index a695c81..0000000 --- a/docs/reference/clients/codex-cli.md +++ /dev/null @@ -1,14 +0,0 @@ ---- -title: Codex CLI -url: https://github.com/openai/codex -logo: https://openai.com/favicon.ico -models: OpenAI GPT-4, Claude, and others -ui_type: CLI -ease_of_use_for_non_technical: 2 -github: https://github.com/openai/codex ---- - -Codex CLI brings OpenAI's reasoning models to the command line. It can create or edit files, -run commands in a sandbox, and track work under version control. Non-technical curators -can prompt tasks and review results directly from the terminal. - diff --git a/docs/reference/clients/gemini-cli.md b/docs/reference/clients/gemini-cli.md deleted file mode 100644 index 120cfaf..0000000 --- a/docs/reference/clients/gemini-cli.md +++ /dev/null @@ -1,13 +0,0 @@ ---- -title: Gemini CLI -url: https://github.com/google-gemini/gemini-cli -logo: https://www.gstatic.com/ai/web/favicons/favicon-32x32.png -models: Gemini 2.5 Pro, Vertex AI models -ui_type: CLI -ease_of_use_for_non_technical: 2 -github: https://github.com/google-gemini/gemini-cli ---- - -Gemini CLI is Google's open-source command-line AI workflow tool that brings the power of Gemini directly into your terminal. It connects to your tools, understands your code, and accelerates your workflows using a "reason and act (ReAct) loop" with built-in tools to complete complex tasks. - -Key features include querying and editing large codebases within Gemini's 1M token context window, generating new applications from PDFs or sketches using multimodal capabilities, and automating operational tasks like querying pull requests or handling complex rebases. The tool supports local and remote Model Context Protocol (MCP) servers and includes built-in tools for grep, terminal operations, file read/write, web search, and web fetch. diff --git a/docs/reference/clients/goose.md b/docs/reference/clients/goose.md deleted file mode 100644 index ba5a864..0000000 --- a/docs/reference/clients/goose.md +++ /dev/null @@ -1,15 +0,0 @@ ---- -title: Goose (CLI and Desktop) -url: https://block.github.io/goose/ -logo: https://block.github.io/goose/img/logo_light.png -models: Any -ui_type: Desktop or CLI -ease_of_use_for_non_technical: 4 -power_rating: 3 -notes: "Good ease-of-use for non-technical users but less powerful than Claude Code for complex agentic tasks" -github: https://github.com/block/goose ---- - -Goose is available as both a Desktop client and a command line client. We roll these into a single entry, because -they share many features and configurations. This means that technical developers, non-technical curators, and github -actions can all share the same confiurations and get similar behavior, just with a different UI. \ No newline at end of file diff --git a/docs/reference/clients/github-copilot.md b/docs/reference/github-copilot.md similarity index 97% rename from docs/reference/clients/github-copilot.md rename to docs/reference/github-copilot.md index 38f836b..448e9b4 100644 --- a/docs/reference/clients/github-copilot.md +++ b/docs/reference/github-copilot.md @@ -154,8 +154,8 @@ While it's possible to disable the firewall entirely, **this is not recommended* - [GitHub Copilot documentation](https://docs.github.com/en/copilot) - [Customizing the agent firewall](https://docs.github.com/en/copilot/how-tos/use-copilot-agents/coding-agent/customize-the-agent-firewall) -- [Set up GitHub actions](../../how-tos/set-up-github-actions.md) -- [Instruct the GitHub agent](../../how-tos/instruct-github-agent.md) +- [Set up GitHub actions](../how-tos/set-up-github-actions.md) +- [Instruct the GitHub agent](../how-tos/instruct-github-agent.md) ## Example: Configuring Copilot for OBO Workflows diff --git a/docs/reference/github-integrations.md b/docs/reference/github-integrations.md index 9c15498..112bedb 100644 --- a/docs/reference/github-integrations.md +++ b/docs/reference/github-integrations.md @@ -12,7 +12,7 @@ This guide documents the different approaches for integrating AI agents with Git ## Dragon-AI Agent -The Dragon-AI Agent approach uses custom GitHub Actions to deploy headless AI coding assistants (Claude Code or Goose) in response to issue/PR comments. +The Dragon-AI Agent approach uses custom GitHub Actions to deploy headless AI coding assistants (Claude Code or Goose) in response to issue/PR comments. New repositories should use a Claude Code Action workflow instead. See [Set up GitHub Actions](../how-tos/set-up-github-actions.md). ### How It Works @@ -153,7 +153,7 @@ Regardless of which approach you use, these files help guide AI behavior: | File | Purpose | |------|---------| | `CLAUDE.md` | System instructions for Claude-based agents | -| `.goosehints` | Instructions for Goose (often symlinked to CLAUDE.md) | +| `.goosehints` | Instructions for Goose. Historic; see [Harnesses](harnesses.md) | | `.github/copilot-instructions.md` | Instructions for GitHub Copilot | | `.github/ai-controllers.json` | Authorized users for Dragon-AI | diff --git a/docs/reference/harnesses.md b/docs/reference/harnesses.md new file mode 100644 index 0000000..b5e675e --- /dev/null +++ b/docs/reference/harnesses.md @@ -0,0 +1,69 @@ +# Harnesses + +A harness is the program that runs the agent loop. It sends your instructions to +a model, runs the tools the model asks for, and feeds the results back. + +The harness matters less than how you set up your repository. A well-configured +repository works with several harnesses. A badly configured one works with none. +Choose from this page, then spend your time on +[patterns](../patterns/index.md). + +## What to use + +| Harness | Where it runs | Instructions file | Use it for | +| --- | --- | --- | --- | +| [Claude Code](https://claude.ai/code) | Terminal, web, desktop, GitHub Actions | `CLAUDE.md` | Most curation work | +| [Codex](https://github.com/openai/codex) | Terminal | `AGENTS.md` | The same work, with OpenAI models | +| [GitHub Copilot coding agent](https://docs.github.com/en/copilot) | GitHub | `.github/copilot-instructions.md` | Issue to pull request inside GitHub | + +All the repositories in our [case studies](../case-studies/index.md) use Claude +Code, Codex, or both. Several also support Copilot. + +## Use a capable model + +[DisMech](../case-studies/dismech.md) states this as a project rule: + +> You should always use best-of-class models, and up to date high quality +> harnesses. Using less powerful models is more likely to generate lower +> quality content. + +Weak model output does not disappear. It arrives at review, gets sent back, and +costs a curator time. The exception is low-risk background work, where a cheaper +model is appropriate. See [Scanners find the work](../patterns/scanners.md). + +## Run in the cloud if installing is hard + +Many curators cannot install command line software, because of institutional IT +rules or because the setup is unfamiliar. You do not have to. + +[Claude Code on the web](https://code.claude.com/docs/en/claude-code-on-the-web) +clones the repository for you, runs the session in a cloud container, and opens +the pull request when you are done. DisMech's +[`CONTRIBUTING.md`](https://github.com/monarch-initiative/dismech/blob/main/CONTRIBUTING.md) +has step-by-step setup instructions, including the one-time environment +configuration that new users find hardest. + +For a whole team at once, a shared hosted environment removes the account and +key problem as well as the installation problem. The Gene Ontology consortium +runs one at [geneontology/go-jupyter](https://github.com/geneontology/go-jupyter). + +## Historic + +**Goose.** Earlier versions of this site recommended +[Goose](https://block.github.io/goose/) for curators who were not comfortable at +the command line. We no longer recommend it for new setups. It is still in use: +[Mondo's](../case-studies/mondo.md) `ai-agent.yml` runs dragon-ai-agent on +Goose, and Goose configuration remains in a few repositories. Keep it working +where it is deployed. Do not start there. + +**dragon-ai-agent.** The agent behind the `@dragon-ai-agent` mention in several +OBO repositories. Still running in Mondo and elsewhere. New repositories should +use a Claude Code Action workflow instead. See +[GitHub integrations](github-integrations.md). + +## Give every session the same tools + +Do not rely on curators configuring MCP servers themselves. Check the +configuration into the repository. [EFO](../case-studies/efo.md) declares its +two servers in `.mcp.json`, so every session gets ontology lookup and literature +access without setup. diff --git a/docs/tutorials/first-curation-session.md b/docs/tutorials/first-curation-session.md new file mode 100644 index 0000000..c8dcb1a --- /dev/null +++ b/docs/tutorials/first-curation-session.md @@ -0,0 +1,139 @@ +# Your first curation session + +This tutorial takes about thirty minutes. At the end you will have asked an +agent to do a real curation task and opened a pull request with the result. + +You do not need to install anything, and you do not need to be able to program. + +## Before you start + +You need: + +* A GitHub account. +* A Claude account on a paid plan. The free plan cannot run Claude Code. +* Write access to a repository that welcomes agent contributions. If you do not + have one, ask on the issue tracker of a project in our + [case studies](../case-studies/index.md). Several add new contributors on + request. + +## 1. Open a session in the cloud + +Go to [claude.ai/code](https://claude.ai/code) and choose to continue on the +web. Connect your GitHub account when you are asked, then select the repository +you want to work in. + +This runs the agent in a cloud container. The repository is cloned for you, and +when you are finished you press **Create PR** and the commits and pull request +are handled for you. + +Curating in the cloud avoids the two problems that stop most curators: you do +not install software, and your institution's IT rules do not apply to a browser +tab. For local installation instead, see [Harnesses](../reference/harnesses.md). + +## 2. Set up your environment once + +An environment holds the network settings, environment variables, and setup +commands your sessions run with. You configure it once, and every later session +reuses it. + +The settings that matter: + +* **Network access.** The default only allows GitHub and package registries. If + your curation reads PubMed, ClinicalTrials.gov, or another external source, + you need broader access. +* **Environment variables.** Any API keys your project's tools need go here, one + `KEY=value` per line, with no quotes. +* **Setup script.** Anything that should be installed before each session + starts, such as a task runner. + +Your project should tell you what to put in each field. DisMech's +[`CONTRIBUTING.md`](https://github.com/monarch-initiative/dismech/blob/main/CONTRIBUTING.md) +has a worked example with screenshots. + +## 3. Ask the agent to explain the project + +Start with this: + +``` +Give me a tour of this project +``` + +Then ask what you actually want to know: + +``` +I want to contribute. How? +``` + +This is not a warm-up exercise. The repository is the current description of how +the project works, and human documentation goes out of date. Asking the agent is +usually faster and more accurate than reading a guide. + +Ask follow-up questions until you understand what a good contribution looks +like. + +## 4. Do one small task + +Pick something narrow. A single term, a single record, one missing definition. + +Many projects give you a command for their main curation job. DisMech has one: + +``` +/curate Parkinson Disease +``` + +If your project has no command, describe the task in plain language and name +the file: + +``` +Add a definition for the term in src/ontology/xxx-edit.obo with id XXX:0000123. +Use the reference cited in issue 456. +``` + +Watch what the agent does. It will run commands, and some of them will fail. +This is normal. Agents try an option, read the error, and try another. Red text +in the output is usually the agent working, not the session breaking. + +## 5. Check the identifiers + +This is the part that needs you. + +Agents produce ontology term identifiers and citations that look right and are +wrong. Most projects on this site have automated checks for this, and you should +run them. Look for a validation command in the project's `CLAUDE.md`, or ask: + +``` +How do I validate what you just changed? +``` + +Then read the result yourself. Check that the term labels match what you would +expect, and that quoted evidence says what the entry claims it says. + +See [Make identifiers hard to fake](../patterns/ground-identifiers.md) for why +this failure mode is so common. + +## 6. Open the pull request + +Press **Create PR**. + +On many projects an automated reviewer now reads your pull request and either +requests changes or marks it ready to merge. If it requests changes, you can ask +your agent to address them in the same session. + +## 7. Start a new session for the next task + +Each session accumulates context: files it has read, commands it has run, output +it has kept. As that grows, the agent starts to lose track of your earlier +instructions. + +The fix is simple. Start a new session for each task, or each small group of +related tasks. Sessions are free to create, and your environment configuration +is reused. + +## What to do next + +* Read the [case study](../case-studies/index.md) for the repository you worked + in, to understand its setup. +* Read [Instruct the GitHub agent](../how-tos/instruct-github-agent.md) to ask + for work from an issue rather than a session. +* Read [Training materials for curators](tutorials-for-curators.md) for recorded + seminars and workshop material. diff --git a/docs/tutorials/ontology-editing-with-ai.md b/docs/tutorials/ontology-editing-with-ai.md deleted file mode 100644 index 8de33d9..0000000 --- a/docs/tutorials/ontology-editing-with-ai.md +++ /dev/null @@ -1,159 +0,0 @@ -# Using local AI tools - -This tutorial walks through how to use AI tools locally. This is aimed at mostly non-technical editors of -[OBO](../glossary.md#obo-format) ontologies that have adopted standardized ODK workflows. There are some technical steps required but these should be straightforward. This will work best with a Mac or Linux. - -## Background - -Most people have by now used web-based chat interfaces such as ChatGPT, or more powerful deep research models (o3, Perplexity Deep Research). These can now take advantage of information online as well as what is already "known" by the model. However, they can't interface with files and tools that you might have locally, e.g. - -- your local copy of the [ontology](../glossary.md#ontology) edit file, checked out from [GitHub](https://github.com) -- [Protege](https://protege.stanford.edu/) -- Reasoners -- Validation/QC workflows - -[Agentic AI](../glossary.md#ai-agent) is a paradigm where an AI application can make use of *tools* to achieve some objective. For ontology editing, this might be tools to edit the ontology or run a reasoner or workflow. These tools could include command line tools (e.g what you have available via ODK), or tools made available via the [Model Context Protocol](../glossary.md#model-context-protocol-mcp) - -There are a growing number of general-purpose applications that allow for easy plug and play of different tools. Many of these are aimed at developers, and hook into existing Integrated Development Environments (IDEs). These are not ideal for non-technical users. - -Two of the main easy-to-use Desktop applications are **[Claude](../glossary.md#claude) Desktop** (not to be confused with Claude Code, or the web interface to Claude) and **Goose**. We focus here on [Goose](../glossary.md#goose), as it allows for easy configurability - -### OBO Academy Seminar -Chris Mungall presented "Using AI Coding Apps for Ontology Development" at [OBO Academy](https://oboacademy.github.io/obook/courses/monarch-obo-training/) on June 10, 2025. You can watch the seminar here: - - - -## Install Goose - -Go to the [install page](https://block.github.io/goose/docs/getting-started/installation/) for Goose. Choose the Desktop app (more ambitious or technical users may want to also install the CLI app) - -## Set up your [LLM](../glossary.md#large-language-model-llm) - -Select settings/advanced, and select a Model. We recommend Anthropic/Claude Sonnet; you can try other models, but this tutorial has not been fully tested with other models. - -You will need an API key. We recommend speaking to your supervisor about getting an API key that is charged to a project you work on. - - -## Try it out - -You should be able to use the UI the same way as a normal AI chat, try asking some questions about your favorite topics. - -However, you can do more here, including with local files. Try giving it a folder full of PDFs and asking it to list the contents. Try finding a particular PDF, and then asking it to summarize it. - -## Install xcode (mac users) - -In order to use most of the useful features of agents, you will need to have certain things installed locally. For Macs, this means installing Xcode: - -* [Xcode](https://apps.apple.com/us/app/xcode/id497799835?mt=12) - -You can just ask Goose to walk you through the installation. - -## Try it out more advanced features - -Some things you can try: - -* `clone the OBO cell ontology repo from github` -* `create a web page for an ontology` -* `create a web app for annotating single-cell experiments, use the OLS API to implement autocomplete over CL` - -## Install OWL-MCP extension - -Next we will try installing an ontology-specific MCP - -You can either install directly from this link: - - * ⬇️ [Install OWL-MCP](goose://extension?cmd=uvx&arg=owl-mcp&id=owl_mcp&name=OWL%20MCP) - - Or to do this manually, in the Extension section of Goose, add a new entry for owlmcp: - - `uvx owl-mcp` - - This video shows how to do this manually: - - - -## Try it out - -You can ask to create an ontology, and add axioms to an ontology. We will outline the steps below. You can also watch this video, which shows an example session. - - - -## Clone the demo repo - -Ask goose to clone [ai4curation/ai-ontology-tutorial](https://github.com/ai4curation/ai-ontology-tutorial). If you have a favorite location on your disk for checked out repos, you can instruct it to place it there. - -Or if you like, you can clone the repo using your favorite git client, e.g. Github Desktop. - -## Navigate to the repo - -In the top of the window for Goose there is a file navigator, you can use this to navigate to the repo in which you checked out the demo repo - -## Open the ontology in Protege - -The demo repo contains one OWL file, `anatomy.ofn`. Open this with Protege. You should see a (highly incomplete!) anatomy ontology, with terms for `digit` and `limb segment` - -## Ask goose to query the ontology - -Try asking questions like `what are all the subclass axioms in the ontology`? - -(remember this is a small test ontology and we don't expect many) - -## Adding terms - -Now try asking for multiple terms to be added in batch: - -`Add terms for fingers, toes, hands, and feet, ensuring part-of relationships between them` - -You should be able to see the AI working - -## Sync Protege - -Switch to your Protege window. Protege will inform you that the file has changed, and ask if you want to see the new content. Say yes. - -The Protege screen should show the changes having been made. - -## Adding more complex axioms - -Now try asking: - -`Add an axiom stating that hands have 5 fingers` - -This should end up putting a cardinality axiom into the ontology. It may do this by using an inverse expression, you can ask it to name the inverse property instead (see video). - -## Other experiments - -I recommend trying experiments on this test repo. Also try experiments like having goose write code that parses the ontology and summarizes the content. - -## Putting this into practice - -- if you want to start using this with an existing OWL ontology like CL, you need to have `robot` on your path, so that the ontology can be normalized prior to commits -- if you are using obo format as source, it's a slightly different workflow, see the odk-ai repo for details -- if your ontology is set up to use [AI agents triggered by github actions](../how-tos/set-up-github-actions.md), then you don't actually need to run any AI tools locally. However, I strongly recommend learning how to do this, as this will empower you to do many more things than are possible with the GitHub agent. - -## Additional training materials - -For more in-depth training, see these workshop materials from 2025-2026: - -- **[Accelerating Ontology Curation with Agentic AI and GitHub](https://go.lbl.gov/icbo-2025-ai)** — ICBO 2025 tutorial with practical guidance, exercises, and social coding patterns. [Recording](https://youtu.be/_9Re39yB7EE), [Zenodo](https://zenodo.org/records/18653147) -- **[Gene Ontology Curators AI Workshop (Part 1)](https://zenodo.org/records/17993529)** — Practical AI support for GO curation workflows -- **[Agentic AI and GO, Oct 2025 (Part 1)](https://zenodo.org/records/18333490)** — Foundational session on using coding agents in ontology workflows -- **[Using Agentic AI to Review GO Standard Annotations (Part 2)](https://zenodo.org/records/17498090)** — Hands-on annotation review, evidence checks, and quality control patterns - -For building the infrastructure around your agents, see [Build your agentic harness](../how-tos/build-agentic-harness.md). - diff --git a/docs/tutorials/tutorials-for-curators.md b/docs/tutorials/tutorials-for-curators.md index 957cb7a..86f94a3 100644 --- a/docs/tutorials/tutorials-for-curators.md +++ b/docs/tutorials/tutorials-for-curators.md @@ -1,6 +1,8 @@ # Training Materials for Curators -This page collects video tutorials and training materials for curators learning to use AI tools. +This page collects video tutorials and training materials for curators learning +to use AI tools. To try an agent yourself first, see +[Your first curation session](first-curation-session.md). ## OBO Academy AI Training @@ -12,7 +14,7 @@ The [OBO Academy Monarch Ontology Training](https://oboacademy.github.io/obook/c Practical introduction to using Claude Code for biocuration tasks. - **Using AI Coding Apps for Ontology Developers** (June 2025) - Chris Mungall - Overview of agentic AI tools for ontology development, focusing on Goose and Claude Code. + Overview of agentic AI tools for ontology development. Covers Goose, which we now treat as [historic](../reference/harnesses.md), alongside Claude Code. @@ -30,3 +32,18 @@ See the full [OBO Academy training calendar](https://oboacademy.github.io/obook/ - [OBO Academy obook](https://oboacademy.github.io/obook/) - Comprehensive semantic engineering training materials - [Claude Code tutorial for ontology developers](https://oboacademy.github.io/obook/tutorial/claude-code-getting-started/) - Step-by-step guide for getting started - [DeepLearning.AI Claude Code course](https://learn.deeplearning.ai/courses/claude-code-a-highly-agentic-coding-assistant/) - General introduction accessible to non-programmers + +## Workshop materials + +- **[Accelerating Ontology Curation with Agentic AI and GitHub](https://go.lbl.gov/icbo-2025-ai)** — + ICBO 2025 tutorial, with exercises and social coding patterns. + [Recording](https://youtu.be/_9Re39yB7EE), + [Zenodo](https://zenodo.org/records/18653147) +- **[Agentic AI and GO, part 1](https://zenodo.org/records/17993529)** and + **[part 2](https://zenodo.org/records/17498090)** — Gene Ontology workshop + series on using coding agents in GO development. +- **[Staying in the Loop: A Biocurator's Guide to Agentic AI Developments](https://zenodo.org/records/18614836)** — + overview talk for biocurators. +- **[geneontology/go-jupyter](https://github.com/geneontology/go-jupyter)** — + the shared cloud environment used to run agent workshops, where each + participant gets a preconfigured workspace and no key of their own. diff --git a/mkdocs.yml b/mkdocs.yml index fbba2d3..9fc54ec 100644 --- a/mkdocs.yml +++ b/mkdocs.yml @@ -36,8 +36,30 @@ plugins: nav: - Home: index.md - - FAQ: faq.md - - Examples: examples.md + - Case studies: + - Overview: case-studies/index.md + - Ontologies: + - GO: case-studies/go-ontology.md + - Uberon: case-studies/uberon.md + - Mondo: case-studies/mondo.md + - Cell Ontology: case-studies/cell-ontology.md + - EFO: case-studies/efo.md + - The mechs: + - DisMech: case-studies/dismech.md + - AI Gene Review: case-studies/ai-gene-review.md + - CommunityMech: case-studies/communitymech.md + - HabitatMech: case-studies/habitatmech.md + - Patterns: + - Overview: patterns/index.md + - Make identifiers hard to fake: patterns/ground-identifiers.md + - Regulate the loop: patterns/human-regulating-the-loop.md + - Break work into skills: patterns/skills-before-automation.md + - One source of instructions: patterns/one-source-of-instructions.md + - Fast and slow validation: patterns/fast-and-slow-validation.md + - Scanners find the work: patterns/scanners.md + - One review checklist: patterns/review-checklists.md + - Shared control: patterns/shared-control.md + - Guard the untrusted surface: patterns/guard-untrusted-input.md - Curator how-tos: - Instruct the GitHub agent: how-tos/instruct-github-agent.md - Make identifiers hallucination-resistant: how-tos/make-ids-hallucination-resistant.md @@ -49,21 +71,17 @@ nav: - Create an agentic curation pipeline: how-tos/create-agentic-curation-pipeline.md - PR reviews for agent improvement: how-tos/pr-reviews-for-agent-improvement.md - Tutorials: + - Your first curation session: tutorials/first-curation-session.md - Training materials for curators: tutorials/tutorials-for-curators.md - - Ontology editing with AI: tutorials/ontology-editing-with-ai.md - Reference: - - Client apps: reference/client-apps.md + - Harnesses: reference/harnesses.md - Agentic tools: reference/agentic-tools.md - GitHub integrations: reference/github-integrations.md + - GitHub Copilot: reference/github-copilot.md - Claude Skills: reference/claude-skills.md - - Clients: - - Claude Desktop: reference/clients/claude-desktop.md - - Claude Code: reference/clients/claude-code.md - - Codex CLI: reference/clients/codex-cli.md - - Gemini CLI: reference/clients/gemini-cli.md - - GitHub Copilot: reference/clients/github-copilot.md - - Goose: reference/clients/goose.md + - Evidence: evidence.md + - FAQ: faq.md - Glossary: glossary.md -site_url: https://ai4curation.github.io/aidocs/ +site_url: https://ai4curation.io/aidocs/ repo_url: https://github.com/ai4curation/aidocs/ From 46ec6be85640ecf5ea24be4a3a9b2a88014828fe Mon Sep 17 00:00:00 2001 From: Claude Date: Tue, 18 Aug 2026 22:57:53 +0000 Subject: [PATCH 2/5] docs: fix the broken links the new link check found The link check on the previous commit found 11 errors out of 456 links. All of them were real except the bot-blocked hosts. Genuine mistakes in the new pages: - HabitatMech's term requests page is at pages/term-requests.html. The link in HabitatMech's own README is missing the pages/ segment and 404s. - geneontology/go-jupyter is not a public repository. Describe the shared JupyterHub setup without linking to it, in three places. Pre-existing errors the check surfaced: - owl-mcp is ai4curation/owl-mcp, not monarch-initiative/owl-mcp. - dismech renamed stale-pr-reassign.yml to pr-shepherd.yml. Excluded from the check: claude.ai and learn.deeplearning.ai return 403 to automated requests. Both links are correct in a browser. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_017tKFksxqHZJNLoH1zJ5WXm --- .github/workflows/docs-checks.yml | 2 ++ docs/case-studies/habitatmech.md | 4 ++-- docs/faq.md | 2 +- docs/how-tos/build-agentic-harness.md | 2 +- docs/how-tos/create-agentic-curation-pipeline.md | 2 +- docs/how-tos/integrate-ai-into-your-kb.md | 7 +++---- docs/reference/harnesses.md | 4 +++- docs/tutorials/tutorials-for-curators.md | 3 --- 8 files changed, 13 insertions(+), 13 deletions(-) diff --git a/.github/workflows/docs-checks.yml b/.github/workflows/docs-checks.yml index 908b780..29e072f 100644 --- a/.github/workflows/docs-checks.yml +++ b/.github/workflows/docs-checks.yml @@ -41,6 +41,8 @@ jobs: --accept 200,206,429 --exclude-path docs/overrides --exclude 'github\.com/ai4curation/agent-watcher' + --exclude 'claude\.ai' + --exclude 'learn\.deeplearning\.ai' 'docs/**/*.md' 'README.md' fail: true diff --git a/docs/case-studies/habitatmech.md b/docs/case-studies/habitatmech.md index f566743..5acac2e 100644 --- a/docs/case-studies/habitatmech.md +++ b/docs/case-studies/habitatmech.md @@ -27,7 +27,7 @@ The same habitat has a different name in every source: | Path | What it is | | --- | --- | | `data/habitats/` | One YAML file per habitat, about 3,200 records | -| [Term requests page](https://culturebotai.github.io/HabitatMech/term-requests.html) | Gaps this project is asking ENVO to fill | +| [Term requests page](https://culturebotai.github.io/HabitatMech/pages/term-requests.html) | Gaps this project is asking ENVO to fill | | `src/habitatmech/schema/` | The LinkML schema | ## What works @@ -56,7 +56,7 @@ pressure to guess. ### Gaps are published as term requests The project renders the terms it needs and cannot find as a public -[term requests page](https://culturebotai.github.io/HabitatMech/term-requests.html) +[term requests page](https://culturebotai.github.io/HabitatMech/pages/term-requests.html) for the ontology community. Curation that finds a gap produces a request rather than a workaround. diff --git a/docs/faq.md b/docs/faq.md index c9aaafe..b559fc3 100644 --- a/docs/faq.md +++ b/docs/faq.md @@ -64,7 +64,7 @@ Use **[ai-blame](https://github.com/ai4curation/ai-blame)**, which extracts prov - **[noctua-mcp](https://github.com/geneontology/noctua-mcp)** — GO-CAM editing via Noctua/Barista - **[oak-mcp](https://github.com/monarch-initiative/oak-mcp)** — ontology search, traversal, and operations via OAK -- **[owl-mcp](https://github.com/monarch-initiative/owl-mcp)** — general OWL ontology operations +- **[owl-mcp](https://github.com/ai4curation/owl-mcp)** — general OWL ontology operations These give agents structured access to domain-specific operations instead of raw file manipulation. diff --git a/docs/how-tos/build-agentic-harness.md b/docs/how-tos/build-agentic-harness.md index c356ae0..26fdff0 100644 --- a/docs/how-tos/build-agentic-harness.md +++ b/docs/how-tos/build-agentic-harness.md @@ -41,7 +41,7 @@ MCP servers give agents structured access to domain-specific operations instead - **[noctua-mcp](https://github.com/geneontology/noctua-mcp)** — GO-CAM editing via Noctua/Barista - **[oak-mcp](https://github.com/monarch-initiative/oak-mcp)** — ontology operations via OAK -- **[owl-mcp](https://github.com/monarch-initiative/owl-mcp)** — OWL ontology operations +- **[owl-mcp](https://github.com/ai4curation/owl-mcp)** — OWL ontology operations Without proper tool access, agents resort to ad-hoc text manipulation of ontology files, which is error-prone. diff --git a/docs/how-tos/create-agentic-curation-pipeline.md b/docs/how-tos/create-agentic-curation-pipeline.md index 84e1b0e..e112f4d 100644 --- a/docs/how-tos/create-agentic-curation-pipeline.md +++ b/docs/how-tos/create-agentic-curation-pipeline.md @@ -486,7 +486,7 @@ Once the baseline validation gate works, add optional GitHub Actions around it: [kgx-release.yaml](https://github.com/monarch-initiative/dismech/blob/main/.github/workflows/kgx-release.yaml). - **Stale review follow-up**: reassigns old PRs with outstanding review feedback back to the agent queue, as in - [stale-pr-reassign.yml](https://github.com/monarch-initiative/dismech/blob/main/.github/workflows/stale-pr-reassign.yml). + [pr-shepherd.yml](https://github.com/monarch-initiative/dismech/blob/main/.github/workflows/pr-shepherd.yml). - **Copilot setup**: if you use GitHub Copilot coding agent, add [copilot-setup-steps.yml](https://github.com/monarch-initiative/dismech/blob/main/.github/copilot-setup-steps.yml) so the agent has the right dependencies and firewall setup. diff --git a/docs/how-tos/integrate-ai-into-your-kb.md b/docs/how-tos/integrate-ai-into-your-kb.md index edf2c2d..29b70bb 100644 --- a/docs/how-tos/integrate-ai-into-your-kb.md +++ b/docs/how-tos/integrate-ai-into-your-kb.md @@ -31,10 +31,9 @@ clones the repository, runs the session in a cloud container, and opens the pull request. A curator needs a paid Claude plan and write access to your repository. For a whole team at once, a shared hosted environment removes the account -problem as well. The Gene Ontology consortium runs one at -[geneontology/go-jupyter](https://github.com/geneontology/go-jupyter), where each -participant gets a workspace with an agent already configured and no key of -their own. +problem as well. The Gene Ontology consortium runs a JupyterHub instance for its +workshops, where each participant gets a workspace with an agent already +configured and no key of their own. Whichever route you choose, write down the setup steps in your repository. DisMech's [`CONTRIBUTING.md`](https://github.com/monarch-initiative/dismech/blob/main/CONTRIBUTING.md) diff --git a/docs/reference/harnesses.md b/docs/reference/harnesses.md index b5e675e..c3efc7c 100644 --- a/docs/reference/harnesses.md +++ b/docs/reference/harnesses.md @@ -45,7 +45,9 @@ configuration that new users find hardest. For a whole team at once, a shared hosted environment removes the account and key problem as well as the installation problem. The Gene Ontology consortium -runs one at [geneontology/go-jupyter](https://github.com/geneontology/go-jupyter). +runs a JupyterHub instance for its workshops, where each participant logs in +with GitHub and gets a workspace with an agent already configured, on a central +key they never see. ## Historic diff --git a/docs/tutorials/tutorials-for-curators.md b/docs/tutorials/tutorials-for-curators.md index 86f94a3..66f0ed4 100644 --- a/docs/tutorials/tutorials-for-curators.md +++ b/docs/tutorials/tutorials-for-curators.md @@ -44,6 +44,3 @@ See the full [OBO Academy training calendar](https://oboacademy.github.io/obook/ series on using coding agents in GO development. - **[Staying in the Loop: A Biocurator's Guide to Agentic AI Developments](https://zenodo.org/records/18614836)** — overview talk for biocurators. -- **[geneontology/go-jupyter](https://github.com/geneontology/go-jupyter)** — - the shared cloud environment used to run agent workshops, where each - participant gets a preconfigured workspace and no key of their own. From b7376f464fb1639121e1554b408334118bd1c478 Mon Sep 17 00:00:00 2001 From: Claude Date: Wed, 19 Aug 2026 00:03:12 +0000 Subject: [PATCH 3/5] docs: add a Bots reference page for GitHub identities "Agent" was covering three different things: the software, the subagent, and the GitHub identity work appears under. Split off the last one and call it a bot, which is the word GitHub already puts in the byline (ai4c-agent[bot]). The frame: the agent is the software that does the work, the bot is the identity it works under, and one agent can work under different bots in different repositories. New docs/reference/bots.md covers: - The bots in use: ai4c-agent, ai4c-reviewer, claude, copilot-swe-agent, github-actions, dragon-ai-agent. - GitHub Apps versus machine accounts, and why Apps are preferable for new setups: per-run tokens, per-repository permissions, no seat. - Why DisMech runs two of its own. GitHub does not let an identity approve its own pull request, so ai4c-agent writes and ai4c-reviewer supplies the approving review that branch protection requires. Without that split, every agent pull request would need a human approval. - Setting the commit author so App commits attribute correctly. - Reading a byline to work out what produced a change. Corrected elsewhere: dragon-ai-agent is semi-retired, not simply "still running". Mondo still runs the machine account on a PAT; DisMech retired the account and keeps @dragon-ai-agent as a trigger keyword only, with the work done by the ai4c-agent App. The same mention means different things in the two repositories, which is worth knowing when reading old issue threads. Updated: harnesses.md, github-integrations.md, the DisMech and Mondo case studies, the untrusted-surface pattern, glossary, and nav. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_017tKFksxqHZJNLoH1zJ5WXm --- docs/case-studies/dismech.md | 15 +++ docs/case-studies/mondo.md | 5 + docs/glossary.md | 20 ++++ docs/patterns/guard-untrusted-input.md | 4 + docs/reference/bots.md | 141 +++++++++++++++++++++++++ docs/reference/github-integrations.md | 2 + docs/reference/harnesses.md | 12 ++- mkdocs.yml | 1 + 8 files changed, 198 insertions(+), 2 deletions(-) create mode 100644 docs/reference/bots.md diff --git a/docs/case-studies/dismech.md b/docs/case-studies/dismech.md index 8d5c054..ad705c3 100644 --- a/docs/case-studies/dismech.md +++ b/docs/case-studies/dismech.md @@ -81,6 +81,21 @@ entries, and move stalled issues and pull requests forward. Low-risk scanners may use cheaper models. A `low_effort` label lets a human assign a task to a cheaper model by hand. See [Scanners](../patterns/scanners.md). +### Two bots, so review counts + +DisMech runs two GitHub Apps of its own. +[`ai4c-agent`](https://github.com/apps/ai4c-agent) does the work. It responds to +mentions and moves stalled pull requests forward. +[`ai4c-reviewer`](https://github.com/apps/ai4c-reviewer) reviews and supplies +the approving review that branch protection requires. + +The split exists because GitHub does not let an identity approve its own pull +request. One bot doing both jobs would mean every agent pull request needed a +human approval, which would undo the whole model. See [Bots](../reference/bots.md). + +The `dragon-ai-agent` machine account is retired here. The +`@dragon-ai-agent please ...` mention survives as a trigger keyword only. + ### The untrusted surface is guarded Pull requests from forks are closed, because GitHub does not give fork workflows diff --git a/docs/case-studies/mondo.md b/docs/case-studies/mondo.md index 553ed94..4ca63f7 100644 --- a/docs/case-studies/mondo.md +++ b/docs/case-studies/mondo.md @@ -44,6 +44,11 @@ subagents under `.claude/agents/` are Claude Code features. An agent triggered by `ai-agent.yml` therefore cannot reach them. They are available in local Claude Code sessions only. +Mondo is also the last repository we track that still runs the +`dragon-ai-agent` machine account. It authenticates with a personal access token +held in the `PAT_FOR_PR` secret. DisMech has moved the same job to a GitHub App +with a token minted per run. See [Bots](../reference/bots.md). + This is not necessarily wrong, but it is undocumented. If you have both, say in `CLAUDE.md` which surface each one serves. If you want the subagents reachable from GitHub, add a workflow that runs Claude Code the way diff --git a/docs/glossary.md b/docs/glossary.md index dafd177..be9f34b 100644 --- a/docs/glossary.md +++ b/docs/glossary.md @@ -153,3 +153,23 @@ A separate agent run started by the main session, with its own context. Useful for jobs that read a great deal and report a little, because the reading does not stay in the main session. See [Break work into skills](patterns/skills-before-automation.md). + +## Bot + +The GitHub identity that agent work appears under: the name in the byline on an +issue, commit, or pull request. Distinct from the agent, which is the software +doing the work. A bot is either a GitHub App or a machine account. See +[Bots](reference/bots.md). + +## GitHub App + +An application installed into a repository or organization, with its own +identity and its own permissions. Workflows mint a short-lived token for it per +run. GitHub shows App bylines with a `[bot]` suffix, as in `ai4c-agent[bot]`. + +## Machine account + +An ordinary GitHub user account that a program logs in as, sometimes called a +machine user. It looks like a person in the interface, uses a seat, and +authenticates with a long-lived personal access token. `dragon-ai-agent` is one. +Prefer a [GitHub App](#github-app) for anything new. diff --git a/docs/patterns/guard-untrusted-input.md b/docs/patterns/guard-untrusted-input.md index dc85911..81792d3 100644 --- a/docs/patterns/guard-untrusted-input.md +++ b/docs/patterns/guard-untrusted-input.md @@ -31,6 +31,10 @@ documenting the keyword does not summon the agent every time someone reads it. production. Give the watcher read access, not write. Instructions are guidance; credentials are the control. +**Pick the identity deliberately.** What a bot may do is set by its permissions, +not by what you told it. A reviewer bot needs to read code and write reviews; it +does not need to push branches. See [Bots](../reference/bots.md). + ## Separate reading from writing Agents that only read are far less risky than agents that write. If your agent diff --git a/docs/reference/bots.md b/docs/reference/bots.md new file mode 100644 index 0000000..c7c7927 --- /dev/null +++ b/docs/reference/bots.md @@ -0,0 +1,141 @@ +# Bots + +A **bot** is the GitHub identity that agent work appears under. It is the name in +the byline on an issue, a comment, a commit, or a pull request. + +Keep this separate from the agent itself: + +* The **agent** is the software that does the work. See + [Harnesses](harnesses.md). +* The **bot** is the identity it works under. + +The same agent can work under different bots in different repositories. Claude +Code runs as `claude[bot]` in one repository and as `ai4c-agent[bot]` in +another. The agent is the same. The permissions are not. + +## The bots you will meet + +| Bot | Kind | What it does | Repositories | +| --- | --- | --- | --- | +| [`ai4c-agent`](https://github.com/apps/ai4c-agent) | GitHub App | Does curation work: responds to mentions, moves stalled pull requests forward | DisMech, AI Gene Review | +| [`ai4c-reviewer`](https://github.com/apps/ai4c-reviewer) | GitHub App | Reviews pull requests and supplies the approving review | DisMech, AI Gene Review | +| [`claude`](https://github.com/apps/claude) | GitHub App | Runs Claude Code from issues and pull requests | Many | +| `copilot-swe-agent` | GitHub App | Works an issue and opens a draft pull request | EFO, Uberon | +| `github-actions` | GitHub App | The built-in identity for workflow runs | Everywhere | +| `dragon-ai-agent` | Machine account | Older curation bot, now semi-retired | Mondo | + +## The two kinds + +### GitHub Apps + +Most of the list above. An App is installed into a repository or an +organization. A workflow mints a short-lived token for it at the start of each +run, and the token expires when the run ends. Nothing long-lived sits in your +secrets except the App's private key. + +You can tell an App by its byline. GitHub appends `[bot]`, so you see +`ai4c-agent[bot]`, not `ai4c-agent`. + +Apps do not use a seat in your organization, and you grant each one only the +permissions it needs. + +### Machine accounts + +A machine account is an ordinary GitHub user account that a program logs in as. +GitHub calls these machine users. `dragon-ai-agent` is one. + +A machine account looks like a person in the interface. There is no `[bot]` +suffix, it can join teams, and it uses a seat. It authenticates with a personal +access token, which is long-lived and sits in your repository secrets. + +Prefer an App for anything new. Shorter-lived credentials and per-repository +permissions are both worth having. + +## Why a project runs more than one bot + +DisMech runs two of its own, and the reason is worth understanding before you +copy the setup. + +**GitHub does not let an identity approve its own pull request.** + +DisMech requires one approving review before a pull request can merge. If the +bot that opened the pull request were also the bot that reviewed it, the +approval would not count, and every agent pull request would need a human +approval. + +So the work is split: + +* `ai4c-agent` writes. It responds to mentions and pushes branches. +* `ai4c-reviewer` reviews. It applies the review rubric and supplies the + approving review that branch protection requires. + +Two identities, one boundary. This is what makes it possible for the project to +say that contributors are not expected to check their own agent's output. See +[One review checklist](../patterns/review-checklists.md). + +## The dragon-ai-agent case + +`@dragon-ai-agent` means two different things depending on which repository you +are in. This trips people up, so it is worth stating plainly. + +**In [Mondo](../case-studies/mondo.md), the machine account is live.** +`.github/workflows/ai-agent.yml` runs `dragon-ai-agent/run-goose-obo` and +authenticates with a personal access token stored as `PAT_FOR_PR`. The account +is doing the work. + +**In [DisMech](../case-studies/dismech.md), the account is retired.** The +workflow says so in a comment. What survives is the *phrase*: writing +`@dragon-ai-agent please ...` in an issue is a trigger keyword that starts a +workflow. The work is then done by the `ai4c-agent` App under a short-lived +token. + +If you are reading an old issue thread, the mention tells you a workflow ran. It +does not tell you which identity did the work. Check the byline on the resulting +pull request. + +## Make commits attribute correctly + +An App token alone does not set the commit author. If you do not configure it, +commits land under whatever Git identity the runner happens to have. + +DisMech sets it explicitly in `pr-shepherd.yml`, using the address form GitHub +recognizes for App commits: + +```bash +APP_USER_ID="$(gh api "/users/${APP_SLUG}[bot]" --jq .id)" +git config --global user.name "${APP_SLUG}[bot]" +git config --global user.email "${APP_USER_ID}+${APP_SLUG}[bot]@users.noreply.github.com" +``` + +Do this. Without it you cannot tell from `git log` which bot wrote a line, and +tools that measure agent contribution have nothing to work from. See +[Evidence](../evidence.md). + +## Restrict who can summon a bot + +A bot with write access, triggered by a comment, is reachable by anyone who can +comment. Both parts need controlling: + +* **Who may trigger it.** DisMech checks the commenter against a list of + authorized controllers before the agent runs. +* **What it may do once triggered.** This is the App's permission set, not a + matter of instructions. Grant each App the narrowest permissions that let it + do its job. + +Separate identities help here too. A reviewer bot needs to read code and write +reviews. It does not need to push branches. + +See [Guard the untrusted surface](../patterns/guard-untrusted-input.md). + +## Working out who did what + +| You see | It means | +| --- | --- | +| `name[bot]` | A GitHub App | +| A plain login | A person, or a machine account | +| `github-actions[bot]` | A workflow, using the built-in token | +| `@dragon-ai-agent` in comment text | A trigger keyword, not necessarily that account | + +A byline tells you the identity. It does not tell you the model, the harness, or +who asked. For that you need the execution trace. See +[Evidence](../evidence.md). diff --git a/docs/reference/github-integrations.md b/docs/reference/github-integrations.md index 112bedb..ca82e0f 100644 --- a/docs/reference/github-integrations.md +++ b/docs/reference/github-integrations.md @@ -14,6 +14,8 @@ This guide documents the different approaches for integrating AI agents with Git The Dragon-AI Agent approach uses custom GitHub Actions to deploy headless AI coding assistants (Claude Code or Goose) in response to issue/PR comments. New repositories should use a Claude Code Action workflow instead. See [Set up GitHub Actions](../how-tos/set-up-github-actions.md). +The `dragon-ai-agent` machine account itself is semi-retired. Mondo still runs it; DisMech keeps the mention as a trigger keyword only. For which identity actually does the work, see [Bots](bots.md). + ### How It Works 1. A controller invokes the agent with `@dragon-ai-agent please` in an issue or PR comment diff --git a/docs/reference/harnesses.md b/docs/reference/harnesses.md index c3efc7c..0aff317 100644 --- a/docs/reference/harnesses.md +++ b/docs/reference/harnesses.md @@ -59,10 +59,18 @@ Goose, and Goose configuration remains in a few repositories. Keep it working where it is deployed. Do not start there. **dragon-ai-agent.** The agent behind the `@dragon-ai-agent` mention in several -OBO repositories. Still running in Mondo and elsewhere. New repositories should -use a Claude Code Action workflow instead. See +OBO repositories. Semi-retired: [Mondo](../case-studies/mondo.md) still runs the +machine account, while [DisMech](../case-studies/dismech.md) has retired it and +keeps only the mention as a trigger keyword, with the work done by the +`ai4c-agent` App. New repositories should use a Claude Code Action workflow +instead. See [Bots](bots.md) and [GitHub integrations](github-integrations.md). +## The agent is not the identity + +The harness is the software. The identity the work appears under is a separate +choice, and it is the one that controls permissions. See [Bots](bots.md). + ## Give every session the same tools Do not rely on curators configuring MCP servers themselves. Check the diff --git a/mkdocs.yml b/mkdocs.yml index 9fc54ec..dd1c3db 100644 --- a/mkdocs.yml +++ b/mkdocs.yml @@ -75,6 +75,7 @@ nav: - Training materials for curators: tutorials/tutorials-for-curators.md - Reference: - Harnesses: reference/harnesses.md + - Bots: reference/bots.md - Agentic tools: reference/agentic-tools.md - GitHub integrations: reference/github-integrations.md - GitHub Copilot: reference/github-copilot.md From c744c3ddec9c0292d6456ff4b8e66c4418380715 Mon Sep 17 00:00:00 2001 From: Claude Date: Wed, 19 Aug 2026 00:32:31 +0000 Subject: [PATCH 4/5] docs: split bots into a page per bot, and correct the GO facts Restructure docs/reference/bots.md into docs/reference/bots/ with an overview plus one page per bot: ai4c-agent, ai4c-reviewer, claude, copilot-swe-agent, dragon-ai-agent. Drop the "bots you will meet" heading. Corrected: GO runs both ai4c-agent and ai4c-reviewer. The previous version listed them for DisMech and AI Gene Review only. GO's two agent workflows turn out to be the best-documented examples of this setup anywhere in these repositories, and their header comments supply several things worth writing down: - A GitHub App cannot be @-mentioned. "@ai4c-agent" renders as plain text, so every bot mention is really a trigger keyword that a workflow matches against comment text. This generalises the dragon-ai-agent case rather than being an exception to it, and it means a keyword can be migrated without touching the identity. GO honours @dragon-ai-agent as a legacy handle during that move. - An App has two different numbers. The app id mints the token; the bot account's user id is what commit attribution needs. Using the wrong one silently unlinks commits from the App. - Review actions skip bot-authored pull requests by default, so the editing bot has to be named in allowed_bots or the highest-value reviews never happen. - GO keeps the review rubric in the pr-review skill rather than the workflow prompt, so humans and bots apply the same criteria. - Authorized triggerers are listed in .github/ai-controllers.json. Also updated the GO case study, which had described the two workflows without naming the identities they run as. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_017tKFksxqHZJNLoH1zJ5WXm --- docs/case-studies/dismech.md | 2 +- docs/case-studies/go-ontology.md | 23 +++- docs/case-studies/mondo.md | 2 +- docs/glossary.md | 2 +- docs/patterns/guard-untrusted-input.md | 2 +- docs/reference/bots.md | 141 --------------------- docs/reference/bots/ai4c-agent.md | 78 ++++++++++++ docs/reference/bots/ai4c-reviewer.md | 74 +++++++++++ docs/reference/bots/claude.md | 53 ++++++++ docs/reference/bots/copilot.md | 58 +++++++++ docs/reference/bots/dragon-ai-agent.md | 62 ++++++++++ docs/reference/bots/index.md | 165 +++++++++++++++++++++++++ docs/reference/github-integrations.md | 2 +- docs/reference/harnesses.md | 4 +- mkdocs.yml | 8 +- 15 files changed, 525 insertions(+), 151 deletions(-) delete mode 100644 docs/reference/bots.md create mode 100644 docs/reference/bots/ai4c-agent.md create mode 100644 docs/reference/bots/ai4c-reviewer.md create mode 100644 docs/reference/bots/claude.md create mode 100644 docs/reference/bots/copilot.md create mode 100644 docs/reference/bots/dragon-ai-agent.md create mode 100644 docs/reference/bots/index.md diff --git a/docs/case-studies/dismech.md b/docs/case-studies/dismech.md index ad705c3..63a4bd7 100644 --- a/docs/case-studies/dismech.md +++ b/docs/case-studies/dismech.md @@ -91,7 +91,7 @@ the approving review that branch protection requires. The split exists because GitHub does not let an identity approve its own pull request. One bot doing both jobs would mean every agent pull request needed a -human approval, which would undo the whole model. See [Bots](../reference/bots.md). +human approval, which would undo the whole model. See [Bots](../reference/bots/index.md). The `dragon-ai-agent` machine account is retired here. The `@dragon-ai-agent please ...` mention survives as a trigger keyword only. diff --git a/docs/case-studies/go-ontology.md b/docs/case-studies/go-ontology.md index 7e09c06..4c312ef 100644 --- a/docs/case-studies/go-ontology.md +++ b/docs/case-studies/go-ontology.md @@ -12,8 +12,9 @@ ontology repository we track. | Path | What it is | | --- | --- | | [`.claude/skills/`](https://github.com/geneontology/go-ontology/tree/master/.claude/skills) | Ten skills, one per recurring job | -| [`.github/workflows/ai-agent.yml`](https://github.com/geneontology/go-ontology/blob/master/.github/workflows/ai-agent.yml) | Runs an agent when someone mentions it in an issue | -| [`.github/workflows/claude-code-review.yml`](https://github.com/geneontology/go-ontology/blob/master/.github/workflows/claude-code-review.yml) | Reviews pull requests | +| [`.github/workflows/ai-agent.yml`](https://github.com/geneontology/go-ontology/blob/master/.github/workflows/ai-agent.yml) | Runs the [ai4c-agent](../reference/bots/ai4c-agent.md) bot on a trigger keyword | +| [`.github/workflows/claude-code-review.yml`](https://github.com/geneontology/go-ontology/blob/master/.github/workflows/claude-code-review.yml) | Runs the [ai4c-reviewer](../reference/bots/ai4c-reviewer.md) bot on pull requests | +| `.github/ai-controllers.json` | Who is allowed to trigger the agent | | [`CLAUDE.md`](https://github.com/geneontology/go-ontology/blob/master/CLAUDE.md) | Repository instructions | The ten skills are `chemical-entity`, `design-pattern`, `external-term-lookup`, @@ -34,6 +35,24 @@ because the detail lives in the skill that needs it. See commands so the agent runs the same build steps an editor runs, instead of inventing its own. +### The workflows explain themselves + +GO's two agent workflows carry long header comments saying what each choice is +for: why the review skips `ontobot` pull requests but not agent-authored ones, +why the app id and the bot user id are different numbers, why the review rubric +lives in a skill rather than the prompt. If you are setting up your own +[bots](../reference/bots/index.md), read these two files before anything else on +this site. + +### Two bots, not one + +GO runs [ai4c-agent](../reference/bots/ai4c-agent.md) for editing and +[ai4c-reviewer](../reference/bots/ai4c-reviewer.md) for review. The split exists +because GitHub does not let an identity approve its own pull request. GO +deliberately reviews the editing bot's own pull requests, which its workflow +calls the highest-value case, since that is where fabricated identifiers get +caught. + ## What to copy first Copy the idea, not the files. List the five tickets your editors handle most diff --git a/docs/case-studies/mondo.md b/docs/case-studies/mondo.md index 4ca63f7..75b954e 100644 --- a/docs/case-studies/mondo.md +++ b/docs/case-studies/mondo.md @@ -47,7 +47,7 @@ Claude Code sessions only. Mondo is also the last repository we track that still runs the `dragon-ai-agent` machine account. It authenticates with a personal access token held in the `PAT_FOR_PR` secret. DisMech has moved the same job to a GitHub App -with a token minted per run. See [Bots](../reference/bots.md). +with a token minted per run. See [Bots](../reference/bots/index.md). This is not necessarily wrong, but it is undocumented. If you have both, say in `CLAUDE.md` which surface each one serves. If you want the subagents reachable diff --git a/docs/glossary.md b/docs/glossary.md index be9f34b..fd53d70 100644 --- a/docs/glossary.md +++ b/docs/glossary.md @@ -159,7 +159,7 @@ not stay in the main session. See The GitHub identity that agent work appears under: the name in the byline on an issue, commit, or pull request. Distinct from the agent, which is the software doing the work. A bot is either a GitHub App or a machine account. See -[Bots](reference/bots.md). +[Bots](reference/bots/index.md). ## GitHub App diff --git a/docs/patterns/guard-untrusted-input.md b/docs/patterns/guard-untrusted-input.md index 81792d3..c4bdbc0 100644 --- a/docs/patterns/guard-untrusted-input.md +++ b/docs/patterns/guard-untrusted-input.md @@ -33,7 +33,7 @@ credentials are the control. **Pick the identity deliberately.** What a bot may do is set by its permissions, not by what you told it. A reviewer bot needs to read code and write reviews; it -does not need to push branches. See [Bots](../reference/bots.md). +does not need to push branches. See [Bots](../reference/bots/index.md). ## Separate reading from writing diff --git a/docs/reference/bots.md b/docs/reference/bots.md deleted file mode 100644 index c7c7927..0000000 --- a/docs/reference/bots.md +++ /dev/null @@ -1,141 +0,0 @@ -# Bots - -A **bot** is the GitHub identity that agent work appears under. It is the name in -the byline on an issue, a comment, a commit, or a pull request. - -Keep this separate from the agent itself: - -* The **agent** is the software that does the work. See - [Harnesses](harnesses.md). -* The **bot** is the identity it works under. - -The same agent can work under different bots in different repositories. Claude -Code runs as `claude[bot]` in one repository and as `ai4c-agent[bot]` in -another. The agent is the same. The permissions are not. - -## The bots you will meet - -| Bot | Kind | What it does | Repositories | -| --- | --- | --- | --- | -| [`ai4c-agent`](https://github.com/apps/ai4c-agent) | GitHub App | Does curation work: responds to mentions, moves stalled pull requests forward | DisMech, AI Gene Review | -| [`ai4c-reviewer`](https://github.com/apps/ai4c-reviewer) | GitHub App | Reviews pull requests and supplies the approving review | DisMech, AI Gene Review | -| [`claude`](https://github.com/apps/claude) | GitHub App | Runs Claude Code from issues and pull requests | Many | -| `copilot-swe-agent` | GitHub App | Works an issue and opens a draft pull request | EFO, Uberon | -| `github-actions` | GitHub App | The built-in identity for workflow runs | Everywhere | -| `dragon-ai-agent` | Machine account | Older curation bot, now semi-retired | Mondo | - -## The two kinds - -### GitHub Apps - -Most of the list above. An App is installed into a repository or an -organization. A workflow mints a short-lived token for it at the start of each -run, and the token expires when the run ends. Nothing long-lived sits in your -secrets except the App's private key. - -You can tell an App by its byline. GitHub appends `[bot]`, so you see -`ai4c-agent[bot]`, not `ai4c-agent`. - -Apps do not use a seat in your organization, and you grant each one only the -permissions it needs. - -### Machine accounts - -A machine account is an ordinary GitHub user account that a program logs in as. -GitHub calls these machine users. `dragon-ai-agent` is one. - -A machine account looks like a person in the interface. There is no `[bot]` -suffix, it can join teams, and it uses a seat. It authenticates with a personal -access token, which is long-lived and sits in your repository secrets. - -Prefer an App for anything new. Shorter-lived credentials and per-repository -permissions are both worth having. - -## Why a project runs more than one bot - -DisMech runs two of its own, and the reason is worth understanding before you -copy the setup. - -**GitHub does not let an identity approve its own pull request.** - -DisMech requires one approving review before a pull request can merge. If the -bot that opened the pull request were also the bot that reviewed it, the -approval would not count, and every agent pull request would need a human -approval. - -So the work is split: - -* `ai4c-agent` writes. It responds to mentions and pushes branches. -* `ai4c-reviewer` reviews. It applies the review rubric and supplies the - approving review that branch protection requires. - -Two identities, one boundary. This is what makes it possible for the project to -say that contributors are not expected to check their own agent's output. See -[One review checklist](../patterns/review-checklists.md). - -## The dragon-ai-agent case - -`@dragon-ai-agent` means two different things depending on which repository you -are in. This trips people up, so it is worth stating plainly. - -**In [Mondo](../case-studies/mondo.md), the machine account is live.** -`.github/workflows/ai-agent.yml` runs `dragon-ai-agent/run-goose-obo` and -authenticates with a personal access token stored as `PAT_FOR_PR`. The account -is doing the work. - -**In [DisMech](../case-studies/dismech.md), the account is retired.** The -workflow says so in a comment. What survives is the *phrase*: writing -`@dragon-ai-agent please ...` in an issue is a trigger keyword that starts a -workflow. The work is then done by the `ai4c-agent` App under a short-lived -token. - -If you are reading an old issue thread, the mention tells you a workflow ran. It -does not tell you which identity did the work. Check the byline on the resulting -pull request. - -## Make commits attribute correctly - -An App token alone does not set the commit author. If you do not configure it, -commits land under whatever Git identity the runner happens to have. - -DisMech sets it explicitly in `pr-shepherd.yml`, using the address form GitHub -recognizes for App commits: - -```bash -APP_USER_ID="$(gh api "/users/${APP_SLUG}[bot]" --jq .id)" -git config --global user.name "${APP_SLUG}[bot]" -git config --global user.email "${APP_USER_ID}+${APP_SLUG}[bot]@users.noreply.github.com" -``` - -Do this. Without it you cannot tell from `git log` which bot wrote a line, and -tools that measure agent contribution have nothing to work from. See -[Evidence](../evidence.md). - -## Restrict who can summon a bot - -A bot with write access, triggered by a comment, is reachable by anyone who can -comment. Both parts need controlling: - -* **Who may trigger it.** DisMech checks the commenter against a list of - authorized controllers before the agent runs. -* **What it may do once triggered.** This is the App's permission set, not a - matter of instructions. Grant each App the narrowest permissions that let it - do its job. - -Separate identities help here too. A reviewer bot needs to read code and write -reviews. It does not need to push branches. - -See [Guard the untrusted surface](../patterns/guard-untrusted-input.md). - -## Working out who did what - -| You see | It means | -| --- | --- | -| `name[bot]` | A GitHub App | -| A plain login | A person, or a machine account | -| `github-actions[bot]` | A workflow, using the built-in token | -| `@dragon-ai-agent` in comment text | A trigger keyword, not necessarily that account | - -A byline tells you the identity. It does not tell you the model, the harness, or -who asked. For that you need the execution trace. See -[Evidence](../evidence.md). diff --git a/docs/reference/bots/ai4c-agent.md b/docs/reference/bots/ai4c-agent.md new file mode 100644 index 0000000..8a22430 --- /dev/null +++ b/docs/reference/bots/ai4c-agent.md @@ -0,0 +1,78 @@ +# ai4c-agent + +The curation bot. It responds to trigger keywords, edits files, pushes branches, +and opens pull requests. + +**Kind**: GitHub App +**Install page**: [github.com/apps/ai4c-agent](https://github.com/apps/ai4c-agent) +**Byline**: `ai4c-agent[bot]` +**Used in**: [GO](../../case-studies/go-ontology.md), +[DisMech](../../case-studies/dismech.md), +[AI Gene Review](../../case-studies/ai-gene-review.md) + +## What it does + +It is the identity that does the work, as opposed to the one that reviews it. +Across the repositories that run it, that means: + +* Responding when an authorized user writes `@ai4c-agent please ...` in an issue + or a pull request comment. +* Moving stalled pull requests forward. +* Editing curated files, pushing branches, and opening pull requests. + +The agent behind it is Claude Code, run through +[`anthropics/claude-code-action`](https://github.com/anthropics/claude-code-action). +The App is the identity, not the agent. + +## How it is wired + +Each job mints its own installation token: + +```yaml +- name: Generate ai4c-agent token + id: ai4c-token + uses: actions/create-github-app-token@v2 + with: + app-id: ${{ secrets.AI4C_AGENT_APP_ID }} + private-key: ${{ secrets.AI4C_AGENT_PRIVATE_KEY }} +``` + +The token lasts for the run. The only long-lived secret is the private key. + +**Secrets you need**: `AI4C_AGENT_APP_ID`, `AI4C_AGENT_PRIVATE_KEY`, and a model +credential, either `ANTHROPIC_API_KEY` or `CLAUDE_CODE_OAUTH_TOKEN`. + +**Where to read it**: +[GO's `ai-agent.yml`](https://github.com/geneontology/go-ontology/blob/master/.github/workflows/ai-agent.yml) +is the best-documented version. Its header comments explain each choice. + +## Triggering it + +`@ai4c-agent` is a keyword, not a mention. Apps cannot be @-mentioned, so the +workflow matches the string in the comment body and checks the author against a +list of authorized users. GO keeps that list in `.github/ai-controllers.json`. + +GO also honours `@dragon-ai-agent` as a legacy keyword while curators move +across. See [dragon-ai-agent](dragon-ai-agent.md). + +## Commit attribution + +GO sets the commit identity from the bot account's numeric user id: + +```yaml +AGENT_LOGIN: ai4c-agent[bot] +AGENT_USER_ID: 242316268 +``` + +That number is the `ai4c-agent[bot]` **account** id, not the app id in +`AI4C_AGENT_APP_ID`. They are different numbers, and using the wrong one +silently unlinks the commits from the App. + +## Notes + +* Pair it with [ai4c-reviewer](ai4c-reviewer.md). A single identity cannot both + open a pull request and supply its approving review. +* List it in your review action's `allowed_bots`, or its pull requests are + skipped as bot noise. +* It cannot run on pull requests from forks, because GitHub withholds secrets + from fork workflows. diff --git a/docs/reference/bots/ai4c-reviewer.md b/docs/reference/bots/ai4c-reviewer.md new file mode 100644 index 0000000..25c971c --- /dev/null +++ b/docs/reference/bots/ai4c-reviewer.md @@ -0,0 +1,74 @@ +# ai4c-reviewer + +The review bot. It reads pull requests, applies the project's review rubric, and +supplies the approving review that branch protection requires. + +**Kind**: GitHub App +**Install page**: [github.com/apps/ai4c-reviewer](https://github.com/apps/ai4c-reviewer) +**Byline**: `ai4c-reviewer[bot]` +**Used in**: [GO](../../case-studies/go-ontology.md), +[DisMech](../../case-studies/dismech.md), +[AI Gene Review](../../case-studies/ai-gene-review.md) + +## Why it is separate from ai4c-agent + +GitHub does not let an identity approve its own pull request. If the bot that +opened a pull request also reviewed it, the approval would not count. + +Splitting the two means an agent pull request can reach a mergeable state +without a human approval, which is what lets these projects treat review as a +process rather than a personal duty. See +[One review checklist](../../patterns/review-checklists.md). + +## How it is wired + +```yaml +- name: Generate ai4c-reviewer token + id: reviewer-token + uses: actions/create-github-app-token@v2 + with: + app-id: ${{ secrets.AI4C_REVIEWER_APP_ID }} + private-key: ${{ secrets.AI4C_REVIEWER_PRIVATE_KEY }} +``` + +**Secrets you need**: `AI4C_REVIEWER_APP_ID`, `AI4C_REVIEWER_PRIVATE_KEY`, and a +model credential. + +**Where to read it**: +[GO's `claude-code-review.yml`](https://github.com/geneontology/go-ontology/blob/master/.github/workflows/claude-code-review.yml). + +## The rubric lives outside the workflow + +GO keeps the review criteria in the `pr-review` skill, not in the workflow +prompt. Its comment gives the reason: + +> The review substance lives in the `pr-review` skill +> (.claude/skills/pr-review/SKILL.md), not in the prompt below, so the same +> criteria apply when a human reviews a PR locally. + +The workflow prompt carries only what is specific to running in CI: what is +installed on the runner, and how to submit the verdict. Copy this split. It +keeps one standard for humans and bots. + +## What it reviews, and what it skips + +GO reviews every same-repo pull request except those from `ontobot`, whose +automated refresh jobs produce large mechanical diffs. It deliberately does +review pull requests opened by `ai4c-agent[bot]`: + +> self-review of agent work is the highest-value case, since it catches +> hallucinated PMIDs and bogus axioms before a curator spends time on them. + +For that to work, the editing bot must be listed in `allowed_bots`. Review +actions skip bot-authored pull requests by default. + +## Notes + +* Fork pull requests cannot be reviewed automatically, because secrets are not + available. GO's documented route for those is `workflow_dispatch`, or a + `/review` comment. +* It finds its own earlier comments by matching + `.user.login == "ai4c-reviewer[bot]"`, which is another reason to get commit + and comment identity right. +* Give it read access to code and write access to reviews. It does not need to + push branches. diff --git a/docs/reference/bots/claude.md b/docs/reference/bots/claude.md new file mode 100644 index 0000000..ad6dacc --- /dev/null +++ b/docs/reference/bots/claude.md @@ -0,0 +1,53 @@ +# claude + +Anthropic's own GitHub App. It runs Claude Code from issues and pull requests. + +**Kind**: GitHub App +**Install page**: [github.com/apps/claude](https://github.com/apps/claude) +**Owner**: Anthropic +**Byline**: `claude[bot]` +**Used in**: many repositories, including this one + +## What it does + +It is the default identity when you add Claude Code to a repository with +`claude install-github-app`. Work appears under `claude[bot]`: review comments, +pull requests, and replies to mentions. + +## How it is wired + +You can authenticate the action two ways. + +**With an App token minted from OIDC.** The workflow passes no `github_token` +and sets `id-token: write`, and the action obtains a Claude App token itself. +DisMech's `claude.yml` does this. Note that job-level `permissions:` then scopes +the unused `GITHUB_TOKEN`, not the App token, so anything the agent needs to +read has to be granted through `additional_permissions`. + +**With your own App token.** Mint a token for a different App and pass it as +`github_token`. This is how GO and DisMech run the review workflow under +[ai4c-reviewer](ai4c-reviewer.md) instead of `claude`. + +The model credential is separate from the identity: either `ANTHROPIC_API_KEY` +or `CLAUDE_CODE_OAUTH_TOKEN`. + +## When to use your own App instead + +Use `claude` when you want the quickest working setup. + +Use your own App when you want: + +* **A name that says what it does.** `ai4c-reviewer[bot]` in a byline is clearer + than `claude[bot]` when several different jobs all run Claude Code. +* **Two identities.** You need a second App if one bot opens pull requests and + another approves them. +* **Your own permission set,** scoped per repository. + +## Notes + +* `CLAUDE_CODE_OAUTH_TOKEN` expires. When it does, every run fails with + `API Error: 401 ... OAuth access token has expired` before reaching the model, + and the job posts "Claude encountered an error" with no detail. Check the job + log rather than the comment. Refresh with `claude setup-token`. +* Bot-authored pull requests are skipped by review actions unless listed in + `allowed_bots`. diff --git a/docs/reference/bots/copilot.md b/docs/reference/bots/copilot.md new file mode 100644 index 0000000..f7f2abd --- /dev/null +++ b/docs/reference/bots/copilot.md @@ -0,0 +1,58 @@ +# copilot-swe-agent + +GitHub's own coding agent. You assign it an issue, and it opens a draft pull +request and works in the background. + +**Kind**: GitHub App +**Byline**: `copilot-swe-agent[bot]` +**Used in**: [EFO](../../case-studies/efo.md), +[Uberon](../../case-studies/uberon.md) + +## What it does + +Assign an issue to Copilot and it works in its own environment, pushing commits +to a draft pull request as it goes. All those commits are authored by +`copilot-swe-agent[bot]`, not by the person who delegated the task. + +It is the only bot here that needs no workflow of your own to run the agent. +The agent runs on GitHub's infrastructure. + +## How it is wired + +Add `.github/workflows/copilot-setup-steps.yml`. The job must be named +`copilot-setup-steps` or GitHub will not pick it up: + +```yaml +jobs: + copilot-setup-steps: + runs-on: ubuntu-latest + permissions: + contents: read +``` + +The job installs whatever the agent needs before it starts. Copilot is given its +own token for its own operations, and it checks out the repository for you if +you do not. + +Instructions go in `.github/copilot-instructions.md`. Keep one authoritative +instructions file and make the others pointers. See +[One source of instructions](../../patterns/one-source-of-instructions.md). + +Uberon's [pull request 3580](https://github.com/obophenotype/uberon/pull/3580) +shows the two files to change. + +## The firewall + +Copilot's agent runs behind a firewall that blocks most external hosts by +default. For ontology work this bites immediately: `purl.obolibrary.org` is +blocked, so PURL-based imports fail. + +Allowlist the hosts you need in the repository's Copilot settings. See +[GitHub Copilot](../github-copilot.md) for the steps and the workarounds. + +## Notes + +* Naming differs by context. The login is `copilot-swe-agent`; API calls that + assign an issue need the exact string `copilot-swe-agent[bot]`. +* It is a good fit for work that starts from a well-specified issue. It is a + poorer fit for exploratory curation, where you want to steer as you go. diff --git a/docs/reference/bots/dragon-ai-agent.md b/docs/reference/bots/dragon-ai-agent.md new file mode 100644 index 0000000..fd95f97 --- /dev/null +++ b/docs/reference/bots/dragon-ai-agent.md @@ -0,0 +1,62 @@ +# dragon-ai-agent + +The original OBO curation bot. It is being retired in favour of +[ai4c-agent](ai4c-agent.md), and `@dragon-ai-agent` now means different things +in different repositories. + +**Kind**: Machine account +**Byline**: `dragon-ai-agent`, with no `[bot]` suffix +**Still live in**: [Mondo](../../case-studies/mondo.md) +**Retired in**: [GO](../../case-studies/go-ontology.md), +[DisMech](../../case-studies/dismech.md) + +## What it was + +A machine account that ran [DRAGON-AI](https://pubmed.ncbi.nlm.nih.gov/39415214/) +against OBO repositories. A curator wrote `@dragon-ai-agent please ...` on an +issue, a workflow picked it up, and the agent produced a pull request. + +Because it is an ordinary user account, it can be @-mentioned like a person, it +shows no `[bot]` suffix, it uses a seat, and it authenticates with a personal +access token stored in repository secrets. + +## Where it stands now + +**In Mondo, the account is live.** `.github/workflows/ai-agent.yml` runs +`dragon-ai-agent/run-goose-obo` and authenticates with a personal access token +held as `PAT_FOR_PR`. The agent is Goose. The account does the work. + +**In DisMech, the account is retired.** The workflow says so: + +> Trigger is a `@dragon-ai-agent please …` mention from an authorized controller +> (a text keyword — the dragon-ai-agent machine account is retired). + +The phrase survives as a trigger keyword. The work is done by `ai4c-agent`. + +**In GO, the keyword is deprecated but still honoured.** The workflow lists it +as a legacy handle, to be dropped once curators have moved to `@ai4c-agent`. + +## Why this matters when reading old threads + +A `@dragon-ai-agent` mention tells you a workflow ran. It does not tell you +which identity did the work, and the answer differs by repository and by date. + +Check the byline on the resulting pull request. If it ends in `[bot]`, an App +did the work. + +## Migrating away from it + +The move is from a machine account with a long-lived token to a GitHub App with +a token minted per run. Steps: + +1. Install [ai4c-agent](ai4c-agent.md), or your own App, and add its app id and + private key as secrets. +2. Change the workflow to mint an installation token with + `actions/create-github-app-token`. +3. Set the commit identity to the bot account, using its numeric user id. +4. Keep the old keyword as a legacy trigger while curators adjust. GO does this + with a separate `AGENT_MENTION_LEGACY` variable. +5. Once nobody uses the old keyword, drop it and remove the account's token. + +You gain shorter-lived credentials, per-repository permissions, and a freed +seat. diff --git a/docs/reference/bots/index.md b/docs/reference/bots/index.md new file mode 100644 index 0000000..4454d10 --- /dev/null +++ b/docs/reference/bots/index.md @@ -0,0 +1,165 @@ +# Bots + +A **bot** is the GitHub identity that agent work appears under. It is the name +in the byline on an issue, a comment, a commit, or a pull request. + +Keep it separate from the agent: + +* The **agent** is the software that does the work. See + [Harnesses](../harnesses.md). +* The **bot** is the identity it works under. + +The same agent can work under different bots in different repositories. Claude +Code runs as `claude[bot]` in one repository and as `ai4c-agent[bot]` in +another. The agent is the same. The permissions are not. + +## Bots in use + +| Bot | Kind | Job | Repositories | +| --- | --- | --- | --- | +| [`ai4c-agent`](ai4c-agent.md) | GitHub App | Does curation work | GO, DisMech, AI Gene Review | +| [`ai4c-reviewer`](ai4c-reviewer.md) | GitHub App | Reviews pull requests | GO, DisMech, AI Gene Review | +| [`claude`](claude.md) | GitHub App | Runs Claude Code from issues and pull requests | Many | +| [`copilot-swe-agent`](copilot.md) | GitHub App | Works an issue, opens a draft pull request | EFO, Uberon | +| [`dragon-ai-agent`](dragon-ai-agent.md) | Machine account | Older curation bot, being retired | Mondo | + +Two more appear in bylines but are not curation bots. `github-actions[bot]` is +the built-in identity for workflow runs, so it authors anything a workflow +commits. `ontobot` runs GO's automated refresh jobs, which produce large +mechanical diffs. + +## The two kinds + +### GitHub Apps + +An App is installed into a repository or an organization. A workflow mints a +token for it at the start of each run, and that token expires when the run +ends. Nothing long-lived sits in your secrets except the App's private key. + +GitHub appends `[bot]` to an App's byline, so you see `ai4c-agent[bot]`. + +Apps do not use a seat in your organization, and you grant each one only the +permissions it needs. + +### Machine accounts + +A machine account is an ordinary GitHub user account that a program logs in as. +GitHub calls these machine users. + +A machine account looks like a person. There is no `[bot]` suffix, it can join +teams, it uses a seat, and it authenticates with a personal access token that is +long-lived and sits in your repository secrets. + +Prefer an App for anything new. + +## You cannot mention an App + +This surprises people, so it is worth stating plainly. + +**GitHub Apps cannot be @-mentioned.** Typing `@ai4c-agent` in a comment does +nothing on its own. It is plain text. GO's workflow says so directly: + +> The handle curators type to invoke the agent. Note this is NOT a real GitHub +> account: the app's login is AGENT_LOGIN (`[bot]`), and apps cannot be +> @-mentioned, so "@ai4c-agent" renders as plain text. + +So every "mention" of a bot is really a **trigger keyword**. A workflow watches +comment text, matches the string, checks that you are authorized, and starts a +run. The App is what the run then authenticates as. + +Two consequences: + +* You can change the keyword without changing the identity. GO honours both + `@ai4c-agent` and the older `@dragon-ai-agent` while curators move across. +* A keyword in an old issue thread tells you a workflow ran. It does not tell + you which identity did the work. Check the byline on the resulting pull + request. + +## Why a project runs two of its own + +GO, DisMech, and AI Gene Review each run both `ai4c-agent` and `ai4c-reviewer`. +The reason matters before you copy the setup. + +**GitHub does not let an identity approve its own pull request.** + +These projects require an approving review before a pull request merges. If the +bot that opened the pull request were also the bot that reviewed it, the +approval would not count, and every agent pull request would need a human. + +So the work is split: + +* `ai4c-agent` writes. It responds to keywords and pushes branches. +* `ai4c-reviewer` reviews. It applies the review rubric and supplies the + approving review. + +Two identities, one boundary. This is what lets a project say that contributors +are not expected to check their own agent's output. See +[One review checklist](../../patterns/review-checklists.md). + +## Make commits attribute correctly + +An App token does not set the commit author. Configure it, or commits land +under whatever Git identity the runner happens to have. + +The address form GitHub recognises for App commits needs the bot account's +numeric user id: + +```bash +APP_USER_ID="$(gh api "/users/${APP_SLUG}[bot]" --jq .id)" +git config --global user.name "${APP_SLUG}[bot]" +git config --global user.email "${APP_USER_ID}+${APP_SLUG}[bot]@users.noreply.github.com" +``` + +!!! warning "The app id is not the user id" + + An App has two different numbers. The **app id** identifies the App and goes + with the private key when you mint a token. The **user id** identifies the + `name[bot]` account and is what commit attribution needs. GO's workflow + warns about this in a comment, because using the wrong one silently breaks + the link between commits and the App. + +Get the user id with `gh api /users/ai4c-agent%5Bbot%5D --jq .id`. + +## Let your reviewer see bot pull requests + +Review actions ignore bot-authored pull requests by default, which is exactly +backwards when the bot is your editing agent. GO has to list it explicitly: + +```yaml +allowed_bots: 'claude,github-actions,ai4c-agent' +``` + +Its comment explains why: + +> ai4c-agent must be listed, or PRs opened by the editing agent -- the main +> reason this workflow exists -- would be ignored as bot noise. + +Reviewing agent work is the highest-value case. It catches fabricated +identifiers before a curator spends time on them. + +## Restrict who can trigger a bot + +A bot with write access, started by a comment, is reachable by anyone who can +comment. Control both halves: + +* **Who may trigger it.** GO keeps a list of authorized usernames in + `.github/ai-controllers.json` and checks the commenter against it before the + agent runs. +* **What it may do.** This is the App's permission set, not a matter of + instructions. Grant each App the narrowest permissions for its job. A reviewer + needs to read code and write reviews. It does not need to push branches. + +See [Guard the untrusted surface](../../patterns/guard-untrusted-input.md). + +## Reading a byline + +| You see | It means | +| --- | --- | +| `name[bot]` | A GitHub App | +| A plain login | A person, or a machine account | +| `github-actions[bot]` | A workflow, using the built-in token | +| `@name` in comment text | A trigger keyword, not necessarily that identity | + +A byline tells you the identity. It does not tell you the model, the harness, or +who asked. For that you need the execution trace. See +[Evidence](../../evidence.md). diff --git a/docs/reference/github-integrations.md b/docs/reference/github-integrations.md index ca82e0f..d9ed14b 100644 --- a/docs/reference/github-integrations.md +++ b/docs/reference/github-integrations.md @@ -14,7 +14,7 @@ This guide documents the different approaches for integrating AI agents with Git The Dragon-AI Agent approach uses custom GitHub Actions to deploy headless AI coding assistants (Claude Code or Goose) in response to issue/PR comments. New repositories should use a Claude Code Action workflow instead. See [Set up GitHub Actions](../how-tos/set-up-github-actions.md). -The `dragon-ai-agent` machine account itself is semi-retired. Mondo still runs it; DisMech keeps the mention as a trigger keyword only. For which identity actually does the work, see [Bots](bots.md). +The `dragon-ai-agent` machine account itself is semi-retired. Mondo still runs it; DisMech keeps the mention as a trigger keyword only. For which identity actually does the work, see [Bots](bots/index.md). ### How It Works diff --git a/docs/reference/harnesses.md b/docs/reference/harnesses.md index 0aff317..bde1691 100644 --- a/docs/reference/harnesses.md +++ b/docs/reference/harnesses.md @@ -63,13 +63,13 @@ OBO repositories. Semi-retired: [Mondo](../case-studies/mondo.md) still runs the machine account, while [DisMech](../case-studies/dismech.md) has retired it and keeps only the mention as a trigger keyword, with the work done by the `ai4c-agent` App. New repositories should use a Claude Code Action workflow -instead. See [Bots](bots.md) and +instead. See [Bots](bots/index.md) and [GitHub integrations](github-integrations.md). ## The agent is not the identity The harness is the software. The identity the work appears under is a separate -choice, and it is the one that controls permissions. See [Bots](bots.md). +choice, and it is the one that controls permissions. See [Bots](bots/index.md). ## Give every session the same tools diff --git a/mkdocs.yml b/mkdocs.yml index dd1c3db..b592e17 100644 --- a/mkdocs.yml +++ b/mkdocs.yml @@ -75,7 +75,13 @@ nav: - Training materials for curators: tutorials/tutorials-for-curators.md - Reference: - Harnesses: reference/harnesses.md - - Bots: reference/bots.md + - Bots: + - Overview: reference/bots/index.md + - ai4c-agent: reference/bots/ai4c-agent.md + - ai4c-reviewer: reference/bots/ai4c-reviewer.md + - claude: reference/bots/claude.md + - copilot-swe-agent: reference/bots/copilot.md + - dragon-ai-agent: reference/bots/dragon-ai-agent.md - Agentic tools: reference/agentic-tools.md - GitHub integrations: reference/github-integrations.md - GitHub Copilot: reference/github-copilot.md From 0b8683d55f07286b9e8eba5f86a6d7f8ec19e8b9 Mon Sep 17 00:00:00 2001 From: Claude Date: Wed, 19 Aug 2026 00:33:42 +0000 Subject: [PATCH 5/5] ci: accept HTTP 203 in the link check The link check failed on one link out of 506: the DRAGON-AI paper at pubmed.ncbi.nlm.nih.gov/39415214/. PubMed answers automated clients with 203 Non-Authoritative Information rather than 200, so lychee rejected it. Verified the link is correct through the NCBI esummary API: PMID 39415214 is Toro S et al., "Dynamic Retrieval Augmented Generation of Ontologies using Artificial Intelligence (DRAGON-AI)", J Biomed Semantics 15, 2024. The fix belongs in the checker, not the citation. 203 is a success code, so accept it generally rather than excluding PubMed. Any future PubMed citation on this site would hit the same thing. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_017tKFksxqHZJNLoH1zJ5WXm --- .github/workflows/docs-checks.yml | 5 ++++- 1 file changed, 4 insertions(+), 1 deletion(-) diff --git a/.github/workflows/docs-checks.yml b/.github/workflows/docs-checks.yml index 29e072f..01ddf58 100644 --- a/.github/workflows/docs-checks.yml +++ b/.github/workflows/docs-checks.yml @@ -35,10 +35,13 @@ jobs: - name: Check links in Markdown uses: lycheeverse/lychee-action@v2 with: + # 203: PubMed answers automated clients with "Non-Authoritative + # Information" rather than 200. The links resolve fine in a browser. + # 429: rate limiting, not a broken link. args: >- --no-progress --max-retries 2 - --accept 200,206,429 + --accept 200,203,206,429 --exclude-path docs/overrides --exclude 'github\.com/ai4curation/agent-watcher' --exclude 'claude\.ai'