Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion .github/workflows/deploy_documentation.yml
Original file line number Diff line number Diff line change
Expand Up @@ -35,7 +35,7 @@ jobs:
run: uv sync --all-extras

- name: Build with MkDocs
run: uv run mkdocs build
run: uv run mkdocs build --strict

- name: Upload artifact
uses: actions/upload-pages-artifact@v3
Expand Down
51 changes: 51 additions & 0 deletions .github/workflows/docs-checks.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,51 @@
name: Docs checks

on:
pull_request:
push:
branches:
- main
schedule:
# Weekly, to catch links that rot in the repos this site points at.
- cron: "0 6 * * 1"
workflow_dispatch:

permissions:
contents: read

jobs:
build:
name: Build site
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: astral-sh/setup-uv@v5
- uses: actions/setup-python@v5
with:
python-version: "3.12"
- run: uv sync --all-extras
- name: Build with MkDocs
run: uv run mkdocs build --strict

links:
name: Check links
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- name: Check links in Markdown
uses: lycheeverse/lychee-action@v2
with:
# 203: PubMed answers automated clients with "Non-Authoritative
# Information" rather than 200. The links resolve fine in a browser.
# 429: rate limiting, not a broken link.
args: >-
--no-progress
--max-retries 2
--accept 200,203,206,429
--exclude-path docs/overrides
--exclude 'github\.com/ai4curation/agent-watcher'
--exclude 'claude\.ai'
--exclude 'learn\.deeplearning\.ai'
'docs/**/*.md'
'README.md'
fail: true
25 changes: 25 additions & 0 deletions AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -68,3 +68,28 @@ curation workflows.
- Follow existing documentation structure instead of creating new sections
unnecessarily.
- Emphasize how AI improves existing workflows rather than replacing curators.

## Site structure

The site has two sections that carry most of its value:

- `docs/case-studies/`: one page per repository that runs agents on real
curation. Each page follows the same shape: what to look at, what works, what
to copy first, and gaps.
- `docs/patterns/`: practices that appear in more than one repository. Each page
names at least two repositories that use it.

Do not centralize material that belongs in another repository. Link to the file
or folder instead. If a page disagrees with the repository it describes, the
repository is right.

Counts of skills, subagents, and workflows come from
[agent-watcher](https://github.com/ai4curation/agent-watcher). Cite the dated
report you took them from, as `docs/evidence.md` does.

## Checks

Run `uv run mkdocs build --strict` before committing. The `docs-checks.yml`
workflow runs the same build and a link check on every pull request, and the
link check also runs weekly to catch links that rot in the repositories this
site points at.
95 changes: 7 additions & 88 deletions CLAUDE.md
Original file line number Diff line number Diff line change
@@ -1,90 +1,9 @@
# CLAUDE.md for aidocs
# CLAUDE.md

This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.
Read [AGENTS.md](AGENTS.md). It is the authoritative guidance for all agents
working in this repository, including Claude Code.

## Project Overview

**aidocs** is a documentation repository for **AI4Curators**, providing practical guides for curators and maintainers of knowledge bases to integrate AI into their workflows. The project focuses on immediate, actionable integration strategies rather than theoretical discussions.

### Core Mission
- Help curators integrate AI agents into existing GitHub-based workflows
- Provide plugins and tools for existing chat UIs
- Support ontology editing and curation workflows with AI assistance

## Repository Structure

```
aidocs/
├── docs/ # MkDocs documentation source
│ ├── how-tos/ # Practical how-to guides
│ ├── tutorials/ # Step-by-step tutorials
│ ├── reference/ # Technical reference materials
│ │ └── clients/ # Documentation for various AI clients
│ └── overrides/ # MkDocs theme customizations
├── src/aidocs/ # Python package source
├── tools/ # Utility scripts and tools
├── mkdocs.yml # MkDocs configuration
└── pyproject.toml # Python project configuration
```

## Technology Stack

- **Documentation**: MkDocs with Material theme
- **Python**: >=3.9 (for tooling and scripts)
- **Deployment**: GitHub Pages
- **Build System**: Hatchling

## Key Files and Their Purpose

- `mkdocs.yml`: Configures the documentation site structure, theme, and plugins
- `pyproject.toml`: Python project metadata and dependencies
- `docs/index.md`: Main landing page content
- `docs/glossary.md`: Terminology definitions for the domain
- `docs/how-tos/`: Practical implementation guides
- `docs/reference/clients/`: Documentation for various AI client applications

## Development Guidelines

### Documentation Standards
- Focus on **practical, immediately actionable** content
- Provide step-by-step guides over theoretical explanations
- Include real-world examples and use cases
- Maintain consistency with the existing Material theme styling

### Content Categories
1. **How-tos**: Task-oriented guides for specific implementation scenarios
2. **Tutorials**: Comprehensive learning-oriented walkthroughs
3. **Reference**: Technical specifications and API documentation
4. **Glossary**: Domain-specific terminology and definitions

### Building and Testing
- Use `mkdocs serve` for local development
- Documentation is automatically deployed to GitHub Pages
- Test all links and examples before committing

## Target Audience

- **Primary**: Curators and maintainers of knowledge bases and ontologies
- **Secondary**: Developers integrating AI into existing curation workflows
- **Focus**: Practitioners who need immediate, working solutions over academic discussions

## AI Agent Integration

This repository serves as both documentation and a practical example of AI agent integration:
- GitHub agents can directly contribute to documentation
- Examples demonstrate real-world AI-assisted curation workflows
- Reference materials guide implementation of similar systems

## Contributing Guidelines

When working on this repository:
1. **Prioritize practical value** - every guide should be immediately applicable
2. **Test all examples** - ensure code snippets and procedures work as documented
3. **Maintain consistency** - follow existing patterns in documentation structure
4. **Focus on integration** - emphasize how AI enhances existing workflows rather than replacing them

# important-instruction-reminders
Do what has been asked; nothing more, nothing less.
NEVER create files unless they're absolutely necessary for achieving your goal.
ALWAYS prefer editing an existing file to creating a new one.
NEVER proactively create documentation files (*.md) or README files. Only create documentation files if explicitly requested by the User.
Do not add instructions here. Keeping one file avoids the drift that happens
when `CLAUDE.md`, `AGENTS.md`, and `.github/copilot-instructions.md` are
maintained separately. See
[One source of instructions](docs/patterns/one-source-of-instructions.md).
4 changes: 2 additions & 2 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,9 +4,9 @@ This repository contains documentation about using AI to assist with curation, p

## Documentation

📖 **[Visit the full documentation](https://ai4curation.github.io/aidocs/)**
📖 **[Visit the full documentation](https://ai4curation.io/aidocs/)**

The complete guides, tutorials, and reference materials are available on our [GitHub Pages site](https://ai4curation.github.io/aidocs/).
The complete guides, tutorials, and reference materials are available on our [GitHub Pages site](https://ai4curation.io/aidocs/).

## About

Expand Down
71 changes: 71 additions & 0 deletions docs/case-studies/ai-gene-review.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,71 @@
# AI Gene Review

A review of existing Gene Ontology annotations, one file per gene, with a public
voting interface for feedback.

**Repository**: [ai4curation/ai-gene-review](https://github.com/ai4curation/ai-gene-review)
**App**: [Browse gene reviews](https://ai4curation.io/ai-gene-review/app/index.html)
**Product**: `genes/<organism>/<GENE>/<GENE>-ai-review.yaml`
**Agent surface**: local sessions, GitHub mention, pull request review

## What to look at

| Path | What it is |
| --- | --- |
| [`.claude/skills/`](https://github.com/ai4curation/ai-gene-review/tree/main/.claude/skills) | Fifteen skills, including `annotation-reviewer`, `core-function-synthesizer`, `gocam-curation`, and `aigr-pr-review` |
| [`.claude/hooks/`](https://github.com/ai4curation/ai-gene-review/tree/main/.claude/hooks) | Hooks that run around agent actions |
| `genes/` | One directory per gene, holding the review and its cached references |
| `src/ai_gene_review/schema/` | The LinkML schema |

## What works

### Every annotation gets a verdict and a reason

A review does not rewrite an annotation. It records an action against the
existing one, such as `ACCEPT`, `MODIFY`, or `REMOVE`, together with the reason:

```yaml
existing_annotations:
- term:
id: GO:0005515
label: protein binding
action: MODIFY
reason: "While evidence is strong, 'protein binding' is uninformative..."
```

This is reviewable in a way that a rewritten file is not. A curator reads the
verdict and the reason, and does not have to reconstruct what changed.

### Identifiers carry their labels

Every term appears as an identifier and a label together. A wrong pair fails
validation, because an agent would have to fabricate both consistently to get
past the check. See
[Make identifiers hard to fake](../patterns/ground-identifiers.md).

### Non-curators can give feedback

The generated pages carry thumbs-up and thumbs-down controls, and there is a
longer [evaluation form](https://go.lbl.gov/gene-eval) for detailed review. A
domain expert can disagree with an agent without learning Git.

This is the part most projects skip. Validation catches fabricated evidence.
Only a human who knows the gene catches a claim that is well-cited and wrong.

### Skills split reviewing from synthesizing

`annotation-reviewer` judges what exists. `core-function-synthesizer` writes what
the gene does. `aigr-pr-review` reviews the pull request. Three jobs, three
skills, three sets of instructions that do not interfere.

## What to copy first

Copy the action-and-reason record shape. If your agents change existing curated
statements, record the verdict and the reason next to the original instead of
replacing it. Review gets much cheaper.

## Gaps

The public voting data is feedback, not yet a measurement. Turning votes into a
number you can track over releases is still open work, and it is the obvious
next step for anyone copying this pattern.
54 changes: 54 additions & 0 deletions docs/case-studies/cell-ontology.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,54 @@
# Cell Ontology

Cell Ontology runs two agent workflows and has no skills, subagents, or
commands. It is the most useful negative example on this site.

**Repository**: [obophenotype/cell-ontology](https://github.com/obophenotype/cell-ontology)
**Product**: `src/ontology/cl-edit.owl`
**Agent surface**: GitHub mention, pull request review

## What to look at

| Path | What it is |
| --- | --- |
| [`CLAUDE.md`](https://github.com/obophenotype/cell-ontology/blob/master/CLAUDE.md) | The whole agent configuration |
| [`.github/workflows/ai-agent.yml`](https://github.com/obophenotype/cell-ontology/blob/master/.github/workflows/ai-agent.yml) | Runs an agent on a GitHub mention |
| [`.github/workflows/clara-review.yml`](https://github.com/obophenotype/cell-ontology/blob/master/.github/workflows/clara-review.yml) | Reviews pull requests |

## What works

There is one instructions file. Most repositories in
[Block A](index.md#block-a-ontologies) keep the same text in both `CLAUDE.md`
and `.github/copilot-instructions.md` and have to edit both. Cell Ontology does
not have that problem.

`CLAUDE.md` is specific about search. It tells the agent that `cl-edit.owl` has
one axiom per line, and gives exact `grep` commands for finding a term by
identifier and by label. Concrete commands beat general advice.

It is also specific about identifiers. New term requests use the `CL_99xxxxx`
range, defined in `src/ontology/cl-idranges.owl`, and the file says never to
guess a term identifier or a PubMed identifier.

## What to copy first

Copy the search section. Three worked `grep` examples against your own edit file
will save more agent time than a page of prose.

## Gaps

Two workflows send real work to the agent, and nothing underneath breaks that
work into parts. Every run starts from the same general instructions.

`CLAUDE.md` contains this line:

> DO NOT bother doing your own greps over the file, or looking for other files,
> unless otherwise asked, you will just waste time.

That is a fix for something that went wrong once, written as a prohibition. A
search skill would do the same job and would also tell the agent what to do
instead of what to avoid. Prohibitions accumulate; skills compose.

Of the repositories we track, this one would gain the most from a small set of
skills. [Uberon's](uberon.md) `identifier-validator` and `metadata-checker` are
a reasonable starting pair.
Loading
Loading