Skip to content

Latest commit

Β 

History

History
371 lines (321 loc) Β· 24.1 KB

File metadata and controls

371 lines (321 loc) Β· 24.1 KB

AGENTS.md

Guidance for AI coding agents working in this repo. The layout is deliberate and the invariants below are easy to break by accident β€” read them before non-trivial changes.

For what raincloud is and how to use it, see README.md. For the manifest schema, sources.schema.md. For step-by-step procedures, SKILLS.md. CLAUDE.md is a symlink to this file.

Start here

On a fresh clone outputs/ is empty. That's expected β€” artifacts are built, not shipped.

python -m raincloud.pipeline.status --fast --missing-only   # read-only; seconds
python -m raincloud.pipeline.validate_manifest              # schema + registry cross-checks
pytest                                                    # needs --extra dev --extra all

Installs are layered. A bare uv sync --inexact gets only the loader; builds need uv sync --extra build --inexact, plus the extra a dataset needs for its format or source: osm, sas, excel, archives, generated (the TPC-H/TPC-DS generators), kaggle, huggingface. --extra all installs everything. Always pass --inexact: without it, syncing one extra silently uninstalls the others, and a later Kaggle/HF build fails.

Query the catalog rather than grepping sources.json or scrolling docs/v2/datasets.md:

python -m raincloud.pipeline.list_datasets --handler uci_default --count
python -m raincloud.pipeline.list_datasets --kaggle-tos         # gated behind a one-time click-through (Kaggle or Hugging Face)
python -m raincloud.pipeline.list_datasets --grep '\bgeo' --long
raincloud describe <slug>                                      # one dataset's columns and types, from the catalog
python -m raincloud.pipeline.list_datasets --columns --column-grep PATTERN   # columns across locally built files
python -m raincloud.pipeline.list_datasets --coverage --source parquet       # type coverage of locally built files
python -m raincloud.pipeline.list_datasets --stale-version      # slugs BUILT under an OLDER schema_version (never-built excluded)

--grep is a regex over slug short_name full_name description joined by spaces, so anchor one slug as '^<slug> '; '^<slug>\b' also matches <slug>-hydrated and other hyphenated siblings.

Filters AND together across --handler, --license, --fetch-type, --reader, --vortex/--no-vortex, --kaggle-tos, --stale-version, --local, --grep. Output: default one bare slug per line (a terminal also marks hydrated ones [hydrated]), --long, --json, --count. In --long, recorded is what the tracked catalog records (in --json, built_version / stale_version) and local is the formats prepared on this install's disk.

raincloud.pipeline.browse is a human-facing TUI β€” it will hang waiting for keystrokes. Don't run it from an agent context; point the user at it instead.

Invariants

  1. sources.json is authoritative. Every row of every derived artifact maps back to a spec here. Never hand-edit docs/*.md or drop a file into outputs/ β€” fix the manifest, rebuild, regenerate docs.
  2. outputs/raw_downloads/<slug>/ is unversioned; outputs/v{n}/<slug>/<format>/ is version-scoped. Raw upstream bytes don't depend on schema version, so they're cached outside the version prefix and shared across versions. Path helpers live in raincloud/pipeline/spec.py β€” use them rather than composing paths by hand.
  3. <scratch_dir>/.recipes/<recipe-hash>/<slug>/ is scratch. Handlers clean up after themselves; build.py --clean-workdir removes only the selected generation. Clearing it forces re-extraction without disturbing another recipe.
  4. Raincloud code, tests and examples open DuckDB through raincloud.duckdb_connect, never duckdb.connect directly. It applies the RAINCLOUD_DUCKDB_* resource limits and storage_compatibility_version=v1.5.0, which persistent VARIANT writes require.
  5. docs/ is split. Top-level docs/*.md is gitignored scratch. docs/v{n}/* is the tracked canonical set. raincloud.pipeline.docs writes to the top level; promoting to docs/v{n}/ is a deliberate manual copy.
  6. A superseded version is frozen. Artifacts under an older outputs/v{n}/ are not rebuilt and must not be wiped β€” nothing regenerates them.
  7. .archive/ is local-only and gitignored. A fresh clone won't have it. Where other docs name it as a fallback, git history is the only one you can rely on.

How a build works

Orchestrated by raincloud.pipeline.build. The pipeline is canonical-Arrow-spined: transform produces Arrow, write_canonical persists the one canonical artifact, and every output format is derived from it by an exporter.

stage module reads writes
fetch fetch.py fetch.* outputs/raw_downloads/<slug>/
extract extract.py extract.* <scratch_dir>/.recipes/<hash>/<slug>/
parse parse.py parse.* in-memory (Path, Table)
transform transform.py transform.* in-memory (slug, Table)
write_canonical canonical.py transform output outputs/v{n}/<slug>/arrow/<slug>.arrow.zstd
validate validate.py expect.* hashes canonical schema, checks rows; [WARN] unless --strict
run_exporters export/ export.formats, export.priority parquet/, vortex/ under outputs/v{n}/<slug>/; the build record
hydrate (named builds only) hydrate.py derive.hydrate a <parent>-hydrated dataset β€” outbound HTTP, safety-filter gated

run_exporters is also invokable on its own, which is the whole job whenever a change touches only the export stage (row-group sizing, a codec, a new cell) β€” the canonical is the input and is left alone:

python -m raincloud.pipeline.export <slug>...                   # re-derive from existing canonicals
python -m raincloud.pipeline.export <slug> --format vortex      # refresh only the Vortex file
python -m raincloud.pipeline.export <slug> --format parquet@rs  # this writer, this run

It refuses a slug with no canonical rather than silently starting a build. A --format parquet@rs override replaces that dataset's file with one from another writer, so its sha256 no longer matches the catalog; --all with it rewrites every Parquet file in the store, which takes hours. Confirm before running either. The file is parquet/<slug>.parquet whichever writer made it, and the build record (<data_dir>/builds.json) records the writer; the catalog learns it only when a maintainer regenerates it (see Regenerating derived docs).

Field-level custom_metadata (the VARIANT_EXT marker, GeoParquet geo metadata) rides through the canonical IPC losslessly. Stamp VARIANT only through variant.attach_variant / attach_variant_schema (the DuckDB bridge does): the stamp also declares the storage struct's metadata, and an unshredded value, non-nullable, as the Parquet VARIANT spec and arrow.parquet.variant require, and checks every row against that. DuckDB's Arrow export declares every field nullable, which a hand-set marker would carry into the Parquet schema.

Streaming handlers write the canonical spine themselves via canonical.open_canonical_writer and return [], so write_canonical is a no-op for them; they share the same validate β†’ run_exporters tail. To find which handlers do this, check the streaming column in docs/v2/handlers.md β€” don't rely on a list here.

In schema_version 2, export.formats is the only declaration of which formats a dataset exports. convert.vortex is v1-only: the schema and validate_manifest reject it in a v2 manifest, while a released v2 catalog that still carries convert.vortex: false (and no export.formats) keeps reading as Parquet-only. A format is one file whichever writer makes it. The writer is the first installed one in export.priority, looked up in the spec, then the catalog's export_priority, then RAINCLOUD_EXPORT_PRIORITY, then the built-in py, rs, java. The spec and catalog levels take a list, which applies to every format and so must name a writer for each one the dataset exports, or a map from format to list ({"parquet": ["rs", "py"]}); a format the map leaves out falls through to the next level. The machine level is a list. The build record, and after regeneration the catalog, records the writer as <fmt>_writer. Sidecar cells (parquet@rs, parquet@java, parquet@hardwood, vortex@rs, vortex@jni) run only where their binary is installed. Compliance measures every writer in scratch, never over the dataset's file.

export.formats lists the formats a dataset wants. When the planned writer cannot produce one for the dataset -- it raises, dies, reports a failed round-trip, or exceeds RAINCLOUD_EXPORT_TIMEOUT -- the previous file comes back and the build records the failure in the build record as that format's unavailable measurement (writer cell, error, toolchain versions, recipe, canonical sha, time), then carries on: the dataset is built with the formats that worked, [unavailable] <slug>/<fmt> is printed and repeated in the summary, and the build exits 0. Never write a writer's limitation into export.notes: docs regen carries the measurement into the snapshot (<fmt>_unavailable), the loader reports it (describe; FormatUnavailable quoting it; auto skips it), a later successful export replaces it, and compliance prints [stale opt-out] once a writer round-trips it. In-process writers run in a forked child so the limits can stop them: RAINCLOUD_EXPORT_TIMEOUT and RAINCLOUD_EXPORT_MEMORY (resident memory, default half of RAM), and the child raises its own oom_score_adj so a machine that runs short loses the writer, not the build. Run an unattended build as a systemd unit with OOMPolicy=continue: the default stop ends the whole unit when the kernel kills one process in it. export without --format behaves like the build; with a bare --format vortex it records the failure and exits 1, and a named cell's failure (--format vortex@rs) exits 1 and records nothing.

Every writer reads back what it writes before its file is promoted. An in-process writer reads its file with the same format's in-process reader and compares it to the canonical (exporters.read_back), streamed window by window (compare.stream_equal: one batch of each side in memory, Parquet read batches sized by bytes), inside the bounded child, so the time and memory limits cover the read too. A mismatch or a read error (Vortex 0.86.1 writes a multi-batch VARIANT column it cannot read back) is a failed round-trip, recorded as above; an in-process writer never reports roundtrip=None, and a read-back that decides neither pass nor fail raises as a bug. A sidecar verifies in its own process and may report roundtrip: null (a comparator gap, or out of memory while verifying): that file is promoted, [unverified] <slug>/<fmt> is printed and repeated in the summary, the build record keeps verified: false and the writer's note as verify_note (every other export records verified: true), docs regen carries them into the snapshot (<fmt>_verified, <fmt>_verify_note), and describe shows the file as UNVERIFIED with the reason. The read-back is one more full read of every file a build writes: on stackoverflow-badges (51M rows, 479 MB canonical) it adds 6.7 s to an 8.1 s Parquet write and 2.0 s to a 2.7 s Vortex write; a streamed compare peaks at 2.9 GiB resident on code-contests' 4.3 GB Parquet (18.5 GiB decoded) and 1.1 GiB on jsonbench's 22.5 GB Vortex file.

A recorded failure is not repeated. Before a writer runs, run_exporters looks up the measurement that applies (this install's build record at the recipe, else the catalog's snapshot; records.recorded_failure). If it names the same writer cell with the same toolchain (writer_toolchain, compared exactly: Python and library versions in process, binary name and sha prefix for a sidecar) and, when it records one, the same canonical sha, the format is skipped: [skip] <slug>/<fmt>: ...; pass --retry-errors to try again, nothing new recorded, the measurement kept. Anything different (an upgraded vortex-data, another writer, a rebuilt canonical) is attempted with a [retry] line naming the difference. A skip is not a new failure: build and export without --format exit 0 and list it in the summary; export --format <fmt|cell> and convert asked for that file, so they exit 1. --retry-errors on build, export and convert (raincloud load --retry-errors, load(..., retry_errors=True), or the retry_errors setting / RAINCLOUD_RETRY_ERRORS, which carries it to a child build) attempts it anyway: a success replaces the measurement, a planned writer's failure records it again. Compliance calls run_bounded directly and measures every write cell regardless, which is how a stale opt-out is found. A write cell's roundtrip is the writer's own read-back, including for an in-process file compliance finds on disk and does not re-encode (bounded.read_back_bounded), and that read-back is also the in-process writer's diagonal read cell (its own reader over its own file), not a second read. Only a sidecar's null is filled from the read matrix's diagonal cell.

Editing sources.json

Large hand-authored JSON with a fixed top-level key order (schema_version, generated_at, audit_cutoff, notes, datasets). Edit structurally, never with sed:

import json
from pathlib import Path
SRC = Path("sources.json")
m = json.loads(SRC.read_text())
for d in m["datasets"]:
    if d["slug"] == "target-slug":
        d["transform"]["handler"] = "new_handler"
        break
SRC.write_text(json.dumps(m, indent=2) + "\n")

Run validate_manifest afterwards. Templates for common edits are in templates/.

Data locations

The build and loader resolve their roots in this order: explicit argument, environment, user config file, system config file, default. raincloud config show prints what is in effect and where each value came from. A source checkout skips the system config files (the user config and RAINCLOUD_CONFIG still apply), so a checkout builds into its own outputs/ rather than a machine's shared store.

env var controls default
RAINCLOUD_HOME when set, forces <home>/outputs and <home>/_workdir checkout root, else the user data dir
RAINCLOUD_OUTPUTS built-artifact root (v{n}/ and raw_downloads/ live directly under it) <checkout>/outputs; outside a checkout, the user data dir itself
RAINCLOUD_RAW_DOWNLOADS cached raw upstream bytes $RAINCLOUD_OUTPUTS/raw_downloads
RAINCLOUD_WORKDIR scratch root holding .recipes/ <checkout>/_workdir; outside a checkout, <user cache>/workdir
RAINCLOUD_MANIFEST / RAINCLOUD_SNAPSHOT a matched local catalog: sources.json and its snapshot checkout copy, else the packaged copy
RAINCLOUD_CATALOG which catalog: auto, active, checkout, bundled, local (the manifest RAINCLOUD_MANIFEST names), a revision or unique prefix, or a bundle/pack directory auto
RAINCLOUD_CATALOG_DIR installed catalog revisions <user data>/catalogs
RAINCLOUD_CATALOG_URL where raincloud catalog update looks when no --source is given unset
RAINCLOUD_CACHE optional separate artifact cache for the loader same as the data root
RAINCLOUD_MIRROR a private artifact store readers fall back to (s3:// needs [s3], https:// needs [http]) unset
RAINCLOUD_OFFLINE 1: read only local files; never contact the mirror unset
RAINCLOUD_RETRY_ERRORS 1: a build attempts a format whose writer, with this toolchain, already failed at the recipe (as --retry-errors) unset
RAINCLOUD_CONFIG / RAINCLOUD_NO_CONFIG select or disable the config file unset
RAINCLOUD_SETTINGS settings JSON the CLI reads with --settings-env; how native readers pass options unset
RAINCLOUD_DUCKDB_MEMORY_LIMIT DuckDB memory ceiling, applied by raincloud.duckdb_connect DuckDB default (~80% RAM)
RAINCLOUD_DUCKDB_THREADS DuckDB thread count DuckDB default
RAINCLOUD_DUCKDB_TEMP_DIRECTORY DuckDB spill directory DuckDB default
RAINCLOUD_FETCH_DEADLINE wall-clock ceiling on one download 6 h (0 disables)
RAINCLOUD_GENERATOR_TIMEOUT ceiling on a generator subprocess 6 h (0 disables)
RAINCLOUD_MAX_TABLE_CELLS rows x columns a whole-table handler may materialize 50,000,000 (0 disables)
RAINCLOUD_MAX_DECOMPRESSED_BYTES one in-memory decompression 4 GiB (0 disables)
RAINCLOUD_ROW_GROUP_TARGET_ENCODED_BYTES Parquet row-group size, in encoded bytes (parquet@java counts compressed pages, so its groups come out larger) 128 MiB
RAINCLOUD_ROW_GROUP_MAX_ROWS row cap per group, used only when a spec omits write.row_group_size_rows; a spec's cap wins in every writer, sidecars included 10,000,000
RAINCLOUD_ROW_GROUP_TARGET_BYTES memory guard: decoded Arrow bytes buffered for one row group 512 MiB
RAINCLOUD_ROW_GROUP_PROBE_ROWS rows the Python Parquet writer samples to size its groups (must be > 0) 262,144
RAINCLOUD_BATCH_ROWS / RAINCLOUD_BATCH_BYTES batch bounds in the streaming ingestion paths (memory only, NOT the row-group size) 4096 rows / 16 MiB
RAINCLOUD_EXPORT_PRIORITY machine writer preference, e.g. rs,py unset (py, rs, java)
RAINCLOUD_EXPORT_TIMEOUT ceiling on one export: an in-process writer (run in a child process) or a sidecar writer; hitting it records the format unavailable 6 h (0 disables)
RAINCLOUD_EXPORT_MEMORY ceiling on one in-process export's resident memory (bytes); the parent stops a writer over it and records the format unavailable half of physical memory (0 disables)
RAINCLOUD_SIDECAR_TIMEOUT ceiling on one sidecar reader call (sidecar writers use RAINCLOUD_EXPORT_TIMEOUT) 30 min (0 disables)
RAINCLOUD_TPCGEN_CLI path to the tpcgen-cli executable the tpcgen-rs TPC-DS generator runs beside the Python executable, else PATH

Individual vars and explicit values win over RAINCLOUD_HOME. A malformed numeric value is an error naming the variable, never a silent default. Loader handles and build entry points freeze these for the duration of an operation. The seven 0 disables ceilings exist because builds are often left to run unattended: each bounds work that is otherwise decided by an upstream file or an external process (a download, a generator, a sidecar), and each can be lifted with 0 for a run that genuinely needs it.

Two directories are named .recipes/. <scratch_dir>/.recipes/<recipe-hash>/<slug>/ is one recipe's extract scratch, so a changed recipe never reuses another's intermediates. <raw_dir>/<slug>/.recipes/<fetch-key>/ holds raw bytes for a catalog or fetch recipe other than the one that owns <raw_dir>/<slug>/ itself. Skills and playbooks that say <recipe-hash> mean the first.

Public loader API

load / load_dataset, slugs(), describe(slug), reader_capabilities(), Config / resolve_config, and the exception hierarchy. Anything under a leading underscore is not it β€” an example that reaches into raincloud._catalog teaches that import to everyone who copies it, and usage is what makes a name public.

Confirm before rebuilding

Large datasets take hours, and a rebuild wipes and redoes existing work. Ask the user before running raincloud.pipeline.build on anything non-trivial. Parquets under ~100 MB are fine to rebuild unprompted.

Regenerating derived docs

python -m raincloud.pipeline.docs    # datasets.md + handlers.md + snapshot.json

This is the only way the catalog changes: builds write the install's build record (<data_dir>/builds.json), never the tracked snapshot, and regenerating takes each built file's sha and writer from that record. In a checkout it writes the gitignored scratch copies under docs/; promote them, then review and commit:

python -m raincloud.pipeline.docs
cp docs/snapshot.json docs/datasets.md docs/handlers.md docs/v2/
git diff docs/v2/

A machine's shared store is released from a commit with python -m raincloud.pipeline.publish <slugs|--all> --store DIR --catalogs DIR.

The snapshot is load-bearing. datasets.md regen reads from disk for locally built slugs and falls back to the snapshot for everything else. A partial regen without that fallback would dash out most rows and destroy ground truth. The no-args form regenerates snapshot and datasets in lockstep β€” prefer it; docs.py datasets alone will not refresh the snapshot.

Conventions

  • One handler per upstream shape. Don't stretch tighten_types or identity β€” add a handler under raincloud/pipeline/handlers/ and declare it in HANDLERS in raincloud/_registry.py (see below). Read docs/v2/handlers.md first to pick precedent and see which extras you'll need.
  • Handlers stay short. Most are under 150 lines; reuse open_canonical_writer, duckdb_connect, workdir_root, spec_field before growing one.
  • No backwards-compat stubs. Remove a handler or slug fully; git history is the fallback.
  • Handlers, exporters and generators are declared in raincloud/_registry.py, and nowhere else. The registries build from it (handlers/__init__.py only resolves names lazily), and the capability list a catalog bundle records derives from it, so adding one is a single edit. That module imports nothing, which is what lets the loader answer "can this be built here?" without pulling in the build toolchain.
  • Version numbers have one home each. The release version is raincloud/__init__.py:__version__. pyproject.toml, the Java client and the C/C++ CMake build read it. Two files cannot and carry a literal: clients/rust/Cargo.toml (Cargo requires one) and CITATION.cff. Bump them with it; tests/test_loader_package.py::test_version_mirrors_agree fails on drift. The set of artifact layouts is the schema_version enum in sources.schema.json, read by Python at runtime; the native clients hold no copy, because the CLI resolves layouts for them. Don't add a second copy, and don't give schema_version a default β€” a wrong one silently selects another layout.
  • Upstream bytes are untrusted input. Archive members get _safe_target + _claim (no escape, no two members on one path); anything written under a final name goes through an atomic temp-then-rename, so [cached] can mean "complete"; and a row the pipeline drops gets counted and printed. Silence is the bug β€” a short table looks exactly like a correct one.
  • Test the observable result, not the mock. Use real small Arrow/Parquet files and real build subprocesses; don't mock the resolver, serializer, checksum or builder that the test exists to verify. Wheel and live-upstream tests are opt-in (--run-wheel, --run-network).

Where things are

  • raincloud/pipeline/ β€” the build stages and CLI entry points
  • raincloud/ β€” the importable loader; clients/ β€” Rust, C/C++, Java readers (clients/README.md). The native readers hold no catalog; they ask the raincloud CLI, so there is nothing of the catalog to keep in step there.
  • sidecars/ β€” reference writers/readers behind the sidecar cells, a PATH-discovered CLI contract. The JVM lanes are a Gradle composite build over a git submodule; a fresh clone needs git submodule update --init.
  • raincloud.pipeline.compliance β€” maintainer-run measurement of the (slug Γ— format Γ— impl) matrix. Never gates a build; absent toolchains skip. --check-oracle runs the additive-only gate: a cell may be added, never removed or mutated. Re-measure the committed oracle whenever the toolchain pins move; until then the gate reports every cell the new toolchain changed as a mutation. See sidecars/README.md.
  • .agents/skills/ β€” invokable skills wrapping the pipeline entry points. .claude β†’ .agents is a symlink. .agents/settings.json is a tracked read-only command allow-list; machine overrides go in the gitignored settings.local.json.

When you're unsure

Read, then grep, then ask β€” don't guess. The pipeline has contracts that aren't visible from any single file: streaming handlers returning [], raw_downloads being unversioned, VARIANT requiring the DuckDB compatibility setting.