Skip to content
maxh213Public

About

Agent harness that makes coding agents loop on a gauntlet of deterministic gates (coverage, mutation, complexity, dead code, Sonar and more)

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

marestail

It's called marestail because working with LLMs reminds me of hacking away at weeds on the allotment. LLMs drift and cause bugs at speed which you have to correct for as you work with them. Marestail pops up each week with a vengance and I need to pull it all out again. The comparison isn't 1:1 but it's good enough for me :)

Marestail is gauntlet of deterministic gates for coding agents, after Uncle Bob's approach: don't tell the agent to be clean, measure cleanliness and make it loop until the measurement passes.

This whole project is very opinionated on what I consider to be clean code / good practices which I want to force an LLM into implementing.

Gates

Gate Python TypeScript Elixir Ruby / Rails C# / .NET Erlang Rust Java
tests, 100% line and branch coverage pytest, coverage.py vitest, v8 (or jest, or playwright + c8) mix test --cover rspec + SimpleCov dotnet test + coverlet eunit + cover cargo llvm-cov, lines and code regions mvn test + JaCoCo
CRAP ≤ 4 per function radon + coverage typescript AST + istanbul elixir AST + cover Ripper AST + SimpleCov Roslyn scanner + coverlet erl_parse AST + cover syn AST + llvm-cov lines javac tree API + JaCoCo
mutation testing, changed files mutmut Stryker muex mutant Stryker.NET built-in operator-swap escript cargo-mutants PIT
dependency direction import-linter dependency-cruiser mix xref cycles Zeitwerk constants vs .ruby-layers.json Roslyn type resolution vs .dotnet-layers.json, cycles beam call-graph cycles use/path resolution vs .rust-layers.json, module cycles import and name resolution vs .java-layers.json, cycles
types and lint mypy strict, ruff tsc strict, eslint mix format, mix compile rubocop Roslyn analyzers via SARIF erlc strong warnings as errors clippy -D warnings -D clippy::pedantic, rustfmt javac -Xlint:all warnings as errors, PMD
no comments, no docstrings tokenizer typescript scanner elixir AST scanner Ripper Roslyn scanner escript scanner proc-macro2 token gaps, doc attributes scanner lexer, javadoc included
no pass-through functions, no imports of private modules ast typescript AST elixir AST Ripper Roslyn scanner (pass-throughs) escript scanner (pass-throughs) syn scanner (pass-throughs) javac tree scanner (pass-throughs)
no unreachable definitions vulture knip BEAM abstract code scan unused private methods unused private members escript scanner rustc dead_code via clippy, unreferenced pub items unused private members
docs match the code: routes ledger, env vars, paths regex over sources regex over sources regex over sources regex over sources regex over sources regex over sources regex over sources regex over sources
the code parses on the interpreter that ships Dockerfile base image vs requires-python, ruff, mypy and shebangs, then ast at that version — — — — — — —
Sonar quality gate, zero issues, zero duplication local SonarQube local SonarQube local SonarQube local SonarQube local SonarQube, SonarScanner for .NET local SonarQube, sonar-erlang plugin local SonarQube, built-in Rust analyzer + LCOV local SonarQube, built-in Java analyzer + JaCoCo XML

The Erlang gates compile and run eunit themselves with erlc and escript (OTP 25+); no rebar3 is required. er.mutation is marestail's own mutation tester: an escript rewrites one operator at a time (comparison, arithmetic, andalso/orelse swaps), recompiles, and runs the eunit suite per mutant — survivors fail the gate, mutation_max caps the mutants checked when a full pass is too slow. For the sonar tier, marestail sonar setup builds the sonar-erlang plugin jar once with docker (pinned to a commit, cached under ~/.config/marestail/, so the build does not recur), mounts it into the SonarQube container's extensions/plugins and restarts the container if the plugin is not loaded yet. The gate imports the coverage er.tests already produced: the eunit run exports .marestail/eunit.coverdata (via cover:export), which the plugin parses into line coverage. It then fails closed when SonarQube shows no Erlang lines or no coverage metric, on top of the usual quality gate, issue, coverage and duplication checks.

The Rust gates need cargo with clippy and rustfmt, plus cargo install --locked cargo-llvm-cov cargo-mutants. marestail's own scanner (marestail/rs/scan, syn and proc-macro2) is built once per repo into .marestail/rs-scan, rebuilt when its source changes, and needs crates.io the first time. rs.tests runs cargo llvm-cov --no-report then writes .marestail/rs-lcov.info (for Sonar and CRAP) and the llvm JSON export; each test binary instruments the library separately, so hits are merged per line and per code region. Stable Rust has no branch coverage, so the gate reports unexecuted code regions, which catch an untaken else or match arm even when it shares a line with covered code. Without rustup, the gate points LLVM_COV/LLVM_PROFDATA at the system LLVM; it must match the LLVM version in rustc -vV. Dead code is two checks: rustc's dead_code lint fails rs.lint, and deadcode reports pub items whose name appears nowhere else in the sources, tests/, examples/ or benches/ (a name match, so a same-named item elsewhere hides one). Items defined in lib.rs count as the crate's API and are never reported. Pass-through detection skips trait impls, because the trait fixes their signature. There are no default layers; rs.deps checks .rust-layers.json ([{"from": "src/domain", "forbid": ["src/web"]}]) and always fails on module cycles. rs.mutation treats a mutant that times out as killed and ignores ones that do not compile. For Sonar, set sonar.rust.lcov.reportPaths=.marestail/rs-lcov.info and sonar.rust.clippy.enabled=false (the scanner container has no cargo, and rs.lint already runs clippy).

The Java gates drive Maven 3.9+ (the repo's ./mvnw when there is one, else mvn, or [java] mvn) on a JDK 21 or newer. [java] root is the module that holds pom.xml; multi-module reactors and Gradle builds are not supported yet. marestail's own scanner (marestail/jvm/Scan.java, built on the JDK's com.sun.source tree API with no dependencies) is compiled with javac once per repo into .marestail/java-scan, and again when its source changes. java.tests runs mvn test with the JaCoCo agent added on the command line, so the pom needs no JaCoCo entry; a pom that sets surefire's <argLine> must start it with @{argLine}. It reads the surefire XML reports and jacoco.xml: a line with no covered instruction is a gap, and so is every missed branch. java.lint compiles the main and test sources together against the Maven test classpath with javac -Xlint:all and fails on any warning, then runs PMD 7 (rulesets/java/quickstart.xml plus CyclomaticComplexity, or your own [java] pmd_ruleset) with the compiled classes on its aux classpath. Maven fetches PMD the first time, from marestail/jvm/tools/pom.xml. @SuppressWarnings and // NOPMD in a source are findings too. java.mutation runs PIT through the project's pitest-maven plugin, because PIT's JUnit 5 support can only be added as a plugin dependency in the pom: copy templates/java-pitest.xml into <build><plugins>. A mutant that times out counts as killed, and a non-viable one is ignored. .java-layers.json (templates/java-layers.json) names layers by package, where com.acme.domain covers its subpackages, or by a path containing /; forbid_external matches imports. java.deps finds a dependency through an import, a simple name in the same package or under an on-demand import, or a fully qualified name, and always fails on cycles between files. Dead code is private methods and fields whose name appears nowhere in the sources (a name match, like C#). It skips annotated members, the fields of annotated classes (Lombok and friends generate accessors for them), and the serialization hooks. Pass-through detection skips @Override methods, because the supertype fixes their signature. For Sonar the gate passes sonar.java.binaries and sonar.coverage.jacoco.xmlReportPaths=.marestail/java-jacoco.xml to the scanner unless sonar-project.properties sets them.

Acceptance is the same in every language: whatever command [qa] cmd names, run from [qa] cwd.

Tiers: fast (tests with coverage, CRAP, lint, the dependency rules, comments, dead code, depth, docs and py.runtime; a language's gates run only when its section is in marestail.toml), sonar (fast plus the Sonar quality gate), full (sonar plus mutation testing), qa (fast plus the acceptance command, with no Sonar and no mutation), all (every gate). marestail gate with no --tier runs fast on the whole repo. The Stop hook runs fast on the diff against [git] base plus the focus paths, and blocks at most five stops per session before letting the agent finish.

py.runtime exists because every other gate runs in the repo's virtualenv, which is not what production runs. It reads the version off the last FROM in the Dockerfile (override with [python] deploy_files, or state it outright with [python] runtime = "3.12"), requires every version the tooling asserts to equal it, and parses each source at that version. Keeping mypy's python_version honest is half the point: typeshed then rejects stdlib names the shipped interpreter does not have. It skips when nothing declares a deployed interpreter.

A Next.js repo that will not move to vitest sets [ts] runner = "jest": the gate runs the repo's own node_modules/.bin/jest with --coverageProvider=babel (v8 cannot express branch arms) and needs [ts] sources so untested files still appear in the report.

A repo whose TypeScript suite is Playwright — including Next.js apps that already run @playwright/test specs, even Node-side unit tests that never open a browser — sets [ts] runner = "playwright". The gate runs the repo's node_modules/.bin/playwright test under c8, writes the same istanbul JSON vitest and jest produce, and needs [ts] sources so files no spec imported still appear. Without @playwright/test or c8 the gate fails closed. [qa] cmd can still be npx playwright test: that is the acceptance command and does not collect coverage. Specs that only drive a browser and never import source files fail closed as uncovered sources.

Scope

--scope changed gates exactly the diff against [git] base: coverage counts only the changed lines and branches, CRAP only the functions the hunks land in, and lint, deps, mutation, comments, deadcode and depth only the changed files. The scope is the diff, so it follows the work — split one file into three or move a function into another existing file and the new hunks are gated wherever they land. --focus PATH (repeatable, or [focus] paths in marestail.toml) adds whole files or folders on top and treats them as fully changed; combined with --scope all it is a usage error. Sonar is filtered, not rescoped: the scanner still sees the whole project and the open issues, duplication and hotspots are cut down to the scope afterwards. docs, qa and py.runtime stay global by nature and print a scope note saying so. --scope all is unchanged and remains the default. marestail run takes the same --scope and --focus flags: every role's verification gate and the gate command each worker is told to run then cover the diff plus the focused paths, with [focus] paths from marestail.toml added; without them a pipeline gates the whole repo. This is a soft scope. It limits what is measured, not what a worker may edit: a worker is free to change any file, and whatever it changes joins the diff and is gated with it. Only freeze restricts edits.

--scope hard gates only the focus paths, each as fully changed, and nothing else in the diff: change one file and only that file goes through coverage, CRAP, lint, deps, mutation, comments, deadcode, depth and Sonar, while a supporting edit elsewhere (a caller, an import, a test) is not measured. It needs at least one focus path, from --focus or [focus] paths. Under marestail run --scope hard, every worker and judge is also told the scope: workers make the change inside the focus paths and only the smallest supporting edits outside them, with no refactoring, renaming, reformatting or cleanup there, and judges bounce a diff that goes further outside them. The runner hands the scope to the agents' Stop hooks through MARESTAIL_SCOPE and MARESTAIL_FOCUS, so the hooks gate the same paths as the pipeline. Tests still run as a whole suite, so a supporting edit that breaks another test still fails the gate, and docs, qa and py.runtime stay global.

Mutation testing always runs on the diff: even without --scope changed, each mutation gate defaults to the files changed against [git] base (an empty diff skips the gate), because a whole-repo mutation pass is too slow to run on every gate. Set [<lang>] mutation_scope = "all" (e.g. [elixir] mutation_scope = "all") to opt back into whole-repo runs; any other value fails the gate. An explicit --scope changed / --focus still wins over the config, and a repo whose [git] base ref does not resolve falls back to a full run, noted in the gate summary.

A mutant that times out counts as killed in every mutation gate, as Stryker, muex, PIT, cargo-mutants and mutant already count it: the mutant changed behaviour enough that the tests never finished. The summary still says how many were killed that way (5 of 1973 mutants not killed (7 killed by timeout)), so a jump in timeouts stays visible. [<lang>] mutation_timeouts = "fail" (e.g. [ts], [dotnet], [python], [elixir]) counts them as survivors instead.

ex.mutation kills any single muex test run that goes past [elixir] muex_run_limit seconds (default 300). muex's own --timeout resets whenever the tests print output, so a mutant that stops the app from starting while it keeps logging would otherwise hang the whole run. muex scores the killed run and moves on.

Use

export PATH="$PATH:/path/to/marestail/bin"
cd your-repo
marestail install .          # marestail.toml, sonar-project.properties, CLAUDE.md / AGENTS.md, PERFORMANCE.md, Stop hooks
marestail install . --gitignore-generated   # also gitignore features/, qa/, tasks/, PERFORMANCE.md, perf/ and the Stop-hook configs, for repos where not everyone runs marestail
marestail sonar setup        # local SonarQube in docker, token in ~/.config/marestail
marestail gate               # fast tier, whole repo
marestail gate --tier full --scope changed
marestail gate --focus app/services   # the diff, plus a whole folder treated as fully changed
marestail gate --scope hard --focus app/services/payments.py   # only that file, nothing else in the diff
marestail route              # the subscription to use now, from dandelion route; marestail route --high for route --high
marestail graph              # module dependency graph, for the architect and for you
marestail depth              # prints, per module, the number of public symbols, the number of statements, and the ratio between them, marking wide-and-thin modules as shallow and files over 300 lines as long.
marestail run tasks/001.md   # Claude (default), or --agent agy|grok|cursor|kilo|kimi / MARESTAIL_AGENT
marestail run tasks/001.md --model claude-opus-5 --effort high   # both are stamped on every commit
marestail run tasks/001.md --scope changed --focus write-to-api/Services   # soft scope: gates cover the diff plus the focused paths; workers may still edit anything, and it joins the diff
marestail run tasks/001.md --scope hard --focus src/render/terminal.ts   # hard scope: gates cover only that file, and roles leave the rest alone
marestail run tasks/001.md --model dandelion/route        # ask dandelion route before every session
marestail run tasks/001.md --model dandelion/route-best   # ask dandelion route --high before every session

marestail run starts at nice 19 with idle I/O and a raised OOM score, so a fleet of pipelines yields the CPU and disk to you and is first in line if the kernel needs memory. Children inherit. [run] nice = 0 or MARESTAIL_NICE=0 turns it off; any other integer 1–19 is the level (MARESTAIL_NICE wins over the file). marestail watch and marestail gate stay at the shell's priority.

--effort names the reasoning effort for the run and every backend carries it in the commit stamp. Claude takes it as --effort (low, medium, high, xhigh, max), agy as --effort (low, medium, high), Grok as --reasoning-effort, Kilo as --variant. Cursor has no flag for it: it goes inside the model, --model 'claude-opus-4-8[context=1m,effort=high]', and --effort there only labels the commits. Kimi has no flag for it either, so --effort only labels the commits. [agent] effort in marestail.toml sets the default; MARESTAIL_GROK_EFFORT and MARESTAIL_KILO_VARIANT still work for those two.

Routing

dandelion watches the usage windows of your AI subscriptions and names the one to use now. marestail route runs dandelion route with the same arguments and exit code (marestail route --high runs dandelion route --high); when dandelion is not installed it says where to get it and exits 127.

--model dandelion/route (or [agent] model = "dandelion/route") asks dandelion route before every agent session, and --model dandelion/route-best asks dandelion route --high. dandelion prints <model> [effort] <account>, and that session runs on the account's backend with that model and effort: claude, claude-work and claude-deepseek run claude (claude-work with CLAUDE_CONFIG_DIR set to DANDELION_CLAUDE_WORK_CONFIG_DIR, default ~/.claude-work; claude-deepseek with CLAUDE_CONFIG_DIR set to DANDELION_CLAUDE_DEEPSEEK_CONFIG_DIR, default ~/.claude-deepseek), and agy, kimi, grok and cursor run their own CLIs. Each session is stamped with the model and effort it ran on, and the log shows the line dandelion printed. A session that hits a rate limit asks dandelion again before retrying, so it can move to another subscription; when dandelion prints none, the runner waits and asks again, like any other rate limit. The backend and effort come from dandelion, so --agent and --effort cannot be combined with these models, and [agent] effort is ignored. MARESTAIL_DANDELION names a different dandelion binary.

Overnight

tools/overnight.sh tasks/000.md tasks/002.md ... runs tasks in order, each to the hardener by default (STOP_AT=qa to include QA), stops at the first failure, waits out rate limits for up to six hours, and writes .marestail/runs/overnight-<stamp>.md with one section per task: exit code, minutes, HEAD, the role and verdict lines, any config proposals, and the performance changes. AGENT, MODEL, EFFORT, SCOPE and FOCUS (a space-separated list of paths) in the environment pass the matching flags through. Start it detached: nohup setsid tools/overnight.sh ... > /dev/null 2>&1 &.

Kilo Code pipeline runs (--agent kilo) use kilo run --auto --format json, prompt on stdin, JSONL on stdout. Default model is StepFun Step 3.7 Flash (free) at variant high; --model overrides. A judge VERDICT: in the JSONL stream still counts. Kilo has no command Stop hook; the runner's four-hour cap is the timeout.

Kimi Code pipeline runs (--agent kimi) use kimi -p --output-format stream-json, JSONL on stdout; the -p text only points at the prompt file under .marestail/runs/<task>/, because a full worker prompt is longer than one argv entry allows (128 KB); -p mode needs no permission flags. --model passes through as -m. A judge VERDICT: in the JSONL stream still counts. Kimi has no command Stop hook; the runner's four-hour cap is the timeout.

install --gitignore-generated exists for repos where not everyone runs marestail: the flag adds the marestail-only working files — features/, qa/, tasks/, PERFORMANCE.md, perf/, the Stop-hook configs — to the target's .gitignore. marestail.toml, sonar-project.properties, CLAUDE.md and AGENTS.md are shared configuration and documentation: they are never gitignored. Workers are told not to git add -f; if they do, the runner untracks those paths after the role (the files stay on disk for the next role).

Watch

marestail watch [paths...] opens a live curses TUI of every repo with a running marestail pipeline — beds whose .marestail directory has a live worker process (--all shows every bed, running or not). With no paths it reads ~/workspace when that exists, else the current directory, and walks nested project folders (so hermes/doin_it is a bed, not only the top-level checkouts). marestail run appends .marestail/runs/<task>/pipeline.log as well as printing; watch tails that, or the newest overnight-*.log, so a run started without overnight.sh still shows its role. Each repo is a bed, and the active worker's row carries its role, elapsed time and a scrolling one-line tail of its latest output, so you can see what the fleet is doing without opening a single log. Between agent sessions the row shows what the runner itself is doing — the gate in progress, or its latest log line such as perf sample filling — falling back to between steps only when nothing is happening. While a claude worker runs, its bed grows up to three dim lines tailed live from the agent's session transcript — its latest thinking, text and tool calls — so progress is visible between step finishes. Enter on a worker opens its full conversation: the prompt it was sent, the handoff it wrote, the result it returned. Arrows or j/k move, Enter opens, q quits. Stdlib curses only, nothing to install. The panels are a registry built to be extended — module graph and coverage views are planned.

Pipeline

Step Kind Gate Does
specifier worker none Gherkin scenarios and a QA procedure from the task
critic judge none judges the spec against the task; bounces to a fresh specifier until it passes; then a human approval pause, skipped by --auto or when stdin is not a terminal
coder worker fast implements; must trace every scenario to a test in its handoff
cleaner worker sonar readability without comments, CRAP, Sonar
architect worker sonar draws module boundaries, moves code, tightens the dependency contracts
practices judge none reviews the diff against the repo's guidance/*.md rulebooks; bounces to a fresh coder only for a cited rule violation in code the task touched; skipped when the repo has no guidance/*.md
perf judge none benchmarks every perf/ bench on the start commit and HEAD; flags degradations and improvements; bounces to a fresh coder only for a fix inside the spec
hardener judge full judges the diff and the full-tier gate report, mutation included; bounces to a fresh coder, or to the specifier when a scenario itself is wrong, until it passes
qa worker qa turns the QA procedure into an executable end-to-end test

The Gate column is the tier a worker is told to run and must pass before its handoff is accepted, and the tier whose report a judge reads before ruling; none means the step runs no gate. Under marestail run every gate follows the run's --scope: a default run gates the whole repo, and --scope changed or --scope hard narrows every tier the same way.

Workers edit and commit. Judges write one verdict file and nothing else (perf may also write perf/**); the runner discards any other edit a judge makes. Handoff and verdict files are runtime state under .marestail/, never committed: when a worker passes verification the runner folds its handoff into that role's commit message, and a judge's verdict becomes an empty commit carrying the verdict. A worker that force-adds a gitignored path (typical: features/, qa/, .marestail/handoffs/) has that path dropped from the commit before the handoff is folded; the working copy is kept. Every commit a run produces starts with the model and effort that produced it, [claude-opus-5 high] coder handoff; the runner rewrites the subject of any commit a worker made without one, so the stamp is deterministic rather than something the agent has to remember. With no --model the backend name stands in for it, and with no effort the stamp is the model alone. git log on the branch is the record, and a role in a fresh clone reads its predecessors from there. When a pipeline completes, the task's handoff files are archived under .marestail/runs/. Every role runs in a fresh session with a short prompt: the role file, the task, the earlier handoffs, and how to finish. Judges also get the gate report.

After every worker the runner checks, deterministically: the handoff exists, the tree is committed, no frozen file changed (one tolerance: a .csproj whose diff is nothing but added PackageReference or InternalsVisibleTo lines is kept, because a test needs its packages and Moq needs InternalsVisibleTo; a removed or changed line, or any other element, is reverted as before), the gate for that tier passes, and for the coder that every scenario in the feature file is traced to a test that exists and that the coder may edit (a trace into a frozen path such as qa/ fails at once, because the QA role writes the end-to-end tests). Anything failing goes back to the same role as feedback until it passes (--retries N caps it; default is unlimited). A worker that gets the same problems back three attempts in a row, ignoring numbers such as timings and attempt counters, is not making progress, and the pipeline stops for a human. A judge's gate failing is a bounce regardless of what the judge wrote. A judge bounces as many times as it takes, with one stop: if it writes the same numbered findings twice in a row, the worker is not making progress and the pipeline stops for a human.

Best practices

The practices judge reviews each task's diff against the per-language rulebooks in the repo's guidance/ folder (guidance/ts.md, guidance/rb.md, …; the file stem names the language). A repo with no guidance/*.md is skipped entirely, so the step only runs where a maintainer has added rulebooks, and [practices] enabled = false turns it off per repo. It bounces to the coder only for a clear violation of a numbered rule in a line the task added or changed, citing the rule id and file:line; violations in pre-existing code are listed informationally under ## Pre-existing on a PASS. React and Next.js rules apply only where that stack is present; Cowboy and pg fan-out rules only to real-time/websocket code; Phoenix and LiveView rules only to Phoenix/LiveView/Ecto code; Hotwire rules only to Rails views and broadcasting models.

marestail install places guidance/ts.md, the curated TypeScript/React/Next.js rulebook (rules numbered TS-1…); in a repo with a .csproj guidance/cs.md, whose one rule so far (CS-1) requires every test to follow Arrange, Act, Assert; in a repo with a *.erl file guidance/er.md, the OTP 29 rulebook (rules numbered ER-1…); in a repo with mix.exs guidance/ex.md, the Elixir 1.20 / Phoenix 1.8 rulebook (rules numbered EX-1…); and in a repo with a Gemfile guidance/rb.md, the Rails 8.1 rulebook (rules numbered RB-1…). Guidance files are maintainer policy: committed, never gitignored, and frozen (guidance/** in freeze.SPEC), so no agent can weaken a rulebook during a run. Other languages get no shipped rulebook; write your own guidance/<lang>.md with one numbered rule per line.

[practices] keys: enabled (default true).

Performance

The perf judge measures how each task changed performance. Before every attempt the runner checks out detached worktrees in temp directories and lists them in .marestail/perf/trees.json and the prompt: baseline at the task's start commit (recorded in .marestail/runs/<task>/start-commit when marestail run starts, or the merge-base with [git] base when that is missing), head (the repo itself), and, on the first perf run in a repo only, pre-marestail, the parent of the commit that added marestail.toml. The worktrees, their databases and trees.json are removed after the attempt.

Benchmarks are executable perf/bench_* scripts that the perf agent writes and extends, committed with every perf verdict. Workers cannot edit perf/** or PERFORMANCE.md, and every gate ignores the root perf/ directory; the Sonar scanner is passed the target's own sonar.exclusions plus perf/**, so existing sonar-project.properties files need no edit. In a repo with a .csproj at its root, that project would compile any .cs file under perf/, so the runner rejects a perf verdict while one exists there. The perf step runs in two phases: in the authoring phase the perf agent edits benches (perf/** is writable), and in the verdict phase perf/** is frozen and the agent only analyses and writes a verdict (VERDICT: AUTHOR sends it back to authoring, up to three times). Between the phases the runner itself fills every missing sample to [perf] min_runs on every tree with the current bench fingerprints — the agent never samples in the verdict phase, so a verdict can no longer die waiting on benchmark rounds its session outlives. Every perf run re-runs every bench on every tree, so each row is a full snapshot. Samples come through marestail perf run perf/bench_<name> --tree <tree> [--samples N] [--db] (the runner calls the same machinery when filling): it runs the script in that tree with MARESTAIL_PERF_TREE, MARESTAIL_PERF_TREE_PATH and MARESTAIL_PERF_SAMPLE set, reads JSON lines {"target": "GET /donations", "unit": "ms", "better": "lower", "values": [0.61, 0.58, ...]} (or {"target": "GET /donations", "absent": true} when a tree has no such target), and appends them to .marestail/perf/samples.jsonl. A sample that reports values must time at least [perf] values_per_sample (200) requests, so p95 describes slow requests rather than slow processes; a one-shot cost such as app startup reports a single "value". The runner, not the agent, pools the values and turns them into p50 and p95 per target and tree. Each sample is stamped with a fingerprint of its bench and every shared file under perf/: taking samples after a bench or a shared helper changed drops that bench's older samples on every tree, and a verdict is rejected while any sample's fingerprint is out of date, so old and new measurements are never averaged together. Files under perf/ whose names start with _, and __pycache__, are scratch: before committing a perf verdict the runner deletes the ones no other perf file refers to.

The runner rejects a perf verdict and retries it, with the reasons in the next prompt, when a bench has no samples on a tree, a target has fewer than [perf] min_runs samples, a column already in PERFORMANCE.md was not re-measured, or a degraded or improved target is missing from the verdict. A target is degraded or improved only when the change clears three bars. It must move by at least the unit's [perf] min_change in absolute terms (1 ms by default, so sub-millisecond wobbles never flag). The whole 95% confidence interval of the change must lie past [perf] threshold_percent; the runner resamples each tree's pooled values [perf] bootstrap times with a fixed seed to get it. With [perf] control = true, the runner also checks out the baseline commit a second time as a control tree, and a change must exceed the difference between baseline and control too. A p95 with fewer than [perf] p95_min_values (200) pooled values per tree is never flagged. Perf bounces to the coder only when a concrete change inside the task and its feature file would recover a degradation without changing a scenario (an N+1 query, repeated work in a loop, a missing batch, an unbounded payload); a degradation the specification makes inherent passes, listed under ## Degradations with the reason, and every improvement is listed under ## Improvements. marestail run ends with a ## Performance changes summary.

On a PASS the runner updates PERFORMANCE.md, which marestail install creates: one row per task (a re-run replaces its row), one column per target and metric, and a pre-marestail row the first time. A cell holds HEAD's value and its change against the start commit: 12.1ms (-2.4%), 15.1ms (+21.8%) ⚠, 9ms (-27.4%) ✓, 0.02ms (new), removed, or 57ms (n=10) for a p95 with too few pooled values to classify. Rows records the performance database size the row was measured with.

A bench that touches a database runs with --db. With [perf.db] migrate set, marestail runs a local Postgres in Docker at the repo's version: [perf.db] image, else the first postgres: image in a root compose file, then in .github/workflows, then .tool-versions, else postgres:18; every tree uses the same image. For each tree schema it builds a golden data directory once: it migrates an empty bench database, runs perf/seed.sql (with the psql variable :rows) or an executable perf/seed* (with MARESTAIL_PERF_ROWS), requires every table to hold [perf.db] rows rows, vacuums and stops cleanly. rows defaults to 50,000,000 per table; 0 builds an empty migrated database with no seed script, and MARESTAIL_PERF_DB_ROWS overrides it for one run. Goldens live in the marestail-perf-pgdata volume, cached by repo, image, schema (the files matching [perf.db] migrations, else the commit), seed and row count. Before every --db sample the tree's container is replaced by a fresh Postgres on a cp --reflink=always copy of the golden, so the Docker data root must be on a reflink-capable filesystem such as btrfs or XFS. For scale, on btrfs with postgres:16, a golden with one 50M-row table took about 140 s to build and 6.9 GB of disk, and a reset about 1.7 s (tools/perf-db-scale.py measures it). A golden build refuses to start when the disk estimate exceeds the free space. marestail perf db golden --tree <tree> builds in the background and status shows progress, url --tree <tree> prints the URL to start a tree's app against (it stays fixed across resets), prune deletes this repo's goldens the current run does not need, and down removes the containers but keeps the volume. The password is generated into ~/.config/marestail/perf-db.json.

[perf] keys: enabled (default true: perf runs in every pipeline unless set to false), threshold_percent (10), min_runs (10), values_per_sample (200), p95_min_values (200), min_change ({ ms = 1 }), bootstrap (500), control (false), sample_timeout (600 seconds per sample, excluding the reset), setup (a command run in each extra worktree, e.g. uv sync). [perf.db] keys: migrate, url_env (DATABASE_URL), rows (50000000), migrations ([]), skip_tables (the usual migration bookkeeping tables), image, port (55432 for head, baseline +1, pre-marestail +2), min_free_gb (50).

Writing tasks

One task is one vertical slice: a user-visible outcome, thin, through every layer it needs. marestail install drops tasks/README.md into the repo with the guidance; the critic bounces a spec that delivers a layer instead of a slice unless the task declares itself a refactor.

Adapting for new languages

Copy the shape, not the tools. Per-language gates live in marestail/gates/; a new language is one file per gate plus a section in marestail.toml.

Also there is a skill in the repo (.claude/skills/add-language/ and .agents/skills/add-language/; Grok scans both) which should make this process relatively (?) trivial.

Inspo

Inspiration for this came from this brilliant interview with uncle bob, would recommend it if you're interested in ideas around delivering quality software in the age of AI! https://www.youtube.com/watch?v=zcLPGC-tvgk&t=1s

About

Agent harness that makes coding agents loop on a gauntlet of deterministic gates (coverage, mutation, complexity, dead code, Sonar and more)

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages