Skip to content

feat(herdr-pond): herdr plugin with sync-on-idle and a read-only session desk - #312

Open
tenequm wants to merge 44 commits into
mainfrom
feat/herdr-pond-desk
Open

tenequm wants to merge 44 commits into
mainfrom
feat/herdr-pond-desk

Conversation

@tenequm

@tenequm tenequm commented Sep 25, 2026 •

Copy link
Copy Markdown
Owner

PR2 of the herdr-pond v1 plan (sections 4-6): a herdr plugin, packages/herdr-pond, with two features. Needs #311's pond serve (/v1/x/sql, --socket) at runtime; there is no compile-time dependency, and the crate does not depend on the pond crate.

  1. Sync-on-idle. When a herdr agent goes idle, its session is synced into the pond store seconds later.
  2. The desk. One herdr overlay lists recent sessions across all harnesses and machines, searches message content, and reads transcripts from the store through pond serve on an owner-only Unix socket. It never reads a harness file.

Changes

  • Workspace: new member packages/herdr-pond (publish = false; release-plz release = false). default-members = ["packages/pond"] keeps bare root cargo and the dist builds pond-only. It has its own moon tasks and is added to the CI gate list.

  • Plugin: herdr-plugin.toml declares plugin pond with:

    • a pond.desk action and an overlay pane
    • a pane.agent_status_changed event hook
    • a startup hook

    bin/herdr-pond is a committed relative symlink to target/release/herdr-pond, for herdr plugin link dev installs.

  • types.rs: the seam. It holds the Api trait the desk reads through, the pond wire mirrors, and every SQL query the desk runs:

    • each query carries an explicit LIMIT
    • JSON getters appear only in page-scoped queries
    • the pager seeks on (timestamp, message_id)
  • hook.rs (sync-on-idle):

    • The hook exits in milliseconds, with no runtime.
    • It hands off to a detached per-adapter worker behind a flock, with stdio redirected so no herdr command slot stays pinned.
    • Trailing-edge coalescing means the last idle event always produces a pond sync <adapter> -q. The sync waits on the per-host sync lock rather than skipping.
    • Enabled by default; sync_on_idle = false in the plugin config turns it off.
  • daemon.rs:

    • One pond serve --socket <state>/serve/<hash>/owner.sock per herdr server, supervised by a detached owner holding a flock. POND_HOST/POND_PORT are stripped from its env, since clap would still parse a malformed value.
    • It publishes a token-stamped endpoint record ({socket, token, pid}) when it spawns the serve (desks treat it as absent until the socket answers a capability probe), then warms the cold FTS path with its full 300s budget. A record naming any socket but the owner's is ignored.
    • It tears down when the herdr server is gone for more than 10s (long enough to survive a live handoff) or when the owner is signalled. It never adopts an endpoint no owner supervises: a live serve left by a killed owner holds pond's socket lock, so the next owner stops it (SIGTERM, then SIGKILL after a grace period) before starting a fresh one. It is signalled only while alive and when it answers on its socket or its command line (/proc on Linux, ps elsewhere) passes --socket <that path> as a whole argument; any other pid is left alone.
    • The serve never passes --with-sync, and the plugin sends only reads.
  • api.rs / serve.rs:

    • An HTTP-over-Unix-socket client with deadlines on every call and a distinct error per failure shape (pond envelope, axum rejection, old pond, timeout, unreachable).
    • Endpoint resolution is lazy: the daemon's endpoint if live, else a desk-owned fallback serve on its own socket, killed on exit. A failed resolution is shared by every queued caller for 1s, so one dead endpoint costs one spawn, not one per lane. Socket paths are checked against the sun_path limit before spawning.
    • Resolution runs in an api-owned task, so a cancelled request never kills a half-open serve.
    • Only a refused or missing socket triggers failover; a timeout is reported as an error.
    • A pond that predates --socket is detected by its clap exit code, and the desk shows the upgrade hint.
    • herdr CLI calls are bounded at 3s.
    • A capability probe (SELECT 1 over /v1/x/sql) detects a pond that is too old.
  • desk/ (ratatui):

    • Pure on_event/apply reducers.
    • Debounced, generation- and epoch-tagged request lanes.
    • A 14-day project-scoped listing with all-projects and all-time toggles. Titles, counts and origin hosts hydrate as three narrow concurrent page-scoped queries, only for rows near the selection, each rendering as it lands.
    • An owner-only disk cache (desk-cache.json) paints the last listing and known titles/hosts/counts before pond answers; only what is missing or stale is fetched.
    • Machine column: the session's origin host as a short name, this for the local machine, local? for unstamped rows.
    • Typed FTS search over the whole store by default (p narrows to this project, t to 14 days) that distinguishes "filters excluded everything" from zero matches.
    • Preview, and a pre-wrapped pager labeled conversation-only.
    • Errors show as a compact toast; the full text of every one goes to desk.log.
    • Jump to live agent panes.
    • Transcript hygiene (ANSI, tabs, CR).
  • README.md: prerequisites, build and link, the keybinding, config, and what the serve is and is not.

Testing

  • moon run herdr-pond:format herdr-pond:lint herdr-pond:test: green, 142 tests, repeated runs with no flakes (a spawn-time ETXTBSY race in the test helper is fixed).
    • HTTP client tests run against a real Unix-socket fake pond fed the golden /v1/x/sql and /v1/search bodies.
    • Hook, worker and daemon tests use fake pond/herdr scripts in a sandbox: pipes reach EOF while a long-running child continues, coalescing, one owner under concurrent starts, token-matched endpoint removal, one spawn per failed resolution.
    • Desk tests use a MockApi plus TestBackend frames and paused-time lane tests, covering stale results, debounce, slow-server cancellation, concurrent hydration lanes, count exactness for windowed listings, and cache restore/invalidation.
  • Dogfooded live inside herdr 0.9.1 (herdr plugin link, prefix+slash) against the S3 store, with feat(serve): add POST /v1/x/sql JSON endpoint and --socket #311 (at f3dd32c) + fix(serve): survive clients that disconnect mid-response (ignore SIGPIPE) #316's pond:
    • The startup hook's daemon published its owner-only socket 1.3s after start.
    • Warm listing 0.9s; a reopen paints fully hydrated rows from the disk cache 0.14s after the keypress.
    • Machine column shows this / short origin hosts; whole-corpus search returned 50 matching sessions, hydrated.
    • Sync-on-idle, the live-agent glyph and jump-to-pane verified earlier in the same live session.
    • Orphan recovery: SIGKILLing the owner left its serve holding the socket lock; the next owner stopped it in 0.5s and its fresh serve answered 1.9s later (re-verified at c64af0f).
  • Not yet verified: daemon behavior across a herdr server handoff, and macOS.

Release note

New herdr plugin (packages/herdr-pond, built from source): syncs each agent session into pond when the agent goes idle, and adds a session desk overlay to list, search and read sessions from every machine. Needs a pond with /v1/x/sql and --socket.

Workspace member (default-members keeps root builds pond-only), manifest,
launcher symlink, moon/CI/release-plz wiring, the Api seam with every desk
SQL query, golden /v1/x/sql and /v1/search bodies, and a fake pond server.
The desk overlay per plan section 6: a hand-built current-thread runtime
owning the terminal (restored on every return path), pure `on_event`/`apply`
reducers emitting effects, request lanes with generation + view epoch +
target identity, the 14-day listing with cached all-projects/all-time
toggles and page-scoped hydration, debounced search and preview, a
pre-wrapped keyset pager, live-row jump, and transcript hygiene.
…pers

The `hook` leg maps an idle/done agent to its pond adapter, stamps
`pending.<adapter>` and detaches a per-adapter worker (setsid, all stdio
to `sync.log`) unless one holds the worker flock, then exits 0 in
milliseconds. The worker consumes the stamp before each
`pond sync <adapter> -q` (no `--no-wait`, so a busy store lock waits
instead of dropping the event) and re-checks it after releasing the
flock, so no idle event is lost. `open` focuses the workspace's existing
desk or opens one; config.toml gates sync and names `pond`.
…ient

`serve-daemon` (startup hook) adopts a live endpoint or detaches an
`--owner` that holds a per-herdr-server flock for life, runs
`pond serve --host 127.0.0.1 --port 0 --port-file`, publishes
`{port, pid, token, pond_version}` after the `/v1/x/sql` capability
probe, warms the 14-day listing and one FTS search, and tears the serve
down (SIGTERM, 10s, SIGKILL) when herdr's socket stops answering; the
endpoint is removed only by the token that wrote it.

`HttpApi` implements the desk's `Api`: endpoint resolution is lazy, a
dead serve fails over once (daemon record, else a desk-owned fallback
serve killed on drop), every SQL request passes its own LIMIT as
`limit`, and responses map to pond envelope / plain rejection / old
pond / unreachable errors. `tui` runs the desk and jumps with
`herdr agent focus` only after the terminal is restored.
Shared helpers replace repeated blocks: SqlResponse::into_rows,
SqlRequest::new / SearchRequest::new().within() owning protocol_version and
the from_date format, ListingScope::recent, serve::live_endpoint,
Config::pond, config::log_stdio and ensure_parent, one runtime builder and
usage string in main.rs, and the test fixtures (fake serve script, dead
port, endpoint, sandbox origin/lines) in fake_pond.

Fields nothing reads are gone: Endpoint.pid and pond_version (and the extra
pond --version spawn), LiveAgent.agent, and the SQL/search response fields
the desk never shows.
Owner daemon: SIGTERM/SIGINT/SIGHUP run the same teardown as herdr leaving;
a live endpoint with a free owner lock is an unsupervised serve, so a fresh
owner replaces its record and never signals it; herdr counts as gone only
after misses spanning a live handoff; the capability probe is retried while
a fresh serve settles and waits past its server budget; the warm-up search
gets the rest of the warm-up budget instead of the SQL deadline; daemon.log
is capped on every liveness check, not only on open.

Desk: serve resolution runs in a task the api owns, so a lane abort no
longer kills a half-open fallback serve; only a refused connection fails
over, a timeout is reported as an error; a success re-arms failover; serve
teardown runs off the async thread, and the api is dropped after the
terminal is restored. A pond that rejects --port-file (clap exit 2) is
reported as too old in the desk and in daemon.log.
A refresh joins requests already in flight for the same scope or query
instead of restarting them, and toggling back to a cached scope lets the
in-flight listing land in the cache. Previews are clipped and sanitized
once when cached (the cache is bounded), the spinner redraws only where it
shows, and a resize re-wraps the pager once at the next draw. Every herdr
CLI call is killed after 3s, and desk exit waits at most 500ms for
blocking work.
…erequisites

The README names the pond change the plugin needs (#311) and says where
the too-old message appears; module docs drop the ambiguous plan section
refs in favor of one pointer in the crate doc; the all-time loading line
and type docs drop internal jargon; the desk's layout breakpoints are named
constants; moon's compiled sources no longer list the manifest or the bin
symlink, neither of which affects the build.
…rejected --port-file

A retiring fallback removed the shared desk.<pid>.port on drop, which could
delete its successor's freshly published file; each spawn now owns
desk.<pid>.<n>.port. clap exits 2 for any usage error, so PondTooOld is now
reported only when this child's log output names --port-file; other early
exits name the log.
resize decided load_more against the old wrap once the rewrap moved to draw
time, so widening could leave the pager short with no fetch until a key.
relayout now returns the post-rewrap load_more, and the loop performs it
before drawing.
The sticky failed_over flag was cleared only by a success, so a failover
whose retry timed out left every later refusal an error without
re-resolving. Each call now retries once on a refused connection.
A process left behind by herdr could hold the pipes open and block the
reader joins forever; the drain now waits only out the call deadline and
abandons a stuck reader. A try_wait error kills and reaps the child.
@tenequm
tenequm marked this pull request as ready for review September 25, 2026 08:26
pond replaces `serve --port-file` with `serve --socket` (#311). The daemon
and the desk fallback now spawn `pond serve --socket` at
serve/<hash>/owner.sock and desk.<pid>.<n>.sock, never --host/--port, with
POND_HOST/POND_PORT stripped (clap counts env values as given, so an
inherited one would conflict with --socket). Readiness is the socket
existing and the SELECT 1 probe passing over it, within the same 180s
deadline; a pond rejecting --socket with exit 2 is PondTooOld.

The client is reqwest's ClientBuilder::unix_socket (no new dependency or
feature), one client per socket, Host: localhost. The endpoint record is
{socket, token}. A missing or refusing socket is is_connect and fails over;
timeouts stay request errors. FakePond now serves on a UnixListener and the
fake pond scripts answer by symlinking --socket to it.
…gin hosts and whole-corpus search

The one-shot hydration read the wide options and search_text columns for
every row of the page's sessions (2.2-6.8s on the live store). It is now
three page-scoped, LIMITed queries on their own lanes, running
concurrently and rendered as each lands:

- titles: first non-empty user message with search_text kept out of WHERE
  (late materialization reads it only for user rows)
- stats: narrow COUNT(*) + MIN(timestamp), the whole-session count and start
- hosts: the JSON getter read only at each session's first timestamp, so
  the machine column is the session's origin host

The listing also returns its in-window COUNT(*) and MIN(timestamp). An
in-window count is shown as the whole-session count only when the window
provably reaches the session's start (all-time listing, or a start already
known), since a session resumed inside the window has an in-window first
row too.

The desk keeps titles, hosts, starts, counts (per last_ts) and the last
listing per scope in desk-cache.json in the plugin state dir (bounded,
atomic write, unreadable = empty and logged to desk.log), paints it before
pond answers, and hydrates only what the cache lacks.

The machine column shows the short host name, `this` for this machine and
fits the widest name in view. Typed search covers the whole corpus by
default; p / t narrow it to this project / the last 14 days, and the
header names the active scope.
…kets, socket-path limit

A failed `pond serve` resolution now answers every caller queued behind it
(and any caller within a second) with its error, instead of each one
spawning another serve and paying another store open. A socket that
refused is used again when `connect` just probed it live, since a
successor owner reuses `owner.sock`. A socket path past the platform's
sun_path limit is refused before pond is spawned, naming the state dir,
rather than failing at bind after a full store open.
…ve env

pond's `--socket` ignores those env vars (#311), so the strip guarded
against nothing.
Correctness:
- Stats carry the `MAX(timestamp)` they counted up to, so a count read
  before a newer listing no longer shows or persists as current.
- `desk-cache.json` is written owner-only (it holds prompt titles, project
  paths and hosts), and a cache of another format version is logged to
  desk.log instead of dropped silently.
- The fatal error screen keys on the failed scope having no listing, not
  on the whole cache being empty.
- A failed hydration un-asks its ids, so their rows are asked again.
- Toggling to a listing restored from disk shows it at once and refetches
  it; listings fetched during this run still toggle back instantly.

Folded in, since they touch the same paths:
- Titles are asked only where they render (listing rows, pager header);
  the pager header reads the title at draw time.
- The hydration window snaps to page-sized blocks, so holding j sends no
  per-keypress requests.
- Listing lookups compare borrowed scope keys and `machine_width` walks
  the rows directly, instead of cloning per row per frame.
- One `Narrowing` type and `scope_for` serve both the listing and search
  filters; one `on_hydration` handler, `Known::set_host`, and `request`
  deriving its lane and ids from the call.
…emon probe outcome

A script another test's fork still holds open for writing fails to exec
with ETXTBSY; `write_script` now execs it in a dry-run mode until that
succeeds. The settle-retry daemon test asserted on log text that no
longer exists; it now checks the probes were retried and never gave up.
…tion lanes and cache

The data plane is HTTP on an owner-only Unix socket, and pond ignores
POND_HOST/POND_PORT beside `--socket`. 6.2 lists every lane and the
view-lane vs data-lane rule; 6.3 describes the three concurrent
page-scoped hydration lanes, desk-cache.json and whole-corpus search.
Only a refused or missing socket fails over; a timeout is an error.
pond ignores them beside --socket, but clap still parses POND_PORT as a u16 when it is set, so a malformed ambient value (a Kubernetes service link such as tcp://10.0.0.1:9797) fails serve with a usage error that never names --socket.
…itle stands

Advancing last_ts from stats invalidated Title::Missing, flipping a no-user-message session back to '...' with no re-ask. last_ts is fed by listings only again; a count holds while its own last_ts is at or past the listing's, and neither a listing row nor a stats read replaces a count that saw newer activity.
open() cannot ask for the title while the titles lane is busy, and if that request fails the header stayed '...'. A successful page reply now re-runs hydration; the hydration error path still does not, since the resolver answers a failed resolution from cache and would loop.
… desk's 25s

On a cold store the 14-day listing takes over 25s, so warm-up failed with pond's timeout and never warmed the serve, and the first desk open, cold too, timed out as well. The listing now sends the whole budget as timeout_seconds (capped at pond's 600) with the matching client deadline; the search keeps what is left.
A failed titles/stats/hosts lookup no longer pops a toast (the user asked for none of it): it becomes an Effect::Log line in desk.log and its ids are still un-asked, so a later move retries. Toasts show a pond envelope error only up to its first ';' or '. ' - the rest is agent-directed recovery advice - and are capped at three lines with an ellipsis. The fatal screen keeps the full error text.
…down

pond serve --socket (#311 at 33a232c) holds a lifetime lock on <socket>.lock and never removes it. Fallback sockets are unique per spawn, so each one left a desk.<pid>.<n>.sock.lock behind for good; a fallback's teardown now removes it once the child has exited. The owner's owner.sock.lock is shared by successive owners and stays. The pre-spawn socket unlink stays too: pond clears a dead serve's socket itself, but an unsupervised orphan's still answers and would pass the fresh child's readiness probe.
…ng its own

pond serve's lifetime lock on owner.sock.lock refuses a fresh serve while a dead owner's orphan still runs, so the owner could no longer replace it. The endpoint record now carries the serve's pid; an owner finding a live record with no owner SIGTERMs that pid, waits the grace period, then SIGKILLs, and only then unlinks the socket and spawns. Only a record whose socket just answered the probe is signalled (the lock makes the answering process the recorded child), and on Linux /proc/<pid>/cmdline must also name the socket. A record without a pid reads as absent; a termination that fails is logged and the owner exits. The fake pond now takes the lock like pond, so the orphan test fails without the termination.
…ing orphans

The owner now writes `{socket, token, pid}` right after spawning, so the
record always names the process holding pond's owner.sock.lock; desks
already treat a record whose socket does not answer as absent. `pond` is
resolved before any orphan is touched, so a working serve is never stopped
for a replacement that cannot spawn. An orphan is stopped when its socket
answers or, on Linux, when its pid is alive with the socket as an argv
element (still opening its store, or wedged); its record is removed once
it is gone. On Linux a zombie or reused pid (argv no longer naming the
socket) counts as gone during the wait and is never SIGKILLed. Off Linux a
non-answering record is never signalled; a fresh serve that then fails to
start names the recorded pid to kill. The warm-up listing sends its budget
unclamped: pond clamps server-side.
…d a stricter readiness

Off Linux the orphan check was a stub that always matched, so a
non-answering record was never stopped there and a refusal hint stood in.
The command line now comes from /proc/<pid>/cmdline on Linux and a bounded
`ps -ww -o args= -p <pid>` elsewhere, and one pure matcher requires the
whole `--socket <path>` argument (never split on whitespace: macOS paths
hold spaces). Non-answering records are handled the same everywhere and
the hint is gone. A record aimed at any socket but this dir's owner.sock
is ignored, so a corrupt one cannot steer a signal. `ready()` checks the
child's exit before probing and accepts an answer only while the child
still runs. Tests: the matcher's edge cases, SIGKILL after a TERM-ignoring
orphan's grace, a record aimed elsewhere, and the macOS-portable orphan
tests no longer gated to Linux.
Toasts are clipped to a first clause and capped at three lines when drawn, which can drop e.g. a 'see <log>' pointer from an unclipped error; desk.log now holds the whole text of every toast.
@tenequm tenequm self-assigned this Sep 25, 2026
@dogebonker

dogebonker commented Oct 5, 2026 •

Copy link
Copy Markdown

@tenequm heads-up: the red build-and-test-linux here (and on #323) is the self-hosted runner, not the code. The log shows No space left on device (os error 28) on /ci-cache/target, then the linker dying with a bus error. main went green later the same day, so clearing /ci-cache and re-running failed jobs should be enough.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants