diff --git a/.gitignore b/.gitignore index bafd54a5..78983498 100644 --- a/.gitignore +++ b/.gitignore @@ -38,6 +38,21 @@ sbom-server.json htmlcov/ .pytest_cache/ +# Local UAT evidence. Keep compact manifests/results/draft issue summaries +# visible for selective review, but never offer raw household media, event +# streams, health snapshots, or container logs for an accidental commit. +/uat-sessions/**/video/ +/uat-sessions/**/clips/ +/uat-sessions/**/logs/ +/uat-sessions/**/ndjson/ +/uat-sessions/**/snapshots/ + +# Generated audit/session outputs. These can contain machine-local paths, +# prompts, or deployment observations; regenerate or curate before publishing. +/AUDIT-REPORT.md +/audit-draft-issues.md +/pi-session-*.html + # Frontend build deps for the dashboard's vendored Tailwind bundle. # package.json + tailwind.config.js are committed; build artifacts and # the npm install state are not. Re-run `npm install && npm run build:css` diff --git a/:memory:.ses b/:memory:.ses new file mode 100644 index 00000000..8e2bcfa4 --- /dev/null +++ b/:memory:.ses @@ -0,0 +1,2 @@ +1789215087076 +80298cc7-4cef-98c0-2eea-5f395ca28db5 diff --git a/AUDIT-REPORT.md b/AUDIT-REPORT.md new file mode 100644 index 00000000..56f6da47 --- /dev/null +++ b/AUDIT-REPORT.md @@ -0,0 +1,1286 @@ +# dotty-stackchan — Project-Wide Audit Report + +## Executive summary + +This audit covered the full `dotty-stackchan` stack — the four server-side containers (xiaozhi-server, dotty-pi, dotty-behaviour, bridge dashboard), the xiaozhi-server drop-in patches, the ASR/TTS providers, the voice LLM providers, the dotty-pi-ext voice tools, the StackChan ESP32-S3 firmware, the top-level scripts, the docs, the monitoring assets, the shipped compose templates, the open GitHub issues, and the open PRs. + +**282 findings** were reported: **4 critical, 24 high, 114 medium, 125 low, 15 info**. By kind: **157 bugs, 47 doc, 44 quality, 19 PR reviews, 15 issue triages**. Four false positives were rejected during verification. + +**Headline risks:** + +1. **A remote-LAN → docker-host-root exploit chain (G9, critical).** Unauthenticated `/xiaozhi/admin/*` routes on a `0.0.0.0:8003` listener + a `/var/run/docker.sock` mount in the same container + host networking on the other three services compose into a single hop from LAN to host root. The individually-reported symptoms (unauth admin routes F16, arbitrary-path play-asset G10, all-root containers G11) are each one link of this chain. +2. **PR #141 is a personal de-safety fork mislabeled as a bug fix (F120–F124, two critical).** It blanks the kid-mode/safety turn suffix, drops all system messages, disables the perception relay, and adds an explicit-sexual-content persona with hardcoded PII to a repo whose default audience is children. **It must be closed, not merged.** +3. **The shipped all-in-one compose is structurally broken (G18, G19, critical/high).** `compose.all-in-one.yml` omits the docker.sock mount the default `PiVoiceLLM` provider requires and the xiaozhi-patches mounts the entire behaviour layer depends on — a fresh deploy following its own Quick start gets a non-functional voice path. +4. **Kid-safety guardrails are thinner than the docs imply.** There is no output content filter on any live path (#138 still open); `memory_lookup` can surface unreviewed `person_pending` minor facts (F251); and the dashboard `say`/`start-story` ingresses bypass filtering entirely (F186). + +## How this was produced + +Multi-agent adversarial audit of the whole repository: parallel area specialists, **3 deepen rounds** to chase compound/cross-file issues, **3-vote adversarial verification** (3 independent lenses per finding), and a completeness critic that swept files never targeted as primary audit areas (monitoring assets, compose templates, the emoji→firmware contract, the vendored `firmware/server/` tree). + +**Caps hit:** verification was capped at **50 of 172 qualifying findings**; the lowest-severity overflow is reported **unverified**. Critical/high findings were prioritized for the 3-lens treatment. Where a finding shows "confirmed N/3 lenses" it was multi-vote verified; "confirmed 1/1 lenses" is a single completeness-critic confirmation; "unverified — over cap" / "low-sev, unverified" means it was reported but not put through the multi-vote gate. + +--- + +## gap: Container privilege/trust boundary (docker.sock + host networking) + +### [CRITICAL] Remote-LAN → docker-host-root exploit chain +**File:** `docker-compose.yml.template` (compose 43-60, 596) + `custom-providers/xiaozhi-patches/http_server.py` (596, 620-672) +**Issue:** The four-container trust boundary collapses into one exploitable chain. xiaozhi-server binds `0.0.0.0:8003` and registers all 11 `/xiaozhi/admin/*` routes with no auth; the same container mounts `/var/run/docker.sock` + `/usr/bin/docker` (the compose comment itself admits this gives "effective root on the docker host"); the other three services run `network_mode: host` with no segmentation; none of the four Dockerfiles set `USER`, so every container is root. An attacker on the LAN/Tailnet has a confirmed write primitive into the socket-holding root container. +**Fix:** Break the chain at >1 link: (1) replace the raw docker.sock with a least-privilege `docker-socket-proxy` exposing only `exec`; (2) add a shared-secret bearer-token aiohttp middleware over `/xiaozhi/admin/*`; (3) bind admin HTTP to `127.0.0.1` and reach it only from sibling containers on loopback; (4) add non-root `USER` to all four Dockerfiles. Do at least the socket-proxy + route auth. +**Verification:** confirmed 1/3 lenses (completeness critic; compound finding assembled from confirmed symptoms). + +### [HIGH] play-asset accepts arbitrary absolute filesystem path with no allowlist +**File:** `custom-providers/xiaozhi-patches/http_server.py:376-411` +**Issue:** `/xiaozhi/admin/play-asset` reads `asset` as a raw absolute path; the only validation is `os.path.exists`, then it is handed to `AudioSegment.from_file` (ffmpeg). The sibling `songs` handler hard-codes a base dir + extension allowlist — play-asset has neither. An unauthenticated LAN caller can point ffmpeg at any readable file (existence-probe via 404-vs-200, exercise libav demuxer CVEs) inside the docker.sock-holding container. +**Fix:** `os.path.realpath(asset)` must start with the canonical songs base before any decode; enforce the `{opus,ogg,wav,mp3}` allowlist; prefer accepting only a basename joined onto the fixed base. +**Verification:** confirmed 1/1 lenses. + +### [HIGH] All four containers run as root (no USER directive) +**File:** `Dockerfile`, `dotty-pi/Dockerfile`, `dotty-behaviour/Dockerfile`, `bridge/Dockerfile` +**Issue:** None of the four images declare `USER`. For xiaozhi-server the root process is the one mounting docker.sock, so any code-exec bug there is already host root. The FastAPI/uvicorn and pi workloads do not need root. +**Fix:** Add a non-root `USER` to each Dockerfile (create app user, chown state dirs). Pair with `no-new-privileges:true` and read-only rootfs where feasible; run the socket-proxy so root-in-container is no longer root-on-host. +**Verification:** confirmed 1/1 lenses. + +### [MEDIUM] Host networking removes network segmentation +**File:** `dotty-pi/docker-compose.yml:17`, `dotty-behaviour/docker-compose.yml:23`, `bridge/docker-compose.yml:31` +**Issue:** Three of four services use `network_mode: host`, so their listeners bind every interface with no bridge isolation and a compromise of any one has unrestricted loopback access to every host service. This is the segmentation half of why the chain finding is critical rather than contained. +**Fix:** Where loopback-to-llama-swap is the only reason for host mode, use a shared user-defined bridge network instead; bind behaviour/bridge to `127.0.0.1` if callers are host-local; at minimum document the host-networking decision as an explicitly accepted risk. +**Verification:** low-sev, unverified. + +--- + +## gap: Shipped compose templates & .env.example (config-drift surface) + +### [CRITICAL] compose.all-in-one.yml omits the docker.sock + docker CLI mounts PiVoiceLLM requires +**File:** `compose.all-in-one.yml:64-79` +**Issue:** The shipped config defaults to `selected_module.LLM: PiVoiceLLM`, which reaches the brain via `docker exec -i dotty-pi pi`. That needs the host docker socket + docker CLI bind-mounted (provided by `docker-compose.yml.template:59-60`). compose.all-in-one.yml mounts the provider dir but NOT the socket or binary — so a new user following the file's own Quick start with the default config gets a container that cannot exec into dotty-pi, and every voice turn fails. The file's own header even says PiVoiceLLM "requires the host docker socket" but never mounts it. +**Fix:** Add the two PiVoiceLLM mounts to compose.all-in-one.yml. Or, if all-in-one is meant to be OpenAICompat-only, change the shipped `selected_module` default and say so in the header. +**Verification:** confirmed 1/1 lenses. + +### [HIGH] compose.all-in-one.yml is missing the xiaozhi-patches mounts +**File:** `compose.all-in-one.yml:64-79` +**Issue:** The canonical template mounts `portal_bridge.py`, `websocket_server.py`, `http_server.py`, and `textMessageHandlerRegistry.py` (admin routes, active_connections registry, perception relay). all-in-one mounts none of them, so a user deploying with this file gets upstream xiaozhi-server with no admin portal and no perception relay — the entire dotty-behaviour layer is inert (consumers fire into 404s). Also missing vs the template: openai_compat, personas, receiveAudioHandle.py, dances.py, textUtils.py, songs, and the kid/smart state mount. +**Fix:** Bring the volume list to parity with `docker-compose.yml.template`, or generate the all-in-one from the template so it can't drift. +**Verification:** confirmed 1/1 lenses. + +### [MEDIUM] compose.all-in-one.yml never sets VISION_BRIDGE_URL/BRIDGE_URL +**File:** `compose.all-in-one.yml:56-58` +**Issue:** The relay resolves its target from `BRIDGE_URL`/`VISION_BRIDGE_URL`; with neither set it logs "dropping perception event" and returns. all-in-one's environment block sets only `TZ`, so even if the relay handler were mounted, every perception event would be dropped. +**Fix:** Add `VISION_BRIDGE_URL=http://:8090` (document `BRIDGE_URL` as its alias) alongside mounting the relay handler. +**Verification:** confirmed 1/1 lenses. + +### [MEDIUM] .env.example omits load-bearing xiaozhi-server vars +**File:** `.env.example:4-21` +**Issue:** Several vars the template declares and the code reads are absent: `VISION_BRIDGE_URL`, `DOTTY_KID_MODE_STATE`, `DOTTY_SMART_MODE_STATE`, `DOTTY_STATE_DIR`, `DOTTY_PI_CONTAINER`, `DOTTY_PI_EXTRA_FLAGS`. The state-dir/kid-mode ones desync the firmware LED pips from the dashboard if left at the dead default. +**Fix:** Add the missing vars with the same explanatory comments already in `docker-compose.yml.template`. +**Verification:** low-sev, unverified. + +### [LOW] Dead '/root/zeroclaw-bridge default' fallback comment in template +**File:** `docker-compose.yml.template:38-40` +**Issue:** The comment warns the reader "falls through to the dead /root/zeroclaw-bridge default"; the real defaults are `/var/lib/dotty-bridge/state/{kid,smart}-mode`. ZeroClaw was retired in #36. +**Fix:** Rewrite the comment to state the real default; grep for other zeroclaw/RPi residue. +**Verification:** low-sev, unverified. + +### [LOW] .env.example documents many dotty-behaviour vars but not SOUND_TURN_YAW_DEG +**File:** `.env.example:23-153` +**Issue:** The file mixes xiaozhi-server and dotty-behaviour vars; `SOUND_TURN_YAW_DEG` (read at `dotty-behaviour/config.py:105`) is absent while sibling perception vars are documented — arbitrary coverage that obscures which container consumes each var. +**Fix:** Split into a per-service `.env.example`, or section-label each block by consuming container; add `SOUND_TURN_YAW_DEG`. +**Verification:** low-sev, unverified. + +--- + +## PR #141 + +### [CRITICAL] #141 is a personal de-safety fork, not a bug fix — close +**File:** PR #141 +**Issue:** The title claims to fix kid-mode constraints and system-prompt leakage; the diff does the opposite — it (a) blanks the safety/format suffix (`_TURN_SUFFIX = ""`), (b) drops all system messages, (c) comments out the perception relay + room_view capture, (d) adds `personas/sweetheart.md`, an explicit-sexual-content persona with the author's hardcoded PII. None is gated, tested, or documented; all of it contradicts the repo's stated kid-safe scope (ages 4-8). +**Fix:** **Close the PR.** Extract only the legitimate piper emoji-strip into a standalone PR. Ask the author to keep the persona, prompt-stripping, and perception-disabling on a private fork. +**Verification:** confirmed 3/3 lenses. + +### [CRITICAL] _TURN_SUFFIX blanked — disables English-only, emoji-prefix, length AND kid-safety constraints +**File:** `custom-providers/openai_compat/openai_compat.py:28` +**Issue:** `_TURN_SUFFIX = build_turn_suffix(KID_MODE)` is replaced with `_TURN_SUFFIX = ""`, stripping the full HARD-CONSTRAINTS block: English-only, the mandatory single-emoji prefix contract, length/no-Markdown rules, AND the entire kid-mode topic blocklist plus jailbreak resistance. `build_turn_suffix` is now dead while still imported. +**Fix:** Revert to `_TURN_SUFFIX = build_turn_suffix(KID_MODE)`. If the suffix was leaking into TTS, fix the echo, do not delete the contract. +**Verification:** confirmed 3/3 lenses. + +### [HIGH] All system messages dropped from dialogue +**File:** `custom-providers/openai_compat/openai_compat.py:108-109` +**Issue:** `if role == "system": continue` skips every system message, removing the top-level `.config.yaml` `prompt:` safety/emoji layer (one of the two load-bearing layers per CLAUDE.md). The title calls this fixing "system prompt leakage"; it deletes the system prompt. +**Fix:** Remove the block. If one system message was being spoken, filter that source specifically. +**Verification:** confirmed 3/3 lenses. + +### [HIGH] Perception event relay and room_view capture disabled by commenting out _spawn calls +**File:** `custom-providers/xiaozhi-patches/textMessageHandlerRegistry.py:153-156, 193` +**Issue:** Both the room_view background capture and the perception-event POST to dotty-behaviour are commented out, silently disabling the entire perception bus relay and the room_view VLM identity feature, leaving dead scaffolding. Looks like personal-fork debugging left in. +**Fix:** Revert both comment-outs. Disable a misbehaving consumer in dotty-behaviour config instead of severing the relay. +**Verification:** confirmed 2/3 lenses. + +### [HIGH] sweetheart.md adds NSFW persona with hardcoded PII to a kid-safe public repo +**File:** `personas/sweetheart.md:1-23` +**Issue:** New persona instructs "flirty"/"sexually explicit stories with detail" and hardcodes the author's real profile. Contradicts `personas/default.md` (ages 4-8, no romantic/adult topics), omits the mandatory emoji-prefix contract, and uses gemma chat-template tags that aren't this repo's persona format. PII + explicit content in a public repo is a privacy and content-safety regression. +**Fix:** Do not merge; remove the file. A personal/adult persona belongs on a private fork. +**Verification:** confirmed 3/3 lenses. + +### [MEDIUM] Full conversation (system+user messages) logged at info level — privacy/PII regression +**File:** `custom-providers/openai_compat/openai_compat.py:150` +**Issue:** `logger.info(f"Sending messages to LLM: {json.dumps(messages,...)}")` dumps persona, system prompt, and every user utterance to persistent docker logs on every turn — including anything a child says. Reads as leftover debugging. +**Fix:** Remove the line, or gate behind a debug flag at DEBUG level with contents redacted. +**Verification:** confirmed 2/3 lenses. + +### [MEDIUM] dotty-behaviour env toggles added with no code/doc reference +**File:** `dotty-behaviour/docker-compose.yml:49-50` +**Issue:** `SOUND_TURN_YAW_DEG=20` / `WAKE_TURN_YAW_DEG=20` added with no consumer change, no docs, no indication they're honored. Dead config if unread; incomplete if read. Also muddies the PR scope. +**Fix:** Confirm the consumers read them; if so ship the consumer code + doc and split into its own PR. If not, drop the lines. +**Verification:** unverified — over cap. + +### [LOW] piper emoji/non-speakable strip is the one reasonable change but is bundled and untested +**File:** `custom-providers/piper_local/piper_local.py:137-140` +**Issue:** Stripping emoji/non-Latin codepoints before Piper synth is defensible, but the regex also strips CJK with no log, has no test, and is bundled into an un-mergeable PR. An all-emoji utterance produces silence with no log. +**Fix:** If salvaged into a standalone PR: add tests for emoji-only/CJK inputs and log at debug when a chunk is fully stripped. +**Verification:** unverified — over cap. + +--- + +## xiaozhi-server drop-in patches (http/ws/ota/admin routes) + +### [HIGH] All /xiaozhi/admin/* routes are completely unauthenticated +**File:** `custom-providers/xiaozhi-patches/http_server.py:620-672` (handlers 34-589) +**Issue:** Every admin endpoint (inject-text, abort, set-head-angles, set-state, set-toggle, set-face-identified, take-photo, play-asset, songs, say, devices) is registered with no auth check, while OTA and WS paths guard themselves. Anyone reaching port 8003 can make the robot speak arbitrary text, move, change state, trigger the camera, and (play-asset) read any absolute path. inject-text is a remote prompt-injection vector into the pi agent. +**Fix:** Gate the routes behind a shared secret (Bearer / `X-Admin-Token` aiohttp middleware, 401 otherwise). At minimum confine play-asset to an allow-listed base via realpath+startswith. +**Verification:** confirmed 3/3 lenses. + +### [MEDIUM] face_detected during talk state silently drops the bridge perception relay +**File:** `custom-providers/xiaozhi-patches/textMessageHandlerRegistry.py:134-138` +**Issue:** The `face_detected`/`talk` branch does an early `return` that also skips the relay POST at the bottom of `handle()`. The talk-gate's intent is only to suppress re-triggering room_view capture, not forwarding to dotty-behaviour, so consumers desync for the duration of a conversation. +**Fix:** Wrap only the capture kick-off in `if cstate != 'talk':` and fall through to the relay. +**Verification:** confirmed 2/3 lenses. + +### [MEDIUM] face_lost does not reset _room_description_in_flight, can wedge room_view capture +**File:** `custom-providers/xiaozhi-patches/textMessageHandlerRegistry.py:92-102, 139-141` +**Issue:** `face_lost` clears the room-view caches but not `_room_description_in_flight`. If a capture task crashed without clearing the flag, every subsequent `face_detected` refuses to re-capture for the life of the connection. +**Fix:** Also set `conn._room_description_in_flight = False` in the `face_lost` branch. +**Verification:** unverified — over cap. + +### [MEDIUM] WebSocket close uses removed `closed` attribute on modern websockets +**File:** `custom-providers/xiaozhi-patches/websocket_server.py:143-151` +**Issue:** The finally block checks `hasattr(websocket,'closed')` — gone in the `ServerConnection` API this code targets — so the branch is dead and the guard is largely pointless. +**Fix:** Drop the `closed` branch; rely on `getattr(websocket,'state',None)` or a try/except `await websocket.close()`. +**Verification:** unverified — over cap. + +### [MEDIUM] Query-param device-id can bypass token auth via the allowlist short-circuit +**File:** `custom-providers/xiaozhi-patches/websocket_server.py:84-108` +**Issue:** Query params are injected into `request.headers` and read back by `_handle_auth`; a `?device-id=` is trusted by the `allowed_devices` whitelist the same as a header, letting a client skip token verification. Headers mutation semantics are also version-fragile. +**Fix:** Treat query-param identity as untrusted — do not let a query-string device-id satisfy the allowlist bypass; use a separate dict rather than mutating `request.headers`. +**Verification:** unverified — over cap. + +### [MEDIUM] OTA _is_higher_version treats pre-release/build-suffix versions as newer +**File:** `custom-providers/xiaozhi-patches/ota_handler.py:24-43, 315` +**Issue:** `_parse_version` does `re.findall(r'\d+', ver)`, so `1.2.3-rc1 → (1,2,3,1)` compares greater than GA `(1,2,3)`, and `1.2` equals `1.2.0`. A device on GA is offered a stale rc; build-suffix versions invert ordering. Verified: `_is_higher_version('1.2.3-rc1','1.2.3')` returns True. +**Fix:** Use a real semver-ish compare (strip leading `v`, parse the numeric core, rank pre-release suffixes BELOW the release, consider only the first 3 segments). Add unit tests for rc/build cases. +**Verification:** confirmed 3/3 lenses. + +### [MEDIUM] play-asset clobbers conn.client_abort / client_is_speaking, corrupting a concurrent voice turn +**File:** `custom-providers/xiaozhi-patches/http_server.py:452-453, 474` +**Issue:** `_dispatch()` unconditionally sets `conn.client_abort=False` then `conn.client_is_speaking=True`, with a finally that marks idle. play-asset is timer-driven on the "first available device", so a play-asset landing mid-turn cancels a user's barge-in and marks the device idle while chat TTS streams. +**Fix:** Refuse with 409 if the conn is mid-turn; save/restore prior flags, or route admin audio through the TTS-priority queue. +**Verification:** confirmed 3/3 lenses. + +### [MEDIUM] Admin handlers call conn.websocket.send() concurrently with the chat-path writer +**File:** `custom-providers/xiaozhi-patches/http_server.py:146, 194, 242, 288, 338, 456-468` +**Issue:** Fire-and-forget `_spawn(conn.websocket.send(...))` admin sends run concurrently with xiaozhi's own send on the same `ServerConnection`. The websockets library forbids interleaved sends; a `ConcurrencyError` is swallowed by the spawned task and the chat frame can be corrupted. play-asset (hundreds of opus frames) maximizes overlap. +**Fix:** Serialize device-bound writes through a per-conn `asyncio.Lock`; at minimum log the `ConcurrencyError` and gate play-asset behind an idle check. +**Verification:** unverified — over cap. + +### [MEDIUM] _dotty_say mutates conn.sentence_id from a worker thread, racing the consumer and pre-empting chat TTS +**File:** `custom-providers/xiaozhi-patches/http_server.py:547-574, 582` +**Issue:** `_enqueue()` runs on a thread and sets `conn.sentence_id` with no busy-check or lock; the greeter fires on perception events that can coincide with a live conversation, so a `/say` mid-turn overwrites `sentence_id` and the consumer drops the chat turn's remaining sentences. Also no `tts_text_queue` guard. +**Fix:** Refuse `/say` with 409 when mid-turn; guard for `tts_text_queue`; set `sentence_id` on the loop thread. +**Verification:** unverified — over cap. + +### [MEDIUM] play-asset omits the tts 'start' lifecycle frame and uses an unvalidated Opus rate +**File:** `custom-providers/xiaozhi-patches/http_server.py:414, 421-423, 456-461` +**Issue:** Playback never sends `{type:tts,state:start}` (firmware-fragile), and feeds `conn.sample_rate` straight into the Opus encoder without checking it is in `{8000,12000,16000,24000,48000}`; a non-legal rate raises, is swallowed by the broad except, and the endpoint still returns 200. +**Fix:** Emit the full lifecycle; validate/resample the rate; propagate decode failure as 500. +**Verification:** unverified — over cap. + +### [MEDIUM] EventTextMessageHandler relays perception events to a stale BRIDGE_URL +**File:** `custom-providers/xiaozhi-patches/textMessageHandlerRegistry.py:67-78, 178` +**Issue:** Events POST to `{BRIDGE_URL or VISION_BRIDGE_URL}/api/perception/event`, but the bus moved to dotty-behaviour (:8090) in #115. If the deploy sets only a behaviour URL var, or BRIDGE_URL points at the dashboard container (which no longer serves that route), every event 404s and is dropped — silently disabling face_greeter/sound_turner/state gating. +**Fix:** Rename/select the var to point at dotty-behaviour explicitly and update the warning text; verify the deploy compose sets the var the handler reads. +**Verification:** unverified — over cap. + +### [MEDIUM] state_changed/face_detected ordering can corrupt the room_view gate +**File:** `custom-providers/xiaozhi-patches/textMessageHandlerRegistry.py:109-115, 133-138` +**Issue:** The room_view gate reads `conn.current_state`, set from `state_changed`. If the firmware emits `face_detected` before the IDLE→TALK `state_changed` it acted on, the gate sees `idle` and fires a capture for a mid-talk flicker — the "stacked Hi NAME" failure the comment claims to prevent. +**Fix:** Gate on a positive "room_view already captured this session" flag rather than `current_state`, or have firmware stamp `face_detected` with its state. +**Verification:** unverified — over cap. + +### [LOW] OTAHandler.handle_post finally-block can swallow a secondary exception +**File:** `custom-providers/xiaozhi-patches/ota_handler.py:359-361` +**Issue:** Unlike `handle_download`, `handle_post`'s finally calls `_add_cors_headers` with no try/except and a bare `return` that masks a propagating exception; an edge case 500s the device with non-JSON. +**Fix:** Wrap `_add_cors_headers` in try/except; initialize `response = None` and short-circuit. +**Verification:** unverified — over cap. + +### [LOW] _dotty_inject_text reads conn.headers without the 'or {}' guard used elsewhere +**File:** `custom-providers/xiaozhi-patches/http_server.py:66` +**Issue:** Line 66 lacks the `(getattr(conn,'headers',{}) or {})` guard every sibling handler uses; if `conn.headers` is None during teardown, inject-text raises after already dispatching chat, returning a misleading 500. +**Fix:** Apply the same guard on line 66. +**Verification:** unverified — over cap. + +### [LOW] OTA mqtt branch ships an empty-password block with no websocket fallback +**File:** `custom-providers/xiaozhi-patches/ota_handler.py:260-279` +**Issue:** When mqtt_gateway is set but the signature key is missing/fails, the handler sends an mqtt block with `password:""` and (being exclusive with the websocket else) no fallback — the device gets a config it cannot authenticate with. +**Fix:** Fall back to the websocket block (or return an error) when the signature is unavailable. +**Verification:** unverified — over cap. + +### [LOW] OTA model-name matching is brittle — updates silently never offered +**File:** `custom-providers/xiaozhi-patches/ota_handler.py:84-90, 200-201, 304` +**Issue:** `files_by_model.get(device_model, [])` requires an exact match between the firmware-reported model and the operator's bin-filename prefix, which rarely align; the device is silently told it's up-to-date and the same log fires for both "unknown model" and "already latest". +**Fix:** Log the parsed model + known keys when candidates is empty; allow a configured fallback bucket or case-insensitive match; distinct log lines for the two cases. +**Verification:** unverified — over cap. + +### [LOW] OTA timezone_offset / firmware_cache_ttl assume numeric YAML +**File:** `custom-providers/xiaozhi-patches/ota_handler.py:62, 226` +**Issue:** `timezone_offset * 60` on a quoted `"8"` yields a 60-copy string in the OTA payload; a non-int ttl makes `int(...)` raise. +**Fix:** Coerce with `int(...)` and validate `firmware_cache_ttl` at construction. +**Verification:** unverified — over cap. + +### [LOW] Firmware bin cache re-scans disk on every OTA POST until a valid bin exists +**File:** `custom-providers/xiaozhi-patches/ota_handler.py:66-72` +**Issue:** The TTL short-circuit requires a non-empty `files_by_model`; with no conforming bins yet that's falsy, so the glob+regex scan runs on every poll and `updated_at` is meaningless. +**Fix:** Track freshness with a separate `_scanned` flag/timestamp independent of emptiness. +**Verification:** unverified — over cap. + +### [LOW] Admin MCP handlers fire-and-forget send() with no guard for a torn-down socket +**File:** `custom-providers/xiaozhi-patches/http_server.py:146, 194, 242, 288, 338` +**Issue:** Between the registry lookup and the spawned send, the WS can close; `send()` on a dead socket raises in the spawned task while the HTTP caller already got `{"ok":true}`. +**Fix:** Check `conn.websocket` state in a small coroutine and catch+log; or await the send before returning. +**Verification:** unverified — over cap. + +### [LOW] set-toggle/set-state have no response correlation — device-side failures invisible; head-angles unclamped +**File:** `custom-providers/xiaozhi-patches/http_server.py:143, 191, 239, 285, 335` +**Issue:** MCP calls use a truncated wall-clock id and never match the device's jsonrpc reply, so a device-side rejection still returns `{"ok":true}`. set-head-angles forwards arbitrary ints (no clamp). +**Fix:** Clamp head-angle yaw/pitch/speed to firmware-legal ranges; optionally correlate by request id with a timeout, or document best-effort. +**Verification:** unverified — over cap. + +### [LOW] play-asset format inference treats bare .opus as Ogg +**File:** `custom-providers/xiaozhi-patches/http_server.py:406-411` +**Issue:** `.opus → format=ogg` fails on raw Opus; unmapped extensions yield `fmt=None` (autodetect) with no distinct log; the songs picker advertises `.opus` so it can silently fail behind an already-sent 200. +**Fix:** Let ffmpeg autodetect for `.opus`; on decode failure return a distinguishable error rather than a swallowed warning. +**Verification:** unverified — over cap. + +### [LOW] OTA mqtt allowlist token logic undocumented coupling +**File:** `custom-providers/xiaozhi-patches/ota_handler.py:283-295` +**Issue:** A whitelisted device gets an empty token and relies on the WS-side allowlist bypass; remove it from the whitelist later and it has no token to fall back on. +**Fix:** Document the coupling, or always issue a token regardless of whitelist membership. +**Verification:** unverified — over cap. + +### [LOW] play-asset resample can feed huge factors to resample_poly; no sample_rate guard +**File:** `custom-providers/xiaozhi-patches/http_server.py:415-420` +**Issue:** Coprime rates give large up/down factors (slow/memory-heavy filter); a 0/None `conn.sample_rate` blows up inside `_decode` and is masked by the broad except. +**Fix:** Validate `conn.sample_rate > 0`; log src/tgt rates on decode failure. +**Verification:** unverified — over cap. + +### [LOW] Admin routes resolve target via next(iter(...)) — nondeterministic device with multiple StackChans +**File:** `custom-providers/xiaozhi-patches/http_server.py:55, 87, 123, 171, 218, 266, 316, 388, 523` +**Issue:** With >1 device and no `device_id`, handlers act on an arbitrary dict entry; abort/set-state can hit the wrong robot. +**Fix:** Require explicit `device_id` (409/400) when multiple devices are connected. +**Verification:** unverified — over cap. + +### [LOW] EventTextMessageHandler attributes events to "unknown" device on header miss +**File:** `custom-providers/xiaozhi-patches/textMessageHandlerRegistry.py:60-65` +**Issue:** A bare try/except leaves `device_id="unknown"` if `conn.headers` is None, so dotty-behaviour accumulates per-device state under a bogus id that never matches the real device — consumers silently never fire. +**Fix:** Resolve device_id like the WS registration does; drop the event with a warning rather than relay under a fabricated id. +**Verification:** unverified — over cap. + +### [QUALITY/LOW] Massive duplication + ms-truncated ids across the MCP admin handlers +**File:** `custom-providers/xiaozhi-patches/http_server.py:101-343` +**Issue:** Five near-identical handlers repeat parse/lookup/envelope/spawn (~250 lines), each using `id = int(time.time()*1000) % 0x7FFFFFFF` which collides within a millisecond. +**Fix:** Extract a `_mcp_call(...)` helper; use a monotonic counter / uuid-derived id. +**Verification:** unverified — over cap. + +### [INFO] play-asset path containment (bridge-callee view) +**File:** `custom-providers/xiaozhi-patches/http_server.py:376-384` +**Issue:** The bridge's own play-song basename-guards, but the xiaozhi callee validates only `os.path.exists`, so the trust boundary is weaker than the bridge layer implies (duplicate of G10, flagged from the bridge scope). +**Fix:** Resolve realpath under the songs base in play-asset. +**Verification:** unverified — over cap. + +### [INFO] _http_response param name does not match the websockets process_request contract +**File:** `custom-providers/xiaozhi-patches/websocket_server.py:157-164` +**Issue:** The param is really a `Request` (accessed as `request_headers.headers`), which misleads readers. +**Fix:** Rename to `request`, access `request.headers`. +**Verification:** unverified — over cap. + +--- + +## Bridge dashboard (FastAPI :8081) + +### [MEDIUM] Dashboard action endpoints always block 8s and report "no reply" — the event producer is dead +**File:** `bridge.py:613-626` (consumed at `bridge/dashboard.py:658-704, 2091-2133`) +**Issue:** `_inject_or_error` subscribes to the bridge event bus and waits 8s for Dotty's reply, but nothing ever enqueues onto `_dashboard_event_listeners` — the producer moved to dotty-pi in #36 and was never re-wired. Every Say/Dance/Mood/Story click blocks the full 8s then renders "Sent — no reply in 8s."; the SSE stream only emits heartbeats. +**Fix:** Either wire a real producer (dotty-pi/xiaozhi POSTs completed turns to a bridge endpoint that calls `put_nowait`), or drop the subscribe/wait and return "Sent" immediately. +**Verification:** confirmed 2/3 lenses. + +### [MEDIUM] Blocking synchronous HTTP inside async dashboard routes stalls the event loop +**File:** `bridge.py:471-504` (`_dotty_behaviour_get`); `bridge/dashboard.py:193-218` (`_fetch_robot_photo`), 1366, 2085, 801, 1925 +**Issue:** The dotty-behaviour getters and `_fetch_robot_photo` do synchronous `requests.get` with no `to_thread`, called from async routes; each uncached call blocks the whole asyncio loop for 1.5-2.0s, freezing all clients. The HTMX 10s poll fans each render into 3-6 such calls. The codebase already uses `to_thread` for device-count/songs, so these are clear omissions. `_build_perception_card_ctx`'s "no I/O" docstring is now false. +**Fix:** Wrap blocking HTTP in `asyncio.to_thread` (or use `httpx.AsyncClient`); fix the stale docstring. +**Verification:** confirmed 3/3 lenses. + +### [MEDIUM] CSRF middleware blocks the localhost /admin/* operator endpoints +**File:** `bridge/csrf.py:33, 69-70, 83`; `bridge.py:800-924` +**Issue:** CSRFMiddleware is app-wide and exempts only `/api/`, `/metrics`, `/health`. The `/admin/*` router (documented as the localhost CLI back-channel) is not exempt and has no CSRF cookie, so curl/script calls are 403'd before reaching `_admin_require_localhost` — the entire documented admin surface is unusable via its intended caller. +**Fix:** Add `/admin/` to `_EXEMPT_PREFIXES` (the router already has its own localhost auth). +**Verification:** confirmed 2/3 lenses. + +### [MEDIUM] SSE /events served under BaseHTTPMiddleware (CSRF) — Starlette buffering hazard +**File:** `bridge/dashboard.py:2091-2133`; `bridge/csrf.py:73-98` +**Issue:** CSRFMiddleware subclasses `BaseHTTPMiddleware`, a documented SSE footgun that buffers chunks and breaks `is_disconnected()`, defeating `X-Accel-Buffering: no`. Today only heartbeats flow, but the channel is structurally degraded once a producer is wired back. +**Fix:** Exempt the SSE path, or migrate CSRFMiddleware to raw ASGI so streaming responses pass through. +**Verification:** unverified — over cap. + +### [MEDIUM] Kid/smart-mode toggles report ok:False while having already persisted+applied the flip +**File:** `bridge.py:630-643, 695-707` +**Issue:** `_dashboard_set_kid_mode` persists + applies the state first, then dispatches the firmware toggle; if the device doesn't ack, it returns `{ok:False}` even though the bridge already flipped. The UI shows "toggle failed", the operator re-toggles, and the state ping-pongs out of sync. +**Fix:** Return `ok:True` with a `device_pushed:false` warning when the bridge write succeeded but the LED push failed (mirror the `/admin/kid-mode` shape). +**Verification:** unverified — over cap. + +### [MEDIUM] network_mode: host binds dashboard + /metrics to 0.0.0.0 with auth unset by default +**File:** `bridge/docker-compose.yml:31, 41-60`; `bridge.py:936`; `bridge/dashboard.py:259-267` +**Issue:** The shipped compose sets no dashboard user/pass/CSRF secret; auth is opt-in and returns immediately when either env var is empty; host networking exposes the full mutation surface and `/metrics` to the whole LAN unauthenticated. +**Fix:** Ship the compose with dashboard user/pass + a persistent CSRF secret from `.env`, and/or bind `127.0.0.1` by default; at minimum document the required auth env. +**Verification:** unverified — over cap. + +### [MEDIUM] Async dashboard routes read + JSON-parse the entire daily convo log synchronously +**File:** `bridge/dashboard.py:377-398, 472-514, 517-555, 1631-1648` +**Issue:** `_stackchan_last_seen`, `alerts_count`, `alerts_detail` read the whole `convo-*.ndjson` and `json.loads` every line with no `to_thread`, from async routes polled every ~10s per tab; a multi-MB log blocks the loop on each poll. +**Fix:** Wrap the read+parse in `asyncio.to_thread`; optionally tail rather than parse the whole day. +**Verification:** unverified — over cap. + +### [MEDIUM] Dashboard /actions/say and /actions/start-story bypass the kid-mode content filter +**File:** `bridge/dashboard.py:706-729, 732-773` +**Issue:** `say()`/`start_story()` sanitise control chars and length but never call `content_filter()` (imported and available), so an operator (or any LAN client when auth is unset) can make Dotty speak arbitrary unfiltered text, and these turns never populate the safety ring. +**Fix:** Run the sanitised text through `content_filter()` when kid-mode is on; record the hit. +**Verification:** unverified — over cap. + +### [QUALITY/MEDIUM] security_watch.py is entirely dead code; only get_recent_cycles() is reachable +**File:** `bridge/security_watch.py:1-630` +**Issue:** Nothing calls `run_security_consumer`, `start_device_timer`, etc.; the perception role moved to dotty-behaviour in #36. ~580 lines run nowhere; `RECENT_CYCLES` is always empty; the docstring still describes firmware tools that "do not exist as of 2026-04-27". +**Fix:** Delete the module + the `_build_security_panel_ctx` route, or repoint the panel at a dotty-behaviour HTTP getter. +**Verification:** unverified — over cap. + +### [LOW] host/footer 'disk' vitals read the container overlay fs, not the host disk +**File:** `bridge/dashboard.py:1614-1620, 1734-1739, 1793-1796` +**Issue:** `shutil.disk_usage('/')` inside the container reports the overlay fs, not the Unraid appdata volume, so the disk gauge is misleading (mem/cpu are host-wide via the shared kernel). +**Fix:** Point at a bind-mounted host path, or relabel as "container disk". +**Verification:** unverified — over cap. + +### [LOW] say/start-story length cap enforced before whitespace collapse +**File:** `bridge/dashboard.py:714-723, 746-753` +**Issue:** The `len > 500` check runs before whitespace collapse, so a pasted message with long whitespace runs but short visible content is wrongly rejected. +**Fix:** Sanitise/collapse first, then length-check the cleaned text. +**Verification:** unverified — over cap. + +### [LOW] play-asset / vision-photo proxy interpolate unencoded device_id into an internal URL +**File:** `bridge/dashboard.py:201, 638, 1924` +**Issue:** `_fetch_robot_photo` builds the behaviour URL with the raw `device_id` path-param (unvalidated, unencoded), allowing query/fragment injection into the internal request. +**Fix:** `urllib.parse.quote(device_id, safe='')` and validate against an id charset, returning 400 on mismatch. +**Verification:** unverified — over cap. + +### [LOW] _dotty_behaviour_get caches the failure fallback for the full TTL +**File:** `bridge.py:492-504` +**Issue:** On any error it writes the empty fallback into the cache for the TTL, so every tile blanks for ~2s after each blip even after recovery. +**Fix:** Cache only on success; return `fallback` on the error path without caching. +**Verification:** unverified — over cap. + +### [LOW] bridge writes to brain.db (RW) primarily owned by dotty-pi — approve/redact can lock/race +**File:** `bridge.py:316, 349` +**Issue:** `_voice_memory_approve/_delete_blocking` open brain.db RW across containers; in non-WAL delete mode a writer blocks readers, and failures are swallowed and surface as a generic "not found". +**Fix:** Set `PRAGMA busy_timeout`, surface lock failures distinctly; longer term route approve/redact through dotty-pi (the canonical writer). +**Verification:** unverified — over cap. + +### [LOW] host_detail 'server' modal hardcodes 'Recent errors today: —' despite the count being computed elsewhere +**File:** `bridge/dashboard.py:1832` +**Issue:** `alerts_count()` already tallies today's errored turns from the convo log; the modal shows a permanent dash. +**Fix:** Factor the tally into a helper and reuse it. +**Verification:** unverified — over cap. + +### [LOW] security_watch defaults use retired RPi/zeroclaw paths and wrong port +**File:** `bridge/security_watch.py:85-101` +**Issue:** `SECURITY_LOG_DIR` defaults `CONVO_LOG_DIR` to `/root/zeroclaw-bridge/logs` (disagreeing with dashboard.py's default) and `_BRIDGE_INTERNAL_URL` defaults to port 8080 while the bridge listens on 8081. +**Fix:** Align the default with dashboard.py; fix the port to 8081 — or delete the dead module. +**Verification:** confirmed 2/3 lenses. + +### [LOW] Voice-tools inventory shows 5 tools but dotty-pi-ext ships 7 +**File:** `bridge/dashboard.py:965-981` +**Issue:** `_VOICE_TOOLS` hardcodes 5 and the docstring says "same five values"; recall_person/remember_person (added in #53) are missing. +**Fix:** Add the two person-memory tools; source from a shared manifest. +**Verification:** unverified — over cap. + +### [LOW] Host-detail 'bridge' modal hardcodes Device: Raspberry Pi +**File:** `bridge/dashboard.py:1773` +**Issue:** The bridge now runs as the `dotty-bridge` Docker container on Unraid; the modal reports stale hardware provenance. +**Fix:** Relabel as "Docker container (Unraid host)" or drop the row. +**Verification:** unverified — over cap. + +### [LOW] CSRF warning + /admin/safety self-rewrite reference retired zeroclaw paths +**File:** `bridge/csrf.py:43`; `bridge.py:872-921` +**Issue:** The CSRF warning points operators at `/etc/default/zeroclaw-bridge` (no effect in a container; CSRF cookies invalidate on every restart). `_admin_safety` self-edits bridge.py on disk — changes nothing live (the consumer was in ZeroClaw) and is wiped on the next image rebuild. +**Fix:** Update the CSRF message to the container env mechanism; remove the `/admin/safety` self-rewrite or move the allowlist to a mounted state file. +**Verification:** unverified — over cap (both). + +### [QUALITY/LOW–INFO] Dead/retired metrics + Tier1Slim narrative +**File:** `bridge/metrics.py:100-160` (128 = `dotty_active_acp_sessions`); `bridge/dashboard.py:847-872, 1113-1126, 1438-1480` +**Issue:** Only `dotty_kid_mode_active` and `dotty_content_filter_hits_total` are ever observed; the rest (incl. the ACP gauge) always export 0. The logger name is `zeroclaw-bridge.metrics`. Smart-mode code carries large dead Tier1Slim/zeroclaw commentary for a card that does nothing. +**Fix:** Remove unproduced metrics (esp. the ACP gauge), rename the logger, collapse the smart-mode card to its real behaviour. +**Verification:** unverified — over cap. + +### [INFO] Dashboard live perception feed (EventSource '/api/perception/feed') 404s — endpoint ripped in #111, consumer left behind +**File:** `bridge/templates/dashboard.html:564` +**Issue:** The page (served from :8081) opens `EventSource('/api/perception/feed')`, but the bridge serves no `/api/*` routes — commit c6df5c5 (#111) deleted the endpoint and left the EventSource. The live feed never streams, every dotty-refresh nudge is dead, and the browser hammers a 404 reconnect every few seconds per open tab. The real SSE lives on dotty-behaviour:8090 (different origin). +**Fix:** Add a bridge SSE passthrough under `/api/`, or point the EventSource at the dotty-behaviour origin with CORS configured. +**Verification:** confirmed 2/3 lenses. + +--- + +## Voice LLM providers (PiVoiceLLM, OpenAICompat) + +### [MEDIUM] Turn timeout is a fixed wall-clock deadline that ignores streaming progress +**File:** `custom-providers/pi_voice/pi_client.py:220-269` +**Issue:** `deadline` is set once (default 120s) and never extended while text streams. A healthy long reply (think_hard escalation, 6-sentence story) is killed mid-stream and the user hears the "(brain offline)" fallback. The timeout should bound silence, not total length. +**Fix:** Reset `deadline` on each `text_delta` so the cap only fires when pi goes quiet. +**Verification:** unverified — over cap. + +### [MEDIUM] new_session failure proceeds with un-reset pi session, leaking prior-turn context +**File:** `custom-providers/pi_voice/pi_voice.py:136-141` +**Issue:** On `PiClientError` from `new_session()`, the code logs and continues and sets `_first_turn = False` unconditionally, so the next turn runs against pi's non-reset session — carrying prior content (or a partially-absorbed jailbreak) into the next, possibly child, interaction. +**Fix:** Force a hard reset (close + respawn) on `new_session` failure, or surface the failure rather than treating a failed reset as a fresh turn. +**Verification:** unverified — over cap. + +### [MEDIUM] _closed flag never reset after close() + reuse, suppressing reader-thread crash logging +**File:** `custom-providers/pi_voice/pi_client.py:171-182, 311-313, 352-354` +**Issue:** `close()` sets `_closed=True`; `_ensure_started()` respawns without clearing it, so the reader threads' `if not self._closed:` guards swallow genuine crashes — turns time out with no diagnostic. +**Fix:** Reset `_closed=False` on respawn; scope the suppression to the process generation. +**Verification:** unverified — over cap. + +### [MEDIUM] Abandoned iter_turn_text leaves pi streaming; a stale agent_end can terminate the next turn +**File:** `custom-providers/pi_voice/pi_client.py:211-269` (also 188-209) +**Issue:** On barge-in/TTS-abort the generator is GC'd without consuming `agent_end`; pi keeps generating and queues frames. `agent_end` is not id-matched, so an abandoned turn's trailing `agent_end` can terminate the NEXT `iter_turn_text`, truncating the next reply to empty/partial. +**Fix:** Send an explicit cancel to pi on abandon (try/finally), drain prior frames by req_id, and id-match `agent_end`. +**Verification:** confirmed 2/3 lenses. + +### [MEDIUM] OpenAICompat can emit two emojis when a valid emoji is not at the start +**File:** `custom-providers/openai_compat/openai_compat.py:218-229` +**Issue:** Enforcement only checks `so_far.startswith(emoji)`. For "Well, 😊 hi there" the check fails, so it prepends `😐` AND yields content still containing `😊` — two emojis, violating the single-emoji HARD CONSTRAINT; firmware face keys on the first, the second garbles TTS. +**Fix:** Strip embedded emojis from the body before yielding, or buffer until the first non-whitespace char and filter subsequent emojis. +**Verification:** unverified — over cap. + +### [MEDIUM] OpenAICompat speaks chain-of-thought when pointed at a reasoning model +**File:** `custom-providers/openai_compat/openai_compat.py:189-229` +**Issue:** Unlike PiClient, OpenAICompat has no thinking filter — it yields `delta.content` verbatim. Inline `...` is spoken aloud and the leading `<` defeats the emoji check; separate `reasoning_content` is dropped but the silent gap can trip the turn timeout. +**Fix:** Strip `...` (stateful across chunks), ignore `reasoning_content`, or send a param to disable thinking. +**Verification:** unverified — over cap. + +### [MEDIUM] OpenAICompat appends the safety/format suffix to the wrong message when no user turn exists +**File:** `custom-providers/openai_compat/openai_compat.py:101-114` +**Issue:** `_TURN_SUFFIX` is appended only at `last_user_idx`; if the final turn is system/assistant/tool (greeter injections), `last_user_idx` is None and the suffix — carrying emoji rule + kid-mode filter — is never injected, so the model runs unconstrained for that turn. +**Fix:** When `last_user_idx` is None, append the suffix to a synthesized trailing user message or the last message regardless of role. +**Verification:** unverified — over cap. + +### [MEDIUM] OpenAICompat yields whitespace-only content chunks to TTS before emoji enforcement +**File:** `custom-providers/openai_compat/openai_compat.py:216-229` +**Issue:** Whitespace-only deltas are yielded before `so_far` is non-empty and the emoji check runs; the first thing TTS/firmware sees is whitespace, not the emoji the contract promises. If the stream is all whitespace, the fallback fires after whitespace already went out. +**Fix:** Buffer until `so_far` is non-empty before yielding anything. +**Verification:** unverified — over cap (duplicate-class with F139). + +### [LOW] Documented extra_pi_flags config key is silently ignored +**File:** `custom-providers/pi_voice/pi_voice.py:29-30, 107-117` +**Issue:** The docstring documents `extra_pi_flags`, but `__init__` reads only `container_name`; `make_default_pi_client()` consults only the `DOTTY_PI_EXTRA_FLAGS` env var. A user setting `extra_pi_flags` in `.config.yaml` gets no effect. +**Fix:** Read `config.get('extra_pi_flags')` and thread it in, or remove it from the docstring and document the env var as the only mechanism. +**Verification:** confirmed 2/3 lenses. + +### [LOW] make_default_pi_client ignores the resolved container_name +**File:** `custom-providers/pi_voice/pi_voice.py:107-117` +**Issue:** The provider resolves and logs `container_name`, but `make_default_pi_client()` re-reads `DOTTY_PI_CONTAINER` itself and ignores it — the configured container name is logged but inert in production. +**Fix:** Pass `self._container` into `make_default_pi_client`. +**Verification:** unverified — over cap. + +### [LOW] PiVoiceLLM fallback strings violate the mandatory emoji-prefix protocol +**File:** `custom-providers/pi_voice/pi_voice.py:130, 150` +**Issue:** `yield "(empty turn)"` and `yield "(brain offline ...)"` begin with `(`; the firmware emotion parser sees `(` and produces no/garbage face. OpenAICompat correctly prefixes its fallbacks. +**Fix:** Prepend an allowed emoji (e.g. `😐`) to both fallbacks. +**Verification:** confirmed 2/3 lenses. + +### [LOW] new_session response match does not verify request id +**File:** `custom-providers/pi_voice/pi_client.py:192-209` +**Issue:** The drain loop matches only `type==response, command==new_session`, not `id==req_id`; a stale `new_session` response can be accepted early. +**Fix:** Add `and frame.get('id') == req_id`. +**Verification:** unverified — over cap. + +### [LOW] PiVoiceLLM _first_turn desyncs from the pi process after a respawn +**File:** `custom-providers/pi_voice/pi_voice.py:136-141` +**Issue:** If pi dies after turn 1, turn 2 runs `new_session()` against a freshly respawned process that has no session — a spurious 10s `new_session timed out` is swallowed, adding dead air. +**Fix:** Have PiClient signal a fresh spawn so the first-turn-after-respawn skips `new_session()`. +**Verification:** unverified — over cap. + +### [LOW] iter_turn_text process-exit detection only runs on queue-empty +**File:** `custom-providers/pi_voice/pi_client.py:222-269` +**Issue:** A crash with frames still queued never hits the Empty branch, so it spins to the 120s turn timeout and reports "turn timed out" instead of "pi process exited", delaying the fallback by up to 2 minutes. +**Fix:** Check `self._proc.poll()` on every iteration. +**Verification:** unverified — over cap. + +### [LOW] _next_id increments shared counter without the lock; reader-thread frame routing races a respawn +**File:** `custom-providers/pi_voice/pi_client.py:283-285, 149-169` +**Issue:** A respawn swaps `_event_queue` under lock while the old daemon reader (never joined) can still `put` stale frames into the new queue, injecting a previous process's frames into the new turn. +**Fix:** Join/stop prior reader threads before swapping the queue; bind each reader to its specific queue/proc; lock the id mutation. +**Verification:** unverified — over cap. + +### [LOW] _completions_url mishandles a base ending in /chat/completions +**File:** `custom-providers/openai_compat/openai_compat.py:125-132` +**Issue:** The endsWith heuristic can double the path (e.g. a gateway route `/chat/completions/stream`). +**Fix:** Make the endpoint construction explicit via a config flag, or document url must be the base. +**Verification:** unverified — over cap. + +### [LOW] OpenAICompat yields leading whitespace before emoji-prefix engages +**File:** `custom-providers/openai_compat/openai_compat.py:218-229` +**Issue:** A whitespace-only first chunk is yielded before enforcement, so the first TTS chunk isn't the emoji. +**Fix:** Buffer until `so_far` is non-empty; strip leading whitespace. +**Verification:** unverified — over cap. + +### [QUALITY/LOW–INFO] Dead/unreachable helpers + unsynchronized stderr ring +**File:** `custom-providers/openai_compat/openai_compat.py:134-140` (`_chunk_sentences`); `custom-providers/pi_voice/pi_client.py:63-81` (`default_subprocess_factory`), 275-277/343-354 (stderr ring) +**Issue:** `_chunk_sentences` + its `_SENTENCE_BOUNDARY` import are dead; the ssh `default_subprocess_factory` is unreachable in production; the stderr ring is mutated/snapshotted across threads without the lock the rest of the class uses. +**Fix:** Remove dead code or wire it; use a `deque(maxlen=...)` and/or guard with `self._lock`. +**Verification:** unverified — over cap. + +--- + +## ASR/TTS providers (FunASR, Whisper, EdgeStream, Piper) + +### [HIGH] Mid-utterance synth failure emits SentenceType.LAST, truncating the rest of the reply +**File:** `custom-providers/edge_stream/edge_stream.py:139-143` (mirror: `piper_local.py:187-191`) +**Issue:** The except handler unconditionally pushes `(SentenceType.LAST, [], None)` regardless of `is_last`. A multi-segment reply where an early segment's synth throws tells xiaozhi/firmware the whole response is finished, so every remaining segment is dropped and the device stops mid-sentence. +**Fix:** Only emit LAST from the failure path when `is_last` is True; otherwise log and continue. +**Verification:** confirmed 2/3 lenses. + +### [MEDIUM] processed_chars over-count in edge_stream and piper_local +**File:** `custom-providers/edge_stream/edge_stream.py:67-74`; `custom-providers/piper_local/piper_local.py:114-121` +**Issue:** `remaining_text = full_text[self.processed_chars:]` (absolute index) then `self.processed_chars += len(full_text)` over-counts; post-condition should be `== len(full_text)`. Masked today by the FIRST reset, but a re-entered LAST drops the final sentence's audio. +**Fix:** `self.processed_chars = len(full_text)` (or `+= len(remaining_text)`) in both; consider a shared mixin. +**Verification:** confirmed 2/3 lenses. + +### [MEDIUM] fun_local crashes in __init__ when output_dir is unset +**File:** `custom-providers/asr/fun_local.py:51-58` +**Issue:** `os.makedirs(None)` raises TypeError when `output_dir` is absent; whisper_local guards this, fun_local does not, so the two ASR providers are inconsistent. +**Fix:** `if self.output_dir: os.makedirs(self.output_dir, exist_ok=True)`. +**Verification:** unverified — over cap. + +### [MEDIUM] ASR model invoked from worker threads with no lock — not thread-safe under concurrent transcription +**File:** `custom-providers/asr/whisper_local.py:107-109, 161-173` (mirror: `fun_local.py:81-88`) +**Issue:** The singleton model is called via `asyncio.to_thread` from separate threads for overlapping sessions; faster-whisper/FunASR aren't safe for concurrent calls on the same model, risking corrupted output or a segfault. +**Fix:** Serialize inference with a per-instance `threading.Lock` inside the blocking call. +**Verification:** unverified — over cap. + +### [MEDIUM] before_stop play files dropped when the final segment's synthesis fails +**File:** `custom-providers/edge_stream/edge_stream.py:139-143` (mirror piper 187-191) +**Issue:** The except handler never calls `_process_before_stop_play_files()`, so a LAST-segment synth failure abandons queued trailing assets (and can leak them into the next utterance). +**Fix:** On the except path, when `is_last`, still process+clear the before_stop play files. +**Verification:** unverified — over cap. + +### [MEDIUM] SentenceType.FIRST emitted once per segment instead of once per reply +**File:** `custom-providers/edge_stream/edge_stream.py:101` (mirror `piper_local.py:146`) +**Issue:** `text_to_speak` puts FIRST at the top of every call; a multi-segment reply pushes multiple FIRSTs, which downstream treats as start-of-response and can re-trigger talk animation mid-sentence. +**Fix:** Track a per-reply first-segment flag; emit FIRST only for the first segment. +**Verification:** unverified — over cap. + +### [LOW] Whisper odd-length PCM buffer raises ValueError instead of degrading +**File:** `custom-providers/asr/whisper_local.py:104` +**Issue:** `np.frombuffer(..., int16)` raises on an odd byte count (truncated final frame); caught by the broad except → empty transcript with no retry. +**Fix:** Trim to an even length before `frombuffer`. +**Verification:** unverified — over cap. + +### [LOW] get_emotion matches the first mapped emoji anywhere, not the required leading emoji; default 🙂 outside the set +**File:** `custom-providers/textUtils.py:146-152` +**Issue:** The protocol requires the FIRST char to be the emotion emoji, but `get_emotion` breaks on the first mapped emoji anywhere, selecting a wrong/secondary face; the default `🙂` is outside `ALLOWED_EMOJIS`/the firmware set. +**Fix:** Resolve emotion from the leading grapheme only; default to `😐` (the in-set neutral). +**Verification:** unverified — over cap. + +### [LOW] piper effective_rate==24000 shortcut skips resampling when native rate isn't 24000 +**File:** `custom-providers/piper_local/piper_local.py:67-74` +**Issue:** The no-resample shortcut conflates "no pitch shift" with "no rate conversion"; a 22050 Hz voice with a pitch_scale that lands on 24000 emits raw native-rate PCM into a 24000 Hz encoder, corrupting playback speed/pitch. +**Fix:** Gate the shortcut on `src_rate == 24000 AND pitch_scale == 1.0`; otherwise always go through the gcd/resample branch. +**Verification:** unverified — over cap. + +### [LOW] piper length_scale not validated for <= 0 (unlike pitch_scale) +**File:** `custom-providers/piper_local/piper_local.py:32-42` +**Issue:** `pitch_scale` raises on `<= 0` but `length_scale` (Piper's speed knob) is unchecked; a 0/negative value yields empty/garbage audio with no early error. +**Fix:** Mirror the pitch_scale guard for length_scale. +**Verification:** unverified — over cap. + +### [LOW] whisper metrics: lang_prob can be None with forced language, raising in the f-string +**File:** `custom-providers/asr/whisper_local.py:132-137` +**Issue:** With a forced language, `language_probability` can be present-but-None; the `getattr(..., 0.0)` default doesn't fire and `{lang_prob:.3f}` raises, swallowed by the except — so the #105 ASR-METRICS line is silently dropped on every such utterance. +**Fix:** `lang_prob = getattr(...) or 0.0`. +**Verification:** unverified — over cap. + +### [LOW] FIRST audio-start marker emitted even when synth yields zero audio +**File:** `custom-providers/edge_stream/edge_stream.py:101-115` (mirror piper 146-159) +**Issue:** A mid-reply empty-synth segment emits FIRST then returns with no frames and no LAST, wedging per-sentence marker accounting. +**Fix:** Emit FIRST only after confirming there is audio, or enqueue a terminating marker on the empty-synth early return. +**Verification:** unverified — over cap. + +### [QUALITY/LOW–INFO] delete_audio_file unused; end_of_stream always True on non-last segments +**File:** `custom-providers/asr/fun_local.py:55` (and `whisper_local.py:46`); `custom-providers/edge_stream/edge_stream.py:121-137` (mirror piper 176-182) +**Issue:** `delete_audio_file` is stored but never honored (utterance WAVs accumulate); the trailing partial frame is flushed with `end_of_stream=True` on every segment, not just the last. +**Fix:** Honor or remove `delete_audio_file`; pass `end_of_stream=is_last`. +**Verification:** unverified — over cap. + +--- + +## dotty-behaviour core (app, routes, perception, dispatch, config) + +### [HIGH] room_view roster recognition silently fails when person id != display_name +**File:** `dotty-behaviour/vision/room_view.py:115-124, 159-167` +**Issue:** `name_choices` is built from `p.display_name`, but `parse_room_view_response` validates against `roster_ids` (which are `p.id` from `roster_ids_with_appearance()`). Whenever `id != display_name.lower()` (the common case), a correct VLM identification falls through to None, the synthetic `face_recognized` never fires, and named greetings regress to bare/no greeting. Masked because the test fake reimplements the helper as `display_name.lower()` with id==display_name fixtures. +**Fix:** Make the prompt vocabulary and the validation set use the same key (build `name_choices` from `p.id`, or resolve the returned display-name back to the id). Fix the test fake to return `p.id` and add an id != display_name fixture. +**Verification:** confirmed 3/3 lenses. + +### [MEDIUM] Consumer task crashes are silently swallowed (no log, no restart) +**File:** `dotty-behaviour/main.py:325-341` +**Issue:** Consumers are launched with `create_task` and only awaited at shutdown via `gather(..., return_exceptions=True)`; a runtime exception is captured and discarded at shutdown, never logged when it happens, and the consumer is never restarted while the daemon still reports "ready". +**Fix:** Attach a done-callback that logs non-cancelled exceptions; optionally supervise/restart each consumer. +**Verification:** unverified — over cap. + +### [MEDIUM] No upload size limit on /api/vision/explain and /api/audio/explain (LAN OOM) +**File:** `dotty-behaviour/routes/vision.py:145` (and `audio.py:87`) +**Issue:** `await file.read()` reads the entire multipart upload into memory then base64-encodes (~1.33x) with no Content-Length cap; both endpoints bind `0.0.0.0:8090`. A single large/malicious POST can exhaust container memory and take down the daemon (and all consumers). The JPEG is also retained per-device in `vision_cache`. +**Fix:** Enforce a max upload size (e.g. 5 MB JPEG / 10 MB audio) up front and reject >N with 413. +**Verification:** unverified — over cap. + +### [MEDIUM] vision_latest lost-wakeup race produces spurious 404s +**File:** `dotty-behaviour/routes/vision.py:331-352` +**Issue:** `vision_latest` pops the cache, registers a waiter, then awaits the Event. If `explain`'s `signal_vision_waiters()` fires between pop and registration, the signal is lost (signal only set()s already-registered Events), so the waiter blocks to the 15s timeout and 404s even though a valid description exists. +**Fix:** Make the wakeup level-triggered: re-check the cache after registering before awaiting, or use a generation counter. Same for the audio waiter. +**Verification:** unverified — over cap. + +### [MEDIUM] vision_latest unconditionally evicts a fresh capture (and its JPEG) written by another producer +**File:** `dotty-behaviour/routes/vision.py:332` +**Issue:** The unconditional `pop` deletes whatever is cached, including a fresh entry another producer wrote microseconds earlier (idle_photographer / room_view), discarding its `jpeg_bytes` that `/api/vision/photo` serves; a subsequent photo GET 404s. +**Fix:** Pop only entries older than the TTL; or snapshot a request-start time and accept any entry newer than it. +**Verification:** unverified — over cap. + +### [MEDIUM] Failed room_view VLM call still arms the 120s cooldown +**File:** `dotty-behaviour/routes/vision.py:188-204` +**Issue:** `last_room_view_capture_t` is set before the VLM call, which never raises (returns sentinels on failure), so a transient VLM blip arms the full 120s cooldown despite no usable description — repeated blips disable named greetings for minutes. +**Fix:** Arm the cooldown only after a successful, non-sentinel response, or use a short failure-cooldown. +**Verification:** unverified — over cap. + +### [LOW] /api/perception/state emits invalid JSON (Infinity) for devices with no last_event_t +**File:** `dotty-behaviour/perception/state.py:328-336` +**Issue:** `_annotate` sets `age = float('inf')` → serialized as bare `Infinity` (not valid JSON), breaking strict parsers/the dashboard card. Reachable for a known device that got a room_view capture but never sent a perception event, and unconditionally for any unknown `device_id`. +**Fix:** Use a JSON-safe value (None / -1) for missing `last_event_t`; add a regression test asserting the body parses as standard JSON. +**Verification:** unverified — over cap. + +### [LOW] Failed/empty weather fetch arms the full 30-min TTL +**File:** `dotty-behaviour/calendar_/cache.py:124-127` (driven by `poll.py:46-48`) +**Issue:** `set_weather` updates `weather_text` only on truthy text but always bumps `weather_fetched_perf`, so a transient weather failure marks the cache fresh for 30 minutes and serves stale (or empty) text. Same class as the room_view cooldown bug, distinct site (reported twice as F206/F246). +**Fix:** Advance `weather_fetched_perf` only on a non-empty fetch; add a test. +**Verification:** unverified — over cap. + +### [LOW] perception/state limit query params unvalidated (negative/zero slice quirks) +**File:** `dotty-behaviour/routes/perception.py:88-99, 124` +**Issue:** `limit` flows into `items[:limit]` / `balances[-limit:]` unbounded; `limit=-3` drops the 3 newest, `limit=0` returns `[]` for one path and the FULL list for `balances[-0:]`. +**Fix:** Clamp `limit = max(0, min(limit, MAX))`; special-case `limit <= 0` for the balance series. +**Verification:** unverified — over cap. + +### [LOW] Client-supplied event ts can be set in the future, defeating staleness gates +**File:** `dotty-behaviour/routes/perception.py:61-68` +**Issue:** `payload.ts` is used verbatim; a future ts makes `now - last_t` negative → clamped to 0, so `get_fresh_face_id` treats a stale identity as permanently fresh. ESP32 RTCs are frequently unsynced. +**Fix:** Reject/clamp implausible future timestamps on ingest; guard `get_fresh_face_id` against negative age. +**Verification:** unverified — over cap. + +### [LOW] VLM error/offline SENTINEL strings cached as a real description and leaked into the perception snapshot +**File:** `dotty-behaviour/routes/vision.py:195-242` +**Issue:** `vision_explain` stores the `VLM_OFFLINE_SENTINEL`/`VLM_NETWORK_ERROR_SENTINEL` verbatim with a fresh `wall_ts`; the snapshot then surfaces `You see: ERROR: the vision service didn't respond...` into Dotty's voice prompt for up to 60s, and the idle photographer can persist it to NDJSON. +**Fix:** Detect the sentinels after `describe_image` and skip the cache write + `signal_vision_waiters`. Same for the audio fallback. +**Verification:** unverified — over cap. + +### [LOW] snapshot.py age-gate constants hardcoded, ignoring env-configurable counterparts +**File:** `dotty-behaviour/perception/snapshot.py:18-20` +**Issue:** `VISION_AGE_GATE_SEC` etc. are literals, while config defines the env-overridable TTLs; raising `VISION_CACHE_TTL_SEC` keeps entries longer but the snapshot still gates at 60s. `config.SCENE_SYNTHESIS_AGE_GATE_SEC` has zero readers. +**Fix:** Import the gates from config (or pass them in); remove the unused config var or document the gates are fixed. +**Verification:** unverified — over cap. + +### [LOW] perception/state stale dance_active never cleared on a missed terminal event +**File:** `dotty-behaviour/perception/state.py:185-199` +**Issue:** `dance_active` has no time-based expiry; a dropped `dance_ended`/`state_changed` latches it True indefinitely, permanently suppressing room_view captures and greetings for that device. +**Fix:** Make `is_dance_active` freshness-bounded against `last_dance_started_t` with a max-dance TTL. +**Verification:** unverified — over cap. + +### [LOW] SSE feed only re-checks client disconnect every 15s, leaking a subscriber queue +**File:** `dotty-behaviour/routes/perception.py:142-153` +**Issue:** `is_disconnected()` is checked once per loop before a 15s `wait_for`, so a disconnect mid-block isn't noticed for up to 15s; the queue stays registered and `broadcast()` keeps filling/dropping it. +**Fix:** Reduce the poll timeout or race the disconnect check against the queue get. +**Verification:** unverified — over cap. + +### [QUALITY/LOW–INFO] Dead prometheus dep; per-request weather subprocess; misplaced imports; None-guard inconsistency +**File:** `dotty-behaviour/requirements.txt:12-14`; `routes/calendar.py:51-62`; `main.py:40`; `consumers/sleep_dreamer.py:167` & `dance_reflector.py:90` +**Issue:** `prometheus-client` is pinned but no `/metrics` or counters exist; `/api/calendar/today` spawns a `curl wttr.in` subprocess on the request path even with no calendars configured; `import re`/`import asyncio` are misplaced; two consumers dereference `event.data` directly while siblings use `(event.data or {})`. +**Fix:** Implement or drop the metrics dep; gate weather refresh on `CALENDAR_IDS`; tidy imports; normalize `event.data` (default_factory or guard uniformly). +**Verification:** unverified — over cap. + +--- + +## dotty-behaviour consumers/greeter/household/calendar/vision + +### [HIGH] Double greeting on every face_recognized — FaceGreeter and ProactiveGreeter both fire +**File:** `dotty-behaviour/consumers/face_greeter.py:133-186` +**Issue:** FaceGreeter's named greet ("Oh, it's {name}!") and ProactiveGreeter's LLM greeting both subscribe to `face_recognized`, keep separate cooldown state, and don't coordinate — one recognition produces two back-to-back utterances. The bare path was deliberately suppressed when the roster is identifiable; the named path was not given the same deference. +**Fix:** Pick one owner (disable FaceGreeter's named path when the proactive greeter is enabled, or share a per-identity cooldown). +**Verification:** confirmed 2/3 lenses. + +### [HIGH] room_view name parser cannot match multi-word display names +**File:** `dotty-behaviour/vision/room_view.py:72-81, 115-124` +**Issue:** `name_choices` injects `display_name`, but the NAME regex is a single whitespace-free token; a reply `NAME: Mary Anne` fails the anchored regex entirely, so any member with a space in their display name is a 100% silent identification failure even on a perfect VLM response. +**Fix:** Allow internal spaces in the NAME group (and normalise), or request a single-token id and map ids→display names. Add a two-word display_name test. +**Verification:** confirmed 3/3 lenses. + +### [HIGH] room_view roster recognition fails when person id != display_name (greeter path) +**File:** `dotty-behaviour/vision/room_view.py:115-124, 159-167` +**Issue:** Same id/display_name name-space mismatch as the core finding; `room_match_person_id` is the VLM's display-name token lowercased, then passed as `identity` to `household.get()` which is keyed by id and also misses. (Reported as both F38 and F201.) +**Fix:** Agree on a single key (prefer `p.id`); fix the test fake and add an id != display_name fixture. +**Verification:** confirmed 3/3 lenses. + +### [HIGH] Greeter calendar lookup never matches an identified person (case mismatch) +**File:** `dotty-behaviour/greeter/greeter.py:257-259` → `calendar_/cache.py:69-73` +**Issue:** `identity` is a lowercased person id, but the calendar `person` is set from the title prefix with no case folding (`[Hudson] → "Hudson"`), so `"Hudson" != "hudson"` drops that person's own events from their greeting. Household-bucket events mask the bug. +**Fix:** Normalise both sides to a common case (lower-case the parsed calendar `person`, or compare case-insensitively); prefer mapping the prefix to the canonical id at fetch time. +**Verification:** confirmed 2/3 lenses. + +### [MEDIUM] FaceGreeter consumes the greet cooldown even when suppressed by an active dance +**File:** `dotty-behaviour/consumers/face_greeter.py:148-168, 105-126` +**Issue:** The cooldown slot is written before the dance-active gate, so every face event during a dance silently burns the cooldown without greeting; the user walks in during a dance, is never greeted, and stays un-greeted for the full cooldown after. `last_face_greet_t` also arms the aborter. (Reported as F160/F207.) +**Fix:** Move the dance-active (and empty-text) checks above the cooldown write in both handlers. +**Verification:** confirmed 2/3 lenses. + +### [MEDIUM] last_chat_t is never written on real conversations — QUIET_AFTER_CHAT gates are dead +**File:** `dotty-behaviour/consumers/face_greeter.py:149-156` + `sound_turner.py:77-79` + `perception/state.py:200-203` +**Issue:** Named-greet and sound-turner suppress themselves for `*_quiet_after_chat_sec` after the last chat by reading `last_chat_t`, but the only writer is `purr_player.py`; nothing records actual conversation activity, so the gate never fires and Dotty name-greets/head-turns over a live conversation. Regression from bridge.py's `/api/message` ingress. +**Fix:** Write `last_chat_t` on conversation activity (set it on `chat_status` listening, or POST a conversation event), or drop the gates as inert. +**Verification:** confirmed 2/3 lenses. + +### [MEDIUM] Greeter day-GC discards every other day's slots; midnight-roll cooldown corruption +**File:** `dotty-behaviour/greeter/greeter.py:224-228` +**Issue:** `_take_slot` derives `today` from wall-now but uses `event_ts` for the cooldown; the GC collapses `_state` to `{today: ...}`, wiping (and persisting away) other days. A pre-midnight event handled just after midnight is filed under the new day with old-day cooldown math, letting a near-instant second greeting slip through. +**Fix:** Key the day on `event_ts`; GC by a retention window (keep today + yesterday). +**Verification:** unverified — over cap. + +### [MEDIUM] SecurityCycle keeps capturing on a missed/early transition +**File:** `dotty-behaviour/consumers/security_cycle.py:238-257` +**Issue:** Capture starts/stops purely off live `state_changed`; if the device is already in `security` when the consumer subscribes the timer never starts, and if the non-security transition frame is dropped the per-device capture loop fires the camera forever. +**Fix:** Seed timers from current `current_state` on subscribe; reconcile each running timer against the live state per cycle. +**Verification:** unverified — over cap. + +### [MEDIUM] calendar fetch time window excludes the final minute of the day +**File:** `dotty-behaviour/calendar_/fetch.py:60-66` +**Issue:** `time_max = 23:59:59` with an exclusive `timeMax` loses the last minute; the correct upper bound is the start of the next local day. +**Fix:** `time_max = (start_of_today + timedelta(days=1)).isoformat()`. +**Verification:** unverified — over cap. + +### [MEDIUM] CalendarCache.flush_for_new_day produces a long empty-context window on day-roll fetch failure +**File:** `dotty-behaviour/calendar_/cache.py:116-122` + `poll.py:53-75` +**Issue:** The day-roll eagerly flushes events, and a failing refresh only bumps `calendar_failures`; with backoff the loop may sleep up to 600s showing an EMPTY calendar in the morning even though yesterday's data was serviceable a moment earlier. +**Fix:** Flush only on a successful fetch, or reset failures / use the base interval for the first retry after a day-roll flush. +**Verification:** unverified — over cap. + +### [MEDIUM] room_view 'no one in view' sentinel uses substring match +**File:** `dotty-behaviour/vision/room_view.py:152-153` +**Issue:** `if ROOM_VIEW_NO_PERSON in cleaned.lower()` discards a valid identification whose DESC merely mentions the phrase, throwing away both the description and the matched identity. +**Fix:** Match the sentinel only when it's the entire reply; try the DESC/NAME regex first. +**Verification:** unverified — over cap. + +### [MEDIUM] FaceIdentifiedRefresher keeps the green ID LED lit when identity came from room_view +**File:** `dotty-behaviour/consumers/face_identified_refresher.py:65-71` +**Issue:** The skip gate only suppresses refresh when the face was both absent AND previously lost; a room_view identity sets `last_face_id` without `face_present` or a `face_lost`, so the LED stays green for the full TTL even after the person walked off. +**Fix:** Treat "identity present but face_present False and never positively detected" as a stop condition. +**Verification:** unverified — over cap. + +### [LOW] FaceLostAborter leaks completed abort tasks in _pending forever +**File:** `dotty-behaviour/consumers/face_lost_aborter.py:101-103, 47-55` +**Issue:** A completed abort task with no subsequent face event stays in `_pending`; the dict never shrinks for departed devices. +**Fix:** Add a done-callback that discards the entry. +**Verification:** unverified — over cap. + +### [LOW] Vision room_view block-reason early return skips stale-cache eviction +**File:** `dotty-behaviour/routes/vision.py:176-186` +**Issue:** A gated room_view request writes a fresh cache entry and returns before the stale-eviction loop, so on frequently-gated deployments stale entries for other devices are never reclaimed. +**Fix:** Run eviction on every exit path, including the block-reason early return. +**Verification:** unverified — over cap. + +### [LOW] Blocking file I/O on the event loop in greeter state persistence +**File:** `dotty-behaviour/greeter/greeter.py:246, 420-435` +**Issue:** `_take_slot` calls `_save_state` synchronously (mkdir + write + os.replace) on the loop on every greet. +**Fix:** `await asyncio.to_thread(self._save_state)` or batch/debounce. +**Verification:** unverified — over cap. + +### [LOW] SleepDreamer / DanceReflector dereference event.data without None-guard, crashing the consumer loop +**File:** `dotty-behaviour/consumers/sleep_dreamer.py:166-168`; `dance_reflector.py:89-92` +**Issue:** Direct `.get` on `event.data`; a `data=None` frame raises AttributeError caught only by the OUTER try/except, which logs "crashed" and EXITS the consumer permanently. +**Fix:** Use `(event.data or {}).get(...)`; wrap per-event handling so one bad frame doesn't terminate the loop. +**Verification:** unverified — over cap. + +### [LOW] HouseholdRegistry.get_by_calendar_prefix only matches with bracketed YAML; Person.days_until_birthday Feb-29 drift +**File:** `dotty-behaviour/household/registry.py:190-200, 293-294` (prefix); `82-92` (birthday) +**Issue:** The prefix index stores the raw YAML value but the query is bracket-normalized, so an unbracketed `calendar_prefix` is unreachable. A Feb-29 birthdate silently shifts to Feb-28 in common years, firing the birthday greeting a day early. +**Fix:** Bracket-normalize both sides; decide and document a deliberate Feb-29 policy. +**Verification:** unverified — over cap. + +### [INFO/QUALITY] Inconsistent None-guarding; convoluted day-GC expression +**File:** `dotty-behaviour/consumers/sleep_dreamer.py:167` / `dance_reflector.py:90`; `greeter/greeter.py:226-227` +**Issue:** None-guard inconsistency across consumers (also flagged above); the day-GC `{today: self._state.get(today, {})}` is provably `{today: {}}` and misleads readers. +**Fix:** Give `PerceptionEvent.data` a default_factory or guard uniformly; simplify the GC expression with a comment. +**Verification:** unverified — over cap. + +--- + +## dotty-pi-ext voice tools (TypeScript) + +### [HIGH] memory_lookup FTS search bypasses the #53 kid-safety namespace gate +**File:** `dotty-pi-ext/src/lib/brain_db.ts:82-90` (leaked into `tools/memory_lookup.ts`) +**Issue:** `fetchPersonMemories` is scoped to `person:` and deliberately never returns `person_pending:` (unreviewed minor facts), but `searchMemories` runs a bare `WHERE memories_fts MATCH ?` over ALL rows, so a pending fact about a minor ("kiddo is allergic to peanuts") is returned to a live turn whenever the query phrase-matches. The oracle (bridge.py) is equally unfiltered, so the gap exists in both. +**Fix:** Add `AND m.namespace NOT LIKE 'person_pending:%'` to the FTS query; mirror into oracle.py; add a regression test. +**Verification:** confirmed 3/3 lenses. + +### [MEDIUM] Admin/behaviour HTTP clients disarm the timeout before reading the body +**File:** `dotty-pi-ext/src/lib/xiaozhi_admin.ts:33-46`; `dotty-pi-ext/src/lib/dotty_behaviour.ts:20-37` +**Issue:** `adminFetch`/`behaviourFetch` clear the AbortController timer in `finally` after `await fetch()` (headers), then callers do an un-timed `.json()`/`.text()`; a server that sends headers then stalls the body hangs the voice tool forever with no exception, so the try/catch fallbacks never run. The sibling `llama_swap.ts` does this correctly. +**Fix:** Keep the timer armed across the body read (move the body read inside the fetch helper, mirroring `llama_swap.ts`). +**Verification:** confirmed 3/3 lenses. + +### [MEDIUM] turn_logger.formatTurnLog omits the final strip the oracle applies +**File:** `dotty-pi-ext/src/lib/turn_logger.ts:98-102` +**Issue:** The oracle stores `content.strip()` of the assembled string; the TS version doesn't, so a user-only turn produces `"user: hello | assistant: "` (trailing space) vs the oracle's stripped form — every user-only turn stores a non-byte-identical row. The integration test that would catch it (`assistant_empty_user_only`) only runs when `DOTTY_BRAIN_DB_SNAPSHOT` is set, which `npm test` never sets. +**Fix:** `return (\`user: ${u} | assistant: ${a}\`).trim();`; wire the Layer-2 pass into `npm test`. +**Verification:** unverified — over cap. + +### [LOW] play_song caches an empty catalogue on transient fetch failure for 60s +**File:** `dotty-pi-ext/src/tools/play_song.ts:83-90` +**Issue:** `getCatalog` caches `fetchSongCatalog`'s `[]` (which means "no songs" OR "fetch failed") for the full 60s TTL, so a momentary xiaozhi blip pins "(catalogue is empty)" for up to a minute. +**Fix:** Cache only non-empty results; treat empty as a non-cacheable miss. +**Verification:** unverified — over cap. + +### [LOW] play_song matchSong: empty stem from a dotfile matches any query, diverging from the oracle +**File:** `dotty-pi-ext/src/tools/play_song.ts:39-74` +**Issue:** A leading-dot catalogue entry yields stem `''`, and `qStem.includes('')` is always true, so the dotfile matches every query and can win — Python's `splitext('.mp3')` keeps `.mp3`, so the implementations disagree, violating the byte-equal contract. +**Fix:** Only strip an extension when the dot index > 0; guard against an empty stem; add a leading-dot fixture. +**Verification:** unverified — over cap. + +### [LOW] think_hard truncates by UTF-16 code units, diverging from the oracle and sibling tools +**File:** `dotty-pi-ext/src/tools/think_hard.ts:72-77` +**Issue:** `.slice(0, 500)` cuts by code units while the oracle and the other tools use codepoint slicing; an emoji-prefixed near-cap reply mis-cuts. think_hard is the lone inconsistent path. +**Fix:** Use `Array.from(...)` codepoint slicing like the sibling tools. +**Verification:** unverified — over cap. + +### [LOW] play_song host-presence guard ignores the _XIAOZHI_HOST fallback +**File:** `dotty-pi-ext/src/tools/play_song.ts:104` +**Issue:** The guard checks only `XIAOZHI_HOST`, but the admin client also honors `_XIAOZHI_HOST`; the two layers disagree on "host configured". +**Fix:** Use the same precedence as the admin client, or drop the guard and let the empty-list fallback drive the message. +**Verification:** unverified — over cap. + +### [LOW] extractTurnText concatenates assistant text blocks with no separator +**File:** `dotty-pi-ext/src/lib/turn_logger.ts:56-62, 79-91` +**Issue:** Multi-message/multi-block assistant turns join with `""`, fusing fragments ("Let me check…The answer is 4") in the logged turn. +**Fix:** Join with a space/newline. +**Verification:** unverified — over cap. + +### [LOW] think_hard default model resolved at module load, diverging from the per-call oracle +**File:** `dotty-pi-ext/src/tools/think_hard.ts:25` +**Issue:** `DEFAULT_MODEL` reads `VOICE_THINKER_MODEL` once at import; the oracle reads it per call, so a later env change diverges the request body. +**Fix:** Resolve the env lazily inside `buildThinkRequest`. +**Verification:** unverified — over cap. + +### [LOW] play_song getCatalog TOCTOU lets concurrent turns both refetch +**File:** `dotty-pi-ext/src/tools/play_song.ts:83-90` +**Issue:** Two interleaved `play_song` calls both see the stale timestamp and both fetch; combined with the empty-cache bug, one can poison the other. +**Fix:** Store an in-flight promise so concurrent callers await the same fetch; cache only non-empty success. +**Verification:** unverified — over cap. + +### [DOC/QUALITY/LOW–INFO] Stale ZeroClaw comments; take_photo 300-char cap claim; turn_logger "background write" wording; brain_db WAL/handle-staleness; take_photo test not run +**File:** `dotty-pi-ext/src/lib/brain_db.ts:5,44-46,176,42-52,32-52`; `src/tools/take_photo.ts:1-9`; `src/lib/turn_logger.ts:104-128`; `package.json:17-26` +**Issue:** brain_db comments still claim ZeroClaw co-writes (now the extension is the sole writer); take_photo's "capped at 300 chars" cap lives server-side, not in the TS client; turn_logger claims a background write but the sqlite INSERT is synchronous inline; brain_db's "WAL concurrency" comment never enables WAL and relies on the default busy_timeout; cached handles go stale if brain.db is swapped under a running process; `tests/take_photo.test.ts` exists but `npm test` never runs it. +**Fix:** Update the comments to the post-#36 reality; add the take_photo test to the aggregate; either enable WAL or fix the comment; document the no-swap-under-running-process contract. +**Verification:** unverified — over cap. + +--- + +## StackChan firmware (ESP32-S3 C++/ESP-IDF) + +### [MEDIUM] Face recognizer reads the camera frame buffer after releasing the arbiter (latent UAF) +**File:** `firmware/firmware/main/stackchan/face/face_detector.cpp:236-292` +**Issue:** `processFrame()` releases the detection arbiter at :236 but still uses the raw V4L2 `frame_data` (captured at :183) afterward — stored into `face_img.data` (:277) and passed to `recognize()` (:292). Once released, the capture path can dequeue/requeue the buffer and overwrite it. The stub ignores the data today, but the planned ESP-DL embedding crop will make this a live use-after-free across two camera tasks. +**Fix:** Hold the detection lock across `recognize()`, or copy the bounded crop while still holding the arbiter and point `face_img.data` at the copy. +**Verification:** confirmed 3/3 lenses. + +### [MEDIUM] SoundLocalizer high-pass filter state goes stale during cooldown, producing a spurious spike +**File:** `firmware/firmware/main/stackchan/sound_localizer.cpp:19-23, 44-51` +**Issue:** `OnStereoFrame()` returns early on the cooldown gate before advancing the HPF state, so the stored x[n-1]/y[n-1] freeze for up to 750 ms; the first post-cooldown frame evaluates across that discontinuity, injecting a transient that can push energy past threshold and fire a spurious direction change. +**Fix:** Advance the HPF state every frame regardless of the gates (or prime `_hp_*_prev_x` to the current input on cooldown). +**Verification:** unverified — over cap. + +### [MEDIUM] ImuEventModifier steals and releases the shared modify-lock it may not own +**File:** `firmware/firmware/main/stackchan/modifiers/imu.h:58-61, 117, 79, 105` +**Issue:** IMU sets the motion lock only if not held, but `restore_state()` always clears it; if `Capture()` holds the lock to keep the head still for the shutter, a shake's restore releases Capture's lock mid-photo, letting the head move. The avatar lock has the same shape. +**Fix:** Give the lock real ownership (refcount/owner-token), or have IMU remember whether it acquired it and only release if so. +**Verification:** unverified — over cap. + +### [MEDIUM] FaceDetector::stop() frees buffer and deletes the stop-semaphore without checking the take succeeded — UAF on overrun +**File:** `firmware/firmware/main/stackchan/face/face_detector.cpp:100-117` +**Issue:** `stop()` ignores the `xSemaphoreTake(_stop_sem, 2000ms)` return, then unconditionally deletes the semaphore and frees `_rgb_buffer`; if the task overruns the 2s timeout it later gives a deleted semaphore and uses freed state. Currently unwired but a real teardown defect. +**Fix:** Capture the take result; only delete/free when it succeeded. +**Verification:** unverified — over cap. + +### [LOW] Seqlock reader can return a torn read on weak memory (missing acquire barrier before the second seq load) +**File:** `firmware/firmware/main/stackchan/face/face_detection_result.h:33-47` +**Issue:** `read()` loads s1 (acquire), copies plain data fields, loads s2 (acquire) — but an acquire load doesn't forbid the earlier data reads from sinking below the s2 load, so a torn read can pass the seq check. The writer side is correct; the reader is missing the symmetric barrier. +**Fix:** Insert `std::atomic_thread_fence(std::memory_order_acquire)` between the data copies and the second seq load. +**Verification:** unverified — over cap. + +### [LOW] Idle-channel proactive reconnect churns close+reopen every ~30 ticks +**File:** `firmware/firmware/patches/xiaozhi-esp32.patch:23-32` +**Issue:** The reconnect fires on `IsTimeout()` (time-since-last-incoming, 120s); a reopened quiet-but-healthy channel keeps `IsTimeout()` true, so it close+reopens every 30 ticks indefinitely, each iteration paying a 200ms delay + full TLS/WS handshake on the main task. +**Fix:** Gate on an actual dead-channel signal or a backoff; reset the incoming-frame clock on `OpenAudioChannel()`. +**Verification:** unverified — over cap. + +### [LOW] CloseAudioChannel() blocks the calling task 200ms on every close +**File:** `firmware/firmware/patches/xiaozhi-esp32.patch:582` +**Issue:** An added `vTaskDelay(200ms)` stalls whatever task calls close (incl. the main Run loop), compounding the reconnect churn. +**Fix:** Move the drain delay to only the reopen path that needs it, or make close async. +**Verification:** unverified — over cap. + +### [LOW] PrivacyLeds::update() non-atomic RMW races the guard's setMicState(Off), latching the listening LED on +**File:** `firmware/firmware/main/stackchan/privacy/privacy_leds.cpp:49-60` +**Issue:** `update()` load→derive→store of `_mic_state` isn't atomic as a unit; the codec-close dtor can set Off in the window and `update()` overwrites it, so the green mic LED keeps showing "listening" after the mic closed (privacy over-indication). +**Fix:** Use a CAS that only updates while non-Off so a concurrent Off isn't resurrected. +**Verification:** unverified — over cap. + +### [LOW] CameraPeripheralGuard ctor-fallback/dtor-normal mismatch underflows the refcount +**File:** `firmware/firmware/main/stackchan/privacy/camera_peripheral_guard.cpp:40-66, 68-86` +**Issue:** The null-mutex ctor fallback doesn't increment `g_refcount`, but the dtor unconditionally `fetch_sub(1)`; a ctor-fallback/dtor-normal mismatch underflows uint32 to UINT32_MAX, so the stream never tears down and the camera LED stays Active indefinitely. +**Fix:** Track per-guard `_incremented`; only `fetch_sub` if it incremented; clamp against underflow. +**Verification:** unverified — over cap. + +### [LOW] MCP take_photo arms a fixed 10s camera-glow timer decoupled from capture duration +**File:** `firmware/firmware/patches/xiaozhi-esp32.patch:523` +**Issue:** `setCameraLedActive(true, 10000)` is cosmetic but hardcoded to 10s with no off-call; the glow lingers after a fast capture or goes dark before a slow one, no longer tracking the real camera lifecycle. +**Fix:** Tie the glow to the capture scope (no timeout; clear after Capture, or drive from the guard refcount). +**Verification:** unverified — over cap. + +### [LOW] Unescaped name/session_id in Protocol::SendEvent JSON construction +**File:** `firmware/firmware/patches/xiaozhi-esp32.patch:537-542` +**Issue:** The frame is built by raw concatenation; neither `name` nor server-supplied `session_id_` is escaped. Not currently exploitable (literal names, pre-escaped data), but any future runtime name or a quoted session id emits malformed JSON the relay drops. +**Fix:** Build with cJSON (escapes both); validate `data_json` parses; or assert+escape. +**Verification:** unverified — over cap. + +### [LOW] Debug log leftovers fire at WARN/ERROR on every WS and LLM frame +**File:** `firmware/firmware/patches/xiaozhi-esp32.patch:65, 590` +**Issue:** An `ESP_LOGW` on every LLM frame and an `ESP_LOGE` on every inbound WS frame bypass log-level filtering, spam serial, and bury real errors. +**Fix:** Drop or demote to `ESP_LOGD`. +**Verification:** unverified — over cap. + +### [LOW] HeadPet/Imu cross-task gesture flags are plain volatile bool, not atomic; fixed Press-before-Release order can invert fast sequences +**File:** `firmware/firmware/main/stackchan/modifiers/head_pet.h:50-52` (and `head_pet.cpp:48-83`, `imu.h:124`) +**Issue:** `volatile bool` flags written from the signal callback and read+cleared on the tick task aren't atomic (a gesture can be lost), and processing press→swipe→release in fixed order can end with `_is_touched=false` on a fast release-then-press, silently never arming the hold-to-listen window. +**Fix:** Use `std::atomic` with `exchange(false)`, or an ordered event queue. +**Verification:** unverified — over cap. + +### [LOW] ParentalGate unlock at millis()==0 treated as "never unlocked" +**File:** `firmware/firmware/main/stackchan/face/parental_gate.cpp:43, 71-72` +**Issue:** `_unlocked_at_ms` uses 0 as the "never" sentinel; an unlock in the first ms after boot stores 0 and is silently dropped. Dormant scaffold but a latent collision. +**Fix:** Use a separate `_is_unlocked` flag, or store `max(now,1)`. +**Verification:** unverified — over cap. + +### [LOW] CameraArbiter::acquireForCapture clears the shared _capture_pending on timeout +**File:** `firmware/firmware/main/stackchan/face/camera_arbiter.cpp` (header :38-44) +**Issue:** `_capture_pending` is a single shared atomic, not per-call; a future second capture path plus a timeout could drop both callers' intent, wedging the other's still-pending capture — the TOCTOU the flag exists to prevent. +**Fix:** Make `_capture_pending` a refcount, or document+assert single-caller. +**Verification:** unverified — over cap. + +### [QUALITY/DOC/LOW–INFO] PrivacyLeds doc/comment drift; dead FaceRecognizer members +**File:** `firmware/firmware/main/stackchan/privacy/privacy_leds.{h,cpp}` (h:15-19,137-141; cpp:64-65,89); `face/face_recognizer.h:172-173` +**Issue:** Header docs describe MIC LEDs as white (actually green) and the self-test as amber→cyan→red (actually green→red→off); inline comments say the hal guard "rejects 6/7" and the camera LED is at "index 7" (actually 6/11, camera at kCameraLedIndex=11); `_last_recognize_ms`/`_last_identity` are dead members. +**Fix:** Update the docs/comments to match the implementation; delete the dead members. +**Verification:** unverified — over cap. + +--- + +## Top-level + scripts (receiveAudioHandle, dances, scripts/, provision) + +### [HIGH] provision.py: voice "General" collides with text "general" and gets moved on every run +**File:** `community/discord/provision.py:93-96, 197-202` +**Issue:** `find_channel` searches the whole guild case-insensitively and ignores channel type, so syncing the VOICE "General" matches the existing TEXT "general" and `existing.edit(category=VOICE)` MOVES #general into the VOICE category (destructive, recurs every run); the real voice channel is never created. +**Fix:** Scope `find_channel` by channel type (compare (name, type) tuples); never move a channel whose type doesn't match. +**Verification:** confirmed 2/3 lenses. + +### [HIGH] dotty_doctor model checks look under data/models instead of repo-root models — false FAIL +**File:** `scripts/dotty_doctor.py:128-155` +**Issue:** With config at `data/.config.yaml`, `root = config_path.parent` is `data/`, so the checks look for `data/models/...` while the models actually live at repo-root `models/` — both checks FAIL and the doctor exits non-zero on a healthy standard deploy. +**Fix:** Don't anchor to `config_path.parent`; if it's a `data/` dir use the parent, or search both `root/models` and `root.parent/models`. +**Verification:** confirmed 2/3 lenses. + +### [MEDIUM] startToChat: missing "language" key throws KeyError and feeds raw JSON to the pipeline +**File:** `receiveAudioHandle.py:1057-1063` +**Issue:** `_language_tag = data["language"]` (subscript) inside a try that swallows KeyError; when speaker+content are present but language is absent, `actual_text` stays the RAW JSON envelope and gets run through ASR corrections/intent/LLM, and `current_speaker` is never set. +**Fix:** `data.get("language")`. +**Verification:** unverified — over cap. + +### [MEDIUM] provision.py: read-only permission overwrites never re-applied to existing channels +**File:** `community/discord/provision.py:197-202` +**Issue:** The existing-channel branch fixes only the category and `continue`s, never applying overwrites, so READ_ONLY channels that already exist keep `@everyone` send_messages — the script isn't idempotent for permissions on a public server. +**Fix:** Compute and `existing.edit(overwrites=ov)` in the existing branch when non-empty. +**Verification:** unverified — over cap. + +### [MEDIUM] dotty_doctor check_http treats 404 (and all <500) as pass for the OTA endpoint +**File:** `scripts/dotty_doctor.py:166-168` +**Issue:** A 404 on `/xiaozhi/ota/` — exactly the misconfiguration the doctor exists to catch — is reported as PASS. +**Fix:** Treat 2xx/3xx as pass; flag 4xx as warn/fail with the status code. +**Verification:** unverified — over cap. + +### [MEDIUM] MCP JSON-RPC request ids collide for calls in the same millisecond +**File:** `receiveAudioHandle.py:281, 319, 349, 373, 399, 456, 559` +**Issue:** Every MCP id is `int(time.time()*1000) % 0x7FFFFFFF`; back-to-back sends (e.g. two `set_toggle`, or HEAD+LED at the same timeline mark) get identical ids, so a relay correlating by id can mismatch/drop responses. +**Fix:** Use a monotonic per-connection counter via a shared `_next_mcp_id(conn)` helper. +**Verification:** unverified — over cap. + +### [LOW] dotty_doctor passes Piper without the required .onnx.json config pair +**File:** `scripts/dotty_doctor.py:146-155` +**Issue:** The check globs only `*.onnx`; a missing/stub `.onnx.json` (the exact corrupt-download failure the SenseVoice check was hardened against) still reports green PASS while the TTS provider crashes at runtime. +**Fix:** For each `*.onnx`, assert the sibling `*.onnx.json` exists above a byte floor. +**Verification:** unverified — over cap. + +### [LOW] render_singing_sinsy crashes with ZeroDivisionError on empty/all-comment lyrics +**File:** `scripts/render_singing_sinsy.py:89` +**Issue:** `syllables[i % len(syllables)]` with `len==0` raises an opaque traceback when the lyrics file is empty/comments-only; the notes list is guarded but syllables isn't. +**Fix:** Add a `if not syllables:` guard mirroring the notes guard. +**Verification:** unverified — over cap. + +### [LOW] _send_led_color swallows all exceptions silently +**File:** `receiveAudioHandle.py:285-286` +**Issue:** A bare `except Exception: pass` with no log; a broken WS during a dance silently no-ops every LED update, making "dance ran but no lights" undebuggable. +**Fix:** Log (warn-once) like `_send_led_multi`. +**Verification:** unverified — over cap. + +### [LOW] _encode_midi_to_opus forces ALL set_tempo events to one value +**File:** `receiveAudioHandle.py:850-857` +**Issue:** Rewriting every tempo event to one value flattens intended tempo changes; benign for the current registry, latent for any song with tempo automation. +**Fix:** Rewrite only the first/global tempo, or scale all tempos by the same ratio. +**Verification:** unverified — over cap. + +### [LOW] _handle_dance: singing stream task untracked, so a new utterance cancels choreography but not the audio +**File:** `receiveAudioHandle.py:789-802, 1092-1096` +**Issue:** Choreography is tracked and cancelled on the next turn; the singing audio task handle is never stored, so on barge-in it stops only if it observes `client_abort` flip between 60ms packets — fragile, no explicit cancel. +**Fix:** Store `conn._singing_task` and cancel it alongside `_dance_task`. +**Verification:** unverified — over cap. + +### [LOW] Idle end-prompt re-enters the full intent pipeline via startToChat +**File:** `receiveAudioHandle.py:1208-1211` +**Issue:** The idle-timeout end-prompt is fed through `startToChat`, which runs state/dance/vision detection; a customised prompt containing "dance"/"sleep"/a vision phrase fires that side-effect at the moment the connection is closing. +**Fix:** Submit the end-prompt directly to the LLM, bypassing intent detection (or add a skip flag for system prompts). +**Verification:** unverified — over cap. + +### [LOW] provision.py create_forum drops the computed permission overwrites +**File:** `community/discord/provision.py:204-213` +**Issue:** The FORUM branch omits `overwrites=ov`; benign today (no forum in READ_ONLY) but silently no-ops a future read-only forum lockdown. +**Fix:** Pass `overwrites=ov` to `guild.create_forum`. +**Verification:** unverified — over cap. + +### [QUALITY/LOW] Dead _write_smart_mode_state; stale ZeroClaw docstrings; redundant JSON re-parse; shared-list in-place sort; unreachable _encode_song_to_opus +**File:** `receiveAudioHandle.py:75-81, 634/644/651, 1054-1063/1114-1119, 761-762/928`; `dances.py:111-112, 302` +**Issue:** `_write_smart_mode_state` is never called (xiaozhi side is read-only by design); `_with_room_view_marker` docstring still describes the retired "zeroclaw `_payload`"; `actual_text` is JSON-re-parsed a second time after it's already unwrapped; `_macarena_moves` returns the shared module-global list that `resolve_timeline` sorts in place; `_encode_song_to_opus` (WAV path) is unreachable because every registry entry is `.mid`, so rendered singing `.wav` never plays. +**Fix:** Delete dead code; fix stale docstrings to name PiVoiceLLM; drop the redundant parse; return `list(_MACARENA_TIMELINE)`; wire a `.wav` registry entry or remove the WAV path. +**Verification:** unverified — over cap. + +--- + +## docs: architecture, internals, setup, operations, policy + +### [HIGH] disable-kid-mode.md references non-existent _content_filter and _ensure_emoji_prefix as active +**File:** `docs/cookbook/disable-kid-mode.md:43-44` +**Issue:** The doc claims `_content_filter` and `_ensure_emoji_prefix` "remain active in both modes"; neither symbol exists anywhere (the only content filter, planned per #138, lives unused in bridge/text.py; `_ensure_emoji_prefix` was ZeroClaw-era). Emoji enforcement is now in `textUtils.build_turn_suffix()` / the `ALLOWED_EMOJIS` fallback in openai_compat.py. +**Fix:** Delete the sentence; replace with the real enforcement points and note there is no content-word filter yet (#138). +**Verification:** confirmed 3/3 lenses. + +### [HIGH] disable-kid-mode.md tells the user to restart the bridge, but kid mode is read by xiaozhi-server +**File:** `docs/cookbook/disable-kid-mode.md:19` +**Issue:** Kid mode drives the live voice path via `pi_voice._read_kid_mode()` (and openai_compat) running inside the xiaozhi-server container; restarting only the bridge won't change the suffix. +**Fix:** Change to `docker compose restart xiaozhi-server`; mention the dashboard's live toggle as the no-restart alternative. +**Verification:** confirmed 3/3 lenses. + +### [MEDIUM] Tool count says 'five voice tools'; there are seven (recall_person/remember_person, #53) +**File:** `docs/brain.md:12,58-68`; `docs/voice-pipeline.md:74`; `docs/llm-backends.md:131-145,22`; `docs/COMPATIBILITY.md:16-18`; `COMPATIBILITY.md:18`; `docs/architecture.md:125,151,239`; `docs/protocols.md:249`; `docs/latent-capabilities.md:66`; `docs/modes.md:175` +**Issue:** Ten+ docs undercount the dotty-pi-ext tools as five and omit recall_person/remember_person (registered in `src/index.ts:21-27`). COMPATIBILITY.md also pins firmware v1.2.4 while the submodule is at fw-v1.3.1. +**Fix:** Update every "five voice tools" reference to seven with the full list; bump the firmware version. +**Verification:** confirmed 2/3 (brain.md, voice-pipeline.md, COMPATIBILITY, architecture, protocols, modes) / 3/3 (latent-capabilities). + +### [MEDIUM] architecture.md claims smart-mode model swap is 'now handled in PiVoiceLLM' — it isn't +**File:** `docs/architecture.md:175` +**Issue:** PiVoiceLLM has no smart-mode/model-swap logic; bridge.py and brain.md/modes.md all say the swap is unwired v2 scope. +**Fix:** Replace with "Smart-mode model swap is v2 scope and NOT wired on the PiVoiceLLM path." +**Verification:** confirmed 2/3 lenses. + +### [MEDIUM] Other architecture/protocol doc drift (shared_llm, /api/voice/escalate, port 8000, ZeroClaw paths) +**File:** `docs/architecture.md:248,172-176`; `docs/protocols.md:292`; `docs/voice-mode-entry.md:40`; `docs/cookbook/add-emoji.md:8-37`; `docs/troubleshooting.md:86`; `docs/proactive-greetings.md:66/83,129-157` +**Issue:** `shared_llm singleton` was removed with Tier1Slim; `/api/voice/escalate` does not exist (not "non-functional"); inject-text is on 8003 not 8000; add-emoji.md edits bridge.py for symbols that live in textUtils.py (and never mentions EMOJI_MAP, the load-bearing edit); troubleshooting.md maps a directory mount where compose mounts a file; proactive-greetings.md points all code refs at retired `bridge/` paths and gives a wrong `GREETER_PER_DAY_MAX` default. +**Fix:** Repoint each to the real symbol/path/port; rewrite add-emoji.md around `textUtils.py` + `EMOJI_MAP`. +**Verification:** confirmed 2/3 (proactive-greetings Files) / others unverified — over cap. + +### [LOW] Misc doc nits (container name `bridge` vs `dotty-bridge`, persona source, missing scripts/backup.sh, README section, SceneSynthesisLoop, taxonomy, /admin table, observability metric) +**File:** `SETUP.md:33`; `docs/troubleshooting.md:135,138`; `docs/COMPATIBILITY.md:60`; `COMPATIBILITY.md:59`; `CLAUDE.md:48`/`SETUP.md:9`/`CONTRIBUTING.md:32,70`; `docs/architecture.md:216`/`modes.md:13,202`; `docs/interaction-map.md:65`; `docs/architecture.md:172-176`; `docs/observability.md:74-84`; `docs/cookbook/change-persona.md:24,47`; `docs/faq.md`/`docs/kid-mode.md` +**Issue:** `docker logs bridge` should be `dotty-bridge`; `scripts/backup.sh` doesn't exist; the "Configuring for your environment" README section is gone (now `docs/quickstart.md`); the consumer class is `SceneSynthesisLoop`; the modes taxonomy is the six-state mutex; the /admin table omits /state and /persona; observability.md omits the one live metric (`dotty_content_filter_hits_total`); change-persona/faq/kid-mode misattribute the persona source (loaded by the pi agent from `/root/.pi`, not the extension). +**Fix:** Apply each correction as listed in the per-finding suggestions. +**Verification:** unverified — over cap. + +--- + +## root + policy docs + +### [LOW] CHANGELOG [Unreleased] presents retired ZeroClaw systemd bridge infra as pending +**File:** `CHANGELOG.md:10` +**Issue:** The [Unreleased] Added entry describes a `zeroclaw-bridge.service.template` + `scripts/install-bridge.sh` that no longer exist; the bridge now ships as the `dotty-bridge` container. +**Fix:** Move the entry to the historical ZeroClaw-era section or strike it. +**Verification:** unverified — over cap. + +### [LOW] SECURITY.md threat model cites a retired /api/message bridge endpoint +**File:** `SECURITY.md:23-24` +**Issue:** `/api/message` was the ZeroClaw voice endpoint; the bridge registers no such route and its `/admin/*` is localhost-gated. The real unauthenticated LAN surface is xiaozhi-server's `/xiaozhi/admin/*` and dotty-behaviour's `/api/*`. +**Fix:** Replace the `/api/message` reference with the real surface. +**Verification:** confirmed 3/3 lenses. + +### [LOW] ROADMAP / FAQ / persona docs carry stale Tier1Slim framing +**File:** `ROADMAP.md:52`; `docs/faq.md:59,73`; `docs/kid-mode.md:50,59,225,297`; `personas/dotty_voice.md` (title) +**Issue:** ROADMAP frames latency as a "two-tier path" (Tier1Slim removed); FAQ/kid-mode imply `personas/dotty_voice.md` is the live agent persona while the only wired persona is `personas/default.md` (OpenAICompat) and the agent loads from `/root/.pi`; `dotty_voice.md` is still titled "(Tier 1)". +**Fix:** Reword to the single PiVoiceLLM + think_hard design; clarify the persona mapping; drop the "Tier 1" label. +**Verification:** unverified — over cap. + +--- + +## monitoring/grafana-dashboard.json (completeness sweep) + +### [HIGH] Six of eight dashboard panels query metrics that are never recorded +**File:** `monitoring/grafana-dashboard.json:107-461, 523-537` +**Issue:** Only `dotty_kid_mode_active` and `dotty_content_filter_hits_total` are ever written; the first-audio-latency, request-rate, error-rate, smart-mode, perception-events, and calendar-failure panels all render empty on import (7 of 8 panels broken). +**Fix:** Wire the metrics at their call sites, or trim the dashboard to live panels and add a Content-Filter-hits panel. +**Verification:** confirmed 1/1 lenses. + +### [HIGH] Grafana description + tag still name the retired ZeroClaw system +**File:** `monitoring/grafana-dashboard.json:18, 629, 660` +**Issue:** The description says "zeroclaw-bridge ... ACP sessions" and the tags array carries a literal `zeroclaw` tag (polluting Grafana tag search). +**Fix:** Rename the description, drop the ACP mention, remove the `zeroclaw` tag. +**Verification:** confirmed 1/1 lenses. + +### [HIGH] Active ACP sessions panel queries a dead metric +**File:** `monitoring/grafana-dashboard.json:266-337` +**Issue:** Panel id 4 queries `dotty_active_acp_sessions` (defined but never observed; ACP retired in #36), with a red/green threshold that shows a permanently RED "ACP child not respawning" for a subsystem that no longer exists. +**Fix:** Delete panel id 4. +**Verification:** confirmed 1/1 lenses. + +### [MEDIUM/LOW] monitoring README/observability.md overstate live coverage; metrics.py identity still 'zeroclaw-bridge' +**File:** `monitoring/README.md:5-7`; `docs/observability.md:8-10,66-84`; `bridge/metrics.py:1,36` +**Issue:** The docs list 9 metrics as emitted while only 2 are, and omit the one live metric (`dotty_content_filter_hits_total`); the metrics module docstring/logger are still named `zeroclaw-bridge`. +**Fix:** Mark unwired metrics as "defined, not yet wired", add the content-filter row, rename the logger to `dotty-bridge.metrics`. +**Verification:** low-sev, unverified. + +--- + +## gap: Emoji-glyph → firmware-emotion contract + +### [MEDIUM] add-emoji.md omits the EMOJI_MAP edit and points at stale bridge.py locations +**File:** `docs/cookbook/add-emoji.md:8-37` +**Issue:** The only edit that makes a new glyph produce a face is adding it to `EMOJI_MAP` in `custom-providers/textUtils.py` — the doc never mentions it, and instead tells the user to edit `bridge.py` `ALLOWED_EMOJIS`/`_BASE_SUFFIX`, which now live in textUtils.py (and bridge.py is off the voice path). Following it verbatim adds a glyph that passes prompt enforcement but never resolves to an emotion frame. +**Fix:** Rewrite around textUtils.py: add to `ALLOWED_EMOJIS`, add a glyph→firmware-name entry to `EMOJI_MAP` (matching a `stackchan_display.cc` strcmp branch), update `_BASE_SUFFIX` + the `.config.yaml` prompt, redeploy xiaozhi-server. +**Verification:** low-sev, unverified. + +### [INFO] Emoji→firmware-emotion contract verified intact (no verb-form fallthrough) +**File:** `custom-providers/textUtils.py:64-90` + `firmware/.../stackchan_display.cc:355-432` +**Issue:** None — the hypothesised `laugh`/`laughing` mismatch is NOT a bug; `EMOJI_MAP` already emits the gerund/past-tense names the firmware expects, and all 9 protocol emojis resolve end-to-end. `laughing≡happy` and `crying≡sad` collapse by design. +**Fix:** No code change; optionally document the two intentional collapses. +**Verification:** low-sev, unverified. + +### [LOW] CLAUDE.md says the firmware parses the emoji glyph; the server does +**File:** `CLAUDE.md:84` +**Issue:** The firmware reads a pre-resolved english emotion string; the glyph→name translation happens server-side in `get_emotion()`/`EMOJI_MAP`. This imprecision is the conceptual root of the add-emoji.md misdirection. +**Fix:** Amend to say xiaozhi-server's `get_emotion()`/`EMOJI_MAP` parses the leading emoji into an emotion name and the firmware renders the face. +**Verification:** low-sev, unverified. + +--- + +## gap: Test-coverage gaps for confirmed bugs + +### [HIGH] TTS processed_chars over-count is untested AND the provider package is excluded from coverage +**File:** `custom-providers/edge_stream/edge_stream.py:67-78`; `custom-providers/piper_local/piper_local.py:114-125` +**Issue:** Both providers carry the confirmed over-count bug, neither is in `[tool.coverage.run] source`, and no test imports them — so the 56% floor is computed over a source set that physically cannot see these lines; a fix can regress green. +**Fix:** Fix the bug, add a pure-logic unit test, add both packages to the coverage source and the ci.yml pytest targets. +**Verification:** confirmed 1/1 lenses. + +### [HIGH] OTA _is_higher_version pre-release ordering bug is outside the coverage source set with no covering test +**File:** `custom-providers/xiaozhi-patches/ota_handler.py:24-43` +**Issue:** The pure version-compare functions are trivially testable but 0% covered and invisible to the CI floor (xiaozhi-patches is only ruff-linted). +**Fix:** Replace the digits-only parser with a semver-aware comparator; add a version-table unit test; add the module to coverage and wire a test dir into ci.yml. +**Verification:** confirmed 1/1 lenses. + +### [HIGH] Firmware C++ concurrency code has no native test harness and is invisible to every CI gate +**File:** `firmware/firmware/main/stackchan/face/face_detection_result.h:19-46` (+ camera path) +**Issue:** The seqlock torn-read and camera UAF findings have no host/native unit-test target; ci.yml has no firmware build-or-test job, so these subtle-ordering bugs can be refactored with zero signal. +**Fix:** Stand up a host-native (TSan-enabled) test target for the hardware-free headers (FaceDetectionResult compiles standalone); add a firmware-host-tests job to ci.yml. +**Verification:** confirmed 1/1 lenses. + +### [MEDIUM] openai_compat emoji-prefix + turn-suffix logic untested and excluded from coverage; 56% floor is blind to the high-sev modules +**File:** `custom-providers/openai_compat/openai_compat.py`; `pyproject.toml:22-33` + `.github/workflows/ci.yml:88-99` +**Issue:** OpenAICompat (the only PiVoiceLLM fallback) is entirely untested and outside the coverage source, and the `--cov-fail-under=56` gate measures only 3 of the shipped Python modules — every module carrying a confirmed high/medium bug (edge_stream, piper_local, openai_compat, asr, ota_handler) is outside the source set, so the floor provides no regression protection for them. +**Fix:** Add a streaming-generator unit test for the emoji/suffix logic; expand `[tool.coverage.run] source` + the ci.yml pytest targets as part of each fix; document that the floor is per-source-set. +**Verification:** low-sev, unverified. + +--- + +## gap: firmware/server/ subtree + +### [INFO] firmware/server/ is vendored upstream m5stack Go backend — dead in Dotty, undocumented as such +**File:** `firmware/server/` (entire subtree) +**Issue:** The hypothesis that this is live OTA/provisioning that could drift from `ota_handler.py` does NOT hold — `firmware/` is a submodule and `firmware/server/` is the stock m5stack GoFrame social-feed backend vendored inside the fork, referenced nowhere in the parent repo, never built by the ESP-IDF flow, and not COPY'd by any container. The flagged config files are upstream placeholders (GoFrame default ports, boilerplate image prefixes), not Dotty values. The only real issue is navigational: CLAUDE.md never says the submodule carries this unrelated tree, so a reviewer can mistake it for live infra. +**Fix:** No code change. Add one line to CLAUDE.md's "Firmware iteration" section noting the firmware submodule vendors upstream `server/`/`app/`/`remote/` trees that are NOT part of Dotty's stack; Dotty's OTA is `custom-providers/xiaozhi-patches/ota_handler.py`. +**Verification:** low-sev, unverified. + +--- + +## Open issues & PRs + +| Ref | Disposition | Action | +|-----|-------------|--------| +| #104 | still-valid (P1) | Keep open. Implement fix (b): give server-pushed audio its own `server_push_sentence_id`; pair with #105. | +| #105 | still-valid (P1) | Keep open. Start with SileroVAD threshold bump; then ASR RMS floor / wake-word gating. Triage with #104. | +| #138 | still-valid (planned) | Keep open. If picked up: lift `content_filter()` into a shared module, gate on kid_mode, wire onto TTS-bound text, red-team on device. | +| #125 | still-valid | Resolved by PR #139 (deploy-dotty-pi.sh). Keep open until merged. | +| #98 / #135 | still-valid (speculative) | Keep open. Largely delivered by PR #140 (SenseVoiceOnnx). No urgency — funasr path works. | +| #60 | still-valid | Keep open. Add docker.sock mount to bridge compose + a CSRF-guarded restart POST. | +| #21 | still-valid | Keep open. Implement firmware fix (b): re-fire `state_changed` on audio-channel (re)open. | +| #77 | still-valid (low) | Keep open. Add a CSRF-guarded `/ui/reload-config` for the reload-safe env subset. | +| #122 | still-valid (tracker) | Keep open as bench tracker. Drain 2-3 checklist items per session; note #39 partially passed. | +| #121 | still-valid (tracker) | Keep open as firmware observation umbrella. No code action until a symptom recurs instrumentably. | +| PR #137 | merge | Docs-only AI transparency policy. Verified sound. Optionally fold in the two cosmetic doc nits. | +| PR #139 | merge | deploy-dotty-pi.sh faithfully mirrors the sibling scripts. Optionally add IMAGE_TAG/node_modules/image-running guards. | +| PR #140 | changes-needed | Sound SenseVoiceOnnx provider. **Blocker:** fix the Makefile `doctor` to SKIP-when-absent (it hard-fails every default FunASR install). | +| PR #141 | **close** | De-safety fork mislabeled as a fix (blanks safety suffix, drops system messages, disables perception, adds NSFW+PII persona). Extract only the piper emoji-strip into a standalone PR. | + +--- + +## Recommended priorities + +1. **Close PR #141** (F120–F124) — it removes the child-safety guardrails and adds NSFW+PII content to a public kid-safe repo. Highest urgency; no code archaeology needed. +2. **Break the LAN→host-root chain** (G9): add a docker-socket-proxy + a shared-secret auth gate on `/xiaozhi/admin/*`, confine play-asset to the songs base (G10), and add non-root `USER` to the four Dockerfiles (G11). +3. **Fix the shipped all-in-one compose** (G18, G19, G21): add the docker.sock + xiaozhi-patches mounts and the VISION_BRIDGE_URL var, or re-scope all-in-one to OpenAICompat-only with an honest header. +4. **Close the kid-safety leaks**: namespace-guard `memory_lookup` against `person_pending` (F251), run dashboard say/start-story through `content_filter()` (F186), and pursue #138. +5. **Fix the room_view id/display_name + multi-word name mismatches** (F38/F201, F161/F202) — the entire named-greet feature silently regresses; also de-duplicate the double greeting (F39). +6. **Fix the TTS truncation bugs** (F148, F26/F27 over-count) and the OTA version comparator (F141), and add them to the coverage source set (G13, G14, G17). +7. **Un-block the event loop in the bridge** (F128, F185) — wrap blocking HTTP and convo-log parsing in `to_thread` — and exempt `/admin/` from CSRF (F129) so the operator back-channel works. +8. **Fix the dotty-pi-ext HTTP client timeout-before-body bug** (F212) — a stalled body wedges the whole voice turn — and the dead bridge perception EventSource (F226). +9. **Repair dotty_doctor** (F257 model path, F51 404-as-pass, F216 Piper config) so `make doctor` is trustworthy on a standard deploy. +10. **Clean the monitoring dashboard + ZeroClaw/tool-count doc drift** (G1–G4, F76/F77, the "five vs seven tools" sweep) so the docs and the Grafana import match reality. diff --git a/CHANGELOG.md b/CHANGELOG.md index 0f22feb5..3ec70318 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -7,6 +7,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/). ## [Unreleased] ### Added +- **Overnight A/V runner can run unattended** (#182) — opt-in `--mic-opener admin_say` opens the microphone through `/xiaozhi/admin/say` before a warm case; playback is also proven by the robot's own recognition; reference-microphone slips and extra speech in the room are `INCONCLUSIVE` rather than failures; a case the evidence cannot settle is set aside after three attempts. `dotty_av_clips.py export --crop` fills the 9:16 frame from a portrait source region. See `docs/dotty-av-tests.md`. - **Kid Mode: blocked-words content filter on the live voice path** (#157, closes the gap tracked in #138) — the pure matcher (three severity tiers + the kid-safe replacement) moved from `bridge/text.py` into the shared `custom-providers/textUtils.py`, the single source of truth both containers import; the bridge keeps its metrics/safety-ring/logging wrapper on top with unchanged behaviour. Both live LLM providers (`PiVoiceLLM`, `OpenAICompat`) now wrap their TTS-bound streams in `filter_tts_stream()`. In Kid Mode it drains and checks the complete response before TTS, then emits the original clean chunks or atomically replaces a blocked turn; outside Kid Mode it remains a transparent streaming passthrough. Full-turn consumption also prevents PiVoiceLLM from abandoning the RPC iterator before `agent_end` and leaking stale frames into the next turn. Honest caveat (see `docs/faq.md`): a word-level blocklist is a weak, bypassable backstop — prompt steering remains the primary defence. **Bench: needs on-device red-team verification before release sign-off.** - **Admin-API auth: `X-Admin-Token` across the whole stack** (#149, #150, #151, #152) — xiaozhi-server's `/xiaozhi/admin/*` routes now accept an `X-Admin-Token` shared secret (`DOTTY_ADMIN_TOKEN`, timing-safe compare, permissive when unset), and all three callers — the bridge dashboard, dotty-behaviour's `XiaozhiAdminClient`, and the dotty-pi voice tools' `adminFetch` — send it. The secret is now **plumbed end-to-end**: `make setup` generates one into `.env`, the compose template / all-in-one pass it to xiaozhi-server, the three service compose files document where their copy lives, and `.env.example` + `SETUP.md` §10 describe the enable-everywhere-or-nowhere semantics. Previously the code shipped with no config path, so every deploy silently stayed permissive. - **PersonResolver — one answer to "who is this?"** (`dotty-behaviour/household/resolver.py`) — identity resolution was smeared across consumers, and the 2026-06-06 audit found four separate identity bugs because of it. All resolution now funnels through one module with `Person.id` as the canonical key space. Fixed by the consolidation: **room_view roster recognition silently failing whenever `id != display_name`** (the VLM echoes display names; validation compared ids — confirmed 3/3, both the core and greeter paths), **multi-word display names never matching** (the NAME parser was single-token, so "Mary Anne" was a 100% silent miss — confirmed 3/3), **the greeter's calendar lookup dropping a person's own events on a case mismatch** (`[Hudson]` ≠ `hudson` — confirmed 2/3), and **bracketless `calendar_prefix:` YAML never matching**. `summarize_for_prompt` now matches person tags case-insensitively and accepts the resolver's tag set; the room_view test fakes were also fixed to carry real ids (the old fakes re-derived ids from display names, which is exactly what masked the bug). @@ -26,6 +27,12 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/). - **`docs/multi-daemon-split.md`, `docs/advanced/multi-host.md`** — both documented ZeroClaw-host topologies that no longer exist. ### Fixed +- **Voice: one abort no longer silences every later reply** (#176) — a device or admin abort set `conn.client_abort` and nothing on the `nointent` chat path cleared it, so ASR and the LLM kept running but no TTS was sent until the robot reconnected. `startToChat` now clears the flag when a new user turn is accepted. Reproduced and verified on the physical robot. +- **Live voice path now has a Dotty system prompt** (#177) — `PiClient` spawned pi with no `--system-prompt`, so pi's built-in coding-assistant prompt was the only identity the model saw ("I'm just a tiny coding assistant…"). `personas/pi_voice.md` is now passed via `--system-prompt` (`DOTTY_PI_SYSTEM_PROMPT_FILE` overrides; a missing file keeps the old behaviour). The persona also tells the 4B model to answer briefly and stop. +- **WhisperLocal drops near-silence hallucinations** (#105) — transcripts whose mean `no_speech_prob` exceeds `no_speech_threshold` (default 0.6, `1.0` disables) are discarded instead of being answered ("Thank you." after the robot stops speaking). +- **Memory: prompt scaffolding is no longer stored or spoken** (#186) — every voice turn was logged to `brain.db` with the per-turn tool-routing and HARD CONSTRAINTS text attached, and an explicit recall ("do you remember…") spoke `memory_lookup`'s raw rows. The turn logger now stores only what the person said, `memory_lookup` cleans older rows at read time, and recall passes the search results to the model as context so the answer is phrased naturally. +- **think_hard no longer fails on a cold start** (#187) — its request timeout defaulted to 30 s, shorter than the reasoner's 30–50 s cold load, so the first think-hard after an idle period was cancelled mid-load. Default is now 90 s (as the pre-cutover bridge unit had it), the direct tool runner allows 105 s, and Dotty says "Let me think hard about that." before the wait. +- **PiVoice RPC turns are serialized** (#175, thanks @0Downtime) — one connection can no longer interleave two pi transactions and consume each other's frames. - **No-GPU ASR path no longer crash-loops on first run** (#124, #136) — `make fetch-models` was requesting two SenseVoiceSmall filenames that don't exist on Hugging Face (`tokens.json` and `chn_jpn_yue_eng_ko_spectral.fbank.conf.yaml`); the real SentencePiece tokenizer is `chn_jpn_yue_eng_ko_spectok.bpe.model`. Both 404s were silently saved as 15-byte "Entry not found" stubs (curl had no `--fail`), so funasr loaded with `bpemodel=None` and `xiaozhi-esp32-server` crash-looped on every GPU-less host. The file list is corrected; **all `fetch-models` downloads now fail loudly** (`curl --fail --retry` + a size floor + delete-on-failure) instead of saving junk; and **`make doctor` now size-checks the required SenseVoice assets** so a corrupt download FAILs instead of passing. Huge thanks to **[@miltieIV2](https://github.com/miltieIV2)** — a meticulous bug report *and* a self-driven root-cause that pinned it on the `HAS_CUDA=0` FunASR switch. A lighter int8 sherpa-onnx SenseVoice runtime (no PyTorch) for Pi-class hosts has landed as the opt-in `SenseVoiceOnnx` provider (#135; see the Added section above). ## [server-v0.1.0] - 2026-05-17 diff --git a/CLAUDE.md b/CLAUDE.md index ee7f4c57..8bfc22d8 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -27,7 +27,7 @@ The voice path runs through a single LLM provider — `PiVoiceLLM`, selected via ``` StackChan hardware → configured persona - │ ESP32-S3, xiaozhi firmware (built from m5stack/StackChan source) + │ ESP32-S3, xiaozhi firmware (pinned BrettKinny/StackChan@dotty fork) │ WiFi / WebSocket (Xiaozhi protocol) ▼ xiaozhi-esp32-server (Docker) @@ -197,7 +197,8 @@ For hardware specs, protocol details, model internals, latent capabilities, and - xiaozhi-esp32-server: https://github.com/xinnan-tech/xiaozhi-esp32-server - xiaozhi-esp32 firmware (upstream): https://github.com/78/xiaozhi-esp32 -- StackChan (hardware + firmware patches): https://github.com/m5stack/StackChan +- StackChan upstream (hardware + base firmware): https://github.com/m5stack/StackChan +- Dotty firmware fork (the pinned build source): https://github.com/BrettKinny/StackChan/tree/dotty - Emotion protocol: https://xiaozhi.dev/en/docs/development/emotion/ ## Agent skills diff --git a/COMPATIBILITY.md b/COMPATIBILITY.md index 84d56c50..c64bc4fb 100644 --- a/COMPATIBILITY.md +++ b/COMPATIBILITY.md @@ -13,11 +13,16 @@ For protocol wire formats see `docs/protocols.md`. | Component | Current Version | Protocol / Interface | Breaking Change Policy | |---|---|---|---| -| StackChan firmware (m5stack/StackChan v1.2.4) | v1.2.4 | Xiaozhi WebSocket protocol, MCP over WS (JSON-RPC 2.0) | Pin firmware to a known-good build; do not OTA-update without verifying server compatibility first | +| StackChan firmware (`BrettKinny/StackChan@dotty`) | `fw-v1.3.3` release; submodule `969c2b2` | Xiaozhi WebSocket protocol, MCP over WS (JSON-RPC 2.0), StateManager event contract | Build the pinned submodule; do not substitute the official upstream tree or OTA-update without coordinated verification | | xiaozhi-esp32-server (local build) | `xiaozhi-esp32-server-piper:local` | Custom LLM provider API, `.config.yaml` schema, Xiaozhi WS server | Rebuild image only after checking upstream changelog for provider API or config schema changes | -| dotty-pi (pi agent) | `dotty-pi:0.1.0` | pi RPC (JSONL over stdio), the five `dotty-pi-ext` voice tools | Pin the image tag; pi-version or model changes need end-to-end cutover testing | +| dotty-pi (pi agent) | `dotty-pi:0.1.0` | pi RPC (JSONL over stdio), the seven `dotty-pi-ext` voice tools | Pin the image tag; pi-version or model changes need end-to-end cutover testing | | dotty-behaviour | `dotty-behaviour:0.1.0` | HTTP API (`/api/perception/*`, `/api/vision/*`, `/api/audio/*`, `/health`) | Endpoint signatures stable; perception event-schema changes require firmware + xiaozhi review | -| bridge.py (dashboard) | unversioned (HEAD) | `/ui` dashboard, `/admin/*`, `/health` | Dashboard/admin service only post-#36; admin route changes require updating dashboard callers | +| bridge.py (dashboard) | `dotty-bridge:0.1.0` image from repo HEAD | `/ui` dashboard, `/admin/*`, `/health` | Dashboard/admin service only post-#36; admin route changes require updating dashboard callers | + +The public `fw-v1.3.3` superproject tag (commit `24a009c`, 2026-07-12) is the +current coordinated release pointer. It pins firmware submodule `969c2b2` and +the matching server-side patches. Draft PRs and a dirty submodule are not a +released compatibility set. ## What counts as a breaking change @@ -42,27 +47,28 @@ Any of the following require coordinated updates across components: ## Versioning strategy -No formal versioning is adopted yet (tracked in -[ROADMAP.md](ROADMAP.md#community-wishlist) under "Firmware/server -compatibility matrix"). When adopted, the plan is: +The repo uses separate tag namespaces: + +- `server-vX.Y.Z` for server-only release milestones. +- `fw-vX.Y.Z` for coordinated firmware release pointers in this superproject. -- Separate tag namespaces: `server-vX.Y.Z` and `fw-vX.Y.Z`. -- This matrix will document which server versions are compatible with which - firmware versions. -- The bridge will carry its own version once it moves to a tagged release - cadence. +Container image tags remain `0.1.0` today and are not sufficient by themselves +to identify the exact source revision; retain the Git commit/tag alongside a +deployment record. ## Upgrade guidance 1. **Check this matrix first.** Confirm the component you are upgrading is compatible with the versions of the other components you are running. -2. **Back up before upgrading.** Run `scripts/backup.sh` (or the equivalent - manual steps) to snapshot config, persona files, and bridge state. +2. **Back up before upgrading.** Manually snapshot `.env`, rendered + `data/.config.yaml`, persona/household files, `brain.db` (including WAL/SHM + companions), and the bridge state directory. This repo does not currently + ship an automated backup script. 3. **Upgrade one component at a time.** Validate with a health check (`curl http://:8090/health` and `:8081/health`) plus a live voice turn before moving to the next component. -4. **Tail logs during validation.** Watch both the xiaozhi-server container - logs and the bridge journal simultaneously to catch mismatches early. +4. **Tail logs during validation.** Watch the xiaozhi-server, dotty-pi, + dotty-behaviour, and bridge container logs together to catch mismatches. 5. **Roll back if broken.** Restore from the backup taken in step 2 and revert to the previous image or binary. @@ -106,4 +112,4 @@ versions work with which firmware versions. --- -Last verified: 2026-05-22. +Last verified against the repository and `fw-v1.3.3` pin: 2026-07-16. diff --git a/Dockerfile b/Dockerfile index 08b9caa9..f9748586 100644 --- a/Dockerfile +++ b/Dockerfile @@ -3,6 +3,11 @@ FROM ghcr.io/xinnan-tech/xiaozhi-esp32-server@sha256:3accd82a7d1a6c01c58f32f6199 RUN pip install --no-cache-dir piper-tts scipy numpy mido faster-whisper sherpa-onnx==1.13.2 +# Patch upstream TTS consumers to treat explicitly registered server-push +# sentence IDs as independent from chat-turn stale-message arbitration (#104). +COPY scripts/patch-tts-server-push.py /tmp/patch-tts-server-push.py +RUN python /tmp/patch-tts-server-push.py /opt/xiaozhi-esp32-server/core/providers/tts + # fluidsynth + General MIDI soundfont for runtime rendering of dance/song MIDI # files to Opus. Installed as the LAST layer so iteration on Python deps above # doesn't invalidate the soundfont download (~141MB). diff --git a/Makefile b/Makefile index 4f798bc0..aa182ebc 100644 --- a/Makefile +++ b/Makefile @@ -221,9 +221,10 @@ setup: _preflight-compose ## Interactive first-run wizard (re-runnable; remember echo -e "$(GREEN)$(BOLD)Setup complete.$(RESET)"; \ echo ""; \ echo "Next steps:"; \ - echo " 1. Flash the StackChan firmware (see SETUP.md or m5stack/StackChan repo)."; \ - echo " 2. In the device's Advanced Options, set the OTA URL to:"; \ - echo " http://$$XIAOZHI_HOST:8003/xiaozhi/ota/"; \ + echo " 1. Build and flash StackChan firmware with this compiled setting:"; \ + echo " CONFIG_OTA_URL=\"http://$$XIAOZHI_HOST:8003/xiaozhi/ota/\""; \ + echo " See SETUP.md. The on-device Settings app has no OTA URL editor."; \ + echo " 2. Provision the robot's 2.4 GHz Wi-Fi using its displayed setup flow."; \ echo " 3. Run 'make doctor' to verify everything is healthy."; \ echo "" diff --git a/README.md b/README.md index df050558..b26d5c1b 100644 --- a/README.md +++ b/README.md @@ -36,7 +36,7 @@ Full policy: [`AI_TRANSPARENCY.md`](./AI_TRANSPARENCY.md). - **Streaming responses** — the bridge streams LLM output to the voice pipeline for lower perceived latency. - **Emoji expressions** — every response starts with an emoji that the firmware maps to a face animation (smile, laugh, sad, surprise, thinking, angry, love, sleepy, neutral). - **Voice tools** — the pi agent can search its memory, escalate hard questions to a bigger model, take a photo, and play songs, all mid-conversation. -- **States, toggles & LEDs** — a six-state mutex (`idle / talk / story_time / security / sleep / dance`) plus two orthogonal toggles (`kid_mode`, `smart_mode`), all owned by the firmware StateManager and surfaced on the 12-pixel LED ring. Shipped on the active firmware fork (commit `d78118b`, 2026-04-27); the `firmware/firmware/` submodule pin in this repo lags, so flash from the active fork to get it. See "States, Toggles & LEDs" below and [`docs/modes.md`](./docs/modes.md). +- **States, toggles & LEDs** — a six-state mutex (`idle / talk / story_time / security / sleep / dance`) plus two orthogonal toggles (`kid_mode`, `smart_mode`), all owned by the firmware StateManager and surfaced on the 12-pixel LED ring. Shipped on the active firmware fork (commit `d78118b`, 2026-04-27) and included in the release pin checked on 2026-09-12 (`969c2b2`). Later fixes may still differ between the release pin and active fork; see [`docs/modes.md`](./docs/modes.md). - **Vision (camera)** — the robot's built-in camera can capture images for multimodal LLM queries. - **Privacy LEDs** — hardware-bound mic (green) and camera (red) indicators on the LED ring. They light from the codec/camera enable signals via RAII guards, so a misbehaving server or model can't capture with the lights off. - **Calendar context** — optional calendar integration feeds upcoming events into the conversation context. @@ -46,7 +46,7 @@ Full policy: [`AI_TRANSPARENCY.md`](./AI_TRANSPARENCY.md). Behaviour is a **six-state mutex** (`idle / talk / story_time / security / sleep / dance`) plus two orthogonal toggles (`kid_mode`, `smart_mode`), all owned by the firmware StateManager (shipped on the active fork in commit `d78118b`, 2026-04-27; bench checks tracked in [#38](https://github.com/BrettKinny/dotty-stackchan/issues/38)). Voice phrases, camera edges, and dashboard controls all flow through it. -> Note: the `firmware/firmware/` submodule pin in this repo deliberately lags the active fork — flashing from the submodule won't give you Phase 4 yet. See the "Firmware iteration" section in [`CLAUDE.md`](./CLAUDE.md) and the submodule-pin caveat in [`docs/modes.md`](./docs/modes.md). +> Note: `firmware/` is a release pointer, not the active development checkout. The historical `35f701a` pin predates Phase 4; the pin checked on 2026-09-12 (`969c2b2`) includes it. Check your actual revision for subsequent fixes. See the "Firmware iteration" section in [`CLAUDE.md`](./CLAUDE.md) and the submodule-pin caveat in [`docs/modes.md`](./docs/modes.md). The 12-pixel LED ring shows the current state at a glance. **Left ring 0-5 is the state arc** — all six pixels paint the state colour, matching the dashboard's state buttons: @@ -56,12 +56,12 @@ The 12-pixel LED ring shows the current state at a glance. **Left ring 0-5 is th | 🟢 | `talk` — conversation engaged. | | 🟠 | `story_time` — long-running interactive story. | | ⚪ | `security` — watching the room (1 Hz white flash). | -| 🔵 | `sleep` — quiescent, mic open for "wake up". | +| 🔵 | `sleep` — privacy sleep: face detection, voice processing, and wake-word detection off; wake by head touch or explicit dashboard state change. | | 🟣 | `dance` — rainbow sweep + choreography. | On the right ring, **indices 8-9 are toggle pips** for kid_mode (salmon pink) and smart_mode (orange), and **index 11 (bottom) lights red while you have the turn** (LISTENING). The `idle → talk` transition fires on `face_detected` from the firmware; VLM identity recognition runs in parallel and feeds the LLM context. -> Heads up: that right-ring layout is the active-fork StateManager. On the firmware **submodule pinned in this repo**, pixels 6 and 11 instead drive the **privacy LEDs** — 6 = mic (green), 11 = camera (red) — and the StateManager pips arrive once the submodule catches up to the active fork. +> Heads up: the historical pre-StateManager build (`35f701a`) used a different right-ring layout: 6 = mic (green), 11 = camera (red). The release pin checked on 2026-09-12 already includes StateManager. Interpret indicators against the actual firmware revision, not the old pre-StateManager layout. Full state taxonomy, colour palette, transition diagram, and per-state backing architecture: [`docs/modes.md`](./docs/modes.md). diff --git a/SETUP.md b/SETUP.md index fee1f2e6..a775bcdb 100644 --- a/SETUP.md +++ b/SETUP.md @@ -48,34 +48,38 @@ There is **no SoftAP captive portal** in stock firmware (some older third-party xiaozhi builds had one; the shipped firmware does not). To run the StackChan **fully self-hosted** (no phone-app account, no vendor cloud, -your own xiaozhi-server as the endpoint), you need to **reflash the device -with firmware built from the open source tree**. - -The upstream firmware lives at **https://github.com/m5stack/StackChan**: -- `firmware/` — M5Stack's patches + ESP-IDF project wrapper -- `firmware/fetch_repos.py` — pulls `78/xiaozhi-esp32` as a dependency and - applies StackChan-specific patches, adding the board target - `CONFIG_BOARD_TYPE_M5STACK_STACK_CHAN` +your own xiaozhi-server as the endpoint), reflash it with Dotty's pinned +`BrettKinny/StackChan@dotty` firmware fork. The official `m5stack/StackChan` +tree is its upstream, but it does not contain Dotty's state, motion, LED, and +perception changes. + +The StackChan's on-device Settings app has no **Advanced Options** or OTA URL +editor. The similarly named **Advanced** tab documented by Xiaozhi belongs to +its browser-based SoftAP captive portal and appears only when that provisioning +mode is active. The current `fw-v1.3.3` prebuilt is compiled for the +maintainer's LAN, so another self-hosted deployment must build from source with +its own `CONFIG_OTA_URL` as shown below. + +The exact firmware source is the `firmware/` submodule of this repo. It pins +`BrettKinny/StackChan@dotty`, whose `firmware/fetch_repos.py` pulls +`78/xiaozhi-esp32` v2.2.4 and applies the StackChan integration patch. --- ## 2. Build and flash open firmware -> **Note — build flow documented from first-pass session findings; not yet -> end-to-end verified. Will be updated after a successful first flash.** - Requires **ESP-IDF v5.5.4**. Easiest path: the official `espressif/idf:v5.5.4` Docker image. ### 2a. Clone and configure ```bash -git clone https://github.com/m5stack/StackChan.git -cd StackChan/firmware +git clone --recursive https://github.com/BrettKinny/dotty-stackchan.git +cd dotty-stackchan/firmware/firmware ``` Point the firmware at your xiaozhi-server for OTA. Edit -`firmware/sdkconfig.defaults` and add (or modify) the line: +`sdkconfig.defaults` and add (or modify) the line: ``` CONFIG_OTA_URL="http://:8003/xiaozhi/ota/" @@ -83,10 +87,20 @@ CONFIG_OTA_URL="http://:8003/xiaozhi/ota/" Trailing slash matters — that's the path the server exposes. +`sdkconfig.defaults` seeds a newly generated `sdkconfig`; it does not overwrite +an existing one. On a repeat build, either make the same change through +`idf.py menuconfig` (**Xiaozhi Assistant → Default OTA URL**) or remove the +generated `sdkconfig` before rebuilding. Verify the effective value after +configuration with: + +```bash +grep '^CONFIG_OTA_URL=' sdkconfig +``` + ### 2b. Build inside the IDF container ```bash -docker run --rm -it -v "$PWD/..":/project -w /project/firmware \ +docker run --rm -it -v "$PWD":/project -w /project \ espressif/idf:v5.5.4 \ bash -c "python3 fetch_repos.py && idf.py build" ``` @@ -101,7 +115,7 @@ Linux (`/dev/cu.usbmodem*` on macOS — adapt the `--device` flag). ```bash docker run --rm -it --device=/dev/ttyACM0 \ - -v "$PWD/..":/project -w /project/firmware \ + -v "$PWD":/project -w /project \ espressif/idf:v5.5.4 \ idf.py -p /dev/ttyACM0 flash ``` @@ -111,14 +125,23 @@ flow), use that instead. ### 2d. First boot after flash -- No pairing-code screen -- No BLE provisioning step -- The device boots, loads WiFi credentials compiled into the firmware - (or, if you left WiFi unconfigured, whatever fallback the upstream - build offers — consult the xiaozhi-esp32 README for the current default - behaviour) -- It POSTs to `http://:8003/xiaozhi/ota/`, gets back the - WebSocket endpoint, and connects +Seeing **“Welcome! Let's get started.” is expected on a fresh flash**. It is +the StackChan app-configuration wizard, controlled by an NVS flag; it does not +mean the compiled `CONFIG_OTA_URL` was ignored. + +For a self-hosted setup: + +1. Tap **Skip** on the welcome screen. This bypasses the M5Stack account/BLE + wizard for the current boot. +2. If the launcher remains visible, open **SETUP**, then use its home control + to exit. Closing SETUP starts the Xiaozhi voice application. +3. With no saved SSID, the pinned build enters its default Xiaozhi hotspot + provisioner. Join the temporary `Xiaozhi-*` Wi-Fi network and open the URL + shown on the robot (normally `http://192.168.4.1`). +4. Select the robot's 2.4 GHz Wi-Fi network. The browser portal also has an + **Advanced** tab where you can confirm or override `ota_url`. +5. After Wi-Fi connects, the device POSTs to the compiled/persisted OTA URL, + receives the WebSocket endpoint, and connects. Tail the server logs while the device boots so you can watch the handshake happen (see step 4 below). @@ -127,18 +150,13 @@ handshake happen (see step 4 below). ## 3. WiFi credentials -Two options, depending on what the upstream xiaozhi-esp32 build exposes at -the version you pulled: +The pinned firmware does **not** define `CONFIG_WIFI_SSID` or +`CONFIG_WIFI_PASSWORD`; adding those names to `sdkconfig.defaults` has no +effect. Wi-Fi credentials live in NVS and are supplied through the Xiaozhi +hotspot portal described above (or retained from an earlier compatible +provisioning). -- **Compile-time WiFi credentials** — set `CONFIG_WIFI_SSID` and - `CONFIG_WIFI_PASSWORD` in `sdkconfig.defaults`. Simplest for a static - home setup; easy to forget they're in the binary. -- **Fallback SoftAP or BLE provisioning** — some upstream builds include a - fallback provisioning flow if no credentials are saved. Check the - upstream README for what your commit supports. - -Either way, the device must land on a **2.4 GHz** network. ESP32-S3 does -not do 5 GHz. +The device must land on a **2.4 GHz** network. ESP32-S3 does not do 5 GHz. --- @@ -234,6 +252,9 @@ Some older xiaozhi builds expose a SoftAP captive portal on first boot: If you have a build where this works, it's the fastest provisioning flow. It just isn't what M5Stack ships today. +This **Advanced** control is a tab in the browser portal at `192.168.4.1`, not +an item in the StackChan's on-device Settings app. + --- ## 9. When it's working: bookmark these diff --git a/audit-draft-issues.md b/audit-draft-issues.md new file mode 100644 index 00000000..afa123a3 --- /dev/null +++ b/audit-draft-issues.md @@ -0,0 +1,3248 @@ +# Draft GitHub issues — dotty-stackchan audit + +Ready-to-file issue drafts for confirmed bugs and all critical/high-severity findings (excluding the issue/PR triage entries, which are tracked in `AUDIT-REPORT.md`). Each block is copy-pasteable into `gh issue create`. + +--- + +## [CRITICAL] Container chain: unauthenticated admin routes + docker.sock + host networking = LAN→host-root + +``` +Summary +An attacker who reaches the LAN (or Tailnet) has a confirmed write primitive into the +container that holds the docker socket, and that container is one `docker run --privileged` +call from full host root. The individually-reported symptoms (unauth admin routes, +arbitrary-path play-asset, all-root containers, host networking) are each one link of a +single exploit chain. + +Location +docker-compose.yml.template:43-60,59-60,596 +custom-providers/xiaozhi-patches/http_server.py:596,620-672 +dotty-pi/docker-compose.yml:17 / dotty-behaviour/docker-compose.yml:23 / bridge/docker-compose.yml:31 + +Details +- xiaozhi-server binds 0.0.0.0:8003 and registers all 11 /xiaozhi/admin/* routes with NO auth. +- The SAME container mounts /var/run/docker.sock + /usr/bin/docker (the compose comment admits + this gives "effective root on the docker host"). +- The other three services run network_mode: host with no segmentation. +- None of the four Dockerfiles set USER — every container runs as root. + +Proposed fix +Break the chain at more than one link: +1. Replace the raw docker.sock bind with a least-privilege docker-socket-proxy + (tecnativa/docker-socket-proxy) exposing ONLY the exec API PiVoiceLLM needs. +2. Add a shared-secret bearer-token aiohttp middleware over the /xiaozhi/admin/* subapp + (constant-time compare, 401 otherwise); have callers present it. +3. Bind admin HTTP to 127.0.0.1 (set server.ip) and reach it only from sibling containers + on loopback instead of publishing 0.0.0.0:8003. +4. Add non-root USER directives to all four Dockerfiles. Pair with no-new-privileges:true. +Do at least the socket-proxy + route auth. + +Suggested labels: security, area:xiaozhi, area:bridge, priority:critical +``` + +--- + +## [CRITICAL] compose.all-in-one.yml omits the docker.sock + docker CLI mounts the default PiVoiceLLM provider requires + +``` +Summary +A fresh deploy following compose.all-in-one.yml's own Quick start with the default config +gets a container that cannot exec into dotty-pi — every voice turn fails. + +Location +compose.all-in-one.yml:64-79 + +Details +The shipped .config.yaml.template defaults to selected_module.LLM: PiVoiceLLM, which reaches +the brain via `docker exec -i dotty-pi pi --mode rpc ...` (pi_client.py:360-368). That needs +the host docker socket + docker CLI bind-mounted (docker-compose.yml.template:59-60). +compose.all-in-one.yml mounts the pi_voice provider dir but NOT the socket or the docker +binary — even though its own header says PiVoiceLLM "requires the host docker socket". + +Proposed fix +Add the two PiVoiceLLM mounts (matching the template): + /var/run/docker.sock:/var/run/docker.sock + /usr/bin/docker:/usr/bin/docker:ro +Or, if all-in-one is meant to be OpenAICompat-only, change the shipped selected_module +default and say so explicitly in the header. + +Suggested labels: bug, area:deploy, priority:high +``` + +--- + +## [HIGH] compose.all-in-one.yml is missing the xiaozhi-patches mounts — admin routes + perception relay dead + +``` +Summary +Deploying with compose.all-in-one.yml yields upstream xiaozhi-server with no admin portal and +no perception relay, so the entire dotty-behaviour layer is inert. + +Location +compose.all-in-one.yml:64-79 + +Details +The canonical template mounts portal_bridge.py, websocket_server.py, http_server.py, and +textMessageHandlerRegistry.py (admin routes, active_connections registry, EventTextMessageHandler). +all-in-one mounts none of them, so consumers fire admin HTTP calls into 404s and no perception +events ever reach the bus. Also missing vs the template: openai_compat, personas, +receiveAudioHandle.py, dances.py, textUtils.py, songs, and the kid/smart state mount. + +Proposed fix +Bring compose.all-in-one.yml's volume list to parity with docker-compose.yml.template, or +generate the all-in-one from the template so it can't drift. At minimum add the four +xiaozhi-patches mounts plus openai_compat, personas, receiveAudioHandle, dances, textUtils, +songs, and VISION_BRIDGE_URL. + +Suggested labels: bug, area:deploy, priority:high +``` + +--- + +## [HIGH] All /xiaozhi/admin/* routes are completely unauthenticated + +``` +Summary +Every admin endpoint is registered with no auth check, so anyone reaching port 8003 can drive +the robot and read arbitrary files. + +Location +custom-providers/xiaozhi-patches/http_server.py:620-672 (handlers 34-589) + +Details +inject-text, abort, set-head-angles, set-state, set-toggle, set-face-identified, take-photo, +play-asset, songs, say, devices — all added with no auth. inject-text runs the full LLM+MCP +pipeline (remote prompt-injection into the pi agent); say is verbatim TTS; play-asset reads any +absolute path (only os.path.exists). The OTA and WS paths guard themselves; these do not. + +Proposed fix +Gate the routes behind a shared secret: an aiohttp middleware requiring a Bearer / X-Admin-Token +header on every /xiaozhi/admin/* request, 401 otherwise. At minimum confine play-asset to an +allow-listed base dir via os.path.realpath + startswith. + +Suggested labels: security, area:xiaozhi, priority:high +``` + +--- + +## [HIGH] play-asset accepts arbitrary absolute filesystem path with no allowlist + +``` +Summary +/xiaozhi/admin/play-asset feeds any readable file to ffmpeg inside the docker.sock-holding +container, with the only validation being os.path.exists. + +Location +custom-providers/xiaozhi-patches/http_server.py:376-411 + +Details +asset is a raw absolute path; no base-dir confinement, no extension allowlist before +AudioSegment.from_file (ffmpeg/libav). The sibling songs handler hard-codes a base dir + a +{opus,ogg,wav,mp3} allowlist; play-asset has neither. An unauthenticated LAN caller can +probe arbitrary host/container paths (404-vs-200) and exercise libav demuxer CVEs against +attacker-staged files — and a decoder RCE lands on the privileged process. + +Proposed fix +Resolve os.path.realpath(asset) and require it to start with the songs base +(/opt/xiaozhi-esp32-server/config/assets/songs + os.sep) before decode; enforce the +{opus,ogg,wav,mp3} allowlist; prefer accepting only a basename joined onto the fixed base. + +Suggested labels: security, area:xiaozhi, priority:high +``` + +--- + +## [HIGH] All four containers run as root (no USER directive) + +``` +Summary +None of the four service images declare USER; for xiaozhi-server the root process is the one +mounting docker.sock, so any code-exec bug there is already host root. + +Location +Dockerfile, dotty-pi/Dockerfile, dotty-behaviour/Dockerfile, bridge/Dockerfile + +Details +xiaozhi-server (upstream base), dotty-pi (node:25.9-alpine), dotty-behaviour + bridge +(python:3.12-slim) all run as root. The FastAPI/uvicorn and pi workloads do not need root. +dotty-behaviour and bridge run root on host networking, widening blast radius. + +Proposed fix +Add a non-root USER to each Dockerfile (create app user, chown app/state dirs, drop to it). +For xiaozhi-server, run the socket-proxy so root-in-container is no longer root-on-host. +Pair with no-new-privileges:true and read-only rootfs where feasible. + +Suggested labels: security, area:deploy, priority:high +``` + +--- + +## [HIGH] room_view roster recognition silently fails when person id != display_name + +``` +Summary +A correct VLM identification is dropped to person_id=None whenever a person's canonical id +differs from display_name.lower() — the common case — so named greetings never fire. + +Location +dotty-behaviour/vision/room_view.py:115-124,159-167 +dotty-behaviour/routes/vision.py:192,256-269 + +Details +name_choices is built from p.display_name, but parse_room_view_response validates against +roster_ids, which household.roster_ids_with_appearance() sources from p.id. For id `hudson_jr` +with display_name `Hudson`, the lowercased returned name doesn't match roster_ids and falls to +None; the synthetic face_recognized broadcast never fires. Masked because the test fake +reimplements the helper as display_name.lower() with id==display_name fixtures. + +Proposed fix +Make the prompt vocabulary and the parser's validation set use the SAME key: build name_choices +from p.id, or resolve the returned display name back to the canonical id before the membership +check. Fix the test fake to return p.id and add an id != display_name fixture. + +Suggested labels: bug, area:behaviour, area:vision, priority:high +``` + +--- + +## [HIGH] Double greeting on every face_recognized — FaceGreeter and ProactiveGreeter both fire + +``` +Summary +A single face recognition produces two back-to-back utterances because two independent +consumers both speak and don't coordinate. + +Location +dotty-behaviour/consumers/face_greeter.py:133-186 +dotty-behaviour/greeter/greeter.py:162 (wired in main.py:309-323) + +Details +FaceGreeter._handle_face_recognized says "Oh, it's {name}!" and ProactiveGreeter._on_face_recognized +independently generates an LLM greeting and also says(). They keep separate cooldown state. The +bare-greet path was deliberately suppressed when the roster is identifiable; the named path was not. + +Proposed fix +Pick one owner for the named greeting: disable FaceGreeter's named path when the proactive +greeter is enabled (empty FACE_NAME_GREET_TEMPLATE), or share a single per-identity cooldown +timestamp between the two so the second consumer backs off. + +Suggested labels: bug, area:behaviour, priority:high +``` + +--- + +## [HIGH] room_view name parser cannot match multi-word display names + +``` +Summary +Any roster member whose display name contains a space is a 100% silent identification failure, +even on a perfect VLM response. + +Location +dotty-behaviour/vision/room_view.py:72-81,115-124 + +Details +name_choices injects display_name and the prompt asks for `NAME: `, but the NAME +capture group is `[A-Za-z_][A-Za-z0-9_-]*` — a single whitespace-free token. A reply +`NAME: Mary Anne` fails the anchored regex, so parse_room_view_response returns (cleaned, None, None). + +Proposed fix +Make the NAME group accept internal spaces (e.g. `[A-Za-z_][\w '\-]*?`) and normalise before +lookup, OR request a single-token id and map ids->display names. Add a two-word display_name test. + +Suggested labels: bug, area:behaviour, area:vision, priority:high +``` + +--- + +## [HIGH] Greeter calendar lookup never matches an identified person (case mismatch) + +``` +Summary +A greeted person's personal calendar items are silently dropped from their greeting unless +their calendar prefix is all-lowercase. + +Location +dotty-behaviour/greeter/greeter.py:257-259 -> calendar_/cache.py:69-73 +dotty-behaviour/calendar_/fetch.py:99 + +Details +summarize_for_prompt is called with person=identity (lowercased id), but the calendar `person` +field is set from the title prefix with no case folding (`[Hudson] -> "Hudson"`). The filter +`if ev["person"] != person` then drops "Hudson" against "hudson". Household-bucket events mask it. + +Proposed fix +Normalise both sides: lower-case the parsed calendar person at fetch time (prefer mapping the +prefix through HouseholdRegistry.get_by_calendar_prefix to the canonical id), or compare +case-insensitively in summarize_for_prompt. + +Suggested labels: bug, area:behaviour, area:calendar, priority:high +``` + +--- + +## [HIGH] memory_lookup FTS search bypasses the #53 kid-safety namespace gate + +``` +Summary +A fact about a minor routed to the person_pending review queue can be returned to a live turn +by memory_lookup whenever the query phrase-matches it. + +Location +dotty-pi-ext/src/lib/brain_db.ts:82-90 (searchMemories SQL); tools/memory_lookup.ts + +Details +fetchPersonMemories is scoped to person: and deliberately never returns person_pending: +(unreviewed minor facts). But searchMemories runs a bare `WHERE memories_fts MATCH ? ORDER BY rank` +over ALL rows in `memories` (external-content FTS over every row), including person_pending:. +The oracle (bridge.py _voice_memory_search_blocking) is equally unfiltered, so the gap exists in both. + +Proposed fix +Add a namespace guard: + ... WHERE memories_fts MATCH ? AND m.namespace NOT LIKE 'person_pending:%' ORDER BY rank LIMIT ? +Mirror the exclusion into oracle.py/bridge.py. Add a test seeding a person_pending row and +asserting memory_lookup never returns it. + +Suggested labels: bug, safety, area:dotty-pi, priority:high +``` + +--- + +## [HIGH] Mid-utterance synth failure emits SentenceType.LAST, truncating the rest of the reply + +``` +Summary +On any synthesis error in a non-last segment, the provider tells xiaozhi/firmware the whole TTS +response is finished, so the device stops talking mid-sentence. + +Location +custom-providers/edge_stream/edge_stream.py:139-143 +custom-providers/piper_local/piper_local.py:187-191 (identical bug) + +Details +text_to_speak's except handler unconditionally pushes (SentenceType.LAST, [], None) regardless +of is_last. A reply is split into multiple segments; if an early/middle segment's synth throws +(edge_tts blip, Piper decode error), the LAST marker drops every remaining segment. + +Proposed fix +Only emit LAST from the failure path when is_last is True; for non-last segments, log and +continue (or emit a segment-level failure sentinel) so subsequent segments still play. + +Suggested labels: bug, area:tts, priority:high +``` + +--- + +## [HIGH] dotty_doctor model checks look under data/models instead of repo-root models — false FAIL + +``` +Summary +In the standard documented layout (config at data/.config.yaml, models at models/), both +model checks report FAIL and dotty_doctor exits non-zero even though the models are present. + +Location +scripts/dotty_doctor.py:128-155 + +Details +_find_config prefers data/.config.yaml; check_models_* then compute root = config_path.parent +(= data/) and check data/models/SenseVoiceSmall and data/models/piper. But the models live at +repo-root models/ (docker-compose mounts ./models/...; Makefile SENSEVOICE_DIR := models/...). + +Proposed fix +Don't anchor to config_path.parent. If config_path is .config.yaml under a data/ dir, use +config_path.parent.parent; or search both root/models and root.parent/models. + +Suggested labels: bug, area:tooling, priority:high +``` + +--- + +## [HIGH] provision.py: voice "General" collides with text "general" and gets moved on every run + +``` +Summary +Syncing the VOICE "General" channel destructively moves the existing TEXT #general into the +VOICE category, recurring on every idempotent re-run; the real voice channel is never created. + +Location +community/discord/provision.py:93-96,197-202 + +Details +find_channel searches the whole guild case-insensitively and ignores channel type, so +find_channel(guild, "General") matches text "general"; the existing-channel branch then calls +existing.edit(category=VOICE), moving #general. Same collision class hits any text/voice name +pair differing only by case. + +Proposed fix +Scope find_channel by the expected channel type (pass kind and filter on isinstance / +ChannelType), or compare (name, type) tuples. Don't move a channel whose type doesn't match. + +Suggested labels: bug, area:community, priority:high +``` + +--- + +## [MEDIUM] Dashboard action endpoints always block 8s and report "no reply" — the event producer is dead + +``` +Summary +Every dashboard Say/Dance/Mood/Story action blocks for the full 8s then renders +"Sent — no reply in 8s." because nothing ever enqueues onto the bridge event bus. + +Location +bridge.py:613-626 (producer stub); bridge/dashboard.py:658-704,2091-2133 + +Details +_inject_or_error subscribes to _dashboard_event_listeners and waits 8s for Dotty's reply, but +the voice-turn producer moved to dotty-pi in #36 and was never re-wired to the bridge SSE bus. +The injection fires (device speaks) but the dashboard can never show the reply. + +Proposed fix +Either wire a real producer (dotty-pi/xiaozhi POSTs completed turns to a bridge endpoint that +calls put_nowait on each subscriber queue), or drop the subscribe/wait and return "Sent" +immediately. Document /ui/events as heartbeat-only until a producer exists. + +Suggested labels: bug, area:dashboard +``` + +--- + +## [MEDIUM] Blocking synchronous HTTP inside async dashboard routes stalls the event loop + +``` +Summary +Uncached dotty-behaviour getters and the robot-photo proxy do synchronous requests.get with no +to_thread, blocking the entire asyncio loop for up to 1.5-2.0s and freezing all dashboard clients. + +Location +bridge.py:471-504 (_dotty_behaviour_get) +bridge/dashboard.py:193-218 (_fetch_robot_photo), 1366, 2085, 801, 1925 + +Details +The perception/vision/audio getters funnel through _dotty_behaviour_get (requests.get timeout=1.5) +and are called from async routes with no await. The HTMX 10s poll fans each render into 3-6 such +calls. The codebase already uses asyncio.to_thread for device-count/songs, so these are clear +omissions. _build_perception_card_ctx's "no I/O" docstring is now false. + +Proposed fix +Wrap the blocking HTTP in asyncio.to_thread at the call sites, or use httpx.AsyncClient. Fix the +stale docstring. + +Suggested labels: bug, area:dashboard +``` + +--- + +## [MEDIUM] CSRF middleware blocks the localhost /admin/* operator endpoints + +``` +Summary +The documented localhost CLI back-channel (/admin/*) is 403'd by the app-wide CSRF middleware +before it ever reaches its own localhost auth, so the entire admin surface is unusable via its +intended caller. + +Location +bridge/csrf.py:33,69-70,83; bridge.py:800-924 + +Details +CSRFMiddleware exempts only ('/api/', '/metrics', '/health'). The /admin router (kid-mode, +smart-mode, state, persona, safety) is not exempt and a curl/script call carries no CSRF cookie, +so it is rejected before _admin_require_localhost runs. + +Proposed fix +Add '/admin/' to _EXEMPT_PREFIXES in bridge/csrf.py (the /admin router already has its own +localhost-only auth, the correct machine-to-machine model). + +Suggested labels: bug, area:dashboard +``` + +--- + +## [MEDIUM] SSE /events served under BaseHTTPMiddleware (CSRF) — Starlette buffering hazard + +``` +Summary +The text/event-stream response is wrapped by BaseHTTPMiddleware (CSRF), a documented Starlette +SSE footgun that buffers chunks and breaks disconnect detection. + +Location +bridge/dashboard.py:2091-2133; bridge/csrf.py:73-98 + +Details +CSRFMiddleware subclasses starlette.middleware.base.BaseHTTPMiddleware, which consumes the +StreamingResponse through an anyio memory stream and can delay/not-flush chunks, defeating the +X-Accel-Buffering: no header. Today only heartbeats flow, but the channel is structurally +degraded once a producer is wired back. + +Proposed fix +Exempt the SSE path from BaseHTTPMiddleware-style wrapping, or migrate CSRFMiddleware to a raw +ASGI middleware so streaming responses pass through untouched. + +Suggested labels: bug, area:dashboard +``` + +--- + +## [MEDIUM] Kid/smart-mode toggles report ok:False while having already persisted+applied the flip + +``` +Summary +When the firmware doesn't ack a toggle, the endpoint returns ok:False even though the bridge has +already flipped and persisted the bit, so the operator sees "failed", re-toggles, and the state +ping-pongs out of sync. + +Location +bridge.py:630-643 (_dashboard_set_kid_mode); 695-707 (_dashboard_set_smart_mode) + +Details +_dashboard_set_kid_mode persists + applies first, then dispatches the firmware set_toggle; if the +device is asleep / WiFi drops, _dispatch_set_toggle returns False and the function returns +{ok:False} even though the bridge state is already flipped. The UI can't distinguish "fully +applied" from "not applied". + +Proposed fix +Return ok:True with a separate device_pushed:false / warning field when the bridge write +succeeded but only the firmware LED push failed (mirror the /admin/kid-mode shape). Render the +persisted new_state and surface the firmware-ack failure as a non-fatal note. + +Suggested labels: bug, area:dashboard +``` + +--- + +## [MEDIUM] network_mode: host binds dashboard + /metrics to 0.0.0.0 with auth unset by default + +``` +Summary +The shipped bridge compose runs host networking with no dashboard auth and no CSRF secret set, so +the full mutation surface and /metrics are reachable unauthenticated from the whole LAN. + +Location +bridge/docker-compose.yml:31,41-60; bridge.py:936; bridge/dashboard.py:259-267 + +Details +DOTTY_DASHBOARD_USER/PASS, DOTTY_CSRF_SECRET, and DOTTY_BRIDGE_HOST are all unset; auth is opt-in +(_verify_dashboard_auth returns immediately when either env var is empty); DOTTY_BRIDGE_HOST +defaults to 0.0.0.0. CSRF alone does not authenticate. An unauthenticated kid-mode-off toggle is +reachable from the LAN. + +Proposed fix +Ship the compose with DOTTY_DASHBOARD_USER/PASS + a persistent DOTTY_CSRF_SECRET from .env, and/or +bind 127.0.0.1 by default with a documented reverse proxy. At minimum document the required auth env. + +Suggested labels: security, area:dashboard +``` + +--- + +## [MEDIUM] Async dashboard routes read + JSON-parse the entire daily convo log synchronously + +``` +Summary +status_strip / host_detail / alerts_count / alerts_detail read the whole convo-*.ndjson and +json.loads every line on the loop thread, blocking all concurrent requests on each ~10s poll. + +Location +bridge/dashboard.py:377-398,472-514,517-555,1631-1648 + +Details +_stackchan_last_seen does path.read_bytes() + per-line json.loads with no to_thread; alerts_count +and alerts_detail do read_bytes().splitlines() + per-line json.loads inline. The log grows all day +and is polled per open tab. + +Proposed fix +Wrap the read+parse bodies in await asyncio.to_thread(...). Optionally tail the file rather than +parsing the whole day. + +Suggested labels: bug, area:dashboard +``` + +--- + +## [MEDIUM] Dashboard /actions/say and /actions/start-story bypass the kid-mode content filter + +``` +Summary +The dashboard say/start-story ingresses hand text straight to inject-text with no content +filtering, a second unfiltered speech path that isn't acknowledged in the kid-safety surface. + +Location +bridge/dashboard.py:706-729,732-773 + +Details +say() and start_story() sanitise control chars and length, then call _inject_or_error without +content_filter() (imported via bridge.text). When DOTTY_DASHBOARD_USER/PASS are unset (default), +any LAN client can make Dotty speak arbitrary unfiltered text; these turns also never populate +the _cf_recent ring powering /ui/safety/recent. + +Proposed fix +Run the sanitised text through content_filter() when kid-mode is on (use kid_mode_getter), +returning the safe replacement / a rejection and recording the hit. + +Suggested labels: bug, safety, area:dashboard +``` + +--- + +## [MEDIUM] face_detected during talk state silently drops the bridge perception relay + +``` +Summary +face_detected/face_lost transitions received while current_state == 'talk' are never forwarded to +dotty-behaviour, desyncing consumers (greeter state, face_lost_aborter) for the whole conversation. + +Location +custom-providers/xiaozhi-patches/textMessageHandlerRegistry.py:134-138 + +Details +The face_detected branch does `if cstate == 'talk': ... return`, which exits handle() entirely so +the relay POST at lines 167-193 never runs. The talk-gate's intent is only to suppress +re-triggering room_view capture, not forwarding the event. + +Proposed fix +Replace the return with control flow that skips only the capture kick-off but still falls through +to the relay (wrap the capture-start in `if cstate != 'talk':`). + +Suggested labels: bug, area:xiaozhi, area:behaviour +``` + +--- + +## [MEDIUM] face_lost does not reset _room_description_in_flight, can wedge room_view capture + +``` +Summary +A leaked _room_description_in_flight flag makes every subsequent face_detected refuse to re-capture +for the life of the connection. + +Location +custom-providers/xiaozhi-patches/textMessageHandlerRegistry.py:92-102,139-141 + +Details +face_detected gates capture on `not _room_description_in_flight`. face_lost clears the other +room-view caches but never clears _room_description_in_flight, so a capture that crashed without +resetting the flag wedges all future captures. + +Proposed fix +In the face_lost branch also set conn._room_description_in_flight = False so a lost-then-reacquired +face always re-arms a capture. + +Suggested labels: bug, area:xiaozhi +``` + +--- + +## [MEDIUM] WebSocket query-param device-id can bypass token auth via the allowlist short-circuit + +``` +Summary +A client can put ?device-id= in the URL and skip token verification, because +query-param identity is injected into request.headers and trusted by the allowlist bypass. + +Location +custom-providers/xiaozhi-patches/websocket_server.py:84-108,216 + +Details +_handle_connection injects device-id/client-id/authorization from query params into +request.headers; _handle_auth reads them back and the allowed_devices whitelist short-circuit +treats a query-param device-id the same as a header one. Headers mutation semantics are also +version-fragile (__setitem__ may append). + +Proposed fix +Treat query-param identity as untrusted: do not let a query-string device-id satisfy the +allowlist bypass; require the token path for query-param connections. Use a separate dict for +derived identity rather than mutating request.headers. + +Suggested labels: security, area:xiaozhi +``` + +--- + +## [MEDIUM] OTA _is_higher_version treats pre-release/build-suffix versions as numerically higher + +``` +Summary +A release-candidate or build-tagged version is seen as HIGHER than its GA, so a device on GA is +offered a stale rc and a device on the rc never sees the GA as newer. + +Location +custom-providers/xiaozhi-patches/ota_handler.py:24-43,315 (also 314-332,94) + +Details +_parse_version does re.findall(r'\d+', ver), so '1.2.3-rc1' -> (1,2,3,1) compares greater than +'1.2.3' -> (1,2,3); '1.2' equals '1.2.0'. Build suffixes ('1.3.2 (build 47)') invert ordering. +Verified: _is_higher_version('1.2.3-rc1','1.2.3') returns True. Drives both the candidate sort +and the update decision. + +Proposed fix +Use a real semver-ish comparator (strip leading v, parse the numeric core, rank pre-release +suffixes BELOW the release, consider only the first 3 segments and ignore build metadata). +Add unit tests for rc/build-suffix cases; add the module to coverage + a ci.yml test dir. + +Suggested labels: bug, area:xiaozhi, area:firmware +``` + +--- + +## [MEDIUM] play-asset clobbers conn.client_abort / client_is_speaking, corrupting a concurrent voice turn + +``` +Summary +A timer-driven play-asset landing mid-turn cancels a user's barge-in and marks the device idle +while chat TTS is still streaming, desyncing the speaking-state machine. + +Location +custom-providers/xiaozhi-patches/http_server.py:452-453,474 + +Details +_dispatch() unconditionally sets conn.client_abort=False then conn.client_is_speaking=True, and in +finally sets client_is_speaking=False. play-asset is fire-and-forget on a ~20s security timer with +non-deterministic device selection, with no guard that the conn is idle. + +Proposed fix +Bail (or queue) if the conn is mid-turn — check client_is_speaking / an in-progress sentence_id +and refuse with 409 if busy. Save/restore prior flags rather than hard-setting, or route admin +audio through the TTS-priority queue. + +Suggested labels: bug, area:xiaozhi +``` + +--- + +## [MEDIUM] Admin handlers call conn.websocket.send() concurrently with the active chat-path writer + +``` +Summary +Fire-and-forget admin sends run concurrently with xiaozhi's own send on the same WS connection; +the websockets library forbids interleaved sends, so the command silently fails or a chat frame +gets corrupted. + +Location +custom-providers/xiaozhi-patches/http_server.py:146,194,242,288,338,456-468 + +Details +Every admin handler and play-asset write directly to conn.websocket.send() via _spawn(). A send +from an admin task while the chat task is mid-send can raise ConcurrencyError (swallowed by the +task done-callback) or interleave frames. play-asset (hundreds of opus frames) maximizes overlap. + +Proposed fix +Serialize all device-bound writes through a single per-conn asyncio.Lock that admin handlers +acquire. At minimum log the ConcurrencyError rather than dropping it, and gate audio pushes behind +an idle check. + +Suggested labels: bug, area:xiaozhi +``` + +--- + +## [MEDIUM] _dotty_say mutates conn.sentence_id from a worker thread, racing the consumer and pre-empting chat TTS + +``` +Summary +A /say landing mid-turn overwrites conn.sentence_id, causing the consumer to drop the chat turn's +remaining sentences — truncating Dotty mid-sentence — and is a data race against the consumer thread. + +Location +custom-providers/xiaozhi-patches/http_server.py:547-574,582 + +Details +_enqueue runs on a thread (asyncio.to_thread) and sets conn.sentence_id with no busy-check or lock +before enqueueing FIRST/MIDDLE/LAST. The greeter fires on perception events that can coincide with +a live conversation. There is also no tts_text_queue guard (AttributeError swallowed → silent no-op). + +Proposed fix +Refuse /say with 409 when the conn is mid-turn; or push the greeting through the queue without +overwriting sentence_id. Guard for tts_text_queue and 503 if absent. If sentence_id must be set, do +it on the loop thread. + +Suggested labels: bug, area:xiaozhi +``` + +--- + +## [MEDIUM] play-asset omits the tts 'start' lifecycle frame and uses an unvalidated Opus rate + +``` +Summary +Playback never sends the {type:tts,state:start} frame (firmware-fragile) and feeds conn.sample_rate +straight into the Opus encoder without checking it is a legal rate; a bad rate silently returns 200. + +Location +custom-providers/xiaozhi-patches/http_server.py:414,421-423,456-461 + +Details +The playback only emits sentence_start then opus frames then stop. tgt = conn.sample_rate is passed +to OpusEncoderUtils; Opus only supports 8/12/16/24/48 kHz, so a non-legal rate raises inside _decode, +is caught by the broad except (logs "decode failed"), and the endpoint still returns 200 ok. + +Proposed fix +Emit the full lifecycle (start -> sentence_start -> frames -> stop). Validate tgt against +{8000,12000,16000,24000,48000} (reject or resample). Propagate decode failure as 500 instead of 200. + +Suggested labels: bug, area:xiaozhi +``` + +--- + +## [MEDIUM] EventTextMessageHandler relays perception events to a stale BRIDGE_URL + +``` +Summary +Perception events are POSTed to {BRIDGE_URL or VISION_BRIDGE_URL}/api/perception/event, but the bus +moved to dotty-behaviour:8090 in #115; if the deploy sets the wrong var, every event 404s and is +dropped, silently disabling face_greeter/sound_turner/state gating. + +Location +custom-providers/xiaozhi-patches/textMessageHandlerRegistry.py:67-78,178 + +Details +The env var names and warning text still say BRIDGE_URL/VISION_BRIDGE_URL. If BRIDGE_URL points at +the dashboard container (which no longer serves /api/perception/event), the relay 404s. + +Proposed fix +Rename/select the env var to point at dotty-behaviour explicitly (e.g. BEHAVIOUR_URL) and update the +warning string; or document that BRIDGE_URL must resolve to dotty-behaviour:8090. Verify the deploy +compose sets the var the handler reads. + +Suggested labels: bug, area:xiaozhi, area:behaviour +``` + +--- + +## [MEDIUM] state_changed/face_detected ordering can corrupt the room_view gate + +``` +Summary +If the firmware emits face_detected before the IDLE->TALK state_changed it acted on, the room_view +gate sees 'idle' and fires a capture for a mid-talk flicker — the stacked "Hi NAME" failure the +comment claims to prevent. + +Location +custom-providers/xiaozhi-patches/textMessageHandlerRegistry.py:109-115,133-138 + +Details +conn.current_state is set from state_changed and read on face_detected. The gate relies on +state_changed always preceding the talk-phase face_detected, which the firmware does not guarantee. + +Proposed fix +Gate on a positive "room_view already captured this session" flag (only capture if +conn._room_description is None and no capture since last face_lost) rather than current_state. Or +have firmware stamp face_detected with its state. + +Suggested labels: bug, area:xiaozhi +``` + +--- + +## [MEDIUM] Turn timeout is a fixed wall-clock deadline that ignores streaming progress + +``` +Summary +A healthy but slow reply (think_hard escalation, 6-sentence story) is killed mid-stream because the +turn deadline is set once and never extended while text streams. + +Location +custom-providers/pi_voice/pi_client.py:220-269 + +Details +iter_turn_text sets deadline = time.time() + self._turn_timeout_sec (default 120s) before the loop +and never resets it. The user hears a partial reply then the "(brain offline)" fallback. The timeout +should bound silence/stall, not total length. + +Proposed fix +Reset deadline = time.time() + self._turn_timeout_sec each time a text_delta (or any progress frame) +is yielded, so the cap only fires when pi goes quiet. + +Suggested labels: bug, area:voice +``` + +--- + +## [MEDIUM] new_session failure proceeds with un-reset pi session, leaking prior-turn context + +``` +Summary +When new_session fails, the next voice turn runs against pi's previous, non-reset session, carrying +prior content (or a partially-absorbed jailbreak) into the next, possibly child, interaction. + +Location +custom-providers/pi_voice/pi_voice.py:136-141 + +Details +response() calls self._client.new_session() but on PiClientError only logs and continues, and sets +self._first_turn = False unconditionally — so a failed/timed-out reset is treated as a successful +fresh turn. + +Proposed fix +On new_session failure force a hard reset before proceeding (close + respawn the pi process), or +surface the failure to the caller. Do not treat a failed reset as a fresh turn. + +Suggested labels: bug, safety, area:voice +``` + +--- + +## [MEDIUM] _closed flag never reset after close() + reuse, suppressing reader-thread crash logging + +``` +Summary +After a close() + reuse, a genuine reader-thread crash is silently swallowed, so turns just time out +with no diagnostic trail. + +Location +custom-providers/pi_voice/pi_client.py:171-182,311-313,352-354 + +Details +close() sets self._closed = True; _ensure_started() respawns the process and starts fresh reader +threads but never resets _closed to False. The stdout/stderr exception handlers guard +logger.exception with `if not self._closed:`, so with _closed stuck True real crashes are hidden. + +Proposed fix +Reset self._closed = False inside _ensure_started() on (re)spawn; better, scope the "expected +shutdown" suppression to the specific process generation rather than a sticky boolean. + +Suggested labels: bug, area:voice +``` + +--- + +## [MEDIUM] Abandoned iter_turn_text leaves pi streaming; a stale agent_end can terminate the next turn + +``` +Summary +After a barge-in/abort, the following voice turn can come back empty or with stale text because an +abandoned turn's trailing agent_end (not id-matched) terminates the next turn. + +Location +custom-providers/pi_voice/pi_client.py:211-269 (also 188-209) + +Details +iter_turn_text sends a prompt then yields deltas until agent_end. On early stop (face_lost_aborter +abort, barge-in) the generator is GC'd without consuming agent_end; pi keeps generating and queues +frames. new_session may race the in-flight turn, and a leftover agent_end satisfies the NEXT +iter_turn_text's terminator (agent_end is matched on type, not req_id). + +Proposed fix +Send an explicit cancel/abort to pi when a turn generator is abandoned (try/finally inside +iter_turn_text). Drain-and-discard the prior turn's frames keyed by the abandoned req_id before the +next prompt. Tag agent_end matching with req_id. Have new_session() block until the prior turn's +agent_end (or a hard reset). + +Suggested labels: bug, area:voice +``` + +--- + +## [MEDIUM] OpenAICompat can emit two emojis when a valid emoji is not at the start + +``` +Summary +When the model's reply contains a valid emoji but not as the first character, the code prepends the +fallback emoji AND yields the original content with its embedded emoji — two emojis, violating the +single-emoji HARD CONSTRAINT. + +Location +custom-providers/openai_compat/openai_compat.py:218-229 + +Details +Enforcement only checks so_far.startswith(emoji). For "Well, 😊 hi there", so_far starts with 'W', +so it prepends 😐 and then yields content still containing 😊. Firmware keys the face on the first +emoji; the second garbles TTS. The check latches after the first content, never stripping embedded emojis. + +Proposed fix +After deciding the leading content lacks a valid emoji prefix, strip embedded emojis from the +streamed body before yielding (reuse textUtils.check_emoji/is_emoji). A robust fix buffers until the +first non-whitespace char, decides the prefix, then filters subsequent emojis. + +Suggested labels: bug, area:voice +``` + +--- + +## [MEDIUM] OpenAICompat speaks chain-of-thought when pointed at a reasoning model + +``` +Summary +OpenAICompat has no thinking filter, so an inline- reasoning model has Dotty narrate its +reasoning aloud, and the leading < defeats the emoji-prefix check. + +Location +custom-providers/openai_compat/openai_compat.py:189-229 + +Details +PiClient filters thinking_delta; OpenAICompat reads only choices[0].delta.content and yields it +verbatim. The local backend serves qwen3.6 reasoning models that emit ... inline +(spoken to TTS) or in delta.reasoning_content (dropped, but the visible answer arrives after a long +silent gap, tripping the turn timeout). + +Proposed fix +Strip ... spans from content before the emoji check/yield (stateful across chunks), +and ignore delta.reasoning_content. Or send a param to disable thinking. Guard against a leading < +defeating the emoji-prefix enforcement. + +Suggested labels: bug, area:voice +``` + +--- + +## [MEDIUM] OpenAICompat appends the safety/format suffix to the wrong message when no user turn exists + +``` +Summary +If the final dialogue turn is a system/assistant/tool message, the HARD CONSTRAINTS suffix +(emoji rule, English-only, length, kid-mode filter) is never injected and the model runs unconstrained. + +Location +custom-providers/openai_compat/openai_compat.py:101-114 + +Details +_build_messages computes last_user_idx and appends _TURN_SUFFIX only there; if last_user_idx is None +(no user role — proactive greeter injections, tool/function turns), the suffix is never added. +pi_voice avoids this by only sending the last user text; OpenAICompat forwards the whole dialogue. + +Proposed fix +If last_user_idx is None, append _TURN_SUFFIX to a synthesized trailing user message (or to the last +message regardless of role) so the safety/format constraints are always present. + +Suggested labels: bug, safety, area:voice +``` + +--- + +## [MEDIUM] OpenAICompat yields whitespace-only content chunks to TTS before emoji enforcement + +``` +Summary +A whitespace-only first chunk is yielded to TTS before the emoji-prefix enforcement engages, so the +first thing TTS/firmware sees is whitespace rather than the emoji the contract promises. + +Location +custom-providers/openai_compat/openai_compat.py:216-229 + +Details +Every non-empty content delta is yielded at :229 unconditionally; enforcement only fires once +so_far = ''.join(full_text).lstrip() is non-empty. Leading whitespace deltas (common with reasoning +models / templated outputs) go out before any emoji; an all-whitespace stream fires the fallback +after whitespace already went to TTS. + +Proposed fix +Buffer content until so_far (lstripped) is non-empty before yielding anything; do not yield raw +whitespace-only deltas. Yield the (optionally emoji-prefixed) accumulated buffer once real content +exists. + +Suggested labels: bug, area:voice +``` + +--- + +## [MEDIUM] processed_chars over-count in edge_stream and piper_local + +``` +Summary +After synthesizing the remaining tail, processed_chars is advanced by len(full_text) on top of an +already-nonzero absolute index, overshooting; a re-entered LAST silently drops the final sentence's audio. + +Location +custom-providers/edge_stream/edge_stream.py:67-74 +custom-providers/piper_local/piper_local.py:114-121 + +Details +remaining_text = full_text[self.processed_chars:] (absolute index), then +self.processed_chars += len(full_text). Post-condition should be processed_chars == len(full_text). +Masked by the per-turn FIRST reset, but a LAST without an intervening FIRST drops the tail. + +Proposed fix +self.processed_chars = len(full_text) (or += len(remaining_text)) in both files. Consider hoisting +the duplicated method into a shared mixin. Add the packages to coverage + a unit test. + +Suggested labels: bug, area:tts +``` + +--- + +## [MEDIUM] fun_local crashes in __init__ when output_dir is unset + +``` +Summary +FunASR fails hard on a config that WhisperLocal tolerates, because os.makedirs(None) raises TypeError. + +Location +custom-providers/asr/fun_local.py:51-58 + +Details +self.output_dir = config.get("output_dir") returns None when absent, then +os.makedirs(self.output_dir, exist_ok=True) is called unconditionally. whisper_local.py:48 guards +with `if self.output_dir:`; fun_local does not. + +Proposed fix +Guard the call: `if self.output_dir: os.makedirs(self.output_dir, exist_ok=True)`. + +Suggested labels: bug, area:asr +``` + +--- + +## [MEDIUM] ASR model invoked from worker threads with no lock — not thread-safe under concurrent transcription + +``` +Summary +faster-whisper / FunASR are not safe for concurrent calls on the same model object; two overlapping +sessions can corrupt output or segfault the container. + +Location +custom-providers/asr/whisper_local.py:107-109,161-173 +custom-providers/asr/fun_local.py:81-88 (mirror) + +Details +speech_to_text offloads inference via asyncio.to_thread on a singleton model shared across +connections; to_thread schedules concurrent calls on separate threads racing the underlying +CTranslate2 / FunASR C++ state. + +Proposed fix +Serialize inference with a per-instance threading.Lock held inside _transcribe_blocking / around +model.generate. Negligible latency for the single-device case, prevents the concurrent crash. + +Suggested labels: bug, area:asr +``` + +--- + +## [MEDIUM] before_stop play files dropped when the final segment's synthesis fails + +``` +Summary +A LAST-segment synth failure abandons any queued before_stop_play_files (e.g. a play_song trailing +asset) — they are neither played nor cleared, and can leak into the next utterance. + +Location +custom-providers/edge_stream/edge_stream.py:139-143 +custom-providers/piper_local/piper_local.py:187-191 (mirror) + +Details +The normal path calls _process_before_stop_play_files() when is_last; the except handler never does. + +Proposed fix +In the except handler, when is_last is True, still call self._process_before_stop_play_files() (and +clear the list) so trailing assets are honored or cleaned up. + +Suggested labels: bug, area:tts +``` + +--- + +## [MEDIUM] SentenceType.FIRST audio marker is emitted once per segment instead of once per reply + +``` +Summary +A multi-segment reply pushes multiple FIRST markers, which downstream treats as start-of-response +and can re-trigger the talk animation / state mid-sentence. + +Location +custom-providers/edge_stream/edge_stream.py:101 +custom-providers/piper_local/piper_local.py:146 (mirror) + +Details +text_to_speak unconditionally puts (SentenceType.FIRST, [], text) at the top of every call; each +segment is a separate text_to_speak call. + +Proposed fix +Gate FIRST so it is emitted only for the first segment of a reply (track a per-reply +first_segment_sent flag, otherwise emit a continuation/MIDDLE marker). + +Suggested labels: bug, area:tts +``` + +--- + +## [MEDIUM] Face recognizer reads the camera frame buffer after releasing the arbiter (latent UAF) + +``` +Summary +processFrame() releases the detection arbiter, then still uses the raw V4L2 frame pointer; once the +ESP-DL embedding crop is wired in, this becomes a live use-after-free / torn-frame read across the +two camera tasks. + +Location +firmware/firmware/main/stackchan/face/face_detector.cpp:236-292 + +Details +arbiter.releaseForDetection() at :236, but frame_data (captured at :183) is stored into face_img.data +(:277) and passed to recognize() (:292). After release, acquireForCapture()/StreamCaptures() can +dequeue/requeue the V4L2 buffer and overwrite the memory frame_data points into. + +Proposed fix +Hold the detection lock across recognize() (move releaseForDetection() after the recognition block), +OR copy the bounded face crop into a heap buffer while still holding the arbiter and point +face_img.data at the copy. Never pass a driver-owned frame pointer to a consumer that runs after release. + +Suggested labels: bug, area:firmware, safety +``` + +--- + +## [MEDIUM] SoundLocalizer high-pass filter state goes stale during cooldown, firing spurious direction changes + +``` +Summary +The HPF state freezes during the 750ms cooldown, so the first post-cooldown frame injects a large +artificial transient that can push energy past threshold and fire a spurious sound_event direction change. + +Location +firmware/firmware/main/stackchan/sound_localizer.cpp:19-23,44-51 + +Details +OnStereoFrame() returns early on the cooldown gate BEFORE the per-sample loop that advances +_hp_*_prev_x/_hp_*_prev_y. The first frame after cooldown evaluates x[n] - x[n-1] across the 750ms +discontinuity. + +Proposed fix +Advance the HPF state on every frame regardless of the gates (move the filter loop above the cooldown +return, keeping only the emit gated), or seed _hp_*_prev_x to the new frame's first sample so the next +post-cooldown frame doesn't see a step. + +Suggested labels: bug, area:firmware +``` + +--- + +## [MEDIUM] ImuEventModifier steals and releases the shared modify-lock it may not own + +``` +Summary +A shake during a photo releases Capture()'s modify-lock while the photo is in progress, letting the +head move during capture and corrupting the still. + +Location +firmware/firmware/main/stackchan/modifiers/imu.h:58-61,117,79,105 + +Details +IMU sets the motion lock only if not held (lines 58-61), but restore_state() ALWAYS clears it (117). +If Capture() already holds the lock when a shake arrives, IMU skips re-locking but its restore +releases Capture's lock. The avatar lock has the same shape. + +Proposed fix +Give the modify-lock real ownership (refcount/owner-token), or have IMU remember whether it actually +acquired the lock (bool _took_motion_lock) and only release in restore_state() if it took it. + +Suggested labels: bug, area:firmware +``` + +--- + +## [MEDIUM] FaceDetector::stop() frees buffer and deletes the stop-semaphore without checking the take succeeded + +``` +Summary +If the detector task overruns the 2s stop timeout, stop() deletes the semaphore and frees the buffer +the still-running task owns — a use-after-free of both. + +Location +firmware/firmware/main/stackchan/face/face_detector.cpp:100-117 + +Details +stop() does xSemaphoreTake(_stop_sem, 2000ms) WITHOUT checking the return, then unconditionally nulls +_task_handle, vSemaphoreDelete(_stop_sem), and heap_caps_free(_rgb_buffer). The live task then gives a +deleted semaphore and uses freed state. Currently unwired (only start()/setEnabled() are used). + +Proposed fix +Capture the take result; only delete the semaphore / free the buffer / null the handle when it +returned pdTRUE. On timeout, log and leak (or loop with a longer bound) rather than tearing down +resources the live task still owns. + +Suggested labels: bug, area:firmware +``` + +--- + +## [MEDIUM] startToChat: missing "language" key throws KeyError and feeds raw JSON to the pipeline + +``` +Summary +For a voiceprint payload lacking a language field, the speaker-extraction silently fails and the raw +JSON envelope is run through ASR corrections / intent detection / the LLM. + +Location +receiveAudioHandle.py:1057-1063 + +Details +The speaker block does _language_tag = data["language"] (subscript) inside a try with +`except (json.JSONDecodeError, KeyError): pass`. When speaker+content are present but language is +absent, line 1059 raises KeyError before actual_text is assigned, so actual_text stays the raw JSON +string and current_speaker is never set. + +Proposed fix +Use _language_tag = data.get("language") so a missing language key is tolerated; only speaker/content +should be required (already gated by the `in` checks). + +Suggested labels: bug, area:xiaozhi +``` + +--- + +## [MEDIUM] provision.py: read-only permission overwrites never re-applied to existing channels + +``` +Summary +The script isn't idempotent for permissions: a READ_ONLY channel that already exists keeps +@everyone send_messages, with the script reporting "= #... (exists)". + +Location +community/discord/provision.py:197-202 + +Details +The existing-channel branch only fixes the category, then continues — it never calls overwrites_for() +or applies overwrites. So missing/incorrect permission overwrites are never reconciled, a +security-relevant gap for a public community server. + +Proposed fix +In the existing branch, compute ov = overwrites_for(...) and, when non-empty, await +existing.edit(overwrites=ov) so permissions converge on every run. + +Suggested labels: bug, area:community +``` + +--- + +## [MEDIUM] dotty_doctor check_http treats 404 (and all <500) as pass for the OTA endpoint + +``` +Summary +A 404 on /xiaozhi/ota/ — exactly the misconfiguration the doctor exists to catch — is reported as PASS. + +Location +scripts/dotty_doctor.py:166-168 + +Details +check_http returns 'pass' for any HTTP status < 500. The OTA path returning 404 (route wrong/missing, +or server up but not serving OTA) passes silently. + +Proposed fix +Treat 2xx (and arguably 3xx) as pass; flag 4xx as warn/fail with the status code in the detail. + +Suggested labels: bug, area:tooling +``` + +--- + +## [MEDIUM] MCP JSON-RPC request ids collide for calls in the same millisecond + +``` +Summary +Two MCP calls dispatched within the same millisecond get identical JSON-RPC ids; any relay that +correlates responses by id will mismatch or drop the colliding pair. + +Location +receiveAudioHandle.py:281,319,349,373,399,456,559 + +Details +Every outbound MCP call computes id = int(time.time()*1000) % 0x7FFFFFFF. This happens routinely: +_sync_toggles_once fires set_toggle(kid_mode) then set_toggle(smart_mode) back-to-back, and +execute_choreography sends a HEAD and a LED at the same t_ms timeline mark. + +Proposed fix +Use a monotonic per-connection counter (itertools.count / conn._mcp_id) via a shared +_next_mcp_id(conn) helper, fixing all seven sites at once. + +Suggested labels: bug, area:xiaozhi +``` + +--- + +## [LOW] dotty_doctor passes Piper without the required .onnx.json config pair + +``` +Summary +A missing/stub .onnx.json (the exact corrupt-download failure the SenseVoice check was hardened +against) still reports a green PASS while the TTS provider crashes at runtime. + +Location +scripts/dotty_doctor.py:146-155 + +Details +check_models_piper only globs *.onnx; Piper requires a companion .onnx.json. The glob *.onnx +does not match *.onnx.json, so the config is never inspected. + +Proposed fix +For each *.onnx found, assert the sibling *.onnx.json exists above a sane byte floor (~200 B), +mirroring the SenseVoice size-floor approach. + +Suggested labels: bug, area:tooling +``` + +--- + +## [LOW] render_singing_sinsy crashes with ZeroDivisionError on empty/all-comment lyrics + +``` +Summary +A lyrics file that is empty or comments-only crashes build_singing_score with an opaque +ZeroDivisionError instead of a clean error. + +Location +scripts/render_singing_sinsy.py:89 (reachable via load_lyrics 48-61) + +Details +load_lyrics skips blank/comment lines and can return []. build_singing_score guards the note list but +not the syllable list; line 89 does syllables[i % len(syllables)] with len==0. + +Proposed fix +After load_lyrics in main() (or at the top of build_singing_score), add a guard: +if not syllables: print error; return 1. Mirror the existing 'has no notes' guard. + +Suggested labels: bug, area:tooling +``` + +--- + +## [LOW] _send_led_color swallows all exceptions silently + +``` +Summary +A broken WS during a dance silently no-ops every LED update with no log, making "dance ran but no +lights" undebuggable. + +Location +receiveAudioHandle.py:285-286 + +Details +_send_led_color wraps the websocket.send in a bare `except Exception: pass` with no logging, unlike +_send_led_multi which warns-once. + +Proposed fix +Log the exception (warn-once, matching _send_led_multi's pattern) rather than a silent pass. + +Suggested labels: bug, area:xiaozhi +``` + +--- + +## [LOW] _encode_midi_to_opus forces ALL set_tempo events to one value + +``` +Summary +Songs whose MIDI encodes legitimate mid-piece tempo changes are flattened to a single constant +tempo, breaking duration/beat alignment for any such file. + +Location +receiveAudioHandle.py:850-857 + +Details +When target_tempo_bpm is set, the loop rewrites every set_tempo in every track to new_tempo. Benign +for the current .mid registry, latent for any added song with tempo automation. + +Proposed fix +Rewrite only the first/global tempo, or scale all tempos by the same ratio relative to the original +first tempo. + +Suggested labels: bug, area:xiaozhi +``` + +--- + +## [LOW] _handle_dance: singing stream task untracked, so a new utterance cancels choreography but not the audio + +``` +Summary +On barge-in the choreography is cancelled but the singing audio keeps streaming Opus to the device, +relying entirely on the abort flag being set synchronously. + +Location +receiveAudioHandle.py:789-802,1092-1096 + +Details +Choreography runs as conn._dance_task (tracked, cancelled on the next turn); the singing audio runs as +a separate fire-and-forget asyncio.create_task whose handle is never stored on conn. + +Proposed fix +Store conn._singing_task and cancel it alongside _dance_task on the new-utterance/abort path. + +Suggested labels: bug, area:xiaozhi +``` + +--- + +## [LOW] Idle end-prompt re-enters the full intent pipeline via startToChat + +``` +Summary +A customised idle end-prompt containing words like "dance"/"sleep"/a vision phrase fires that +side-effect at the moment the connection is closing. + +Location +receiveAudioHandle.py:1208-1211 + +Details +no_voice_close_connect feeds end_prompt.prompt through startToChat, which runs _detect_state_phrase / +_is_vision_request / _is_dance_request. close_after_chat is already set, so the behaviour competes with +shutdown. + +Proposed fix +Submit the end-prompt directly to the LLM (_submit_chat / conn.chat) bypassing the intent pipeline, or +add a flag to startToChat to skip state/dance/vision detection for system-originated prompts. + +Suggested labels: bug, area:xiaozhi +``` + +--- + +## [LOW] provision.py create_forum drops the computed permission overwrites + +``` +Summary +The FORUM branch omits overwrites=ov, silently no-opping a future read-only forum lockdown. + +Location +community/discord/provision.py:204-213 + +Details +overwrites_for is computed for every new channel and the text branch passes overwrites=ov, but +guild.create_forum(...) omits it. Benign today (no forum in READ_ONLY). + +Proposed fix +Pass overwrites=ov to guild.create_forum(...), matching the text-channel branch. + +Suggested labels: bug, area:community +``` + +--- + +## [LOW] /api/perception/state emits invalid JSON (Infinity) for devices with no last_event_t + +``` +Summary +A device with no last_event_t serializes sensor_age_s as the bare token Infinity, which is not valid +JSON and breaks strict parsers / the dashboard perception card. + +Location +dotty-behaviour/perception/state.py:328-336 (reachable via routes/perception.py:73-85, routes/vision.py:188) + +Details +_annotate sets age = float('inf') when last_event_t is absent; Starlette's JSONResponse serializes +with allow_nan=True. Reachable for a known device that got a room_view capture but never sent a +perception event, and unconditionally for any unknown device_id. + +Proposed fix +Replace age = float('inf') with a JSON-safe value (None / -1) and rely on sensor_stale=True; add a +regression test that the route body parses as standard JSON. + +Suggested labels: bug, area:behaviour +``` + +--- + +## [LOW] Failed/empty weather fetch arms the full 30-min TTL + +``` +Summary +A transient weather failure marks the cache fresh for the full TTL, so no retry happens and prompts +keep using stale (or empty) weather text. + +Location +dotty-behaviour/calendar_/cache.py:124-127 (set_weather); driven by poll.py:46-48 + +Details +fetch_weather returns "" on any failure; set_weather only overwrites weather_text on truthy text but +UNCONDITIONALLY bumps weather_fetched_perf, so ttl_expired stays false for the next full TTL window. + +Proposed fix +Advance weather_fetched_perf only on a successful (non-empty) fetch (guard the bump behind `if text:`), +or use a short retry interval on empty. Add a test asserting an empty fetch leaves the timestamp +unchanged. + +Suggested labels: bug, area:behaviour +``` + +--- + +## [LOW] No upload size limit on /api/vision/explain and /api/audio/explain (LAN OOM) + +``` +Summary +Both endpoints read the entire multipart upload into memory (+ ~1.33x for base64) with no size cap, +so a single large/malicious POST can exhaust container memory and take down the daemon. + +Location +dotty-behaviour/routes/vision.py:145 (and audio.py:87) + +Details +await file.read() reads the whole upload; no Content-Length cap, no streaming. Both endpoints bind +0.0.0.0:8090. The JPEG is also retained per-device in vision_cache. + +Proposed fix +Check Content-Length up front and reject >N (e.g. 5 MB JPEG, 10 MB audio) with 413, or read in +bounded chunks. Apply to both endpoints. + +Suggested labels: bug, security, area:behaviour +``` + +--- + +## [LOW] Consumer task crashes are silently swallowed (no log, no restart) + +``` +Summary +A consumer's run() raising at runtime dies silently and the daemon keeps reporting "ready" while the +consumer is dead. + +Location +dotty-behaviour/main.py:325-341 + +Details +Consumers are launched with asyncio.create_task and only awaited at shutdown via +gather(..., return_exceptions=True), which captures and discards exceptions. No done-callback surfaces +them at failure time; no restart. + +Proposed fix +Attach a done-callback that logs non-cancelled exceptions at failure time; optionally wrap each +consumer in a supervising restart loop. + +Suggested labels: bug, area:behaviour +``` + +--- + +## [LOW] vision_latest lost-wakeup race + unconditional fresh-capture eviction + +``` +Summary +A concurrent explain that signals between the cache pop and waiter registration is missed (spurious +404), and the unconditional pop discards a fresh capture another producer just wrote. + +Location +dotty-behaviour/routes/vision.py:331-352,332 + +Details +vision_latest pops the cache, registers a waiter, then awaits the Event; signal_vision_waiters only +set()s already-registered Events, so a signal in the pop→register window is lost and the waiter 404s +at the 15s timeout. The pop has no freshness check, so it deletes a fresh idle-photographer/room_view +entry (and its jpeg_bytes that /api/vision/photo serves). + +Proposed fix +Make the wakeup level-triggered (re-check the cache after registering before awaiting, or use a +generation counter). Pop only entries actually older than the TTL. + +Suggested labels: bug, area:behaviour, area:vision +``` + +--- + +## [LOW] Failed room_view VLM call still arms the 120s cooldown + +``` +Summary +A transient VLM blip arms the full 120s room_view cooldown despite never producing a usable +description, disabling named greetings for minutes. + +Location +dotty-behaviour/routes/vision.py:188-204 + +Details +last_room_view_capture_t is set BEFORE vlm.describe_image, which never raises (returns +VLM_NETWORK_ERROR_SENTINEL / VLM_OFFLINE_SENTINEL on failure). So an unreachable VLM still records the +timestamp and gates further captures for 120s. + +Proposed fix +Set last_room_view_capture_t only after a successful, non-sentinel VLM response, or use a much shorter +failure-cooldown. + +Suggested labels: bug, area:behaviour, area:vision +``` + +--- + +## [LOW] VLM error/offline SENTINEL strings cached and leaked into the perception snapshot + +``` +Summary +A single VLM outage poisons Dotty's voice prompt for up to 60s on subsequent turns by caching the +loud-error sentinel as a real description. + +Location +dotty-behaviour/routes/vision.py:195-242 (sentinels in dispatch/vlm.py:22-29) + +Details +vision_explain stores VLM_OFFLINE_SENTINEL / VLM_NETWORK_ERROR_SENTINEL verbatim as +vision_cache[device_id]["description"] with a fresh wall_ts. perception/snapshot.py:142-147 then +surfaces it as `You see: ERROR: the vision service didn't respond...`, and the idle photographer can +persist it to NDJSON. + +Proposed fix +Detect the sentinel set after describe_image and skip the cache write + signal_vision_waiters on +sentinel returns. Same guard for the audio fallback string. + +Suggested labels: bug, area:behaviour, area:vision +``` + +--- + +## [LOW] perception/state limit query params unvalidated (negative/zero slice quirks) + +``` +Summary +A client-supplied limit flows unbounded into list slices; limit=-3 drops the 3 newest events and +balances[-0:] returns the FULL series. + +Location +dotty-behaviour/routes/perception.py:88-99,124 + +Details +perception_recent forwards limit into items[:limit] (state.py:152) with no >=0 bound; sound_balance_series +ends with balances[-limit:] where limit=0 returns the whole list. + +Proposed fix +Clamp limit = max(0, min(limit, MAX)); special-case limit <= 0 (return []) before the balance slice. + +Suggested labels: bug, area:behaviour +``` + +--- + +## [LOW] Client-supplied event ts can be set in the future, defeating staleness gates + +``` +Summary +A future ts makes age negative → clamped to 0, so get_fresh_face_id treats a stale identity as +permanently fresh and staleness never trips. ESP32 RTCs are frequently unsynced. + +Location +dotty-behaviour/routes/perception.py:61-68 + +Details +perception_event uses payload.ts verbatim. Downstream freshness logic computes age = now - last_t with +max(0.0, ...) clamps or direct > ttl comparisons. + +Proposed fix +Reject or clamp implausible future timestamps on ingest (substitute time.time() if >few seconds +ahead). Guard get_fresh_face_id against negative age (treat as stale). + +Suggested labels: bug, area:behaviour +``` + +--- + +## [LOW] perception/state stale dance_active never cleared on a missed terminal event + +``` +Summary +A dropped dance_ended/state_changed latches dance_active True indefinitely, permanently suppressing +room_view captures and greetings for that device. + +Location +dotty-behaviour/perception/state.py:185-199 + +Details +dance_active is set on dance_started / state_changed->dance and cleared only on a terminal event; the +bus drops events on a full subscriber queue, so a lost terminal frame strands the flag. No time-based +expiry. + +Proposed fix +Make is_dance_active freshness-bounded against last_dance_started_t with a max-dance TTL. + +Suggested labels: bug, area:behaviour +``` + +--- + +## [LOW] Greeter day-GC discards other days' slots; midnight-roll cooldown corruption + +``` +Summary +The greeter persists never more than one day of slots, and a perception event spanning midnight can +slip a near-instant second greeting through. + +Location +dotty-behaviour/greeter/greeter.py:224-228 + +Details +_take_slot derives today from wall-now but uses event_ts for cooldown; the GC collapses _state to +{today: ...} (then _save_state persists the truncated dict). A pre-midnight event handled just after +midnight is filed under the new day with old-day cooldown math. + +Proposed fix +Key the day on event_ts; GC by dropping only days older than a retention window (keep today + yesterday). + +Suggested labels: bug, area:behaviour +``` + +--- + +## [LOW] SecurityCycle keeps capturing on a missed/early transition + +``` +Summary +A dropped non-security state_changed leaves the per-device capture loop firing the camera forever, and +a security session already live when the consumer subscribes never starts the timer. + +Location +dotty-behaviour/consumers/security_cycle.py:238-257 + +Details +Capture is started/stopped purely off live state_changed events with no reconciliation against +current_device_state. + +Proposed fix +On subscribe, seed timers from current per-device current_state. Reconcile each running timer against +current_device_state per cycle and self-cancel when no longer 'security'. + +Suggested labels: bug, area:behaviour, safety +``` + +--- + +## [LOW] calendar fetch time window excludes the final minute of the day + +``` +Summary +Late-evening events near midnight can be silently missing because timeMax is exclusive and set to 23:59:59. + +Location +dotty-behaviour/calendar_/fetch.py:60-66 + +Details +time_max = now.replace(hour=23,minute=59,second=59,microsecond=0). The Calendar API timeMax is +exclusive; the conventional correct bound is the start of the next local day. + +Proposed fix +time_max = (start_of_today + timedelta(days=1)).isoformat(). + +Suggested labels: bug, area:behaviour, area:calendar +``` + +--- + +## [LOW] CalendarCache.flush_for_new_day produces a long empty-context window on day-roll fetch failure + +``` +Summary +A failing day-roll refresh leaves a guaranteed-empty calendar and, with backoff, can show no events +for up to 10 minutes in the morning even though yesterday's data was serviceable a moment earlier. + +Location +dotty-behaviour/calendar_/cache.py:116-122 + poll.py:53-75 + +Details +refresh_if_stale eagerly flushes events then attempts the fetch; on failure the except only bumps +calendar_failures, and _sleep_seconds escalates to up to 600s. + +Proposed fix +Flush only on a successful fetch, or reset calendar_failures / use the base interval for the first +retry after a day-roll flush. + +Suggested labels: bug, area:behaviour, area:calendar +``` + +--- + +## [LOW] room_view 'no one in view' sentinel uses substring match + +``` +Summary +A valid identification whose DESC merely mentions "no one in view" is discarded as an empty frame, +throwing away both the description and the matched identity. + +Location +dotty-behaviour/vision/room_view.py:152-153 + +Details +parse_room_view_response does `if ROOM_VIEW_NO_PERSON in cleaned.lower(): return None, None, None` — a +substring test against the whole reply. + +Proposed fix +Match the sentinel only when it is the entire reply (modulo whitespace/punctuation); try the DESC/NAME +regex first and only fall back to the sentinel check on a format miss. + +Suggested labels: bug, area:behaviour, area:vision +``` + +--- + +## [LOW] FaceIdentifiedRefresher keeps the green ID LED lit when identity came from room_view + +``` +Summary +A room_view identity keeps the green identified pip lit for the full TTL even after the person walked +off, because face_present is never set. + +Location +dotty-behaviour/consumers/face_identified_refresher.py:65-71 + +Details +The skip gate `if not face_present and last_lost:` only suppresses when the face was both absent AND +previously lost. A room_view identification sets last_face_id without face_present and without a +face_lost, so the gate is never entered and the LED re-fires every interval. + +Proposed fix +Treat "identity present but face_present False and never positively detected" as a stop condition +(skip refresh when not face_present, or require a positive face_present before keeping the pip lit). + +Suggested labels: bug, area:behaviour +``` + +--- + +## [LOW] FaceGreeter consumes the greet cooldown even when suppressed by an active dance + +``` +Summary +A face event during a dance silently burns the per-identity/per-device cooldown without greeting, so a +person who arrives during a dance is denied their greeting both during the dance and for the full +cooldown after. + +Location +dotty-behaviour/consumers/face_greeter.py:148-168,105-126 + +Details +The cooldown slot is written BEFORE the is_dance_active gate in both _handle_face_recognized and +_handle_face_detected. last_face_greet_t also arms face_lost_aborter. + +Proposed fix +Move the is_dance_active (and empty-text / mic-only) checks ABOVE the cooldown-slot write in both +handlers. + +Suggested labels: bug, area:behaviour +``` + +--- + +## [LOW] last_chat_t is never written on real conversations — QUIET_AFTER_CHAT gates are dead + +``` +Summary +The named-greet and sound-turner "don't fire while the user is mid-conversation" gates never fire, +because nothing records actual conversation activity. + +Location +dotty-behaviour/consumers/face_greeter.py:149-156 + sound_turner.py:77-79 + perception/state.py:200-203 + +Details +Both gates read dev_state['last_chat_t'], but the only writer is purr_player.py (which pokes it into +the future). update_state handles chat_status by setting 'listening', not last_chat_t. Regression from +bridge.py's /api/message ingress. + +Proposed fix +Write last_chat_t on conversation activity (set it on a chat_status 'listening' event and/or the +relayed user-utterance path), or drop the gates as inert. + +Suggested labels: bug, area:behaviour +``` + +--- + +## [LOW] FaceLostAborter leaks completed abort tasks in _pending forever + +``` +Summary +A completed abort task with no subsequent face event stays in _pending, so the dict never shrinks for +departed devices. + +Location +dotty-behaviour/consumers/face_lost_aborter.py:101-103,47-55 + +Details +_pending entries are only removed by a subsequent face_detected or a newer face_lost; a simple +fire-and-complete leaves a done task in the dict. + +Proposed fix +Add a done-callback that discards the entry (matching the _tasks.discard pattern other consumers use). + +Suggested labels: bug, area:behaviour +``` + +--- + +## [LOW] Vision room_view block-reason early return skips stale-cache eviction + +``` +Summary +On frequently-gated deployments, stale vision_cache entries for other devices are never reclaimed +because the block-reason early return bypasses the eviction loop. + +Location +dotty-behaviour/routes/vision.py:176-186 + +Details +A gated room_view request writes a fresh entry and returns at :186, before the stale-eviction loop at +:273-279 that runs only on the normal path. + +Proposed fix +Factor the stale-eviction loop into a helper and call it on every exit path, including the block-reason +early return. + +Suggested labels: bug, area:behaviour +``` + +--- + +## [LOW] Blocking file I/O on the event loop in greeter state persistence + +``` +Summary +Every successful greet stalls the single event loop on synchronous disk I/O. + +Location +dotty-behaviour/greeter/greeter.py:246,420-435 + +Details +_take_slot calls _save_state() synchronously (mkdir, tmp.write_text, os.replace) from the bus event +handler on the asyncio loop. + +Proposed fix +await asyncio.to_thread(self._save_state) (the handler is already async), or batch/debounce persistence. + +Suggested labels: bug, area:behaviour +``` + +--- + +## [LOW] SleepDreamer / DanceReflector dereference event.data without None-guard, crashing the consumer loop + +``` +Summary +A state_changed/dance_ended frame with data=None raises AttributeError caught only by the OUTER +try/except, which logs "crashed" and EXITS the consumer permanently. + +Location +dotty-behaviour/consumers/sleep_dreamer.py:166-168; dance_reflector.py:89-92 + +Details +Both do `.get` directly on event.data; sibling consumers use `(event.data or {})`. The except is +outside the while loop, so one bad frame kills the consumer for the process lifetime. + +Proposed fix +Use (event.data or {}).get(...); wrap per-event handling so a single bad frame doesn't terminate the +loop. Consider giving PerceptionEvent.data a default_factory=dict. + +Suggested labels: bug, area:behaviour +``` + +--- + +## [LOW] HouseholdRegistry.get_by_calendar_prefix only matches with bracketed YAML + +``` +Summary +A household member whose calendar_prefix is configured without brackets is unreachable by +get_by_calendar_prefix. + +Location +dotty-behaviour/household/registry.py:190-200 vs 293-294 + +Details +_reload indexes by_prefix[person.calendar_prefix.strip().lower()] verbatim, but get_by_calendar_prefix +bracket-normalizes the query (`if not key.startswith('['): key = f'[{key}]'`). The two agree only when +the YAML value already has brackets. (No non-test callers today, so latent.) + +Proposed fix +Normalize both sides identically (bracket-normalize calendar_prefix when building by_prefix, or strip +brackets on both). Add a test for both YAML forms. + +Suggested labels: bug, area:behaviour +``` + +--- + +## [LOW] Person.days_until_birthday shifts Feb-29 birthdays to Feb-28, firing the greeting a day early + +``` +Summary +In non-leap years a Feb-29 birthday is treated as Feb-28, so the greeter says "It is X's birthday +today" a day early. + +Location +dotty-behaviour/household/registry.py:82-92 + +Details +self.birthdate.replace(year=ref.year) raises for Feb-29 in a common year and the except falls back to +date(ref.year, month, 28); days_until_birthday()==0 then fires on Feb-28. + +Proposed fix +Decide and document a deliberate Feb-29 policy (roll to Mar-1 in common years, or keep Feb-28 but make +it intentional and consistent across both branches). + +Suggested labels: bug, area:behaviour +``` + +--- + +## [LOW] Admin/behaviour HTTP clients disarm the timeout before reading the response body + +``` +Summary +A server that sends headers then stalls the body hangs the voice tool forever — no timeout fires, no +exception is thrown, so the try/catch fallbacks never run and the whole voice turn wedges. + +Location +dotty-pi-ext/src/lib/xiaozhi_admin.ts:33-46; dotty-pi-ext/src/lib/dotty_behaviour.ts:20-37 + +Details +adminFetch/behaviourFetch arm an AbortController timer, await fetch() (resolves on headers), then clear +the timer in finally — before any caller reads the body (fetchSongCatalog .json(), playAsset .text(), +fetchTakePhoto/fetchPersonReviewStatus .json()). The sibling llama_swap.ts:42-71 does this correctly. + +Proposed fix +Keep the timer armed across the body read: move the body read inside adminFetch/behaviourFetch (return +parsed JSON/text) so the existing try/finally covers it, mirroring llama_swap.ts. + +Suggested labels: bug, area:dotty-pi +``` + +--- + +## [LOW] play_song caches an empty catalogue on transient fetch failure for 60s + +``` +Summary +A single transient xiaozhi hiccup pins an empty catalogue for the full 60s TTL, so play_song returns +"(song catalogue is empty)" for up to a minute after recovery. + +Location +dotty-pi-ext/src/tools/play_song.ts:83-90 + +Details +getCatalog unconditionally caches fetchSongCatalog's result, which returns [] on ANY failure — it +cannot distinguish "no songs" from "fetch failed". + +Proposed fix +Only cache non-empty results: if fresh.length > 0, cache; else return fresh without caching so the next +turn retries. + +Suggested labels: bug, area:dotty-pi +``` + +--- + +## [LOW] play_song matchSong: empty stem from a dotfile matches any query, diverging from the oracle + +``` +Summary +A leading-dot catalogue entry matches every query and can win the best-candidate race, diverging from +the Python oracle and violating the byte-equal contract. + +Location +dotty-pi-ext/src/tools/play_song.ts:39-74 + +Details +stemLower computes the stem via lastIndexOf('.'); for '.mp3', dot===0 so the stem is '', and +qStem.includes('') is always true. Python's os.path.splitext('.mp3') keeps '.mp3'. + +Proposed fix +Match splitext semantics: only strip an extension when the dot index > 0 (`dot > 0 ? lower.slice(0,dot) : lower`). +Guard the substring branch against an empty stem. Add a leading-dot fixture to play_song.test.ts. + +Suggested labels: bug, area:dotty-pi +``` + +--- + +## [LOW] think_hard truncates by UTF-16 code units, diverging from the oracle and sibling tools + +``` +Summary +think_hard caps the reply with .slice (UTF-16 code units) while the oracle and the other tools use +codepoint slicing, so an emoji-prefixed near-cap reply mis-cuts. + +Location +dotty-pi-ext/src/tools/think_hard.ts:72-77 + +Details +content.trim().slice(0, MAX_OUTPUT_CHARS). bridge.py oracle uses Python codepoint slicing [:500]. +memory_lookup/recall_person/remember/remember_person all use Array.from codepoint slicing; think_hard +is the lone inconsistent path. + +Proposed fix +Use codepoint-aware truncation like the sibling tools: +const cp = Array.from(content.trim()); cp.length > MAX ? cp.slice(0, MAX).join('') : content.trim(); + +Suggested labels: bug, area:dotty-pi +``` + +--- + +## [LOW] play_song host-presence guard ignores the _XIAOZHI_HOST fallback + +``` +Summary +play_song refuses early with "(can't reach xiaozhi-server)" in an environment where only _XIAOZHI_HOST +is set, even though the admin client would reach the server. + +Location +dotty-pi-ext/src/tools/play_song.ts:104 + +Details +The guard checks `!opts.host && !process.env.XIAOZHI_HOST`, but DEFAULT_HOST resolves +XIAOZHI_HOST ?? _XIAOZHI_HOST ?? 'localhost'. (No setter for _XIAOZHI_HOST today, so latent.) + +Proposed fix +Use the same precedence as the admin client, or drop the guard and let fetchSongCatalog's empty-list +fallback drive the result. + +Suggested labels: bug, area:dotty-pi +``` + +--- + +## [LOW] extractTurnText concatenates assistant text blocks with no separator + +``` +Summary +Multi-message/multi-block assistant turns are joined with the empty string, fusing fragments +("Let me check…The answer is 4") in the logged turn. + +Location +dotty-pi-ext/src/lib/turn_logger.ts:56-62 (join("")), 79-91 + +Details +extractTurnText joins assistantParts and TextContent items with "". No oracle exists for this +extraction (pure pi-ext logic). + +Proposed fix +Join multi-message assistant text and TextContent items with a separator (a space, or "\n"). + +Suggested labels: bug, area:dotty-pi +``` + +--- + +## [LOW] think_hard default model resolved at module load, diverging from the per-call oracle + +``` +Summary +DEFAULT_MODEL reads VOICE_THINKER_MODEL once at import; the oracle reads it per call, so a later env +change diverges the request body. + +Location +dotty-pi-ext/src/tools/think_hard.ts:25 + +Details +DEFAULT_MODEL = process.env.VOICE_THINKER_MODEL ?? 'qwen3.6:27b-think' is evaluated at import. The +Python oracle resolves os.environ.get('VOICE_THINKER_MODEL', ...) on every call. + +Proposed fix +Resolve the env lazily inside buildThinkRequest/runThinkHard at call time. + +Suggested labels: bug, area:dotty-pi +``` + +--- + +## [LOW] play_song getCatalog TOCTOU lets concurrent turns both refetch + +``` +Summary +Two interleaved play_song calls both see the stale timestamp and both issue a fetch; combined with the +empty-cache bug, one can poison the other. + +Location +dotty-pi-ext/src/tools/play_song.ts:83-90 + +Details +getCatalog reads _catalogFetchedAt, awaits fetchSongCatalog(), then writes the cache; the pi agent can +run tool calls concurrently across the await. + +Proposed fix +Store an in-flight Promise so concurrent callers await the same fetch, and cache only on a non-empty +successful result (also resolves the empty-cache bug). + +Suggested labels: bug, area:dotty-pi +``` + +--- + +## [LOW] turn_logger.formatTurnLog omits the final strip the oracle applies (user-only turn divergence) + +``` +Summary +Every user-only turn stores a row that is not byte-identical to the tuned-against oracle (trailing +space after "assistant:"). + +Location +dotty-pi-ext/src/lib/turn_logger.ts:98-102 + +Details +formatTurnLog returns `user: ${u} | assistant: ${a}` directly; the oracle stores content.strip(). When +the assistant text is empty (a logged user-only turn), TS produces "...assistant: " vs the oracle's +"...assistant:". The integration test that would catch it (assistant_empty_user_only) only runs when +DOTTY_BRAIN_DB_SNAPSHOT is set, which npm test never sets. + +Proposed fix +return (`user: ${u} | assistant: ${a}`).trim(); wire the Layer-2 integration pass into npm test. + +Suggested labels: bug, area:dotty-pi +``` + +--- + +## [LOW] Seqlock reader can return a torn read on weak memory (missing acquire barrier) + +``` +Summary +A torn read can pass the seqlock check because the data reads can sink below the second seq acquire-load. + +Location +firmware/firmware/main/stackchan/face/face_detection_result.h:33-47 + +Details +read() loads s1 (acquire), copies the plain data fields, loads s2 (acquire), accepts when s1==s2. An +acquire load forbids later ops moving before it, not earlier plain reads from sinking below it. The +writer side is correct; the reader is missing the symmetric barrier. (Low real-world hit-rate on +Xtensa LX7.) + +Proposed fix +Insert an acquire fence between the data copies and the second seq load: +read the fields, then std::atomic_thread_fence(std::memory_order_acquire); then s2 = seq.load(relaxed). + +Suggested labels: bug, area:firmware +``` + +--- + +## [LOW] Idle-channel proactive reconnect churns close+reopen every ~30 ticks + +``` +Summary +A quiet-but-healthy idle channel close+reopens every ~30s indefinitely (a ~30s reconnect loop), each +iteration paying a 200ms delay + full TLS/WS handshake on the main task. + +Location +firmware/firmware/patches/xiaozhi-esp32.patch:23-32 + +Details +The reconnect fires when clock_ticks_ % 30 == 0 && state==Idle && protocol_->IsTimeout(). IsTimeout() +is time-since-last-INCOMING; a reopened channel the server never sends on keeps IsTimeout() true. No +backoff, no recently-reconnected suppression. + +Proposed fix +Gate on an actual dead-channel signal (IsAudioChannelOpened() false) or a last-reconnect-timestamp +backoff, or reset the incoming-frame clock on OpenAudioChannel() so a successful reopen clears IsTimeout(). + +Suggested labels: bug, area:firmware +``` + +--- + +## [LOW] CloseAudioChannel() blocks the calling task 200ms on every close + +``` +Summary +An added vTaskDelay(200ms) stalls whatever task calls close (incl. the main Run loop) on every channel +close, compounding the reconnect churn. + +Location +firmware/firmware/patches/xiaozhi-esp32.patch:582 + +Details +The patch adds vTaskDelay(pdMS_TO_TICKS(200)) to WebsocketProtocol::CloseAudioChannel() after +websocket_.reset(). The intent (drain the FIN) is reasonable but should not be a hard sleep on a +shared task. + +Proposed fix +Move the drain delay to only the reopen path that needs it (or poll socket state), or make close async, +rather than unconditionally sleeping 200ms on the caller's task. + +Suggested labels: bug, area:firmware +``` + +--- + +## [LOW] PrivacyLeds::update() non-atomic RMW races the guard, latching the listening LED on + +``` +Summary +The green mic privacy LED can keep showing "listening" after the mic ADC closed, lying about live +capture state until the next repaint. + +Location +firmware/firmware/main/stackchan/privacy/privacy_leds.cpp:49-60 + +Details +update() (tick task) load→derive→store of _mic_state is not atomic as a unit; MicPeripheralGuard's +dtor (codec-close task) setMicState(Off) between the load and store is overwritten. Fail-safe in +direction (over-indicates) but a privacy-indicator correctness bug. + +Proposed fix +Make the reconciliation a CAS that only updates while non-Off: load s; if s==Off return; compute +desired; compare_exchange_strong(s, desired). + +Suggested labels: bug, area:firmware +``` + +--- + +## [LOW] CameraPeripheralGuard ctor-fallback/dtor-normal mismatch underflows the refcount + +``` +Summary +A ctor that took the null-mutex fallback (no increment) followed by a normal dtor (fetch_sub) underflows +g_refcount to UINT32_MAX, so the stream never tears down and the camera LED stays Active indefinitely. + +Location +firmware/firmware/main/stackchan/privacy/camera_peripheral_guard.cpp:40-66,68-86 + +Details +The null-mutex ctor fallback sets CameraState::Active and returns WITHOUT incrementing g_refcount; the +dtor unconditionally fetch_sub(1) on the normal path. + +Proposed fix +Track per-guard whether the ctor incremented (bool _incremented) and only fetch_sub in the dtor when +true. Clamp/guard against fetch_sub when g_refcount==0. + +Suggested labels: bug, area:firmware +``` + +--- + +## [LOW] MCP take_photo arms a fixed 10s camera-glow timer decoupled from capture duration + +``` +Summary +The cosmetic camera-activity glow is hardcoded to 10s with no off-call, so it lingers after a fast +capture or goes dark before a slow one. + +Location +firmware/firmware/patches/xiaozhi-esp32.patch:523 + +Details +The take_photo patch adds GetHAL().setCameraLedActive(true, 10000) before camera->Capture(). The proper +privacy LED is driven by CameraPeripheralGuard; this glow no longer tracks the real lifecycle. + +Proposed fix +Tie the glow to the capture scope: set it active without a timeout and clear it after Capture() returns, +or drive it from the CameraPeripheralGuard refcount transitions like the privacy pixel. + +Suggested labels: bug, area:firmware +``` + +--- + +## [LOW] Unescaped name / session_id in Protocol::SendEvent JSON construction + +``` +Summary +SendEvent builds the perception frame by raw string concatenation; neither the event name nor the +server-supplied session_id is escaped, so any future runtime name or a quoted session id emits +malformed JSON the relay drops. + +Location +firmware/firmware/patches/xiaozhi-esp32.patch:537-542 + +Details +Not currently exploitable (literal names, pre-escaped data, server-controlled session id), but SendEvent +is the documented thread-safe perception-emit boundary with no escaping/validation. + +Proposed fix +Build the frame with cJSON (escapes both name and session_id; attach pre-validated data_json as a parsed +sub-object), or assert+escape and validate data_json parses. + +Suggested labels: bug, area:firmware +``` + +--- + +## [LOW] Debug log leftovers fire at WARN/ERROR on every WS and LLM frame + +``` +Summary +An ESP_LOGW on every LLM frame and an ESP_LOGE on every inbound WS frame bypass log-level filtering, +spam serial, and bury real errors. + +Location +firmware/firmware/patches/xiaozhi-esp32.patch:65,590 + +Details +Patch line 65 logs WARN on every LLM frame; line 590 logs ERROR on every inbound WS frame (hello, tts, +stt, llm, mcp, event acks). + +Proposed fix +Drop these two lines from the patch, or demote to ESP_LOGD. + +Suggested labels: bug, area:firmware +``` + +--- + +## [LOW] HeadPet/Imu gesture flags are plain volatile bool, not atomic; fixed order can invert fast sequences + +``` +Summary +A gesture can be lost (non-atomic read-then-clear across cores), and a fast release-then-press processed +in fixed press→release order ends with _is_touched=false, silently never arming the hold-to-listen window. + +Location +firmware/firmware/main/stackchan/modifiers/head_pet.h:50-52 (and head_pet.cpp:48-83, imu.h:124) + +Details +_event_press/_event_swipe/_event_release/_event_shake are volatile bools written from the HAL signal +callback (possibly a different core) and read+cleared on the tick task in a hardcoded order. + +Proposed fix +Use std::atomic with exchange(false) to consume the flags, or post gestures through a FreeRTOS +queue so they are consumed in arrival order. + +Suggested labels: bug, area:firmware +``` + +--- + +## [LOW] ParentalGate unlock at millis()==0 treated as "never unlocked" + +``` +Summary +A successful unlock in the first millisecond after boot stores 0, which isUnlocked() treats as the +"never" sentinel, so the unlock is silently dropped. + +Location +firmware/firmware/main/stackchan/face/parental_gate.cpp:43,71-72 + +Details +_unlocked_at_ms uses 0 as the sentinel; tryUnlockByPIN/tryUnlockByLongPress store now_ms() which can be +0 right after boot. (Dormant scaffold.) + +Proposed fix +Use a separate atomic _is_unlocked alongside the timestamp, or store max(now,1) and treat 0 +strictly as "never". + +Suggested labels: bug, area:firmware +``` + +--- + +## [LOW] CameraArbiter::acquireForCapture clears the shared _capture_pending on timeout + +``` +Summary +_capture_pending is a single shared atomic, not per-call; a future second capture path plus a timeout +could drop both callers' intent, wedging the other's still-pending capture. + +Location +firmware/firmware/main/stackchan/face/camera_arbiter.cpp (header camera_arbiter.h:38-44) + +Details +acquireForCapture() sets _capture_pending = true on entry and clears it to false on mutex-take timeout. +Latent today (one Capture() caller) but contradicts the header's "two concurrent consumers compose +cleanly". + +Proposed fix +Make _capture_pending a refcount (fetch_add on entry, fetch_sub on timeout/release), or document+assert +single-caller. + +Suggested labels: bug, area:firmware +``` + +--- + +## [LOW] Whisper odd-length PCM buffer raises ValueError instead of degrading + +``` +Summary +A truncated/short final VAD frame (odd byte count) discards the whole utterance as an empty transcript +with no retry. + +Location +custom-providers/asr/whisper_local.py:104 + +Details +np.frombuffer(artifacts.pcm_bytes, dtype=np.int16) raises ValueError on an odd byte count, caught by the +broad except → ('', None). The failure is deterministic so the retry loop can't help. + +Proposed fix +Trim to an even length before frombuffer: +buf = artifacts.pcm_bytes; buf = buf[: len(buf) - (len(buf) % 2)]. + +Suggested labels: bug, area:asr +``` + +--- + +## [LOW] get_emotion matches the first mapped emoji anywhere, not the leading emoji; default 🙂 outside the set + +``` +Summary +get_emotion can select a wrong/secondary face when a non-prefix emoji leaks into the reply, and its +default 🙂 is outside the firmware-recognized set. + +Location +custom-providers/textUtils.py:146-152 + +Details +The protocol requires the FIRST char to be the emotion emoji, but get_emotion breaks on the first +EMOJI_MAP char wherever it appears; the default emoji='🙂'/emotion='happy' masks a missing prefix. + +Proposed fix +Resolve emotion from the leading grapheme only (inspect text[0]), falling back to FALLBACK_EMOJI/😐 +rather than 🙂/happy. + +Suggested labels: bug, area:tts +``` + +--- + +## [LOW] piper effective_rate==24000 shortcut skips resampling when native rate isn't 24000 + +``` +Summary +A non-24000 Hz voice whose pitch_scale lands effective_rate on 24000 emits raw native-rate PCM into a +24000 Hz encoder, playing back at the wrong speed/pitch. + +Location +custom-providers/piper_local/piper_local.py:67-74 + +Details +The no-resample shortcut (self._up = self._down = 1) fires whenever effective_rate == 24000, conflating +"no pitch shift" with "no rate conversion". The common default voice en_GB-cori-medium is 22050 Hz. + +Proposed fix +Gate the shortcut on the SOURCE rate: only _up=_down=1 when src_rate==24000 AND pitch_scale==1.0. +Otherwise always go through the gcd/resample branch (compute g=gcd(effective_rate,24000), up=24000//g, +down=effective_rate//g unconditionally). + +Suggested labels: bug, area:tts +``` + +--- + +## [LOW] piper length_scale not validated for <= 0 (unlike pitch_scale) + +``` +Summary +A configured length_scale <= 0 yields empty/garbage audio with no early error, while pitch_scale raises +a clear ValueError. + +Location +custom-providers/piper_local/piper_local.py:32-42 + +Details +pitch_scale is validated positive (33-36); length_scale is taken straight from config into +SynthesisConfig(length_scale=float(length_scale)) with no bounds check. + +Proposed fix +Mirror the pitch_scale guard: if length_scale is not None and float(length_scale) <= 0, raise +ValueError. Construct self.syn_config only after the check. + +Suggested labels: bug, area:tts +``` + +--- + +## [LOW] whisper metrics: lang_prob can be None with forced language, dropping the ASR-METRICS line + +``` +Summary +With a forced language, the #105 ASR-METRICS line is silently dropped on every utterance where +language_probability is None. + +Location +custom-providers/asr/whisper_local.py:132-137 + +Details +getattr(_info, "language_probability", 0.0) only fires when the attribute is ABSENT, not present-but-None; +{lang_prob:.3f} then raises TypeError, swallowed by the surrounding try/except. + +Proposed fix +lang_prob = getattr(_info, "language_probability", 0.0) or 0.0 (guards both absent and present-but-None). + +Suggested labels: bug, area:asr +``` + +--- + +## [LOW] FIRST audio-start marker emitted even when synth yields zero audio + +``` +Summary +A mid-reply empty-synth segment emits a FIRST marker with no frames and no LAST, wedging per-sentence +marker accounting. + +Location +custom-providers/edge_stream/edge_stream.py:101-115; piper_local.py:146-159 (mirror) + +Details +text_to_speak puts (SentenceType.FIRST, [], text) before synthesis; on empty audio for a non-last +segment it returns without enqueueing any frames or a terminator. + +Proposed fix +Emit FIRST only after confirming there is audio, or on the empty-synth early return also enqueue a +terminating marker so every FIRST is balanced. Apply to both files. + +Suggested labels: bug, area:tts +``` + +--- + +## [LOW] security_watch defaults use retired RPi/zeroclaw paths and wrong port + +``` +Summary +security_watch's CONVO_LOG_DIR default disagrees with dashboard.py, and _BRIDGE_INTERNAL_URL points at +the wrong port (8080 vs the bridge's 8081). + +Location +bridge/security_watch.py:85-101 + +Details +SECURITY_LOG_DIR defaults CONVO_LOG_DIR to /root/zeroclaw-bridge/logs (the retired RPi path) while +dashboard.py defaults it to /var/lib/dotty-bridge/logs; _BRIDGE_INTERNAL_URL defaults to +http://127.0.0.1:8080 while the bridge listens on 8081. + +Proposed fix +Align the CONVO_LOG_DIR default with dashboard.py and fix _BRIDGE_INTERNAL_URL to :8081. Better: since +the whole consumer is unreachable, delete the module or gate it behind a clear 'not wired' guard. + +Suggested labels: bug, area:dashboard +``` + +--- + +## [LOW] Dashboard live perception feed (EventSource '/api/perception/feed') 404s — endpoint ripped in #111 + +``` +Summary +The dashboard's live perception EventSource 404s on the bridge (the endpoint was deleted in #111, the +consumer left behind), so the feed never streams, every dotty-refresh nudge is dead, and the browser +hammers a 404 reconnect every few seconds per open tab. + +Location +bridge/templates/dashboard.html:564 + +Details +The page (served from :8081) opens EventSource('/api/perception/feed'), but the bridge serves no /api/* +routes. Commit 30d9113 added the endpoint; c6df5c5 (#111) deleted it and left the EventSource. The real +SSE lives on dotty-behaviour (routes/perception.py:127) at :8090 — a different origin. + +Proposed fix +Add a bridge SSE passthrough route proxying dotty-behaviour's /api/perception/feed (under the /api/ +CSRF-exempt prefix) so the relative URL resolves, OR point the EventSource at the dotty-behaviour origin +explicitly with CORS configured. + +Suggested labels: bug, area:dashboard +``` + +--- + +## [LOW] play-asset / vision-photo proxy interpolate unencoded device_id into an internal URL + +``` +Summary +_fetch_robot_photo builds the dotty-behaviour URL with the raw device_id path-param (unvalidated, +unencoded), allowing query/fragment injection into the internal request. + +Location +bridge/dashboard.py:201,638,1924 + +Details +f"{_DOTTY_BEHAVIOUR_URL}/api/vision/photo/{device_id}" with device_id from the URL; FastAPI path +matching blocks literal /, but query/fragment chars still alter the request target. + +Proposed fix +urllib.parse.quote(device_id, safe='') before interpolation, and validate device_id against an expected +id charset (e.g. re.fullmatch(r'[A-Za-z0-9:_-]+', device_id)), returning 400 on mismatch. + +Suggested labels: bug, area:dashboard +``` + +--- + +## [LOW] _dotty_behaviour_get caches the failure fallback for the full TTL + +``` +Summary +A single transient dotty-behaviour failure pins an empty {}/[] result for the cache TTL, so every +dashboard tile blanks for ~2s after each blip even after recovery. + +Location +bridge.py:492-504 + +Details +On any timeout/connection/HTTP/JSON error the helper sets value = fallback and unconditionally writes it +into _dotty_behaviour_cache with the full TTL. + +Proposed fix +Only populate the cache on success (move the write inside the try after r.json()), returning fallback on +the exception path without caching it. + +Suggested labels: bug, area:dashboard +``` + +--- + +## [LOW] bridge writes to brain.db (RW) primarily owned by dotty-pi — approve/redact can lock or race + +``` +Summary +The bridge's memory approve/redact open brain.db RW across containers; in non-WAL delete mode a writer +blocks readers, and failures are swallowed and surface as a generic "not found". + +Location +bridge.py:316,349 + +Details +_voice_memory_approve_blocking / _voice_memory_delete_blocking sqlite3.connect(str(_VOICE_MEMORY_DB)) +RW (timeout=5) against brain.db, which dotty-pi-ext actively writes. The list path correctly uses +?mode=ro. + +Proposed fix +Open with WAL-aware busy handling (PRAGMA busy_timeout); surface lock failures distinctly from +'not found'. Longer term, route approve/redact through dotty-pi (the canonical writer). + +Suggested labels: bug, area:dashboard, area:dotty-pi +``` + +--- + +## [LOW] host_detail 'server' modal hardcodes 'Recent errors today: —' despite the count being computed elsewhere + +``` +Summary +The server host-detail modal shows a permanent dash for today's error count, even though the bridge +already tallies it for the alerts chip. + +Location +bridge/dashboard.py:1832 + +Details +The modal renders ('Recent errors today', '—') with a TODO, but alerts_count() (485-500) already counts +rec.get('error') from the convo log. + +Proposed fix +Factor the per-day error tally out of alerts_count() into a helper and reuse it. If the intent is +xiaozhi-container errors specifically, note that in the label rather than a bare '—'. + +Suggested labels: bug, area:dashboard +``` + +--- + +## [LOW] PiVoiceLLM fallback strings violate the mandatory emoji-prefix protocol + +``` +Summary +PiVoiceLLM's server-generated fallback strings begin with '(', so the firmware emotion parser sees '(' +and produces no/garbage face state. + +Location +custom-providers/pi_voice/pi_voice.py:130,150 + +Details +yield "(empty turn)" and yield "(brain offline — try again in a moment)" bypass the persona and lack an +emoji prefix. OpenAICompat correctly prefixes all its fallbacks with FALLBACK_EMOJI. + +Proposed fix +Prepend an allowed emoji to both fallbacks (e.g. "😐 (empty turn)", "😐 (brain offline — try again ...)"), +or import ensure_emoji_prefix/FALLBACK_EMOJI from the shared textUtils. + +Suggested labels: bug, area:voice +``` + +--- + +## [LOW] new_session response match does not verify request id + +``` +Summary +new_session accepts a stale response from a previously-timed-out call because the drain loop never +checks the request id. + +Location +custom-providers/pi_voice/pi_client.py:192-209 + +Details +new_session sends req_id but the drain loop matches only type=='response' and command=='new_session', +unlike iter_turn_text which matches the accept-ack by id. + +Proposed fix +Add `and frame.get('id') == req_id` to the match condition. + +Suggested labels: bug, area:voice +``` + +--- + +## [LOW] PiVoiceLLM _first_turn desyncs from the pi process after a respawn + +``` +Summary +If pi dies after turn 1, the next turn runs new_session() against a freshly respawned process that has +no session, producing a spurious 10s timeout and dead air. + +Location +custom-providers/pi_voice/pi_voice.py:136-141 + +Details +_first_turn is tracked in PiVoiceLLM but the spawn decision lives in PiClient._ensure_started(); after a +respawn _first_turn is False so new_session runs and waits up to 10s for an ack a clean process may +never send. + +Proposed fix +Have PiClient signal whether _ensure_started actually spawned a fresh process and skip new_session() in +that case (or move the freshly-spawned suppression into PiClient.new_session()). + +Suggested labels: bug, area:voice +``` + +--- + +## [LOW] iter_turn_text process-exit detection only runs on queue-empty + +``` +Summary +A pi crash with frames still queued is reported as a generic 120s turn timeout instead of "pi process +exited", delaying the user-facing fallback by up to 2 minutes. + +Location +custom-providers/pi_voice/pi_client.py:222-269 + +Details +The "pi process exited mid-turn" check is only reached inside the `except Empty` branch; if pi emits +some frames then dies, the loop drains them without hitting Empty and spins to the turn timeout. + +Proposed fix +Check self._proc.poll() is not None on every loop iteration and raise "pi process exited mid-turn" +immediately when the process is gone and no agent_end has been seen. + +Suggested labels: bug, area:voice +``` + +--- + +## [LOW] _next_id increments shared counter without the lock; reader-thread frame routing races a respawn + +``` +Summary +A respawn swaps _event_queue while an old, never-joined reader thread can still put stale frames into +the new queue, injecting a previous process's frames into the new turn. + +Location +custom-providers/pi_voice/pi_client.py:283-285,149-169 + +Details +_next_id mutates _next_req_id outside _lock; _ensure_started() replaces _event_queue on every respawn; +the stdout reader routes frames by reading self._event_queue live without the lock. + +Proposed fix +Join (or signal-stop) the prior reader/stderr threads before swapping _event_queue, bind each reader to +its specific Queue/proc instance (pass as args), and take the lock around _next_req_id mutation. + +Suggested labels: bug, area:voice +``` + +--- + +## [LOW] _completions_url mishandles a base ending in /chat/completions + +``` +Summary +A base_url whose path contains but doesn't end in /chat/completions (e.g. a gateway route +/chat/completions/stream) gets a second /chat/completions appended. + +Location +custom-providers/openai_compat/openai_compat.py:125-132 + +Details +The endsWith heuristic on an already-rstripped base appends /chat/completions when the base path +contains the literal but isn't exactly it. + +Proposed fix +Make the endpoint construction explicit via a config flag (url is base vs full endpoint), or document +that url must be the base and drop the endswith special-case. + +Suggested labels: bug, area:voice +``` + +--- + +## [LOW] OpenAICompat yields leading whitespace before emoji-prefix engages + +``` +Summary +A whitespace-only first chunk is yielded to TTS before enforcement, so the first TTS chunk isn't the +emoji the contract promises. + +Location +custom-providers/openai_compat/openai_compat.py:218-229 + +Details +Enforcement fires only once so_far = ''.join(full_text).lstrip() is non-empty; a leading-whitespace +content chunk is still emitted via yield content while emoji_checked is False. + +Proposed fix +Buffer (do not yield) content until so_far is non-empty and the emoji check has run; strip leading +whitespace from the first emitted chunk. + +Suggested labels: bug, area:voice +``` + +--- + +## [LOW] OTAHandler.handle_post finally-block can swallow a secondary exception + +``` +Summary +handle_post's finally calls _add_cors_headers with no try/except and a bare return that masks a +propagating exception, so an edge case 500s the device with non-JSON. + +Location +custom-providers/xiaozhi-patches/ota_handler.py:359-361 + +Details +handle_download wraps _add_cors_headers in try/except; handle_post does not, and the bare `return` in +finally masks any exception from the except branch. + +Proposed fix +Mirror handle_download: wrap _add_cors_headers in try/except; initialize response = None and +short-circuit if still None. + +Suggested labels: bug, area:xiaozhi +``` + +--- + +## [LOW] _dotty_inject_text reads conn.headers without the 'or {}' guard used by sibling handlers + +``` +Summary +inject-text raises AttributeError (returning a misleading 500 after already dispatching chat) if +conn.headers is present-but-None. + +Location +custom-providers/xiaozhi-patches/http_server.py:66 + +Details +Line 66 does getattr(conn,'headers',{}).get('device-id',''); every other admin handler uses +(getattr(conn,'headers',{}) or {}).get(...). + +Proposed fix +Apply the same (getattr(conn,'headers',{}) or {}) guard on line 66. + +Suggested labels: bug, area:xiaozhi +``` + +--- + +## [LOW] OTA mqtt branch ships an empty-password block with no websocket fallback + +``` +Summary +A misconfigured MQTT signature key yields an mqtt config with a blank password and no websocket +fallback — the device gets a connection config it cannot authenticate with. + +Location +custom-providers/xiaozhi-patches/ota_handler.py:260-279 + +Details +When mqtt_gateway is configured but mqtt_signature_key is missing (or generate_password_signature +fails), the handler sends mqtt with password:"" and, being exclusive with the websocket else, no +alternate. + +Proposed fix +If signature is required but unavailable, fall back to the websocket block (or return an error) rather +than shipping an mqtt config with an empty password. + +Suggested labels: bug, area:xiaozhi +``` + +--- + +## [LOW] OTA model-name matching is brittle — updates silently never offered + +``` +Summary +device_model must exactly match the operator's bin-filename prefix; when they differ (the common case) +the device is silently told it is up-to-date. + +Location +custom-providers/xiaozhi-patches/ota_handler.py:84-90,200-201,304 + +Details +files_by_model.get(device_model, []) requires an exact string match; the StackChan board reports +board.type, which rarely equals the human-chosen filename prefix. The same log fires for "model unknown" +and "already latest". + +Proposed fix +Log the parsed device_model + known model keys when candidates is empty; allow a configured fallback +bucket or case-insensitive match; emit distinct log lines for the two cases. + +Suggested labels: bug, area:xiaozhi +``` + +--- + +## [LOW] OTA timezone_offset / firmware_cache_ttl assume numeric YAML + +``` +Summary +A quoted YAML timezone_offset ("8") produces a 60-copy string in the OTA payload; a non-int +firmware_cache_ttl makes int(...) raise. + +Location +custom-providers/xiaozhi-patches/ota_handler.py:62,226 + +Details +server_config.get("timezone_offset", 8) * 60 and int(self._bin_cache.get("ttl", 30)) trust the YAML +type. + +Proposed fix +Coerce with int(server_config.get("timezone_offset", 8)) * 60 and validate firmware_cache_ttl at +construction (int() with try/except defaulting to 30). + +Suggested labels: bug, area:xiaozhi +``` + +--- + +## [LOW] Firmware bin cache re-scans disk on every OTA POST until a valid bin exists + +``` +Summary +With no conforming bins yet, the TTL short-circuit never considers the cache warm, so the glob+regex +scan + makedirs runs on every OTA POST. + +Location +custom-providers/xiaozhi-patches/ota_handler.py:66-72 + +Details +The short-circuit requires a non-empty files_by_model; {} is falsy, so updated_at is bumped meaninglessly +each pass and the scan never caches. + +Proposed fix +Track freshness with a separate _scanned flag/updated_at sentinel independent of emptiness, so an +empty-but-recent scan is honored for the TTL window. + +Suggested labels: bug, area:xiaozhi +``` + +--- + +## [LOW] Admin MCP handlers fire-and-forget send() with no guard for a torn-down socket + +``` +Summary +A single late send to an already-closed WS fails invisibly while the HTTP caller already got {"ok":true}. + +Location +custom-providers/xiaozhi-patches/http_server.py:146,194,242,288,338 + +Details +Each direct-MCP route does _spawn(conn.websocket.send(msg), ...) after the registry lookup; between +lookup and send executing, the WS can close (the registry pop is in the WS handler's finally), so +send() raises in the spawned task. + +Proposed fix +Wrap the send in a coroutine that checks conn.websocket state and catches+logs the exception, and/or +await the send before returning so the HTTP response reflects whether the frame was queued. + +Suggested labels: bug, area:xiaozhi +``` + +--- + +## [LOW] set-toggle/set-state have no response correlation; set-head-angles unclamped + +``` +Summary +A device-side MCP rejection still returns {"ok":true} (no response correlation), and set-head-angles +forwards arbitrary ints with no clamp. + +Location +custom-providers/xiaozhi-patches/http_server.py:143,191,239,285,335 + +Details +MCP calls use a truncated wall-clock id and the handler returns before any response frame; set_head_angles +validates only type, so a 100000 speed or 999 yaw is sent verbatim. + +Proposed fix +Clamp set-head-angles yaw/pitch/speed to firmware-legal ranges before sending. Optionally correlate by +request id with a timeout, or document the endpoints as best-effort. + +Suggested labels: bug, area:xiaozhi +``` + +--- + +## [LOW] play-asset format inference treats bare .opus as Ogg + +``` +Summary +A .opus asset shown in the songs picker can fail to play (forced format=ogg fails on raw Opus) with only +a swallowed warning behind an already-sent 200. + +Location +custom-providers/xiaozhi-patches/http_server.py:406-411 + +Details +fmt is derived from the extension; {"opus":"ogg",...}.get(ext) forces .opus to ogg, and unmapped +extensions yield fmt=None (autodetect) with no distinct log. + +Proposed fix +Let ffmpeg autodetect for .opus (pass no explicit format); on decode failure return a distinguishable +error code rather than a swallowed warning behind a 200. + +Suggested labels: bug, area:xiaozhi +``` + +--- + +## [LOW] play-asset resample can feed huge factors to resample_poly; no sample_rate guard + +``` +Summary +Coprime rates give large up/down factors (slow/memory-heavy filter), and a 0/None conn.sample_rate blows +up inside _decode masked by the broad except. + +Location +custom-providers/xiaozhi-patches/http_server.py:415-420 + +Details +g = gcd(src_rate, tgt); resample_poly(pcm, tgt//g, src_rate//g). No guard that conn.sample_rate is a +positive int; failures are logged as 'decode failed', hiding the cause. + +Proposed fix +Validate conn.sample_rate > 0 (return 400/503 if not); log the actual src/tgt rates on decode failure; +consider capping or special-casing common rates. + +Suggested labels: bug, area:xiaozhi +``` + +--- + +## [LOW] Admin routes resolve target via next(iter(...)) — nondeterministic device with multiple StackChans + +``` +Summary +With more than one device connected and no device_id, admin handlers act on an arbitrary dict entry, so +abort/set-state can hit the wrong robot. + +Location +custom-providers/xiaozhi-patches/http_server.py:55,87,123,171,218,266,316,388,523 + +Details +next(iter(_dotty_active_connections.values()), None) picks the first dict entry; the caller has no way +to know which device was targeted. + +Proposed fix +When multiple devices are connected and no device_id is given, return 409/400 requiring an explicit +device_id (or target all). Single-device deployments are unaffected. + +Suggested labels: bug, area:xiaozhi +``` + +--- + +## [LOW] EventTextMessageHandler attributes events to "unknown" device on header miss + +``` +Summary +If conn.headers is None, the relayed perception event is attributed to device "unknown", which never +matches the real device id used by the admin routes, so per-device consumers silently never fire. + +Location +custom-providers/xiaozhi-patches/textMessageHandlerRegistry.py:60-65 + +Details +handle() does device_id = conn.headers.get("device-id","unknown") inside a bare try/except that swallows +all exceptions, leaving device_id="unknown". + +Proposed fix +Resolve device_id like the WS registration ((getattr(conn,"headers",None) or {}).get("device-id")); if +it resolves falsy/"unknown", drop the event with a warning rather than relay under a fabricated id. + +Suggested labels: bug, area:xiaozhi, area:behaviour +``` diff --git a/bridge.py b/bridge.py index b4783d23..5bbf2ad5 100644 --- a/bridge.py +++ b/bridge.py @@ -15,12 +15,14 @@ """ import asyncio +import hmac import json import logging import os import re import sys import time +from collections import deque from contextlib import asynccontextmanager from pathlib import Path from typing import Any @@ -596,25 +598,19 @@ def _dashboard_last_user_line_getter(device_id: str) -> dict | None: return None -def _dashboard_sound_balance_series() -> list[float]: - """Live sound_event balance series from dotty-behaviour. +def _dashboard_open_perception_feed() -> Any: + """Open dotty-behaviour's perception SSE stream for the dashboard proxy. - Picks the first device with perception state (single-robot - deployment heuristic — matches the device-picking strategy used - for the perception card) and reads its sound-balance ring via - /api/perception/sound-balance/{device_id}. Same cache + timeout + - circuit-breaker contract as the other dotty-behaviour-backed - getters — returns ``[]`` on any failure.""" - state = _dashboard_perception_state_getter() - device_id = next(iter(state), None) if isinstance(state, dict) else None - if not device_id: - return [] - result = _dotty_behaviour_get( - f"/api/perception/sound-balance/{device_id}", - {"limit": 30}, - [], + Returns a streaming ``requests.Response`` (caller closes it). The read + timeout sits above the upstream's 15 s keepalive so an idle-but-healthy + stream never trips it. Raises on connect failure / non-2xx.""" + r = requests.get( + f"{DOTTY_BEHAVIOUR_URL}/api/perception/feed", + stream=True, + timeout=(_DOTTY_BEHAVIOUR_TIMEOUT_SEC, 30.0), ) - return result if isinstance(result, list) else [] + r.raise_for_status() + return r def _dashboard_vision_failures_last_hour() -> dict[str, int]: @@ -631,10 +627,30 @@ def _identity_display_name(identity: str) -> str | None: return None -# Stub SSE plumbing — kept so /ui/events still wakes the queue handler -# (sends heartbeats) when the browser subscribes. No producer is wired in -# bridge.py post-#111; convo turns are owned by dotty-pi now. +# Turn stream — /ui/events subscribers plus the recent-turns ring behind +# the Turns tab, Errors count and error report. Voice turns are owned by +# dotty-pi since #111, so the PiVoiceLLM provider reports each completed +# turn to /api/voice/turn (below). In-memory ONLY: the ring is lost on +# restart and transcripts are never written to disk. _dashboard_event_listeners: list[asyncio.Queue] = [] +_DASHBOARD_TURNS_MAX = 200 +_dashboard_recent_turns: "deque[dict]" = deque(maxlen=_DASHBOARD_TURNS_MAX) + + +def _dashboard_publish_turn(record: dict) -> None: + """Record a completed turn and fan it out to /ui/events subscribers. + A slow subscriber's full queue drops the event rather than blocking.""" + _dashboard_recent_turns.append(record) + for q in list(_dashboard_event_listeners): + try: + q.put_nowait(record) + except asyncio.QueueFull: + pass + + +def _dashboard_recent_turns_getter() -> list[dict]: + """Recent turns, oldest first.""" + return list(_dashboard_recent_turns) def _dashboard_subscribe_events() -> asyncio.Queue: @@ -764,6 +780,7 @@ async def _dashboard_memory_redact(mem_id: str) -> dict: abort_device=_dashboard_abort_device, subscribe_events=_dashboard_subscribe_events, unsubscribe_events=_dashboard_unsubscribe_events, + recent_turns_getter=_dashboard_recent_turns_getter, perception_state_getter=_dashboard_perception_state_getter, perception_recent_getter=_dashboard_perception_recent_getter, memory_records_getter=_dashboard_memory_records, @@ -771,11 +788,74 @@ async def _dashboard_memory_redact(mem_id: str) -> dict: memory_redact=_dashboard_memory_redact, identity_display_name=_identity_display_name, last_user_line_getter=_dashboard_last_user_line_getter, - sound_balance_getter=_dashboard_sound_balance_series, + perception_feed_opener=_dashboard_open_perception_feed, vision_failures_getter=_dashboard_vision_failures_last_hour, ) +# --------------------------------------------------------------------------- +# /api/voice/* — turn + filter-hit reports from the voice provider +# --------------------------------------------------------------------------- +# PiVoiceLLM runs inside the xiaozhi container, so the bridge only learns +# about a voice turn if the provider tells it. Guarded by the same shared +# X-Admin-Token as /xiaozhi/admin/* when DOTTY_ADMIN_TOKEN is set; open on +# the LAN otherwise, matching the rest of the stack's unset-token posture. + +_VOICE_TEXT_MAX = 2000 + + +def _voice_require_admin_token(request: Request) -> None: + if not _ADMIN_TOKEN: + return + supplied = request.headers.get("X-Admin-Token", "") + if not hmac.compare_digest(supplied.encode(), _ADMIN_TOKEN.encode()): + raise HTTPException(status_code=401, detail="bad admin token") + + +class _VoiceTurnIn(BaseModel): + request_text: str = "" + response_text: str = "" + latency_ms: int | None = None + error: str | None = None + channel: str = "dotty" + + +class _VoiceFilterHitIn(BaseModel): + tier: str + rule: str = "" + prefix: str = "" + + +_voice_router = APIRouter( + prefix="/api/voice", dependencies=[Depends(_voice_require_admin_token)], +) + + +@_voice_router.post("/turn") +async def _voice_turn(payload: _VoiceTurnIn) -> dict: + _dashboard_publish_turn({ + "ts": time.time(), + "channel": payload.channel[:32] or "dotty", + "request_text": payload.request_text[:_VOICE_TEXT_MAX], + "response_text": payload.response_text[:_VOICE_TEXT_MAX], + "latency_ms": payload.latency_ms, + "error": (payload.error or "")[:500] or None, + }) + return {"ok": True} + + +@_voice_router.post("/filter-hit") +async def _voice_filter_hit(payload: _VoiceFilterHitIn) -> dict: + from bridge.text import record_content_filter_hit + record_content_filter_hit( + payload.tier[:16], payload.rule[:64], payload.prefix, + ) + return {"ok": True} + + +app.include_router(_voice_router) + + # --------------------------------------------------------------------------- # /admin/* — localhost-only runtime configuration mutations # --------------------------------------------------------------------------- diff --git a/bridge/dashboard.py b/bridge/dashboard.py index 6ff97548..a3382191 100644 --- a/bridge/dashboard.py +++ b/bridge/dashboard.py @@ -46,6 +46,7 @@ "abort_device": None, "subscribe_events": None, "unsubscribe_events": None, + "recent_turns_getter": None, "perception_state_getter": None, "perception_recent_getter": None, "identity_display_name": None, @@ -53,7 +54,7 @@ "memory_records_getter": None, "memory_approve": None, "memory_redact": None, - "sound_balance_getter": None, + "perception_feed_opener": None, "vision_failures_getter": None, } @@ -67,6 +68,7 @@ def configure(*, send_message: Any = None, vision_cache_getter: Any = None, inject_to_device: Any = None, abort_device: Any = None, subscribe_events: Any = None, unsubscribe_events: Any = None, + recent_turns_getter: Any = None, perception_state_getter: Any = None, perception_recent_getter: Any = None, identity_display_name: Any = None, @@ -74,7 +76,7 @@ def configure(*, send_message: Any = None, vision_cache_getter: Any = None, memory_records_getter: Any = None, memory_approve: Any = None, memory_redact: Any = None, - sound_balance_getter: Any = None, + perception_feed_opener: Any = None, vision_failures_getter: Any = None) -> None: """Register bridge state with the dashboard. Idempotent.""" if send_message is not None: @@ -105,6 +107,8 @@ def configure(*, send_message: Any = None, vision_cache_getter: Any = None, _state["subscribe_events"] = subscribe_events if unsubscribe_events is not None: _state["unsubscribe_events"] = unsubscribe_events + if recent_turns_getter is not None: + _state["recent_turns_getter"] = recent_turns_getter if perception_state_getter is not None: _state["perception_state_getter"] = perception_state_getter if perception_recent_getter is not None: @@ -119,8 +123,8 @@ def configure(*, send_message: Any = None, vision_cache_getter: Any = None, _state["memory_approve"] = memory_approve if memory_redact is not None: _state["memory_redact"] = memory_redact - if sound_balance_getter is not None: - _state["sound_balance_getter"] = sound_balance_getter + if perception_feed_opener is not None: + _state["perception_feed_opener"] = perception_feed_opener if vision_failures_getter is not None: _state["vision_failures_getter"] = vision_failures_getter @@ -298,7 +302,6 @@ def _xiaozhi_admin_headers() -> dict[str, str]: is set (matches the xiaozhi-server middleware); empty otherwise.""" return {"X-Admin-Token": _ADMIN_TOKEN} if _ADMIN_TOKEN else {} LOG_DIR = Path(os.environ.get("CONVO_LOG_DIR", "/var/lib/dotty-bridge/logs")) -VOICE_CHANNELS = ("dotty", "stackchan") _START_TIME = time.time() @@ -340,12 +343,27 @@ def _humanize_age(seconds: float) -> str: return f"{s // 86400}d" -def _today_log_path() -> Path: - return _log_path_for(datetime.now().strftime("%Y-%m-%d")) +def _recent_turns() -> list[dict]: + """Recent voice turns from the bridge's in-memory ring, oldest first. + Empty after a bridge restart — nothing is persisted.""" + getter = _state.get("recent_turns_getter") + try: + return list(getter()) if getter else [] + except Exception: + return [] -def _log_path_for(date_str: str) -> Path: - return LOG_DIR / f"convo-{date_str}.ndjson" +def _errored_turns_today() -> list[dict]: + """Today's (local date) errored turns, newest first.""" + today = datetime.now().date() + out: list[dict] = [] + for rec in reversed(_recent_turns()): + ts = rec.get("ts") + if not rec.get("error") or not isinstance(ts, (int, float)): + continue + if datetime.fromtimestamp(ts).date() == today: + out.append(rec) + return out def _clean_request_text(s: str) -> str: @@ -372,37 +390,34 @@ def _clean_request_text(s: str) -> str: return after -def _parse_ts(ts: str) -> float | None: - if not ts: - return None +def _stackchan_last_seen() -> float | None: + """Timestamp of the most recent voice activity. + + Prefers the newest reported turn; falls back to dotty-behaviour's + per-device `last_chat_t`, which also covers the window after a bridge + restart when the turn ring is empty.""" + return _last_turn_ts() or _perception_last_chat_ts() + + +def _perception_last_chat_ts() -> float | None: + getter = _state.get("perception_state_getter") try: - return datetime.fromisoformat(ts.replace("Z", "+00:00")).timestamp() + pstate = (getter() if getter else None) or {} except Exception: return None + stamps = [ + d["last_chat_t"] for d in pstate.values() + if isinstance(d, dict) and isinstance(d.get("last_chat_t"), (int, float)) + ] + return max(stamps) if stamps else None -def _stackchan_last_seen() -> float | None: - """Timestamp of the most recent voice-channel turn in today's log.""" - path = _today_log_path() - if not path.exists(): - return None - try: - data = path.read_bytes() - except OSError: - return None - last_voice_ts: float | None = None - for line in data.splitlines(): - if not line.strip(): - continue - try: - rec = json.loads(line) - except Exception: - continue - if rec.get("channel") in VOICE_CHANNELS: - ts = _parse_ts(rec.get("ts", "")) - if ts is not None: - last_voice_ts = ts - return last_voice_ts +def _last_turn_ts() -> float | None: + for rec in reversed(_recent_turns()): + ts = rec.get("ts") + if isinstance(ts, (int, float)): + return float(ts) + return None @router.get("", response_class=HTMLResponse, include_in_schema=False) @@ -479,7 +494,7 @@ async def device_status(request: Request) -> Any: @router.get("/alerts/count", response_class=HTMLResponse, include_in_schema=False) async def alerts_count(request: Request, chip: int = 0) -> Any: - """Q6: count today's errored turns from the convo log. + """Q6: count today's errored turns from the recent-turns ring. Two render modes share the same count so the dashboard polls one URL: - default: legacy floating ``alerts_badge.html`` (kept for @@ -490,22 +505,7 @@ async def alerts_count(request: Request, chip: int = 0) -> Any: the chip. innerHTML swap target is the chip itself, so the script runs against its own parent element. """ - today = datetime.now().strftime("%Y-%m-%d") - path = _log_path_for(today) - n = 0 - if path.exists(): - try: - for line in path.read_bytes().splitlines(): - if not line.strip(): - continue - try: - rec = json.loads(line) - except Exception: - continue - if rec.get("error"): - n += 1 - except OSError: - pass + n = len(_errored_turns_today()) if chip: # Don't tint the chip red — selection (`btn-primary` via setFeedFilter) # is the only "selected" cue; the chip just shows a warning glyph + @@ -526,37 +526,15 @@ async def alerts_count(request: Request, chip: int = 0) -> Any: async def alerts_detail(request: Request) -> Any: """F13: render today's errored turns. Opened via the alerts-badge modal in dashboard.html.""" - today = datetime.now().strftime("%Y-%m-%d") - path = _log_path_for(today) entries: list[dict[str, Any]] = [] - if path.exists(): - try: - lines = path.read_bytes().splitlines() - except OSError: - lines = [] - for line in reversed(lines): - if not line.strip(): - continue - try: - rec = json.loads(line) - except Exception: - continue - if not rec.get("error"): - continue - ts = rec.get("ts", "") - try: - time_str = datetime.fromisoformat( - ts.replace("Z", "+00:00") - ).astimezone().strftime("%H:%M:%S") - except Exception: - time_str = ts[-8:] if ts else "?" - entries.append({ - "time": time_str, - "channel": rec.get("channel") or "?", - "request": _clean_request_text(rec.get("request_text") or "")[:400], - "response": (rec.get("response_text") or "")[:300], - "error": str(rec.get("error"))[:500], - }) + for rec in _errored_turns_today(): + entries.append({ + "time": datetime.fromtimestamp(rec["ts"]).astimezone().strftime("%H:%M:%S"), + "channel": rec.get("channel") or "?", + "request": _clean_request_text(rec.get("request_text") or "")[:400], + "response": (rec.get("response_text") or "")[:300], + "error": str(rec.get("error"))[:500], + }) return templates.TemplateResponse( request, "alerts_detail.html", {"entries": entries}, @@ -663,17 +641,13 @@ def _post() -> dict: return templates.TemplateResponse(request, "say_result.html", result) -_INJECT_WAIT_SEC = 8.0 # Q4: how long to wait for Dotty's reply before - # showing "no response in time" fallback. - - async def _inject_or_error(request: Request, text: str, label: str) -> Any: """Helper for action endpoints that fire text into xiaozhi-server's pipeline so the device actually speaks/emotes/runs MCP tools. - Q4: subscribes to the bridge's event stream BEFORE injecting, then - waits up to ~8s for the next turn so the dashboard can show what Dotty - actually said (not just "Sent…").""" + Returns as soon as the inject is accepted. It used to wait ~8 s for the + reply on the bridge's turn stream, but turns are owned by dotty-pi since + #111 and nothing publishes them here — the wait only ever timed out.""" inject = _state.get("inject_to_device") if inject is None: return templates.TemplateResponse( @@ -681,38 +655,22 @@ async def _inject_or_error(request: Request, text: str, label: str) -> Any: {"ok": False, "error": "Inject path not configured (xiaozhi admin patch missing)."}, ) - subscribe = _state.get("subscribe_events") - unsubscribe = _state.get("unsubscribe_events") - queue = subscribe() if subscribe else None try: - try: - result = await inject(text=text) - except Exception as exc: - log.exception("dashboard inject failed") - return templates.TemplateResponse( - request, "say_result.html", - {"ok": False, "error": f"Bridge error: {exc.__class__.__name__}"}, - ) - if not result.get("ok"): - return templates.TemplateResponse( - request, "say_result.html", - {"ok": False, "error": result.get("error", "unknown injection failure")}, - ) - # Wait for the next completed turn (likely ours — single device). - response_text = "Sent — no reply in 8s." - if queue is not None: - try: - event = await asyncio.wait_for(queue.get(), timeout=_INJECT_WAIT_SEC) - response_text = event.get("response_text") or "(no text)" - except asyncio.TimeoutError: - pass + result = await inject(text=text) + except Exception as exc: + log.exception("dashboard inject failed") return templates.TemplateResponse( request, "say_result.html", - {"ok": True, "sent": label, "response": response_text}, + {"ok": False, "error": f"Bridge error: {exc.__class__.__name__}"}, ) - finally: - if queue is not None and unsubscribe is not None: - unsubscribe(queue) + if not result.get("ok"): + return templates.TemplateResponse( + request, "say_result.html", + {"ok": False, "error": result.get("error", "unknown injection failure")}, + ) + return templates.TemplateResponse( + request, "say_result.html", {"ok": True, "sent": label}, + ) def _kid_blocked(text: str) -> bool: @@ -954,29 +912,6 @@ async def kid_mode_partial(request: Request) -> Any: } -def _sound_balance_sparkline() -> dict | None: - """Build the State-tile sound-balance sparkline (#69) — an SVG - polyline over the last 60 s of localizer `balance` samples. Returns - None when there's no recent sound data (fewer than 2 points).""" - getter = _state.get("sound_balance_getter") - series = getter() if getter else [] - if not series or len(series) < 2: - return None - width, height = 100.0, 24.0 - n = len(series) - pts: list[str] = [] - for i, bal in enumerate(series): - x = (i / (n - 1)) * width - y = (1.0 - max(0.0, min(1.0, float(bal)))) * height - pts.append(f"{x:.1f},{y:.1f}") - return { - "points": " ".join(pts), - "last": series[-1], - "width": width, - "height": height, - } - - def _vision_failures_count() -> int: """Total vision-capture failures in the last hour (#74) — summed across error kinds. 0 when the getter is unwired or the window is @@ -1002,21 +937,25 @@ async def state_partial(request: Request) -> Any: "states": _STATES, "labels": _STATE_LABELS, "descriptions": _STATE_DESCRIPTIONS, - "sound_spark": _sound_balance_sparkline(), }, ) # #65 tier 1 — voice tool inventory. Hardcoded list of the tools shipped # by dotty-pi-ext (see dotty-pi-ext/package.json). Per the issue body the -# static list is intentional: "hardcoded; same five values until a sixth -# tool ships". Tier 2 (call counts via PiClient stdout parsing) and Tier 3 -# (safety-denial counts) are deferred follow-ups. +# static list is intentional — keep it in step with dotty-pi-ext/src/tools/ +# (seven tools since #53 added recall_person / remember_person). Tier 2 +# (call counts via PiClient stdout parsing) and Tier 3 (safety-denial +# counts) are deferred follow-ups. _VOICE_TOOLS: list[dict[str, str]] = [ {"name": "memory_lookup", "description": "Search Dotty's long-term memory by keyword"}, {"name": "remember", "description": "Save a new memory (\"Brett's birthday is …\")"}, + {"name": "recall_person", + "description": "Look up what Dotty knows about a named person"}, + {"name": "remember_person", + "description": "Save a fact about a named person (kid-mode: pending review)"}, {"name": "think_hard", "description": "Hand a hard reasoning task to the bigger 27B model"}, {"name": "take_photo", @@ -1815,7 +1754,7 @@ async def host_detail(request: Request, slug: str) -> Any: smart_on = bool(smart_getter()) if smart_getter else None active_llm = _host_detail_llm_label(smart_on) facts = [ - ("Device", "Raspberry Pi"), + ("Device", "Docker container (dotty-bridge)"), ("Status", "online"), ("Version", BRIDGE_VERSION), ("Uptime", _humanize_age(time.time() - _START_TIME)), @@ -2099,6 +2038,54 @@ async def security_recent(request: Request, device_id: str) -> Any: return templates.TemplateResponse(request, "security_panel.html", ctx) +# --- Perception feed proxy ------------------------------------------------ + +@router.get("/perception/feed", include_in_schema=False) +async def perception_feed_proxy(request: Request) -> StreamingResponse: + """Relay dotty-behaviour's perception SSE stream to the browser. + + The perception bus moved to dotty-behaviour in #36, but the page can + only open same-origin EventSources — so the bridge relays the stream + line-for-line. If the upstream is down or drops, the stream just ends + and EventSource reconnects after the `retry` interval (a non-200 would + stop it retrying for good).""" + opener = _state.get("perception_feed_opener") + + async def gen(): + yield b"retry: 5000\n\n" + if opener is None: + return + try: + upstream = await asyncio.to_thread(opener) + except Exception as exc: + log.warning("perception feed upstream unavailable: %s", exc) + return + lines = upstream.iter_lines(chunk_size=1) + try: + while True: + if await request.is_disconnected(): + break + try: + line = await asyncio.to_thread(next, lines, None) + except Exception as exc: + log.warning("perception feed upstream dropped: %s", exc) + break + if line is None: + break + yield line + b"\n" + finally: + upstream.close() + + return StreamingResponse( + gen(), + media_type="text/event-stream", + headers={ + "Cache-Control": "no-cache", + "X-Accel-Buffering": "no", + }, + ) + + # --- P13 + P12: SSE event stream for live log + error toasts ------------- @router.get("/events", include_in_schema=False) @@ -2106,8 +2093,8 @@ async def events_stream(request: Request) -> StreamingResponse: """Server-Sent Events stream of completed conversation turns. Each event is one JSON object: {ts, channel, request_text, response_text, - latency_ms, error, emoji_used}. The bridge's ConvoLogger broadcasts on - every turn. Heartbeats every 15s keep proxies / browsers awake. + latency_ms, error}. Published when the voice provider reports a turn to + /api/voice/turn. Heartbeats every 15s keep proxies / browsers awake. """ subscribe = _state.get("subscribe_events") unsubscribe = _state.get("unsubscribe_events") diff --git a/bridge/templates/dashboard.html b/bridge/templates/dashboard.html index 075eceeb..89913d15 100644 --- a/bridge/templates/dashboard.html +++ b/bridge/templates/dashboard.html @@ -561,7 +561,7 @@

Latest vision capture

} (function () { if (typeof EventSource === 'undefined') return; - var psrc = new EventSource('/api/perception/feed'); + var psrc = new EventSource('/ui/perception/feed'); psrc.onmessage = function (ev) { var data; try { data = JSON.parse(ev.data); } catch (e) { return; } diff --git a/bridge/templates/say_result.html b/bridge/templates/say_result.html index a5e61f35..d99f54db 100644 --- a/bridge/templates/say_result.html +++ b/bridge/templates/say_result.html @@ -1,10 +1,12 @@ {% if ok %}
-
you said
+
{% if response %}you said{% else %}sent to Dotty{% endif %}
{{ sent }}
+ {% if response %}
Dotty replied
{{ response }}
+ {% endif %}
{% else %} diff --git a/bridge/templates/state.html b/bridge/templates/state.html index be57f940..0724261f 100644 --- a/bridge/templates/state.html +++ b/bridge/templates/state.html @@ -44,21 +44,6 @@ {% endif %} - {% if sound_spark %} - {# #69 — sound-localizer balance over the last 60 s. 0.5 is centred; - a flat line near 1.0 is the stuck-left bias tracked in #27. #} -
- Sound balance - - {{ "%.2f"|format(sound_spark.last) }} -
- {% endif %} -
+ + + + + + + + + + + diff --git a/receiveAudioHandle.py b/receiveAudioHandle.py index 7747b430..d819def2 100644 --- a/receiveAudioHandle.py +++ b/receiveAudioHandle.py @@ -3,7 +3,6 @@ import time import json import asyncio -from difflib import SequenceMatcher from typing import TYPE_CHECKING if TYPE_CHECKING: @@ -76,6 +75,25 @@ def _read_smart_mode_state() -> bool: return False +def _camera_access_denied() -> bool: + """Voice-camera access requires explicit adult policy in the shared file. + + This camera-only reader intentionally fails closed on missing/invalid + state, even when a startup environment flag says adult mode. Re-read on + every access; keep in sync with dotty-behaviour/routes/voice.py. + """ + try: + with open(_KID_MODE_STATE_FILE, "r", encoding="utf-8") as file: + value = file.read().strip().lower() + except (OSError, UnicodeError): + return True + return value not in ("false", "0", "no") + + +class _VoiceCameraDenied(PermissionError): + """Distinct from a stale/missing image; never treat denial as a photo.""" + + def _write_smart_mode_state(enabled: bool) -> None: try: os.makedirs(os.path.dirname(_SMART_MODE_STATE_FILE), exist_ok=True) @@ -120,88 +138,67 @@ def _write_smart_mode_state(enabled: bool) -> None: ) -# ---------- Fuzzy phrase corrections ---------- -# Each entry: (canonical_phrase, minimum_similarity_ratio) -# The canonical phrase is what we want. If the ASR text (or a window of it) -# fuzzy-matches above the threshold, we substitute the canonical form. -# Threshold 0.7 is conservative — avoids false positives on short utterances. -_PHRASE_CORRECTIONS: list[tuple[str, float]] = [ +# ---------- Known-phrase punctuation normalization ---------- +# Never infer intent from edit-distance similarity: correct "one short joke" +# was rewritten into "a story", causing an unintended state transition. +_CANONICAL_PHRASES: tuple[str, ...] = ( # Vision triggers - ("take a photo", 0.7), - ("take a picture", 0.7), - ("take a photo of me", 0.7), - ("take a picture of me", 0.7), + "take a photo", + "take a picture", + "take a photo of me", + "take a picture of me", # Common kid requests - ("tell me a story", 0.7), - ("sing a song", 0.7), - ("sing the macarena", 0.7), - ("dance", 0.8), - ("do the macarena", 0.7), - # Song-name fuzzy hits - ("play tetris", 0.7), - ("hall of the mountain king", 0.7), - ("star wars", 0.75), - ("pirates of the caribbean", 0.7), - ("super mario", 0.7), - ("play music", 0.75), + "tell me a story", + "sing a song", + "sing the macarena", + "dance", + "do the macarena", + # Song names + "play tetris", + "hall of the mountain king", + "star wars", + "pirates of the caribbean", + "super mario", + "play music", # Identity questions - ("what's your name", 0.7), - ("what is your name", 0.7), - ("who are you", 0.75), + "what's your name", + "what is your name", + "who are you", # Greetings - ("good morning", 0.7), - ("good night", 0.7), -] + "good morning", + "good night", +) + +# Paired quotation marks only. Apostrophes inside words (don't / Dotty's) +# are not quote delimiters. Quoted spans remain verbatim for the LLM. +_QUOTED_SPAN_RE = re.compile( + r'"[^"]*"|“[^”]*”|(? str: - """Fuzzy-match ASR text against known phrases and substitute if close enough. - - Uses a sliding window: for each canonical phrase of N words, we check every - contiguous N-word window in the ASR text. If the best window exceeds the - similarity threshold, we replace that window with the canonical phrase. + """Normalize one known phrase's punctuation without changing its words. - Only the single best match (highest ratio) is applied per call to avoid - cascading replacements on short utterances. + Explicit observed name aliases remain in _apply_asr_corrections. Unknown + near-matches stay intact for ordinary conversation instead of manufacturing + camera, music, or state commands. Prefer the longest exact word sequence. """ - lower = text.lower().strip() - words = lower.split() - if len(words) < 2: - return text # too short to fuzzy-match phrases - - best_ratio = 0.0 - best_phrase = "" - best_start = 0 - best_length = 0 - - for canonical, threshold in _PHRASE_CORRECTIONS: + # Do not erase the distinction between a command and a quoted mention. + if _QUOTED_SPAN_RE.search(text): + return text + original_words = text.split() + words = [word.lower().strip(".,!?;:\"'()[]{}‘’“”") for word in original_words] + for canonical in sorted(_CANONICAL_PHRASES, key=lambda p: len(p.split()), reverse=True): canon_words = canonical.split() window_size = len(canon_words) - if window_size > len(words): - continue - for i in range(len(words) - window_size + 1): - window = " ".join(words[i : i + window_size]) - ratio = SequenceMatcher(None, window, canonical).ratio() - if ratio >= threshold and ratio > best_ratio: - best_ratio = ratio - best_phrase = canonical - best_start = i - best_length = window_size - - if best_ratio > 0: - # Rebuild using original-case words outside the match window, - # substituting the canonical phrase for the matched span. - original_words = text.split() - # Map word indices from lower-cased split back to original split. - # They should align since we only called .lower() without changing - # word boundaries, but guard against edge cases. - if len(original_words) >= best_start + best_length: - before = " ".join(original_words[:best_start]) - after = " ".join(original_words[best_start + best_length :]) - parts = [p for p in (before, best_phrase, after) if p] - return " ".join(parts) - + # Punctuation cleanup must not concatenate separate sentences into + # a command, e.g. "Tell me. A story is something I dislike." + if any(word.rstrip("\"'‘’“”)]}").endswith((".", "!", "?")) + for word in original_words[i:i + window_size - 1]): + continue + if words[i:i + window_size] == canon_words: + return " ".join(original_words[:i] + [canonical] + original_words[i + window_size:]) return text @@ -246,17 +243,35 @@ def _repl(m): _NON_CONVERSATIONAL_STATES = ("sleep", "security", "story_time") +def _has_state_command(text: str, phrase: str) -> bool: + """Recognize a literal state phrase, excluding explicit mentions/negation. + + This is deliberately not NLU: paired quotes and immediately preceding + do-not/don't/never (with common politeness/adverb fillers) are protected. + Indirect negation, hypothetical/reporting language, and unmatched quotes + remain outside this guard. Wake and entry shortcuts share the same rule. + """ + lower = _QUOTED_SPAN_RE.sub(" [quoted] ", text.lower()).strip() + for match in re.finditer(r"\b" + re.escape(phrase) + r"\b", lower): + prefix = lower[:match.start()] + negated = re.search( + r"\b(?:do\s+not|don['’]t|never)\b" + r"(?:[\s,]+(?:please|ever|just|you))*[\s,]*$", prefix, + ) + if not negated: + return True + return False + + def _detect_state_phrase(text: str) -> tuple[str, str] | None: - lower = text.lower().strip() for phrase, state, ack in _STATE_TRIGGER_PHRASES: - if phrase in lower: + if _has_state_command(text, phrase): return (state, ack) return None def _is_wake_phrase(text: str) -> bool: - lower = text.lower().strip() - return any(phrase in lower for phrase in _WAKE_PHRASES) + return any(_has_state_command(text, phrase) for phrase in _WAKE_PHRASES) _HELP_PHRASES = ( @@ -395,6 +410,8 @@ def _is_vision_request(text: str) -> bool: async def _handle_vision(conn: "ConnectionHandler", text: str) -> str | None: + if _camera_access_denied(): + raise _VoiceCameraDenied("Camera access is disabled in Kid Mode") if not VISION_BRIDGE_URL: conn.logger.bind(tag=TAG).warning("VISION_BRIDGE_URL not set, skipping vision") return None @@ -1050,11 +1067,15 @@ async def startToChat(conn: "ConnectionHandler", text): try: if text.strip().startswith("{") and text.strip().endswith("}"): data = json.loads(text) - if "speaker" in data and "content" in data: - speaker_name = data["speaker"] - _language_tag = data["language"] - actual_text = data["content"] - conn.logger.bind(tag=TAG).info(f"解析到说话人信息: {speaker_name}") + if isinstance(data, dict) and "content" in data: + # ASR may omit speaker/language metadata. Decode content before + # corrections, noise filtering, or intent routing; otherwise a + # fuzzy replacement can corrupt the JSON envelope itself. + actual_text = data["content"] if isinstance(data["content"], str) else "" + speaker_name = data.get("speaker") + _language_tag = data.get("language") + if speaker_name: + conn.logger.bind(tag=TAG).info(f"解析到说话人信息: {speaker_name}") except (json.JSONDecodeError, KeyError): pass @@ -1097,6 +1118,10 @@ async def startToChat(conn: "ConnectionHandler", text): return await send_stt_message(conn, actual_text) + # DOTTY-PATCH: an abort cancels the turn in flight, not the ones after it. + # Nothing else on the `nointent` path clears this, so a single abort frame + # left conn.chat() breaking out of every later reply until reconnect. + conn.client_abort = False thinking_frame = json.dumps({ "type": "llm", @@ -1162,7 +1187,15 @@ async def startToChat(conn: "ConnectionHandler", text): if _is_vision_request(user_text): conn.logger.bind(tag=TAG).info(f"Vision intent detected: {user_text[:60]}") - description = await _handle_vision(conn, user_text) + try: + description = await _handle_vision(conn, user_text) + except _VoiceCameraDenied: + conn.logger.bind(tag=TAG).info("Voice camera denied by Kid Mode policy") + _submit_chat(conn, + "[CAMERA_DISABLED] Camera access is disabled in Kid Mode. " + "This request did not take a photo or access a camera view. " + "Briefly tell the user that camera access is disabled; do not use tools.") + return if description: vision_prompt = ( f"[You just used your camera and took a photo. " diff --git a/scripts/dotty-av-cases.json b/scripts/dotty-av-cases.json new file mode 100644 index 00000000..3b27d966 --- /dev/null +++ b/scripts/dotty-av-cases.json @@ -0,0 +1,24 @@ +{ + "ai_assistance": "OpenAI Codex (GPT-6), human-directed overnight session", + "cases": [ + {"id":"V01-identity","feature":"warm conversation identity/recovery","prompt":"What is your name? Please answer in one short sentence.","asr":[["name"]],"reply":[["Dotty","Dottie","Dotti"]],"seconds":35,"recovery_policy":"ready_for_followup","soak_safe":true,"max_response_words":40}, + {"id":"V02-repeat","feature":"ASR/constrained reply","prompt":"Repeat these five words: purple, robot, seven, window, Brisbane.","asr":[["purple"],["robot"],["seven","7"],["window"],["Brisbane"]],"reply":[["purple"],["robot"],["seven","7"],["window"],["Brisbane"]],"seconds":35,"recovery_policy":"ready_for_followup","soak_safe":true,"max_response_words":25}, + {"id":"V03-arithmetic","feature":"comprehension/recovery","prompt":"What is twelve plus seven? Just say the answer.","asr":[["twelve","12"],["seven","7"]],"reply":[["nineteen","19"]],"seconds":35,"recovery_policy":"ready_for_followup","soak_safe":true,"max_response_words":10}, + {"id":"V04-tiny-review","feature":"persona/short response","prompt":"Give yourself a very tiny performance review in two funny sentences.","asr":[["performance","review"]],"reply":[],"seconds":45,"recovery_policy":"ready_for_followup","soak_safe":true,"max_response_words":60}, + {"id":"V05-toaster","feature":"creative explanation","prompt":"Explain gravity to a sleepy toaster in two short sentences.","asr":[["gravity"],["toaster"]],"reply":[["gravity","fall","pull","ground"]],"seconds":45,"recovery_policy":"ready_for_followup","soak_safe":true,"max_response_words":60}, + {"id":"V06-robot-joke","feature":"expressive speech","prompt":"Tell me one short joke about being a very small robot.","asr":[["joke"],["robot"]],"reply":[],"seconds":45,"recovery_policy":"ready_for_followup","soak_safe":true,"max_response_words":60}, + {"id":"W01-cold-wake","feature":"cold wake word","prompt":"Hi E S P. What is your name?","asr":[["name"]],"reply":[["Dotty","Dottie"]],"seconds":35,"require_wake_event":true,"recovery_policy":"ready_for_followup","required_passes":10,"enabled":false}, + {"id":"V07-hard-thinking","feature":"think_hard","prompt":"Hi E S P. Think hard: if all wugs are blue and some blue things are round, must some wugs be round? Give one sentence.","asr":[["blue"],["round"]],"reply":[["no","not necessarily","cannot","can't"]],"seconds":110,"enabled":false}, + {"id":"S01-sleep","feature":"sleep entry","prompt":"Hi E S P. Goodnight Dotty. Go to sleep.","asr":[["goodnight","good night","sleep"]],"reply":[],"resting_states":["sleep"],"seconds":50,"enabled":false}, + {"id":"S02-wake","feature":"wake from sleep","prompt":"Hi E S P. Wake up. What is your name?","asr":[["wake"],["name"]],"reply":[["Dotty","Dottie"]],"seconds":65,"enabled":false} + ], + "pending_coverage": [ + {"feature":"physical touch, face tracking/loss, spatial sound localization","status":"BLOCKED","reason":"requires physical stimuli; synthetic events only prove downstream consumers"}, + {"feature":"screen, LEDs, movement and Shorts framing","status":"BLOCKED","reason":"preflight framing crops body and shows screen obliquely"}, + {"feature":"memory/person tools","status":"PENDING","reason":"isolate fixture writes and preserve household records"}, + {"feature":"think_hard and cached take_photo","status":"PENDING","reason":"need explicit tool/result assertions and deployed capability check"}, + {"feature":"states, toggles, songs, reminders, interruption","status":"PENDING","reason":"run after acoustic baseline; per-feature cleanup required"}, + {"feature":"dashboard, enabled ambient consumers, restart recovery","status":"PENDING","reason":"inventory live configuration and establish baseline first"}, + {"feature":"story/security full backing and smart-model swap","status":"PENDING","reason":"reconcile code with live deployment; do not implement missing features"} + ] +} diff --git a/scripts/dotty-av-test.sh b/scripts/dotty-av-test.sh index a6a5a386..4bffad36 100755 --- a/scripts/dotty-av-test.sh +++ b/scripts/dotty-av-test.sh @@ -6,12 +6,42 @@ set -euo pipefail VIDEO_DEVICE="${DOTTY_AV_VIDEO_DEVICE:-/dev/v4l/by-id/usb-046d_HD_Pro_Webcam_C920_7B90DC9F-video-index0}" AUDIO_DEVICE="${DOTTY_AV_AUDIO_DEVICE:-hw:C920,0}" +AUDIO_BACKEND="${DOTTY_AV_AUDIO_BACKEND:-pulse}" +AUDIO_SOURCE="${DOTTY_AV_AUDIO_SOURCE:-alsa_input.usb-046d_HD_Pro_Webcam_C920_7B90DC9F-02.analog-stereo}" SINK="${DOTTY_AV_SINK:-@DEFAULT_SINK@}" VIDEO_SIZE="${DOTTY_AV_VIDEO_SIZE:-1280x720}" VIDEO_FPS="${DOTTY_AV_VIDEO_FPS:-15}" VOLUME="${DOTTY_AV_VOLUME:-20}" RESPONSE_SECONDS="${DOTTY_AV_RESPONSE_SECONDS:-20}" OUT_DIR="${DOTTY_AV_OUT_DIR:-uat-sessions/$(date +%F)/av}" +SCRIPT_DIR="$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")" && pwd)" +MAX_SECONDS="${DOTTY_AV_MAX_SECONDS:-180}" +LOSSLESS_AUDIO="${DOTTY_AV_LOSSLESS_AUDIO:-0}" + +cleanup() { + if [[ -n "${capture_pid:-}" ]]; then + kill -INT "$capture_pid" 2>/dev/null || true + wait "$capture_pid" 2>/dev/null || true + fi + [[ -z "${speech:-}" ]] || rm -f -- "$speech" + [[ -z "${tone:-}" ]] || rm -f -- "$tone" +} +trap cleanup EXIT +trap 'exit 130' INT +trap 'exit 143' TERM + +lock_capture() { + # Shared by all invocations, regardless of output directory. + exec 9>"${XDG_RUNTIME_DIR:-/tmp}/dotty-av-${UID}.lock" + flock -n 9 || { echo 'ERROR: another Dotty capture owns the camera' >&2; exit 1; } +} + +bounded() { + [[ "$1" =~ ^[0-9]+$ && "$MAX_SECONDS" =~ ^[0-9]+$ ]] && + (( $1 > 0 && $1 <= MAX_SECONDS )) || { + echo "ERROR: capture must be 1..${MAX_SECONDS} seconds" >&2; exit 2; + } +} usage() { cat <<'EOF' @@ -25,6 +55,9 @@ Usage: Environment overrides: DOTTY_AV_VIDEO_DEVICE, DOTTY_AV_AUDIO_DEVICE, DOTTY_AV_SINK + DOTTY_AV_AUDIO_BACKEND (pulse default, alsa fallback), DOTTY_AV_AUDIO_SOURCE + DOTTY_AV_PROMPT_WAV (optional pre-rendered local prompt) + DOTTY_AV_LOSSLESS_AUDIO=1 (optional same-capture float32 PCM .capture.wav) DOTTY_AV_VIDEO_SIZE, DOTTY_AV_VIDEO_FPS, DOTTY_AV_VOLUME DOTTY_AV_RESPONSE_SECONDS, DOTTY_AV_OUT_DIR @@ -66,14 +99,49 @@ default_output() { capture_args() { printf '%s\n' \ - -f v4l2 -input_format mjpeg -video_size "$VIDEO_SIZE" -framerate "$VIDEO_FPS" -i "$VIDEO_DEVICE" \ - -f alsa -ac 2 -ar 32000 -i "$AUDIO_DEVICE" \ - -c:v libx264 -preset veryfast -pix_fmt yuv420p -c:a aac -b:a 128k + -thread_queue_size 512 -f v4l2 -input_format mjpeg -video_size "$VIDEO_SIZE" -framerate "$VIDEO_FPS" -i "$VIDEO_DEVICE" + case "$AUDIO_BACKEND" in + alsa) printf '%s\n' -thread_queue_size 512 -f alsa -ac 2 -ar 32000 -i "$AUDIO_DEVICE" ;; + pulse) + printf '%s\n' -thread_queue_size 512 -f pulse -ac 2 -ar 32000 + # Pulse defaults to signed 16-bit. Request float input explicitly + # for this experiment so >1.0 samples survive into the PCM sidecar. + [[ "$LOSSLESS_AUDIO" != 1 ]] || printf '%s\n' -c:a pcm_f32le + printf '%s\n' -i "$AUDIO_SOURCE" + ;; + *) echo "ERROR: unknown audio backend $AUDIO_BACKEND" >&2; exit 2 ;; + esac +} + +check_capture_outputs() { + local output="$1" sidecar="${1%.mp4}.capture.wav" + [[ "$LOSSLESS_AUDIO" == 0 || "$LOSSLESS_AUDIO" == 1 ]] || { + echo 'ERROR: DOTTY_AV_LOSSLESS_AUDIO must be 0 or 1' >&2; exit 2; + } + [[ ! -e "$output" && ! -L "$output" ]] || { + echo "ERROR: output already exists: $output" >&2; exit 1; + } + if [[ "$LOSSLESS_AUDIO" == 1 && ( -e "$sidecar" || -L "$sidecar" ) ]]; then + echo "ERROR: output already exists: $sidecar" >&2; exit 1 + fi +} + +capture_output_args() { + local seconds="$1" output="$2" + printf '%s\n' -map 0:v:0 -map 1:a:0 \ + -c:v libx264 -preset veryfast -pix_fmt yuv420p -c:a aac -b:a 128k -t "$seconds" "$output" + if [[ "$LOSSLESS_AUDIO" == 1 ]]; then + # One input and one process: no second microphone reader. -t is an + # output option and must be repeated to bound the audio-only output. + printf '%s\n' -map 1:a:0 -vn -c:a pcm_f32le -ar 32000 -ac 2 \ + -t "$seconds" "${output%.mp4}.capture.wav" + fi } verify() { local file="$1" [[ -s "$file" ]] || { echo "ERROR: missing or empty recording: $file" >&2; exit 1; } + python "$SCRIPT_DIR/dotty_av_media.py" verify "$file" echo "Streams:" ffprobe -v error \ -show_entries stream=codec_type,codec_name,width,height,r_frame_rate,sample_rate,channels,duration,nb_frames \ @@ -89,10 +157,16 @@ record() { echo "ERROR: duration must be a positive whole number" >&2 exit 2 } - [[ ! -e "$output" ]] || { echo "ERROR: output already exists: $output" >&2; exit 1; } + check_capture_outputs "$output" + bounded "$seconds" + lock_capture mkdir -p "$(dirname "$output")" mapfile -t args < <(capture_args) - ffmpeg -hide_banner -loglevel warning "${args[@]}" -t "$seconds" "$output" + mapfile -t outputs < <(capture_output_args "$seconds" "$output") + ffmpeg -nostdin -n -hide_banner -loglevel warning "${args[@]}" "${outputs[@]}" & + capture_pid=$! + wait "$capture_pid" + capture_pid='' verify "$output" } @@ -122,10 +196,11 @@ speaker-test) need pactl set_volume "$VOLUME" tone="$(mktemp --suffix=.wav)" - trap 'rm -f "$tone"' EXIT ffmpeg -hide_banner -loglevel error -y -f lavfi -i 'sine=frequency=440:duration=0.5' -af 'volume=-18dB' "$tone" echo "Playing a quiet half-second calibration tone at ${VOLUME}%..." - pw-play "$tone" + target_sink="$SINK" + [[ "$target_sink" != '@DEFAULT_SINK@' ]] || target_sink="$(pactl get-default-sink)" + pw-play --target "$target_sink" "$tone" ;; record) need ffmpeg @@ -147,23 +222,38 @@ run) exit 2 } output="${4:-$(default_output response)}" - [[ ! -e "$output" ]] || { echo "ERROR: output already exists: $output" >&2; exit 1; } + check_capture_outputs "$output" + lock_capture mkdir -p "$(dirname "$output")" speech="$(mktemp --suffix=.wav)" - trap 'rm -f "$speech"' EXIT - espeak-ng -v en-au -s 145 -w "$speech" "$prompt" + if [[ -n "${DOTTY_AV_PROMPT_WAV:-}" ]]; then + [[ -s "$DOTTY_AV_PROMPT_WAV" ]] || { echo 'ERROR: missing prompt WAV' >&2; exit 2; } + cp -- "$DOTTY_AV_PROMPT_WAV" "$speech" + else + espeak-ng -v en-au -s 145 -w "$speech" "$prompt" + fi speech_seconds="$(ffprobe -v error -show_entries format=duration -of csv=p=0 "$speech")" total_seconds="$(awk -v speech="$speech_seconds" -v response="$response_seconds" 'BEGIN { printf "%d", speech + response + 3.999 }')" + bounded "$total_seconds" set_volume "$VOLUME" + target_sink="$SINK" + [[ "$target_sink" != '@DEFAULT_SINK@' ]] || target_sink="$(pactl get-default-sink)" + cp -- "$speech" "${output%.mp4}.prompt.wav" mapfile -t args < <(capture_args) - ffmpeg -hide_banner -loglevel warning "${args[@]}" -t "$total_seconds" "$output" & + mapfile -t outputs < <(capture_output_args "$total_seconds" "$output") + # Launch-time anchor; first encoded sample may lag device initialization. + date -u +%FT%T.%NZ > "${output%.mp4}.recording-start.txt" + ffmpeg -nostdin -n -hide_banner -loglevel warning "${args[@]}" "${outputs[@]}" & capture_pid=$! - trap 'kill "$capture_pid" 2>/dev/null || true; rm -f "$speech"' EXIT echo "Recording. Prompt plays in 2 seconds; Dotty then has ${response_seconds}s to answer." sleep 2 - pw-play "$speech" + kill -0 "$capture_pid" 2>/dev/null || { echo 'ERROR: capture failed before playback' >&2; exit 1; } + [[ -s "$output" ]] || { echo 'ERROR: capture produced no file before playback' >&2; exit 1; } + date -u +%FT%T.%NZ > "${output%.mp4}.playback-start.txt" + pw-play --target "$target_sink" "$speech" + date -u +%FT%T.%NZ > "${output%.mp4}.playback-end.txt" wait "$capture_pid" - trap 'rm -f "$speech"' EXIT + capture_pid='' verify "$output" echo "Saved: $output" ;; diff --git a/scripts/dotty_av_clips.py b/scripts/dotty_av_clips.py new file mode 100644 index 00000000..dcb8c2f6 --- /dev/null +++ b/scripts/dotty_av_clips.py @@ -0,0 +1,242 @@ +#!/usr/bin/env python3 +"""Local review exports from Dotty case evidence. AI-assisted: OpenAI Codex (GPT-6).""" + +import argparse +from datetime import datetime, timezone +import hashlib +import html +import json +import math +import os +from pathlib import Path +import re +import subprocess +from urllib.parse import quote + + +GATES = {"visual", "privacy", "captions", "music", "context"} + + +def identifier(value): + if not re.fullmatch(r"[A-Za-z0-9][A-Za-z0-9_.-]{0,159}", value): + raise ValueError("case/export name must be a simple identifier") + return value + + +def inside(path, root): + path = path.resolve() + if not path.is_relative_to(root.resolve()): + raise ValueError("path escapes session directory") + return path + + +def load(path): + return json.loads(path.read_text(encoding="utf-8")) + + +def write_new(path, text): + with path.open("x", encoding="utf-8") as output: + output.write(text) + + +def execute(command): + return subprocess.run(command, check=True, capture_output=True, text=True, timeout=600) + + +def media_duration(source): + data = json.loads(execute(["ffprobe", "-v", "error", "-show_streams", "-show_format", + "-of", "json", str(source)]).stdout) + if not {"video", "audio"}.issubset({s["codec_type"] for s in data["streams"]}): + raise ValueError("source requires video and audio") + duration = float(data["format"]["duration"]) + if not math.isfinite(duration) or duration <= 0: + raise ValueError("invalid source duration") + return duration + + +def stamp(seconds): + milliseconds = round(seconds * 1000) + hours, rest = divmod(milliseconds, 3600000) + minutes, rest = divmod(rest, 60000) + seconds, milliseconds = divmod(rest, 1000) + return f"{hours:02}:{minutes:02}:{seconds:02},{milliseconds:03}" + + +def captions(transcript, offset, start, end): + """Convert response-relative segments to clip-relative SRT, retaining provenance.""" + if not all(math.isfinite(n) for n in (offset, start, end)) or offset < 0 or start < 0 or end <= start: + raise ValueError("invalid caption timing") + cues = [] + for segment in transcript.get("segments", []): + source_start = float(segment["start"]) + offset + source_end = float(segment["end"]) + offset + if not all(math.isfinite(n) for n in (source_start, source_end)) or source_end <= source_start: + raise ValueError("invalid transcript segment timing") + begin, finish = max(start, source_start), min(end, source_end) + text = " ".join(str(segment["text"]).split()) + if finish > begin and text: + cues.append({"start": begin - start, "end": finish - start, "text": text, + "source_start": source_start, "source_end": source_end}) + srt = "\n\n".join(f"{i}\n{stamp(c['start'])} --> {stamp(c['end'])}\n{c['text']}" + for i, c in enumerate(cues, 1)) + return srt + ("\n" if srt else ""), cues + + +def file_digest(path): + with path.open("rb") as source: + return hashlib.file_digest(source, "sha256").hexdigest() + + +def url(path, base): + return quote(os.path.relpath(path, base), safe="/") + + +def render_index(manifests, output): + rows = [] + for path in manifests: + record = load(path) + folder = path.parent + evidence = record["source"] + root = Path(record["session"]) + source = root / evidence["file"] + result = root / evidence["result_file"] + links = [("Video", folder / "clean.mp4"), ("Thumbnail", folder / "thumbnail.jpg"), + ("Source", source), ("Case evidence", result), ("Provenance", path)] + if (folder / "captions.srt").exists(): + links.append(("SRT captions", folder / "captions.srt")) + anchor = " · ".join(f'{label}' for label, p in links) + rows.append("" + "".join(f"{html.escape(str(value))}" for value in [ + record["title"], record["case_id"], record["verdict"], record["status"], + ", ".join(record["pending_reviews"]) or "all gates acknowledged", + f'{evidence["start_seconds"]:.3f}–{evidence["end_seconds"]:.3f}s', + record["commit"], record["export_state"]]) + f"{anchor}") + page = ('Dotty local clip review' + '' + '

Dotty local clip review

AI-assisted exports by OpenAI Codex (GPT-6). ' + 'Human review required before posting. Test verdicts are separate from clip readiness. ' + 'Captions are provisional unless their review gate is acknowledged. Source recordings ' + 'retain the full exchange; trimming and audio normalization apply only to exports.

' + '' + '' + + "".join(rows) + '
Title suggestionCaseVerdictReview groupPending reviewsSource intervalCommitExportFiles
') + write_new(output, page) + return output + + +def crop_region(value): + """`W:H:X:Y` in source pixels. Portrait only: the region is scaled to fill + 1080x1920, so a landscape region would be stretched.""" + if not re.fullmatch(r"\d+:\d+:\d+:\d+", value or ""): + raise ValueError("crop must be W:H:X:Y in source pixels") + width, height, _x, _y = (int(part) for part in value.split(":")) + if width < 2 or height < 2 or abs(width / height - 9 / 16) > .02: + raise ValueError("crop region must be 9:16 portrait") + return value + + +def export(session, case_id, *, start=0, end=None, name=None, title=None, reviewed=(), + include_captions=True, crop=None): + session = session.resolve(strict=True) + case = inside(session / "cases" / identifier(case_id), session / "cases") + source = inside(case / "raw.mp4", session) + result_path = inside(case / "result.json", session) + result = load(result_path) + duration = media_duration(source) + end = duration if end is None else end + if not all(math.isfinite(n) for n in (start, end)) or not 0 <= start < end <= duration: + raise ValueError("clip interval must be within source duration") + if crop is not None: + crop = crop_region(crop) + acknowledged = set(reviewed) + if acknowledged - GATES: + raise ValueError("unknown review gate") + pending = sorted(GATES - acknowledged) + verdict = str(result.get("verdict", "INCONCLUSIVE")) + status = "failures" if verdict == "FAIL" else "needs-review" if pending else "ready-for-review" + name = identifier(name or case_id + "-" + datetime.now(timezone.utc).strftime("%H%M%S%f")) + destination = inside(session / "clips" / status / name, session) + # Exclusive directory reservation makes every export immutable, including failed attempts. + destination.mkdir(parents=True, exist_ok=False) + record = {"ai_assistance": "OpenAI Codex (GPT-6)", "session": str(session), + "case_id": case_id, "feature": result.get("case_id", case_id), + "title": title or result.get("case", {}).get("prompt", case_id), + "verdict": verdict, "status": status, "pending_reviews": pending, + "reviewed": sorted(acknowledged), "commit": result.get("commit", "unknown"), + "created": datetime.now(timezone.utc).isoformat(), "export_state": "incomplete", + "source": {"file": str(source.relative_to(session)), + "result_file": str(result_path.relative_to(session)), + "sha256": file_digest(source), "start_seconds": start, "end_seconds": end, + "case_started": result.get("started"), "duration_seconds": duration}, + "layout": (f"1080x1920 cropped from source region {crop}" if crop + else "1080x1920 fitted; entire source frame retained"), + "audio": "original audio with loudness normalization; no replacement soundtrack", + "captions": {"state": "absent", "source": "local response.json transcription"}} + try: + execute(["ffmpeg", "-nostdin", "-n", "-v", "error", "-ss", str(start), "-i", str(source), + "-t", str(end - start), "-map", "0:v:0", "-map", "0:a:0", "-vf", + (f"crop={crop},scale=1080:1920:flags=lanczos,setsar=1" if crop else + "scale=1080:1920:force_original_aspect_ratio=decrease:force_divisible_by=2," + "pad=1080:1920:(ow-iw)/2:(oh-ih)/2,setsar=1"), + "-af", "loudnorm=I=-16:TP=-1.5:LRA=11", + "-c:v", "libx264", "-preset", "fast", "-crf", "20", "-pix_fmt", "yuv420p", + "-c:a", "aac", "-b:a", "160k", "-movflags", "+faststart", str(destination / "clean.mp4")]) + execute(["ffmpeg", "-nostdin", "-n", "-v", "error", "-ss", str(min(2, (end-start)/2)), + "-i", str(destination / "clean.mp4"), "-frames:v", "1", str(destination / "thumbnail.jpg")]) + transcript = inside(case / "response.json", session) + if include_captions and transcript.exists(): + if "response_offset" not in result: + record["captions"]["state"] = "blocked: missing response_offset; cannot align captions" + else: + offset = float(result["response_offset"]) + srt, cues = captions(load(transcript), offset, start, end) + write_new(destination / "captions.srt", srt) + record["captions"].update(state="reviewed" if "captions" in acknowledged else "provisional", + response_offset=offset, cues=cues) + record["export_state"] = "complete" + except Exception as exc: + record["error"] = str(exc) + raise + finally: + write_new(destination / "manifest.json", json.dumps(record, ensure_ascii=False, indent=2) + "\n") + render_index([destination / "manifest.json"], destination / "index.html") + return destination + + +def index(session): + session = session.resolve(strict=True) + clips = inside(session / "clips", session) + clips.mkdir(exist_ok=True) + manifests = [p for group in ("ready-for-review", "needs-review", "failures") + for p in sorted(clips.glob(f"{group}/*/manifest.json"))] + output = clips / ("index-" + datetime.now(timezone.utc).strftime("%Y%m%dT%H%M%S%f") + ".html") + return render_index(manifests, output) + + +def main(): + parser = argparse.ArgumentParser(description=__doc__) + sub = parser.add_subparsers(dest="command", required=True) + p = sub.add_parser("export") + p.add_argument("session", type=Path) + p.add_argument("case_id") + p.add_argument("--start", type=float, default=0) + p.add_argument("--end", type=float) + p.add_argument("--name") + p.add_argument("--title") + p.add_argument("--reviewed", choices=sorted(GATES), action="append", default=[]) + p.add_argument("--no-captions", action="store_true") + p.add_argument("--crop", help="W:H:X:Y portrait source region to fill the 9:16 frame " + "(default: fit the whole frame with padding)") + p = sub.add_parser("index") + p.add_argument("session", type=Path) + args = parser.parse_args() + if args.command == "index": + print(index(args.session)) + else: + print(export(args.session, args.case_id, start=args.start, end=args.end, name=args.name, + title=args.title, reviewed=args.reviewed, include_captions=not args.no_captions, + crop=args.crop)) + + +if __name__ == "__main__": + main() diff --git a/scripts/dotty_av_media.py b/scripts/dotty_av_media.py new file mode 100644 index 00000000..9c4ad661 --- /dev/null +++ b/scripts/dotty_av_media.py @@ -0,0 +1,229 @@ +#!/usr/bin/env python3 +"""Local A/V evidence helpers. AI-assisted: OpenAI Codex (GPT-6). + +Verification never infers a robot response from a nonempty audio track. +""" +import argparse +from array import array +import json +import math +from pathlib import Path +import subprocess +import sys + + +def probe(path): + return json.loads(subprocess.check_output([ + "ffprobe", "-v", "error", "-show_streams", "-show_format", "-of", "json", str(path) + ], text=True, timeout=20)) + + +def verify(path, expected=None): + data = probe(path) + streams = data.get("streams", []) + video = [s for s in streams if s.get("codec_type") == "video"] + audio = [s for s in streams if s.get("codec_type") == "audio"] + if len(video) != 1 or len(audio) != 1: + raise ValueError("expected exactly one video and one audio stream") + vd, ad = [float(s.get("duration", 0)) for s in (video[0], audio[0])] + if not all(math.isfinite(d) and d > 0 for d in (vd, ad)): + raise ValueError("empty/nonfinite stream duration") + if abs(vd - ad) > 1: + raise ValueError("audio/video duration mismatch") + if expected is not None and min(vd, ad) < expected - 1: + raise ValueError("truncated capture") + continuity = audio_continuity(path, int(audio[0]["sample_rate"])) + if continuity["decoded_seconds"] < ad - max(.25, ad * .02): + raise ValueError(f"audio sample loss: {continuity['decoded_seconds']:.3f}s decoded across {ad:.3f}s timeline") + if continuity["max_gap_seconds"] > .15: + raise ValueError(f"audio discontinuity: {continuity['max_gap_seconds']:.3f}s gap") + return {"video_seconds": vd, "audio_seconds": ad, + "width": video[0]["width"], "height": video[0]["height"], + "sample_rate": audio[0]["sample_rate"], "channels": audio[0]["channels"], + **continuity, "audio_quality": audio_metrics(path)} + + +def pcm_metrics(pcm): + """Measure decoded float32 PCM without inferring speech or interaction success. + + Near-full-scale samples include lossy-decoder overshoots; this is a clipping + warning, not proof that the microphone's analogue input clipped. + """ + if len(pcm) % 4: + raise ValueError("incomplete float32 PCM sample") + samples = array("f") + samples.frombytes(pcm) + if sys.byteorder != "little": + samples.byteswap() + if not samples: + raise ValueError("empty decoded PCM") + peak = squared_sum = clipped = 0 + for sample in samples: + if not math.isfinite(sample): + raise ValueError("nonfinite decoded PCM sample") + magnitude = abs(sample) + peak = max(peak, magnitude) + squared_sum += sample * sample + clipped += magnitude >= .999 + rms = math.sqrt(squared_sum / len(samples)) + ratio = clipped / len(samples) + silent = rms < .001 and peak < 10 ** (-45 / 20) + warnings = [] + if ratio > .001: + warnings.append("severe_clipping: over 0.1% of decoded samples near or above full scale") + if silent: + warnings.append("near_silence: RMS below -60 dBFS and peak below -45 dBFS") + return {"pcm_samples": len(samples), "rms": rms, "peak": peak, + "rms_dbfs": 20 * math.log10(rms) if rms else None, + "peak_dbfs": 20 * math.log10(peak) if peak else None, + "clipped_samples": clipped, "clipping_ratio": ratio, + "clipping_threshold": .999, "near_silence": silent, "warnings": warnings, + "note": "Decoded PCM only; signal presence does not verify speech or a robot response. " + "Lossy decoding can produce full-scale overshoots."} + + +def audio_metrics(path, start=0, duration=180): + # Preserve native channels/rate: downmixing or resampling can mask saturation. + # Bound decoding to the maximum permitted case duration; never play audio. + if (any(isinstance(v, bool) or not isinstance(v, (int, float)) or not math.isfinite(v) + for v in (start, duration)) or start < 0 or duration <= 0 or start + duration > 180): + raise ValueError("audio metric window must lie within 0..180 seconds") + pcm = subprocess.check_output([ + "ffmpeg", "-nostdin", "-v", "error", "-i", str(path), "-ss", str(start), "-t", str(duration), + "-map", "0:a:0", "-vn", "-sn", "-dn", "-c:a", "pcm_f32le", "-f", "f32le", "pipe:1" + ], timeout=45) + return {**pcm_metrics(pcm), "window_start_seconds": start, + "analysis_limit_seconds": duration, "channel_policy": "native; no downmix or resampling"} + + +def audio_continuity(path, rate): + frames = json.loads(subprocess.check_output([ + "ffprobe", "-v", "error", "-select_streams", "a:0", "-show_frames", + "-show_entries", "frame=pts_time,nb_samples", "-of", "json", str(path) + ], text=True, timeout=30)).get("frames", []) + samples, end, gap = 0, None, 0.0 + for frame in frames: + count = int(frame.get("nb_samples", 0)) + start = float(frame["pts_time"]) + if end is not None: + gap = max(gap, start - end) + end = start + count / rate + samples += count + return {"decoded_samples": samples, "decoded_seconds": samples / rate, "max_gap_seconds": gap} + + +def analysis_gain(quality): + """Raise quiet analysis audio towards -20 dBFS RMS, preserving 1 dB headroom.""" + if quality["near_silence"] or not quality["rms"] or not quality["peak"]: + return 0.0 + return max(0.0, min(20.0, -20.0 - quality["rms_dbfs"], -1.0 - quality["peak_dbfs"])) + + +def prepare_analysis(path, output): + """Persist a separate normalized derivative; originals are never rewritten.""" + derivative = Path(output).with_suffix(".analysis.wav") + if derivative.exists() or derivative.resolve() == Path(path).resolve(): + raise FileExistsError("analysis derivative already exists or matches input") + # Measure the exact mono/16k signal supplied to Whisper before choosing gain. + # No normalization, VAD or transcript-derived conditioning is applied here. + pcm = subprocess.check_output([ + "ffmpeg", "-nostdin", "-v", "error", "-i", str(path), "-t", "180", + "-map", "0:a:0", "-vn", "-ac", "1", "-ar", "16000", + "-c:a", "pcm_f32le", "-f", "f32le", "pipe:1" + ], timeout=45) + before = pcm_metrics(pcm) + gain = analysis_gain(before) + samples = array("f") + samples.frombytes(pcm) + if sys.byteorder != "little": + samples.byteswap() + multiplier = 10 ** (gain / 20) + normalized = array("f", (sample * multiplier for sample in samples)) + if sys.byteorder != "little": + normalized.byteswap() + rendered = normalized.tobytes() + after = pcm_metrics(rendered) + subprocess.run([ + "ffmpeg", "-nostdin", "-n", "-v", "error", "-f", "f32le", "-ar", "16000", + "-ac", "1", "-i", "pipe:0", "-c:a", "pcm_f32le", str(derivative) + ], input=rendered, check=True, capture_output=True, timeout=45) + return derivative, {"enabled": True, "source": str(Path(path).resolve()), + "derivative": str(derivative.resolve()), "gain_db": gain, + "maximum_gain_db": 20, "target_rms_dbfs": -20, "headroom_db": 1, + "filter": "mono 16000 Hz float32 decode, constant amplitude gain; no limiter", + "original_quality": before, "analysis_quality": after, + "analysis_limit_seconds": 180} + + +def validate_vad_threshold(value): + try: + threshold = float(value) + except (TypeError, ValueError): + raise ValueError("VAD threshold must be finite and between 0 and 1") from None + if isinstance(value, bool) or not math.isfinite(threshold) or not 0 <= threshold <= 1: + raise ValueError("VAD threshold must be finite and between 0 and 1") + return threshold + + +def transcribe(path, model_path, output, normalize=False, vad_filter=True, vad_threshold=.5): + # Lower thresholds are explicit prospective experiments, never inferred from + # expected words or used to reclassify existing case evidence. + vad_parameters = {"threshold": validate_vad_threshold(vad_threshold)} + if Path(output).exists() or Path(output).resolve() == Path(path).resolve(): + raise FileExistsError("transcription output already exists or matches input") + analysis_path, analysis = (path, {"enabled": False, "source": str(Path(path).resolve())}) + if normalize: + analysis_path, analysis = prepare_analysis(path, output) + if normalize and analysis["original_quality"]["near_silence"]: + result = {"model": str(model_path), "language_probability": None, "segments": [], + "analysis": analysis, "vad_filter": vad_filter, "vad_parameters": vad_parameters, + "evidence_status": "INCONCLUSIVE", "reason": "original_audio_near_silent"} + with Path(output).open("x") as file: + json.dump(result, file, indent=2) + file.write("\n") + return result + # Explicit disk path + local_files_only prevent network model/media access. + from faster_whisper import WhisperModel + model = WhisperModel(str(model_path), device="cpu", compute_type="int8", + cpu_threads=4, local_files_only=True) + segments, info = model.transcribe(str(analysis_path), language="en", beam_size=3, + vad_filter=vad_filter, vad_parameters=vad_parameters, + condition_on_previous_text=False, + word_timestamps=True) + result = {"model": str(model_path), "language_probability": info.language_probability, + "analysis": analysis, "vad_filter": vad_filter, "vad_parameters": vad_parameters, + "evidence_status": "UNVERIFIED_TRANSCRIPT", + "note": "Transcription is evidence only, including after gain. Confirm response timing, " + "service TTS and content independently; noise can hallucinate words.", + "segments": [{"start": s.start, "end": s.end, "text": s.text, + "no_speech_prob": s.no_speech_prob, "avg_logprob": s.avg_logprob, + "words": [{"start": w.start, "end": w.end, "word": w.word, + "probability": w.probability} for w in (s.words or [])]} + for s in segments]} + with Path(output).open("x") as file: + json.dump(result, file, indent=2) + file.write("\n") + return result + + +if __name__ == "__main__": + parser = argparse.ArgumentParser(description=__doc__) + sub = parser.add_subparsers(dest="command", required=True) + p = sub.add_parser("verify") + p.add_argument("file", type=Path) + p.add_argument("--expected", type=float) + p = sub.add_parser("transcribe") + p.add_argument("file", type=Path) + p.add_argument("--model", required=True, type=Path) + p.add_argument("--output", required=True, type=Path) + p.add_argument("--normalize", action="store_true", help="create bounded-gain analysis WAV; preserve original") + p.add_argument("--no-vad", action="store_true", help="diagnostic transcription only; increases hallucination risk") + p.add_argument("--vad-threshold", type=validate_vad_threshold, default=.5, + help="explicit VAD speech threshold 0..1 (default: 0.5); recorded in provenance") + args = parser.parse_args() + if args.command == "verify": + print(json.dumps(verify(args.file, args.expected))) + else: + print(json.dumps(transcribe(args.file, args.model, args.output, + normalize=args.normalize, vad_filter=not args.no_vad, + vad_threshold=args.vad_threshold))) diff --git a/scripts/dotty_overnight.py b/scripts/dotty_overnight.py new file mode 100644 index 00000000..2370b153 --- /dev/null +++ b/scripts/dotty_overnight.py @@ -0,0 +1,871 @@ +#!/usr/bin/env python3 +"""Bounded overnight physical Dotty tests. AI-assisted: OpenAI Codex (GPT-6). + +This runner does not edit, deploy, restart or flash product code. The supervising +agent owns controlled repair experiments. All evidence and transcription stay local. +""" +import argparse +import contextlib +from datetime import datetime, timedelta, timezone +import fcntl +import hashlib +import html +import json +import math +import os +from pathlib import Path +import random +import re +import shlex +import shutil +import signal +import subprocess +import time +import uuid +from zoneinfo import ZoneInfo + +from dotty_av_media import audio_metrics, verify + +ROOT = Path(__file__).resolve().parents[1] +SERVICES = ("xiaozhi-esp32-server", "dotty-behaviour", "dotty-bridge", "dotty-pi") +TZ = ZoneInfo("Australia/Brisbane") +STOP = False + + +def now(): + return datetime.now(timezone.utc).isoformat() + + +def write_json(path, data): + path = Path(path) + path.parent.mkdir(parents=True, exist_ok=True) + temporary = path.with_name(path.name + ".tmp") + with temporary.open("w") as f: + json.dump(data, f, indent=2, ensure_ascii=False) + f.write("\n") + f.flush() + os.fsync(f.fileno()) + temporary.replace(path) + + +def journal(session, event): + with (session / "events.jsonl").open("a") as f: + f.write(json.dumps({"at": now(), **event}, ensure_ascii=False) + "\n") + f.flush() + os.fsync(f.fileno()) + + +def load(path, default=None): + return json.loads(path.read_text()) if path.exists() else default + + +def command(args, timeout=30, **kwargs): + return subprocess.run([str(a) for a in args], capture_output=True, text=True, + timeout=timeout, check=True, **kwargs).stdout + + +def ssh(host, argv, timeout=30, **kwargs): + return command(["ssh", "-o", "BatchMode=yes", "-o", "ConnectTimeout=5", host, + shlex.join([str(a) for a in argv])], timeout=timeout, **kwargs) + + +def admin(host, route, body=None): + if route not in {"devices", "songs", "abort", "set-state", "set-toggle", + "set-head-angles", "take-photo", "play-asset", "say", "inject-text"}: + raise ValueError("unsupported admin route") + # Token stays in the container: never in argv, logs, evidence or shell expansion. + code = '''import json,os,sys,urllib.request +route, body = json.loads(sys.argv[1]) +headers={"Content-Type":"application/json", "X-Admin-Token":os.environ.get("DOTTY_ADMIN_TOKEN", "")} +req=urllib.request.Request("http://127.0.0.1:8003/xiaozhi/admin/"+route, + data=None if body is None else json.dumps(body).encode(), headers=headers) +with urllib.request.urlopen(req, timeout=10) as r: print(r.read().decode()) +''' + return json.loads(ssh(host, ["docker", "exec", "-i", "xiaozhi-esp32-server", + "python", "-", json.dumps([route, body])], input=code)) + + +def snapshot(host): + result = {"at": now(), "errors": {}} + for name, url in { + "bridge": "http://localhost:8081/health", + "behaviour": "http://localhost:8090/health", + "perception": "http://localhost:8090/api/perception/state", + }.items(): + try: + result[name] = json.loads(ssh(host, ["curl", "-fsS", "--max-time", "5", url])) + except Exception as exc: + result["errors"][name] = type(exc).__name__ + try: + result["devices"] = admin(host, "devices")["devices"] + raw = ssh(host, ["docker", "inspect", "--format", + '{{json .Name}} {{json .State.Running}} {{json .RestartCount}} {{json .Image}}', + *SERVICES]) + result["containers"] = raw.strip().splitlines() + result["services_running"] = len(result["containers"]) == 4 and all( + ' true ' in line for line in result["containers"]) + except Exception as exc: + result["errors"]["containers"] = type(exc).__name__ + return result + + +def normalize(text): + return " ".join(re.findall(r"[a-z0-9]+", text.lower())) + + +def matches(text, groups): + normalized = normalize(text) + # Token boundaries matter: "no" must not pass merely because "know" occurs. + return all(any(re.search(r"(?= 0 and event_at >= chat_at and age >= 0 else None + + +def wake_events(log_text): + events = [] + for line in log_text.splitlines(): + # Confirmed deployed textMessageProcessor receive logger. A user saying + # "wake_word_detected", ASR JSON, or an admin state change cannot count. + match = re.search(r"\[core\.handle\.textMessageProcessor\]-INFO-收到(listen|event)消息[::]\s*(\{.*)", line) + if not match: + continue + try: + frame, _ = json.JSONDecoder().raw_decode(match[2]) + except (ValueError, TypeError): + continue + if frame.get("type") != match[1]: + continue + if ((frame.get("type") == "listen" and frame.get("state") == "detect") + or (frame.get("type") == "event" and frame.get("name") == "wake_word_detected")): + events.append(frame) + return events + + +def evaluate(case, transcript, log_text, playback_end, after, device, before=None): + # Only response-window words count. Prompt and response transcribed separately + # to prevent Whisper concatenating both speakers into a single segment. + segments = transcript.get("segments", []) + response = " ".join(s["text"].strip() for s in segments + if s.get("end", 0) > s.get("start", 0) + and s.get("no_speech_prob", 1) < .8 + and s.get("avg_logprob", -10) > -1.2) + asr = re.findall(r"结果: (.*)", log_text) + # 识别文本 is logged only for recognitions handed to the chat path, so + # empty and ASR-REJECTed results never appear here. + recognised = [item.strip() for item in re.findall(r"识别文本: (.*)", log_text) if item.strip()] + service_tts = " ".join(text.strip() for text in re.findall( + r"发送音频消息: SentenceType\.(?:FIRST|MIDDLE), (.*)", log_text) if text.strip() != "None") + state = after.get("perception", {}).get(device, {}) + tts = "SentenceType.FIRST" in log_text or "发送第一段语音:" in log_text + tts_edges = re.findall(r"SentenceType\.(FIRST|LAST)\b|(发送第一段语音:)", log_text) + tts_completed = bool(tts_edges) and tts_edges[-1][0] == "LAST" + # One recognition must contain the request; unrelated log lines cannot pool + # their words into evidence that Dotty understood this prompt. + asr_ok = any(matches(item, case.get("asr", [])) for item in asr) + response_ok = bool(response) and matches(response, case.get("reply", [])) + # The reference microphone can mishear a correct reply ("Dadi" for Dotty). + # Service text is not acoustic proof, so agreement there cannot pass the + # case, but it does stop a transcription slip being scored as a robot fault. + service_ok = bool(service_tts) and bool(case.get("reply")) and matches(service_tts, case["reply"]) + matched = next((item for item in recognised if matches(item, case.get("asr", []))), None) + # The first recognition in the capture is the prompt, whether or not it was + # heard correctly; a mishearing is an asr_mismatch, not someone else talking. + extra_speech = list(recognised) + if matched in extra_speech: + extra_speech.remove(matched) + elif extra_speech: + extra_speech.pop(0) + word_count = len(re.findall(r"\b\w+(?:['’]\w+)*\b", response)) + word_limit = case.get("max_response_words") + if word_limit is not None and (type(word_limit) is not int or word_limit < 1): + raise ValueError("max_response_words must be a positive integer") + within_limit = word_limit is None or word_count <= word_limit + expected_tool = case.get("expected_tool") + if expected_tool is not None and (not isinstance(expected_tool, str) or not expected_tool.strip()): + raise ValueError("expected_tool must name one tool") + # This concrete PiClient marker records model tool-call construction. It + # cannot prove execution or a successful result, and a name mentioned in + # an ASR transcript or assistant answer must not satisfy it. + tool_calls = re.findall(r"PiClient: tool call name=([a-z_][a-z0-9_]*) id=([^\s]+)", log_text) + observed_tools = [name for name, ident in tool_calls if ident != "unknown"] + tool_observed = expected_tool in observed_tools if expected_tool else None + observed_wakes = wake_events(log_text) + requires_wake = case.get("require_wake_event", False) + wake_ready = wake_precondition(before, device) if requires_wake else None + recovered = (not after.get("errors") and after.get("services_running", False) + and not state.get("sensor_stale", True) + and state.get("current_state") in case.get("resting_states", ["idle"]) + and state.get("listening") is False) + if case.get("recovery_policy") == "ready_for_followup": + # The deployed voice path intentionally leaves automatic listening open. + # Count readiness only with an ended TTS turn and a fresh live state. + recovered = (not after.get("errors") and after.get("services_running", False) + and not state.get("sensor_stale", True) + and tts_completed + and ((state.get("current_state") == "idle" and state.get("listening") is False) + or (state.get("current_state") == "talk" and state.get("listening") is True))) + failure = None + if requires_wake and not wake_ready: + failure = "wake_precondition" + elif requires_wake and not observed_wakes: + failure = "no_wake_event" + elif not asr: + failure = "no_asr_or_no_wake" + elif not asr_ok: + failure = "asr_mismatch" + elif not tts: + failure = "no_tts" + elif not response: + # VAD/transcription can miss quiet physical replies despite completed + # TTS. Missing words alone cannot establish acoustic silence. + failure = "response_transcription_unavailable" + elif not response_ok: + failure = "response_transcript_disagrees_with_tts" if service_ok else "response_mismatch" + elif not within_limit: + failure = "response_too_long" + elif not recovered: + failure = "not_ready_for_followup" if case.get("recovery_policy") == "ready_for_followup" else "no_idle_recovery" + if extra_speech and failure not in ("wake_precondition", "no_wake_event"): + # Someone else spoke (or the robot answered room noise) inside the + # capture; neither a pass nor a failure can be attributed to the prompt. + failure = "extra_speech_in_capture" + interaction = "FAIL" if failure else "PASS" + verdict = "FAIL" if failure else "INCONCLUSIVE" + if failure == "wake_precondition": + interaction = verdict = "BLOCKED" + elif failure in ("response_transcription_unavailable", "response_transcript_disagrees_with_tts", + "extra_speech_in_capture"): + interaction = verdict = "INCONCLUSIVE" + if failure is None and expected_tool and not tool_observed: + # Absence of a marker may mean the deployed provider lacks this logger; + # do not turn uncertain instrumentation into a product failure or pass. + interaction = "INCONCLUSIVE" + return {"asr": asr, "asr_match": asr_ok, "response_transcript": response, + "response_match": response_ok, "tts_evidence": tts, "tts_completed": tts_completed, + "service_tts_text": service_tts, "service_tts_match": service_ok, + "extra_speech": extra_speech, + "response_word_count": word_count, "max_response_words": word_limit, + "response_length": "PASS" if within_limit else "FAIL", + "semantic_quality": "INCONCLUSIVE", + "semantic_note": "Keyword/length checks do not establish creative quality or complete semantic correctness.", + "observed_tool_calls": observed_tools, + "wake_precondition": "PASS" if wake_ready else "BLOCKED" if requires_wake else "NOT_REQUIRED", + "observed_wake_events": observed_wakes, + "wake_event": "PASS" if observed_wakes else "FAIL" if requires_wake else "NOT_REQUIRED", + "expected_tool_call": "PASS" if tool_observed else "INCONCLUSIVE", + "tool_execution": "INCONCLUSIVE", + "tool_note": "Expected call unverified; logger availability and execution/result require review." + if expected_tool and not tool_observed else "Call markers alone do not prove tool execution/result.", + "expression": "INCONCLUSIVE", + "recovery": "PASS" if recovered else "FAIL", "failure": failure, + "interaction": interaction, + "visual": "INCONCLUSIVE", "verdict": verdict, + "note": "Visual assertion requires frame review; automated acoustic checks are provisional."} + + +def playback_verdict(heard, asr_ok, groups): + """The prompt was audible if the reference microphone transcript contains + it, or if the robot's own recogniser did. The second is the stronger + evidence: the reference transcript can mishear a prompt the robot got right.""" + if heard.strip() and matches(heard, groups): + return "PASS", "reference_mic" + if asr_ok: + return "PASS", "robot_asr" + return "INCONCLUSIVE", None + + +MIC_OPEN_POLLS = 90 +MIC_SETTLE_SECONDS = (4, 14) + + +def warm_listening(snap, device, max_age=30, min_age=0): + dev = (snap or {}).get("perception", {}).get(device, {}) + age = listening_status_age(dev) + fresh = (dev.get("listening") is True and not dev.get("sensor_stale", True) + and age is not None and min_age <= age <= max_age) + return fresh, age + + +def open_mic(config): + """Opt-in (`mic_opener: admin_say`): make the robot say one word through the + admin route; its firmware then opens the microphone. Wait until the mic has + been open a few seconds so a transcript of the trailing silence cannot + collide with the prompt. Returns the settled snapshot, or None.""" + admin(config["host"], "say", {"device_id": config["device"], "text": "Ready."}) + low, high = MIC_SETTLE_SECONDS + for _ in range(MIC_OPEN_POLLS): + snap = snapshot(config["host"]) + if warm_listening(snap, config["device"], max_age=high, min_age=low)[0]: + return snap + time.sleep(.5) + return None + + +def space_ok(path): + usage = shutil.disk_usage(path) + return usage.free >= max(10 * 1024**3, usage.total * .1) + + +def guarded_process(argv, logfile, session, timeout, env=None, deadline=None): + if deadline is not None and time.time() >= deadline: + raise TimeoutError("session deadline reached before process launch") + with logfile.open("w") as log: + p = subprocess.Popen([str(a) for a in argv], stdout=log, stderr=subprocess.STDOUT, + env=env, start_new_session=True) + started = time.monotonic() + heartbeat = started + try: + while p.poll() is None: + if STOP or (session / "STOP").exists(): + raise InterruptedError("operator stopped session") + if time.monotonic() - started > timeout: + raise TimeoutError(f"process exceeded {timeout}s") + if deadline is not None and time.time() >= deadline: + raise TimeoutError("session deadline reached") + if not space_ok(session): + raise RuntimeError("disk pressure") + if time.monotonic() - heartbeat >= 5: + status = load(session / "status.json", {}) + write_json(session / "status.json", {**status, "heartbeat": now(), + "stage": logfile.stem, "pid": os.getpid()}) + heartbeat = time.monotonic() + time.sleep(.25) + if p.returncode: + raise RuntimeError(f"process exit {p.returncode}; see {logfile.name}") + finally: + if p.poll() is None: + os.killpg(p.pid, signal.SIGTERM) + try: + p.wait(timeout=8) + except subprocess.TimeoutExpired: + os.killpg(p.pid, signal.SIGKILL) + p.wait() + + +def read_log_window(host, start, end, directory): + errors = [] + for service in SERVICES: + # Bounded history. Separate stdout/stderr merged by docker logs on remote. + result = subprocess.run(["ssh", "-o", "BatchMode=yes", "-o", "ConnectTimeout=5", host, + shlex.join(["docker", "logs", "--timestamps", "--since", start, + "--until", end, "--tail", "3000", service])], + capture_output=True, text=True, timeout=20) + (directory / f"{service}.log").write_text(result.stdout + result.stderr) + if result.returncode: + errors.append(service) + if errors: + raise RuntimeError("log_collection_failed:" + ",".join(errors)) + + +def playback_window(directory, media_seconds): + def stamp(name): + return datetime.fromisoformat((directory / name).read_text().strip().replace("Z", "+00:00")) + launched = stamp("raw.recording-start.txt") + started = stamp("raw.playback-start.txt") + ended = stamp("raw.playback-end.txt") + start = (started - launched).total_seconds() + end = (ended - launched).total_seconds() + if not 0 <= start < end < end + .4 < media_seconds: + raise ValueError("invalid_or_truncated_playback_window") + # Launch precedes the first camera sample, so this response cut is late, + # never early. A small tail margin excludes speaker reverberation. + return start, end + .4 + + +def capture_window_metrics(media, media_seconds, prompt_end, response_start): + """Keep short loud prompts visible instead of diluting them across silence. + + Measure the native channel/rate signal, not the mono transcription WAV. + This evidence does not alter interaction verdicts or claim ADC clipping. + """ + result = {"timing_note": "Approximate recording-launch-relative windows, not sample-aligned. " + "Prompt includes pre-roll. WAV begins at first audio sample; MP4 has its own " + "stream timestamps. Capture startup delay can shift either window."} + sources = [("container_audio", Path(media))] + sidecar = Path(media).with_suffix(".capture.wav") + if sidecar.exists(): + sources.append(("lossless_capture", sidecar)) + for label, source in sources: + result[label] = {"source": str(source), + "prompt": audio_metrics(source, start=0, duration=prompt_end), + "response": audio_metrics(source, start=response_start, + duration=min(180, media_seconds) - response_start)} + return result + + +def host_clock(host): + started = time.time() + remote = datetime.fromisoformat(ssh(host, ["date", "-u", "+%FT%T.%NZ"]).strip().replace("Z", "+00:00")).timestamp() + ended = time.time() + return {"offset_seconds": remote - (started + ended) / 2, + "uncertainty_seconds": (ended - started) / 2} + + +def host_timestamp(local_iso, clock, margin=0): + stamp = datetime.fromisoformat(local_iso).timestamp() + clock["offset_seconds"] + margin + return datetime.fromtimestamp(stamp, timezone.utc).isoformat() + + +def prepare_prompt(session, directory, config, case, env): + """Select local prompt evidence only; relative paths are session-relative. + + Snapshot supplied waveforms per attempt so a later render cannot replace + the file between validation and the shell harness's playback copy. + """ + env.pop("DOTTY_AV_PROMPT_WAV", None) + entries = config.get("prompt_wavs", {}) + if not isinstance(entries, dict): + raise ValueError("prompt_wavs_invalid_mapping") + if case["id"] not in entries: + return {"source": "harness_default", "renderer": "espeak-ng", "text": case["prompt"]} + entry = entries[case["id"]] + if not isinstance(entry, dict) or any( + not isinstance(entry.get(key), str) or not entry[key].strip() + for key in ("path", "text", "sha256", "renderer")): + raise ValueError("prompt_wav_invalid_entry") + if not re.fullmatch(r"[0-9a-f]{64}", entry["sha256"]): + raise ValueError("prompt_wav_invalid_sha256") + if entry["text"] != case["prompt"]: + raise ValueError("prompt_wav_text_mismatch") + source = Path(entry["path"]) + if not source.is_absolute(): + source = session / source + source = source.resolve() + if not source.is_file() or source.stat().st_size == 0: + raise ValueError("prompt_wav_missing_or_empty") + selected = (directory / "selected-prompt.wav").resolve() + shutil.copyfile(source, selected) + with selected.open("rb") as audio: + digest = hashlib.file_digest(audio, "sha256").hexdigest() + if digest != entry["sha256"]: + raise ValueError("prompt_wav_sha256_mismatch") + env["DOTTY_AV_PROMPT_WAV"] = str(selected) + return {**entry, "source": "configured_wav", "source_path": str(source), "path": str(selected)} + + +def response_vad_threshold(config): + value = config.get("response_vad_threshold", .5) + # The bounded comparison also rejects NaN and infinities. JSON booleans + # are not numerical settings, despite bool being an int subclass. + if isinstance(value, bool) or not isinstance(value, (int, float)) or not 0 <= value <= 1: + raise ValueError("response_vad_threshold_must_be_finite_number_between_0_and_1") + return float(value) + + +def run_case(session, config, case): + ident = datetime.now(timezone.utc).strftime("%Y%m%dT%H%M%S") + "-" + case["id"] + "-" + uuid.uuid4().hex[:6] + directory = session / "cases" / ident + directory.mkdir(parents=True) + result = {"id": ident, "case_id": case["id"], "case": case, "started": now(), + "commit": command(["git", "rev-parse", "HEAD"], cwd=ROOT).strip(), + "capture": "INCONCLUSIVE", "playback": "INCONCLUSIVE", + "interaction": "INCONCLUSIVE", "recovery": "INCONCLUSIVE", + "visual": "INCONCLUSIVE", "verdict": "INCONCLUSIVE"} + write_json(directory / "result.json", result) + write_json(session / "active.json", {"case": ident, "case_id": case["id"], "started": result["started"]}) + journal(session, {"event": "case_started", "id": ident}) + try: + result["response_vad_threshold"] = response_vad_threshold(config) + require_warm = case.get("require_warm_listening", + case.get("recovery_policy") == "ready_for_followup" + and not case.get("require_wake_event")) + if not isinstance(require_warm, bool): + raise ValueError("require_warm_listening_must_be_boolean") + if require_warm and case.get("require_wake_event"): + raise ValueError("warm_listening_and_cold_wake_preconditions_conflict") + result["warm_listening_precondition"] = "NOT_REQUIRED" + before = snapshot(config["host"]) + write_json(directory / "before.json", before) + if before["errors"] or not before.get("services_running") or config["device"] not in before.get("devices", []): + raise RuntimeError("infrastructure_down_or_device_disconnected") + if case.get("require_wake_event") and not wake_precondition(before, config["device"]): + raise CaseBlocked("wake_precondition") + if require_warm: + fresh, age = warm_listening(before, config["device"]) + if not fresh and config.get("mic_opener") == "admin_say": + # Explicit operator opt-in: the default remains "never change + # the robot's state to satisfy a prerequisite". + result["mic_opened_by"] = "admin_say" + opened = open_mic(config) + if opened: + before = opened + write_json(directory / "before.json", before) + fresh, age = warm_listening(before, config["device"]) + result["warm_listening_status_age_seconds"] = age + result["warm_listening_max_age_seconds"] = 30 + # A prerequisite for a new trial, not a claim that an older open + # conversation is dead. Recovery after a long capture is separate. + if not fresh: + raise CaseBlocked("warm_listening_precondition") + result["warm_listening_precondition"] = "PASS" + sink = command(["pactl", "get-default-sink"]).strip() + if sink != config["sink"]: + raise RuntimeError("speaker_sink_changed") + volume = command(["pactl", "get-sink-volume", sink]) + channels = re.findall(r"(\d+)%", volume) + if not channels or any(int(p) != config["volume"] for p in channels): + raise RuntimeError("speaker_volume_changed") + if command(["pactl", "get-sink-mute", sink]).strip() != "Mute: no": + raise RuntimeError("speaker_muted_or_unknown") + env = dict(os.environ, DOTTY_AV_VOLUME=str(config["volume"]), + DOTTY_AV_ALLOW_HIGH_VOLUME="1", DOTTY_AV_SINK=sink, + DOTTY_AV_MAX_SECONDS="180") + result["prompt_provenance"] = prepare_prompt(session, directory, config, case, env) + write_json(directory / "result.json", result) + media = directory / "raw.mp4" + result["host_clock"] = host_clock(config["host"]) + result["capture_started"] = now() + guarded_process([ROOT / "scripts/dotty-av-test.sh", "run", case["prompt"], + case.get("seconds", 50), media], directory / "capture.log", session, 185, env, + datetime.fromisoformat(config["stop_prompts_at"]).timestamp()) + result["capture_ended"] = now() + # Snapshot before transcription: evaluator runtime must not give a stuck + # robot extra minutes to recover unnoticed. + after = snapshot(config["host"]) + write_json(directory / "after.json", after) + clock = result["host_clock"] + read_log_window(config["host"], + host_timestamp(result["capture_started"], clock, -clock["uncertainty_seconds"]), + host_timestamp(result["capture_ended"], clock, clock["uncertainty_seconds"]), directory) + result["media"] = verify(media, expected=case.get("seconds", 50)) + result["capture"] = "PASS" + result["playback_process"] = "PASS" + # Compute actual playback offset from persisted timestamps and recording + # duration. Crop conservatively after playback, never use prompt audio as answer. + prompt_offset, offset = playback_window(directory, result["media"]["audio_seconds"]) + result["response_offset"] = offset + result["timing_note"] = "Recording-launch timestamp precedes first sample; response crop is conservatively late. Early reply/latency may be lost." + result["native_audio_windows"] = capture_window_metrics( + media, result["media"]["audio_seconds"], offset - .4, offset) + for label, options in (("response", ["-ss", str(offset)]), + ("prompt-heard", ["-ss", "0", "-t", str(offset - .4)])): + wav = directory / f"{label}.wav" + command(["ffmpeg", "-nostdin", "-n", "-v", "error", *options, "-i", media, + "-vn", "-ar", "16000", "-ac", "1", wav]) + guarded_process([config["python"], ROOT / "scripts/dotty_av_media.py", "transcribe", wav, + "--model", config["model"], "--output", directory / f"{label}.json", + *(["--normalize", "--vad-threshold", str(result["response_vad_threshold"])] + if label == "response" else [])], + directory / f"{label}-transcribe.log", session, 120, + deadline=datetime.fromisoformat(config["finish_at"]).timestamp()) + analysis = load(directory / f"{label}.json") + result.setdefault("transcription_analysis", {})[label] = { + key: analysis[key] for key in ("model", "analysis", "vad_filter", "vad_parameters") + if key in analysis} + heard = " ".join(s["text"] for s in load(directory / "prompt-heard.json")["segments"]) + result["prompt_heard"] = heard + result.update(evaluate(case, load(directory / "response.json"), + (directory / "xiaozhi-esp32-server.log").read_text(), + offset, after, config["device"], before=before)) + result["playback"], result["playback_evidence"] = playback_verdict( + heard, result["asr_match"], case.get("asr", [])) + if result["playback"] != "PASS" and result["interaction"] == "PASS": + result.update(interaction="INCONCLUSIVE", failure="prompt_playback_unverified") + for label, seconds in (("prompt", 3), ("response", offset + 2), + ("final", max(0, result["media"]["video_seconds"] - 2))): + command(["ffmpeg", "-nostdin", "-n", "-v", "error", "-ss", str(seconds), + "-i", media, "-frames:v", "1", directory / f"frame-{label}.jpg"]) + except CaseBlocked as exc: + result.update(verdict="BLOCKED", interaction="BLOCKED", failure=str(exc), + exception=type(exc).__name__) + result[str(exc)] = "BLOCKED" + except Exception as exc: + result.update(verdict="FAIL", failure=str(exc), exception=type(exc).__name__) + finally: + result["ended"] = now() + write_json(directory / "result.json", result) + # Caller clears active only after checkpoint fsync. A crash between these + # writes otherwise replays a potentially side-effectful successful case. + journal(session, {"event": "case_finished", "id": ident, "verdict": result["verdict"], + "failure": result.get("failure")}) + return result + + +def report(session): + results = [load(p) for p in sorted((session / "cases").glob("*/result.json"))] + rows = [] + for r in results: + rel = "cases/" + r["id"] + rows.append(f'{html.escape(r["case_id"])}{r["verdict"]}' + f'{html.escape(r.get("failure") or r.get("response_transcript", ""))}' + f'Video · Evidence') + (session / "index.html").write_text('Dotty overnight review' + '' + '

Dotty overnight review

AI-assisted: OpenAI Codex (GPT-6). ' + 'Automated acoustic verdicts are provisional. Missing visual/physical coverage is not a pass.

' + '

Coverage · Status

' + '' + ''.join(rows) + '
CaseOverallFindingReview
') + write_json(session / "summary.json", {"at": now(), "attempts": len(results), + "interaction_passes": sum(r.get("interaction") == "PASS" for r in results), + "failures": sum(r["verdict"] == "FAIL" for r in results), + "blocked": sum(r["verdict"] == "BLOCKED" for r in results), + "inconclusive": sum(r["verdict"] == "INCONCLUSIVE" for r in results)}) + + +def preflight(session, config): + threshold = response_vad_threshold(config) + data = {"at": now(), "response_vad_threshold": threshold, + "disk_ok": space_ok(session), "snapshot": snapshot(config["host"])} + for binary in ("ffmpeg", "ffprobe", "espeak-ng", "pw-play", "pactl", "flock"): + data[binary] = shutil.which(binary) + data["model_exists"] = (Path(config["model"]) / "model.bin").is_file() + data["camera_exists"] = Path("/dev/v4l/by-id/usb-046d_HD_Pro_Webcam_C920_7B90DC9F-video-index0").exists() + write_json(session / "preflight.json", data) + return data + + +def lock(session): + f = (session / "runner.lock").open("w") + fcntl.flock(f, fcntl.LOCK_EX | fcntl.LOCK_NB) + return f + + +@contextlib.contextmanager +def device_lock(config): + # All sessions targeting this physical robot serialize, including different + # output directories. The shell separately locks the shared camera. + key = hashlib.sha256(config["device"].encode()).hexdigest()[:20] + path = Path("/tmp") / f"dotty-overnight-device-{key}.lock" + with path.open("a") as handle: + fcntl.flock(handle, fcntl.LOCK_EX | fcntl.LOCK_NB) + yield + + +def restore_checkpoint(session): + checkpoint = load(session / "checkpoint.json", {}) + for key, value in {"completed": [], "quarantined": [], "unresolved": [], "blocked": {}, "consecutive_failures": 0, + "soak": 0, "counts": {}}.items(): + checkpoint.setdefault(key, value) + active = load(session / "active.json", {}) + if active: + if checkpoint.get("last_attempt") == active["case"]: + write_json(session / "active.json", {}) + return checkpoint + path = session / "cases" / active["case"] / "result.json" + result = load(path, {}) + case_id = active.get("case_id") or result.get("case_id") + if not case_id: + raise RuntimeError("interrupted case identity unknown; manual evidence review required") + if case_id not in checkpoint["quarantined"]: + checkpoint["quarantined"].append(case_id) + if result: + result.update(verdict="INCONCLUSIVE", failure="interrupted_case_requires_review", + ended=now()) + write_json(path, result) + journal(session, {"event": "interrupted_case_requires_review", **active}) + # Persist quarantine before clearing active: a second crash cannot lose it. + write_json(session / "checkpoint.json", checkpoint) + write_json(session / "active.json", {}) + return checkpoint + + +def record_attempt(checkpoint, case, result): + if result.get("id"): + checkpoint["last_attempt"] = result["id"] + counts = checkpoint["counts"].setdefault(case["id"], {"attempts": 0, "streak": 0, "failures": 0}) + counts["attempts"] += 1 + if result.get("verdict") == "BLOCKED": + checkpoint.setdefault("blocked", {})[case["id"]] = result.get("failure", "prerequisite_missing") + return + fields = [result.get(k) for k in ("capture", "playback", "interaction", "recovery")] + acoustic_pass = all(value == "PASS" for value in fields) + # An attempt that proved nothing (reference-mic slip, extra speech in the + # room) breaks the streak but is not evidence of a robot fault, so it must + # not push a healthy case toward quarantine or pause the run. + inconclusive = (not acoustic_pass and result.get("verdict") != "FAIL" + and "FAIL" not in fields and "INCONCLUSIVE" in fields) + counts["streak"] = counts["streak"] + 1 if acoustic_pass else 0 + if not acoustic_pass and not inconclusive: + counts["failures"] += 1 + # A case the evidence can never settle (for example a word the reference + # microphone always mishears) must not hold the queue for ever. + counts["inconclusive"] = 0 if acoustic_pass else counts.get("inconclusive", 0) + (1 if inconclusive else 0) + if (counts["inconclusive"] >= case.get("max_inconclusive", 3) + and case["id"] not in checkpoint.setdefault("unresolved", [])): + checkpoint["unresolved"].append(case["id"]) + if acoustic_pass: + checkpoint["consecutive_failures"] = 0 + elif not inconclusive: + checkpoint["consecutive_failures"] += 1 + # Acoustic completion remains provisional while the visual verdict is open. + if counts["streak"] >= case.get("required_passes", 3) and case["id"] not in checkpoint["completed"]: + checkpoint["completed"].append(case["id"]) + if counts["failures"] >= case.get("max_failures", 3) and case["id"] not in checkpoint["quarantined"]: + checkpoint["quarantined"].append(case["id"]) + if result.get("exception") in {"InterruptedError", "TimeoutError"} and case["id"] not in checkpoint["quarantined"]: + checkpoint["quarantined"].append(case["id"]) + + +def eligible_cases(bank, checkpoint): + return [c for c in bank if c["id"] not in checkpoint["quarantined"] + and c["id"] not in checkpoint["blocked"] + and c["id"] not in checkpoint.get("unresolved", [])] + + +def run(session, config, catalogue, selected=None, once=False): + with lock(session), device_lock(config): + checkpoint = restore_checkpoint(session) + deadline = datetime.fromisoformat(config["stop_prompts_at"]).timestamp() + rng = random.Random(config["seed"] + checkpoint["soak"]) + bank = [c for c in catalogue if c.get("enabled", True) and (not selected or c["id"] in selected)] + if not bank: + raise ValueError("no selected executable cases") + previous = time.monotonic() + try: + while not STOP and not (session / "STOP").exists() and time.time() < deadline: + write_json(session / "status.json", {"state": "running", "heartbeat": now(), "pid": os.getpid(), + "checkpoint": checkpoint, "stop_prompts_at": config["stop_prompts_at"]}) + if (session / "PAUSE").exists(): + time.sleep(1) + continue + if checkpoint["consecutive_failures"] >= 3: + write_json(session / "PAUSE", {"reason": "three consecutive failures require agent diagnosis"}) + write_json(session / "checkpoint.json", checkpoint) + continue + eligible = eligible_cases(bank, checkpoint) + pending = [c for c in eligible if c["id"] not in checkpoint["completed"]] + if pending: + case = pending[0] + else: + soak_bank = [c for c in eligible if c.get("soak_safe", False)] + if not soak_bank: + journal(session, {"event": "no_eligible_soak_cases"}) + break + due = previous + config.get("interval_seconds", 600) + if time.monotonic() < due: + time.sleep(min(1, due - time.monotonic())) + continue + case = soak_bank[0] if checkpoint["soak"] % 6 == 0 else rng.choice(soak_bank) + checkpoint["soak"] += 1 + # Reserve enough time for capture and both transcriptions before cutoff. + if deadline - time.time() < 360: + break + if not space_ok(session): + raise RuntimeError("disk pressure") + result = run_case(session, config, case) + previous = time.monotonic() + record_attempt(checkpoint, case, result) + write_json(session / "checkpoint.json", checkpoint) + write_json(session / "active.json", {}) + report(session) + print(json.dumps({"case": case["id"], "verdict": result["verdict"], "failure": result.get("failure")}), flush=True) + if once: + break + # Quiet gap between coverage cases; interruptible, no prompt overlap. + for _ in range(20): + if STOP or (session / "STOP").exists(): + break + time.sleep(1) + finally: + write_json(session / "status.json", {"state": "stopped", "at": now(), "checkpoint": checkpoint}) + report(session) + + +def main(): + parser = argparse.ArgumentParser(description=__doc__) + parser.add_argument("command", choices=["init", "preflight", "run", "resume", "status", "stop", "report", "admin"]) + parser.add_argument("--session", required=True, type=Path) + parser.add_argument("--host", default=os.environ.get("DOTTY_TEST_HOST"), + help="init only: explicit SSH user@host (or DOTTY_TEST_HOST)") + parser.add_argument("--device") + parser.add_argument("--cases", help="comma-separated case IDs") + parser.add_argument("--once", action="store_true") + parser.add_argument("--route") + parser.add_argument("--body", help="JSON object; tokens are resolved remotely") + parser.add_argument("--response-vad-threshold", type=lambda value: response_vad_threshold( + {"response_vad_threshold": float(value)}), + help="init only: response transcription VAD threshold, 0..1 (default 0.5)") + parser.add_argument("--mic-opener", choices=["none", "admin_say"], + help="init only: 'admin_say' lets warm cases open the microphone through " + "/xiaozhi/admin/say instead of blocking (default none)") + args = parser.parse_args() + if args.mic_opener is not None and args.command != "init": + parser.error("--mic-opener is init-only; existing sessions use config.json") + if args.response_vad_threshold is not None and args.command != "init": + parser.error("--response-vad-threshold is init-only; existing sessions use config.json") + if args.command == "init" and (not args.host or not args.host.strip()): + parser.error("init requires --host or DOTTY_TEST_HOST; no deployment host is assumed") + session = args.session.resolve() + session.mkdir(parents=True, exist_ok=True) + if args.command == "init": + if (session / "config.json").exists(): + raise ValueError("session already initialized") + morning = datetime.now(TZ).replace(hour=8, minute=0, second=0, microsecond=0) + if morning <= datetime.now(TZ): + morning += timedelta(days=1) + devices = admin(args.host, "devices")["devices"] + device = args.device or (devices[0] if len(devices) == 1 else None) + if not device: + raise ValueError("select exactly one live device") + config = {"host": args.host, "device": device, + "sink": command(["pactl", "get-default-sink"]).strip(), "volume": 100, + "python": str(session / "evaluator-venv/bin/python"), + "model": str(session / "models/whisper-small.en-ct2"), + "response_vad_threshold": .5 if args.response_vad_threshold is None else args.response_vad_threshold, + "mic_opener": args.mic_opener or "none", + "seed": 20260912, "interval_seconds": 600, + "stop_prompts_at": (morning - timedelta(minutes=30)).isoformat(), + "finish_at": morning.isoformat(), "created": now(), "framing": "BLOCKED"} + write_json(session / "config.json", config) + for folder in ("clips/ready-for-review", "clips/needs-review", "clips/failures", "experiments", "logs"): + (session / folder).mkdir(parents=True, exist_ok=True) + write_json(session / "coverage.json", load(ROOT / "scripts/dotty-av-cases.json")) + print(json.dumps(config, indent=2)) + return + config = load(session / "config.json") + if args.command == "admin": + print(json.dumps(admin(config["host"], args.route, json.loads(args.body) if args.body else None))) + elif args.command == "preflight": + print(json.dumps(preflight(session, config), indent=2)) + elif args.command in {"run", "resume"}: + if args.command == "resume": + # Explicit invocation is the agent's decision to resume after review. + with lock(session): + for name in ("STOP", "PAUSE"): + (session / name).unlink(missing_ok=True) + checkpoint = restore_checkpoint(session) + checkpoint["consecutive_failures"] = 0 + write_json(session / "checkpoint.json", checkpoint) + run(session, config, load(ROOT / "scripts/dotty-av-cases.json")["cases"], + args.cases.split(",") if args.cases else None, args.once) + elif args.command == "stop": + write_json(session / "STOP", {"at": now()}) + elif args.command == "status": + print(json.dumps(load(session / "status.json", {}), indent=2)) + else: + report(session) + + +if __name__ == "__main__": + def stop_signal(signum, frame): + global STOP + STOP = True + signal.signal(signal.SIGINT, stop_signal) + signal.signal(signal.SIGTERM, stop_signal) + main() diff --git a/scripts/patch-tts-server-push.py b/scripts/patch-tts-server-push.py new file mode 100644 index 00000000..73d0bf99 --- /dev/null +++ b/scripts/patch-tts-server-push.py @@ -0,0 +1,42 @@ +"""Patch pinned xiaozhi TTS consumers for explicit server-push arbitration.""" + +from pathlib import Path +import sys + + +OLD = "if message.sentence_id != self.conn.sentence_id:\n" +NEW = ( + "if (\n" + " message.sentence_id != self.conn.sentence_id\n" + " and message.sentence_id not in getattr(\n" + " self.conn, \"_dotty_server_push_sentence_ids\", set()\n" + " )\n" + " ):\n" +) + + +def patch_source(source: str) -> str: + if NEW in source: + return source + if OLD not in source: + raise ValueError("sentence-id predicate anchor not found") + return source.replace(OLD, NEW, 1) + + +def main(root: Path) -> None: + patched = 0 + for path in root.glob("*.py"): + source = path.read_text(encoding="utf-8") + if OLD not in source and NEW not in source: + continue + updated = patch_source(source) + if updated != source: + path.write_text(updated, encoding="utf-8") + patched += 1 + if patched == 0: + raise SystemExit("no xiaozhi TTS consumers contained the expected predicate") + print(f"patched {patched} TTS consumer(s) for server-push arbitration") + + +if __name__ == "__main__": + main(Path(sys.argv[1])) diff --git a/tests/test_abort_flag_cleared_on_new_turn.py b/tests/test_abort_flag_cleared_on_new_turn.py new file mode 100644 index 00000000..004911eb --- /dev/null +++ b/tests/test_abort_flag_cleared_on_new_turn.py @@ -0,0 +1,52 @@ +"""A device/admin abort must not silence every later turn on the connection. + +AI-assisted (Claude). Observed on the physical robot 2026-10-05: after one +`abort` frame, ASR and the LLM kept running but no reply was ever spoken, +because nothing on the `nointent` chat path cleared `conn.client_abort`. +""" +import asyncio +import json +import types +from unittest.mock import AsyncMock, MagicMock, patch + +from tests.test_asr_name_corrections import _module + + +def _start_turn(text, **conn_overrides): + conn = types.SimpleNamespace( + logger=MagicMock(), need_bind=False, max_output_size=0, + client_is_speaking=False, current_state="idle", session_id="fixture", + websocket=types.SimpleNamespace(send=AsyncMock()), + executor=MagicMock(), chat=MagicMock(), + _dotty_toggles_synced=True, client_abort=False, + ) + for key, value in conn_overrides.items(): + setattr(conn, key, value) + flag_at_submit = [] + conn.executor.submit.side_effect = lambda *a, **k: flag_at_submit.append(conn.client_abort) + with patch.object(_module, "handle_user_intent", new=AsyncMock(return_value=False)), \ + patch.object(_module, "send_stt_message", new=AsyncMock()), \ + patch.object(_module, "_mcp_call_tool", new=AsyncMock()): + asyncio.run(_module.startToChat(conn, json.dumps({"content": text}))) + return conn, flag_at_submit + + +def test_new_turn_after_abort_is_not_submitted_with_abort_still_set(): + conn, flag_at_submit = _start_turn("what is your name", client_abort=True) + assert flag_at_submit == [False] + assert conn.client_abort is False + + +def test_turn_handled_by_intent_does_not_touch_the_flag(): + conn = types.SimpleNamespace( + logger=MagicMock(), need_bind=False, max_output_size=0, + client_is_speaking=False, current_state="idle", session_id="fixture", + websocket=types.SimpleNamespace(send=AsyncMock()), + executor=MagicMock(), chat=MagicMock(), + _dotty_toggles_synced=True, client_abort=True, + ) + with patch.object(_module, "handle_user_intent", new=AsyncMock(return_value=True)), \ + patch.object(_module, "send_stt_message", new=AsyncMock()): + asyncio.run(_module.startToChat(conn, json.dumps({"content": "goodbye"}))) + assert conn.client_abort is True + conn.executor.submit.assert_not_called() diff --git a/tests/test_asr_name_corrections.py b/tests/test_asr_name_corrections.py index 2d358f9d..3934b08b 100644 --- a/tests/test_asr_name_corrections.py +++ b/tests/test_asr_name_corrections.py @@ -1,11 +1,14 @@ """Regression tests for wake-name corrections on the live ASR text path.""" import importlib.util +import asyncio +import json import pathlib import sys import types import unittest from contextlib import contextmanager +from unittest.mock import AsyncMock, MagicMock, patch _ROOT = pathlib.Path(__file__).parent.parent @@ -97,5 +100,67 @@ def test_alias_matching_is_word_bounded(self): self.assertEqual(_module._apply_asr_corrections(text), text) +class TestAsrEnvelope(unittest.TestCase): + """AI-assisted GPT-6: exercise the actual post-ASR call boundary.""" + + def test_content_only_json_reaches_intent_as_text(self): + for envelope in ( + {"content": "Purple robot seven window Brisbane."}, + {"content": "Purple robot seven window Brisbane.", "speaker": "Fixture"}, + {"content": "Purple robot seven window Brisbane.", "speaker": "Fixture", "language": "en"}, + ): + with self.subTest(envelope=envelope): + conn = types.SimpleNamespace(logger=MagicMock(), need_bind=False, + max_output_size=0, client_is_speaking=False) + intent = AsyncMock(return_value=True) + with patch.object(_module, "_sync_toggles_once", new=AsyncMock()), \ + patch.object(_module, "handle_user_intent", new=intent): + asyncio.run(_module.startToChat(conn, json.dumps(envelope))) + intent.assert_awaited_once_with(conn, envelope["content"]) + self.assertEqual(conn.current_speaker, envelope.get("speaker")) + + def test_empty_or_nontext_content_is_not_a_spoken_request(self): + for content in ("", "...", None, [], {"nested": "say hello"}): + with self.subTest(content=content): + conn = types.SimpleNamespace(logger=MagicMock(), need_bind=False, + max_output_size=0, client_is_speaking=False) + sync = AsyncMock() + with patch.object(_module, "_sync_toggles_once", new=sync), \ + patch.object(_module, "handle_user_intent", new=AsyncMock(return_value=True)): + asyncio.run(_module.startToChat(conn, json.dumps({"content": content}))) + sync.assert_not_awaited() + + +class TestPhraseIntentPreservation(unittest.TestCase): + def test_correct_ordinary_requests_are_not_rewritten_as_commands(self): + for text in ( + "Tell me one short joke about being a very small robot.", + "Tell me a short joke.", + "What is your game?", + "Can you sing a long note?", + "Please take a phota of the toy.", + "Tell me. A story is something I dislike.", + ): + with self.subTest(text=text): + self.assertEqual(_module._apply_phrase_corrections(text), text) + + def test_known_phrase_punctuation_still_normalizes(self): + self.assertEqual(_module._apply_phrase_corrections("Good night, Dotty."), + "good night Dotty.") + self.assertEqual(_module._apply_phrase_corrections("Please take a photo, now."), + "Please take a photo now.") + + def test_actual_json_joke_turn_reaches_routing_unchanged(self): + text = "Tell me one short joke about being a very small robot." + conn = types.SimpleNamespace(logger=MagicMock(), need_bind=False, + max_output_size=0, client_is_speaking=False) + intent = AsyncMock(return_value=True) + with patch.object(_module, "_sync_toggles_once", new=AsyncMock()), \ + patch.object(_module, "handle_user_intent", new=intent): + asyncio.run(_module.startToChat(conn, json.dumps({"content": text}))) + intent.assert_awaited_once_with(conn, text) + self.assertIsNone(_module._detect_state_phrase(intent.await_args.args[1])) + + if __name__ == "__main__": unittest.main() diff --git a/tests/test_bridge_routes.py b/tests/test_bridge_routes.py index 75f45276..94d80477 100644 --- a/tests/test_bridge_routes.py +++ b/tests/test_bridge_routes.py @@ -321,92 +321,5 @@ def test_scene_synthesis_cache_getter_degrades_on_timeout(self): self.assertEqual(result, {}) -# --------------------------------------------------------------------------- -# Sound-balance series getter — rewired in #115 Tile 5 -# --------------------------------------------------------------------------- - - -class SoundBalanceGetterTests(unittest.TestCase): - """Smoke tests for the dotty-behaviour-backed sound-balance series - getter. Resolves the device id from the perception state getter - (single-robot heuristic) and pulls the sparkline series.""" - - def setUp(self) -> None: - bridge_app._dotty_behaviour_cache.clear() - - def _patch_get(self, by_path): - """Monkeypatch requests.get to dispatch on URL path. ``by_path`` - is a dict mapping path-substring → response (or Exception).""" - original = bridge_app.requests.get - - def fake_get(url, *args, **kwargs): - for needle, response in by_path.items(): - if needle in url: - if isinstance(response, Exception): - raise response - return response - raise AssertionError(f"unexpected URL fetched: {url}") - - bridge_app.requests.get = fake_get - self.addCleanup(lambda: setattr(bridge_app.requests, "get", original)) - - def test_sound_balance_series_returns_fetched_list(self): - class StateResp: - def raise_for_status(self): - return None - - def json(self): - return {"dev-1": {"face_present": True}} - - class BalanceResp: - def raise_for_status(self): - return None - - def json(self): - return [0.1, 0.5, 0.9] - - self._patch_get( - { - "/api/perception/state": StateResp(), - "/api/perception/sound-balance/": BalanceResp(), - } - ) - result = bridge_app._dashboard_sound_balance_series() - self.assertEqual(result, [0.1, 0.5, 0.9]) - - def test_sound_balance_series_no_device_returns_empty(self): - class EmptyStateResp: - def raise_for_status(self): - return None - - def json(self): - return {} - - self._patch_get({"/api/perception/state": EmptyStateResp()}) - result = bridge_app._dashboard_sound_balance_series() - self.assertEqual(result, []) - - def test_sound_balance_series_degrades_on_connection_error(self): - import requests as _requests - - class StateResp: - def raise_for_status(self): - return None - - def json(self): - return {"dev-1": {"face_present": True}} - - self._patch_get( - { - "/api/perception/state": StateResp(), - "/api/perception/sound-balance/": _requests.exceptions.ConnectionError( - "simulated" - ), - } - ) - result = bridge_app._dashboard_sound_balance_series() - self.assertEqual(result, []) - - if __name__ == "__main__": unittest.main() diff --git a/tests/test_dashboard_feed_and_actions.py b/tests/test_dashboard_feed_and_actions.py new file mode 100644 index 00000000..ac1806ce --- /dev/null +++ b/tests/test_dashboard_feed_and_actions.py @@ -0,0 +1,242 @@ +"""Tests for the perception-feed relay, the no-wait inject helper, the +last-seen fallback, the voice-tool inventory, and the voice-turn ingest. + +The dashboard's live pieces all pointed at producers that moved out of the +bridge in #36 / #111: the activity feed opened an SSE route the bridge never +served, and every inject action sat for 8 s on a turn stream nothing publishes +to. Handlers are invoked directly (not via TestClient) so neither CSRF nor a +live dotty-behaviour is in the path. +""" +from __future__ import annotations + +import asyncio +import importlib.util +import os +import re +import sys +import tempfile +import unittest +from contextlib import asynccontextmanager +from pathlib import Path + +from starlette.requests import Request + +_state_dir = Path(tempfile.mkdtemp(prefix="dotty-feed-state-")) +os.environ.setdefault("DOTTY_KID_MODE_STATE", str(_state_dir / "kid-mode")) +os.environ.setdefault("DOTTY_SMART_MODE_STATE", str(_state_dir / "smart-mode")) + +_repo_root = Path(__file__).resolve().parents[1] +_spec = importlib.util.spec_from_file_location("bridge_app", _repo_root / "bridge.py") +assert _spec is not None and _spec.loader is not None +bridge_app = importlib.util.module_from_spec(_spec) +sys.modules["bridge_app"] = bridge_app +_spec.loader.exec_module(bridge_app) + + +@asynccontextmanager +async def _noop_lifespan(_app): + yield + + +bridge_app.app.router.lifespan_context = _noop_lifespan + +import bridge.dashboard as dash # noqa: E402 + + +def _req(path: str = "/ui/perception/feed", method: str = "GET") -> Request: + async def _receive(): + # Never signals http.disconnect — the stream ends on upstream EOF. + await asyncio.sleep(3600) + + return Request( + {"type": "http", "method": method, "path": path, + "headers": [], "query_string": b""}, + _receive, + ) + + +class _FakeUpstream: + def __init__(self, lines): + self._lines = lines + self.closed = False + + def iter_lines(self, chunk_size=None): + return iter(self._lines) + + def close(self): + self.closed = True + + +async def _drain(resp) -> bytes: + return b"".join([chunk async for chunk in resp.body_iterator]) + + +class _StateSaver(unittest.TestCase): + def setUp(self): + saved = dict(dash._state) + self.addCleanup(lambda: dash._state.update(saved)) + + +class PerceptionFeedProxyTests(_StateSaver): + + def test_relays_upstream_lines_and_closes(self): + upstream = _FakeUpstream([b'data: {"name":"face_detected"}', b"", b": keepalive", b""]) + dash._state["perception_feed_opener"] = lambda: upstream + resp = asyncio.run(dash.perception_feed_proxy(_req())) + self.assertEqual(resp.media_type, "text/event-stream") + body = asyncio.run(_drain(resp)) + self.assertEqual( + body, + b'retry: 5000\n\ndata: {"name":"face_detected"}\n\n: keepalive\n\n', + ) + self.assertTrue(upstream.closed) + + def test_upstream_down_ends_stream_without_error(self): + def _boom(): + raise ConnectionError("simulated") + + dash._state["perception_feed_opener"] = _boom + resp = asyncio.run(dash.perception_feed_proxy(_req())) + self.assertEqual(resp.status_code, 200) + self.assertEqual(asyncio.run(_drain(resp)), b"retry: 5000\n\n") + + def test_page_subscribes_to_the_relay_route(self): + html = (_repo_root / "bridge/templates/dashboard.html").read_text() + self.assertIn("new EventSource('/ui/perception/feed')", html) + self.assertNotIn("EventSource('/api/perception/feed')", html) + + +class InjectNoWaitTests(_StateSaver): + + def test_returns_without_waiting_for_a_turn(self): + async def _inject(text): + return {"ok": True} + + dash._state["inject_to_device"] = _inject + + async def _run(): + return await asyncio.wait_for( + dash._inject_or_error(_req("/ui/actions/say", "POST"), "hi", label="hi"), + timeout=1.0, + ) + + body = asyncio.run(_run()).body.decode("utf-8") + self.assertIn("sent to Dotty", body) + self.assertNotIn("no reply", body) + + def test_inject_failure_surfaces_error(self): + async def _inject(text): + return {"ok": False, "error": "no device connected"} + + dash._state["inject_to_device"] = _inject + resp = asyncio.run( + dash._inject_or_error(_req("/ui/actions/say", "POST"), "hi", label="hi") + ) + self.assertIn("no device connected", resp.body.decode("utf-8")) + + +class LastSeenFallbackTests(_StateSaver): + + def setUp(self): + super().setUp() + dash._state["recent_turns_getter"] = lambda: [] + + def test_prefers_newest_reported_turn(self): + dash._state["recent_turns_getter"] = lambda: [{"ts": 300.0}, {"ts": 400.0}] + dash._state["perception_state_getter"] = lambda: {"dev-1": {"last_chat_t": 100.0}} + self.assertEqual(dash._stackchan_last_seen(), 400.0) + + def test_falls_back_to_perception_last_chat(self): + dash._state["perception_state_getter"] = lambda: { + "dev-1": {"last_chat_t": 100.0}, "dev-2": {"last_chat_t": 250.0}, + } + self.assertEqual(dash._stackchan_last_seen(), 250.0) + + def test_none_when_no_chat_recorded(self): + dash._state["perception_state_getter"] = lambda: {"dev-1": {}} + self.assertIsNone(dash._stackchan_last_seen()) + + +class VoiceToolInventoryTests(unittest.TestCase): + + def test_matches_dotty_pi_ext_tools(self): + shipped = set() + for src in (_repo_root / "dotty-pi-ext/src/tools").glob("*.ts"): + shipped.update(re.findall(r'^\s*name: "(\w+)"', src.read_text(), re.M)) + self.assertEqual({t["name"] for t in dash._VOICE_TOOLS}, shipped) + + +class VoiceTurnIngestTests(unittest.TestCase): + """PiVoiceLLM reports turns to /api/voice/turn; the bridge keeps them in + an in-memory ring that feeds /ui/events, the Errors count and report.""" + + def setUp(self): + from fastapi.testclient import TestClient + bridge_app._dashboard_recent_turns.clear() + self.addCleanup(bridge_app._dashboard_recent_turns.clear) + saved_token = bridge_app._ADMIN_TOKEN + self.addCleanup(lambda: setattr(bridge_app, "_ADMIN_TOKEN", saved_token)) + bridge_app._ADMIN_TOKEN = "" + # Each test module loads its own copy of bridge.py, and the shared + # bridge.dashboard is wired to whichever loaded last — pin it to ours. + saved_getter = dash._state.get("recent_turns_getter") + self.addCleanup(lambda: dash._state.update(recent_turns_getter=saved_getter)) + dash._state["recent_turns_getter"] = bridge_app._dashboard_recent_turns_getter + self.client = TestClient(bridge_app.app) + + def test_turn_lands_in_ring_and_reaches_subscribers(self): + q = bridge_app._dashboard_subscribe_events() + self.addCleanup(lambda: bridge_app._dashboard_unsubscribe_events(q)) + r = self.client.post("/api/voice/turn", json={ + "request_text": "hello", "response_text": "😊 Hi!", "latency_ms": 1200, + }) + self.assertEqual(r.status_code, 200) + turns = bridge_app._dashboard_recent_turns_getter() + self.assertEqual(len(turns), 1) + self.assertEqual(turns[0]["request_text"], "hello") + self.assertIsNone(turns[0]["error"]) + self.assertIsInstance(turns[0]["ts"], float) + self.assertEqual(q.get_nowait()["response_text"], "😊 Hi!") + + def test_errored_turn_shows_in_count_and_report(self): + self.client.post("/api/voice/turn", json={"request_text": "a", "response_text": "ok"}) + self.client.post("/api/voice/turn", json={ + "request_text": "what is the weather", "response_text": "😐 (brain offline)", + "error": "pi rpc timeout", + }) + self.assertIn("Errors (1)", self.client.get("/ui/alerts/count?chip=1").text) + detail = self.client.get("/ui/alerts/detail").text + self.assertIn("pi rpc timeout", detail) + self.assertIn("what is the weather", detail) + + def test_no_errors_leaves_plain_chip(self): + self.client.post("/api/voice/turn", json={"request_text": "a", "response_text": "ok"}) + self.assertEqual(self.client.get("/ui/alerts/count?chip=1").text, "Errors") + + def test_admin_token_enforced_when_set(self): + bridge_app._ADMIN_TOKEN = "s3cret" + body = {"request_text": "a", "response_text": "b"} + self.assertEqual(self.client.post("/api/voice/turn", json=body).status_code, 401) + ok = self.client.post("/api/voice/turn", json=body, headers={"X-Admin-Token": "s3cret"}) + self.assertEqual(ok.status_code, 200) + self.assertEqual(len(bridge_app._dashboard_recent_turns_getter()), 1) + + def test_long_text_is_truncated(self): + self.client.post("/api/voice/turn", json={"request_text": "x" * 5000, "response_text": ""}) + self.assertEqual(len(bridge_app._dashboard_recent_turns_getter()[0]["request_text"]), 2000) + + def test_filter_hit_lands_in_safety_ring(self): + import bridge.text as btext + btext._cf_recent.clear() + self.addCleanup(btext._cf_recent.clear) + r = self.client.post("/api/voice/filter-hit", json={ + "tier": "redirect", "rule": "badword", "prefix": "a longer prefix", + }) + self.assertEqual(r.status_code, 200) + hit = btext.recent_content_filter_hits()[0] + self.assertEqual((hit["tier"], hit["rule"], hit["prefix"]), ("redirect", "badword", "a longer")) + self.assertIn("badword", self.client.get("/ui/safety/recent").text) + + +if __name__ == "__main__": + unittest.main() diff --git a/tests/test_device_command.py b/tests/test_device_command.py index 07007e9f..e95dd825 100644 --- a/tests/test_device_command.py +++ b/tests/test_device_command.py @@ -10,10 +10,11 @@ import importlib.util as _ilu import json import pathlib +import queue import sys import types import unittest -from unittest.mock import MagicMock +from unittest.mock import MagicMock, patch _PATCHES = pathlib.Path(__file__).parent.parent / "custom-providers" / "xiaozhi-patches" @@ -231,6 +232,121 @@ def test_resolve_conn_falls_back_to_first_device_and_503s_when_empty(self): self.assertIsNone(got) self.assertEqual(err.status, 503) + def test_server_say_survives_chat_sentence_id_change(self): + """#104: a chat turn must not invalidate queued server speech.""" + class SentenceType: + FIRST = "first" + MIDDLE = "middle" + LAST = "last" + + class ContentType: + TEXT = "text" + + class DTO: + def __init__(self, **kwargs): + self.__dict__.update(kwargs) + + dto_module = types.SimpleNamespace( + ContentType=ContentType, + SentenceType=SentenceType, + TTSMessageDTO=DTO, + ) + saved = { + name: sys.modules.get(name) + for name in ( + "core.providers", "core.providers.tts", + "core.providers.tts.dto", "core.providers.tts.dto.dto", + ) + } + try: + for name in ("core.providers", "core.providers.tts", "core.providers.tts.dto"): + sys.modules[name] = types.ModuleType(name) + sys.modules["core.providers.tts.dto.dto"] = dto_module + + async def go(): + conn = _FakeConn() + conn.headers = {"device-id": "dev-say"} + conn.tts = types.SimpleNamespace(tts_text_queue=queue.Queue()) + type(self).active["dev-say"] = conn + async def run_inline(func): + return func() + + with patch.object(self.mod.asyncio, "to_thread", side_effect=run_inline): + await self._server()._dotty_say( + self._request({"text": "Hello from the server."}) + ) + + conn.sentence_id = "new-chat-turn" + accepted = [] + while not conn.tts.tts_text_queue.empty(): + message = conn.tts.tts_text_queue.get_nowait() + if ( + message.sentence_id != conn.sentence_id + and message.sentence_id not in getattr( + conn, "_dotty_server_push_sentence_ids", set() + ) + ): + continue + accepted.append(message) + + self.assertEqual( + [m.sentence_type for m in accepted], + [SentenceType.FIRST, SentenceType.MIDDLE, SentenceType.LAST], + ) + + asyncio.run(go()) + finally: + for name, module in saved.items(): + if module is None: + sys.modules.pop(name, None) + else: + sys.modules[name] = module + + def test_server_push_exemption_does_not_keep_stale_chat_ids(self): + conn = types.SimpleNamespace( + sentence_id="current-chat", + _dotty_server_push_sentence_ids={"server-push"}, + ) + + def accepted(sentence_id): + return not ( + sentence_id != conn.sentence_id + and sentence_id not in conn._dotty_server_push_sentence_ids + ) + + self.assertFalse(accepted("stale-chat")) + self.assertTrue(accepted("server-push")) + + def test_client_abort_still_drops_server_push(self): + conn = types.SimpleNamespace( + sentence_id="current-chat", + client_abort=True, + _dotty_server_push_sentence_ids={"server-push"}, + ) + + def accepted(sentence_id): + if conn.client_abort: + return False + return not ( + sentence_id != conn.sentence_id + and sentence_id not in conn._dotty_server_push_sentence_ids + ) + + self.assertFalse(accepted("server-push")) + + def test_build_patch_rewrites_consumer_predicate_explicitly(self): + script = pathlib.Path(__file__).parent.parent / "scripts" / "patch-tts-server-push.py" + spec = _ilu.spec_from_file_location("tts_push_patch_under_test", script) + module = _ilu.module_from_spec(spec) # type: ignore[arg-type] + spec.loader.exec_module(module) # type: ignore[union-attr] + source = ( + "if message.sentence_id != self.conn.sentence_id:\n" + " continue\n" + ) + updated = module.patch_source(source) + self.assertIn("_dotty_server_push_sentence_ids", updated) + self.assertIn("continue", updated) + def test_no_ms_truncated_ids_left_anywhere(self): # The collision-prone id pattern must not reappear in either # patch surface (the audit found it at twelve sites). diff --git a/tests/test_docs_onboarding.py b/tests/test_docs_onboarding.py new file mode 100644 index 00000000..fcf79f77 --- /dev/null +++ b/tests/test_docs_onboarding.py @@ -0,0 +1,84 @@ +"""Regression guards for the documented StackChan provisioning path.""" + +from pathlib import Path + + +ROOT = Path(__file__).resolve().parents[1] +ONBOARDING_FILES = ( + ROOT / "docs" / "quickstart.md", + ROOT / "SETUP.md", + ROOT / "docs" / "troubleshooting.md", + ROOT / "Makefile", +) + + +def test_docs_do_not_claim_ota_url_is_in_on_device_advanced_options() -> None: + combined = "\n".join(path.read_text(encoding="utf-8") for path in ONBOARDING_FILES) + stale_directions = ( + "device's Advanced Options", + "robot's Advanced Options", + "Settings > Advanced Options", + ) + for stale in stale_directions: + assert stale not in combined + + +def test_docs_identify_compiled_ota_url_and_missing_device_editor() -> None: + quickstart = (ROOT / "docs" / "quickstart.md").read_text(encoding="utf-8") + assert "CONFIG_OTA_URL" in quickstart + assert "does **not** contain an Advanced" in quickstart + assert "fw-v1.3.3" in quickstart + + +def test_source_build_uses_the_pinned_dotty_firmware_fork() -> None: + setup = (ROOT / "SETUP.md").read_text(encoding="utf-8") + gitmodules = (ROOT / ".gitmodules").read_text(encoding="utf-8") + + assert "github.com/BrettKinny/StackChan.git" in gitmodules + assert "git clone --recursive https://github.com/BrettKinny/dotty-stackchan.git" in setup + assert "cd dotty-stackchan/firmware/firmware" in setup + assert "git clone https://github.com/m5stack/StackChan.git" not in setup + + +def test_docs_only_recommend_real_firmware_kconfig_symbols() -> None: + setup = (ROOT / "SETUP.md").read_text(encoding="utf-8") + + assert "CONFIG_OTA_URL" in setup + assert "set `CONFIG_WIFI_SSID`" not in setup + assert "set `CONFIG_WIFI_PASSWORD`" not in setup + assert "does **not** define `CONFIG_WIFI_SSID`" in setup + + +def test_docs_explain_the_real_first_boot_path() -> None: + setup = (ROOT / "SETUP.md").read_text(encoding="utf-8") + quickstart = (ROOT / "docs" / "quickstart.md").read_text(encoding="utf-8") + + assert "Skip" in setup + assert "Welcome!" in setup + assert "Xiaozhi" in setup and "hotspot" in setup.lower() + assert "Skip" in quickstart + + +def test_current_docs_do_not_claim_the_firmware_submodule_lacks_state_manager() -> None: + current_docs = ( + (ROOT / "README.md").read_text(encoding="utf-8") + + (ROOT / "docs" / "modes.md").read_text(encoding="utf-8") + + (ROOT / "CLAUDE.md").read_text(encoding="utf-8") + ) + stale_claims = ( + "submodule pin in this repo lags", + "submodule pin lags", + "flashing from the submodule won't give you Phase 4", + "does **not** include StateManager", + "xiaozhi firmware (built from m5stack/StackChan source)", + ) + for stale in stale_claims: + assert stale not in current_docs + + +def test_compatibility_doc_matches_the_current_public_contract() -> None: + compatibility = (ROOT / "COMPATIBILITY.md").read_text(encoding="utf-8") + assert "`fw-v1.3.3`" in compatibility + assert "the seven `dotty-pi-ext` voice tools" in compatibility + assert "No formal versioning is adopted yet" not in compatibility + assert "scripts/backup.sh" not in compatibility diff --git a/tests/test_dotty_av_audio_metrics.py b/tests/test_dotty_av_audio_metrics.py new file mode 100644 index 00000000..7172faed --- /dev/null +++ b/tests/test_dotty_av_audio_metrics.py @@ -0,0 +1,200 @@ +"""Decoded PCM quality checks. AI-assisted: OpenAI Codex (GPT-6).""" +import importlib.util +import math +from pathlib import Path +import struct +import sys +from types import SimpleNamespace + +import pytest + + +spec = importlib.util.spec_from_file_location("dotty_av_media", Path(__file__).parents[1] / "scripts/dotty_av_media.py") +media = importlib.util.module_from_spec(spec) +spec.loader.exec_module(media) + + +def pcm(values): + return struct.pack(f"<{len(values)}f", *values) + + +def test_silent_recording_warns_without_claiming_interaction(): + result = media.pcm_metrics(pcm([0] * 100)) + assert result["near_silence"] is True + assert result["rms_dbfs"] is None + assert result["clipping_ratio"] == 0 + assert "near_silence" in result["warnings"][0] + assert "verdict" not in result and "interaction" not in result + + +def test_known_amplitude_metrics(): + result = media.pcm_metrics(pcm([.5, -.5] * 100)) + assert result["rms"] == .5 + assert result["peak"] == .5 + assert result["rms_dbfs"] == pytest.approx(-6.0205999) + assert result["warnings"] == [] + + +def test_severe_clipping_strictly_above_point_one_percent(): + threshold = media.pcm_metrics(pcm([1.] + [.1] * 999)) + assert threshold["clipping_ratio"] == .001 + assert threshold["warnings"] == [] + severe = media.pcm_metrics(pcm([1., -1.] + [.1] * 998)) + assert severe["clipping_ratio"] == .002 + assert "severe_clipping" in severe["warnings"][0] + + +def test_lossy_overshoot_retained_and_explained(): + result = media.pcm_metrics(pcm([1.1, -1.2])) + assert result["peak_dbfs"] > 0 + assert result["clipped_samples"] == 2 + assert "Lossy" in result["note"] + + +@pytest.mark.parametrize("data", [b"", b"abc", pcm([math.nan]), pcm([math.inf])]) +def test_rejects_invalid_pcm(data): + with pytest.raises(ValueError): + media.pcm_metrics(data) + + +def test_decode_is_bounded_and_preserves_channels(monkeypatch): + commands = [] + + def decode(command, **kwargs): + commands.append(command) + assert kwargs["timeout"] == 45 + return pcm([.1, -.1]) + + monkeypatch.setattr(media.subprocess, "check_output", decode) + result = media.audio_metrics(Path("capture.mp4")) + assert commands[0][commands[0].index("-t") + 1] == "180" + assert "-ac" not in commands[0] and "-ar" not in commands[0] + assert commands[0][-1] == "pipe:1" + assert result["analysis_limit_seconds"] == 180 + + +def test_native_window_metrics_do_not_downmix_or_resample(monkeypatch): + commands = [] + monkeypatch.setattr(media.subprocess, "check_output", + lambda argv, **kwargs: commands.append(argv) or pcm([1.1, .1])) + result = media.audio_metrics(Path("capture.mp4"), start=2, duration=4.5) + assert commands[0][commands[0].index("-ss") + 1] == "2" + assert commands[0][commands[0].index("-t") + 1] == "4.5" + assert "-ac" not in commands[0] and "-ar" not in commands[0] + assert result["window_start_seconds"] == 2 + assert result["analysis_limit_seconds"] == 4.5 + assert result["clipping_ratio"] == .5 + assert "verdict" not in result + + +@pytest.mark.parametrize("start, duration", [(-1, 1), (0, 0), (0, 181), (179, 2), + (math.nan, 2), (0, math.inf), (True, 2)]) +def test_invalid_audio_metric_windows_fail_before_decode(monkeypatch, start, duration): + def unexpected(*args, **kwargs): + pytest.fail("invalid window reached decoder") + monkeypatch.setattr(media.subprocess, "check_output", unexpected) + with pytest.raises(ValueError, match="window"): + media.audio_metrics("unused", start=start, duration=duration) + + +def test_verify_adds_quality_without_replacing_continuity(monkeypatch): + monkeypatch.setattr(media, "probe", lambda _: {"streams": [ + {"codec_type": "video", "duration": "10", "width": 640, "height": 480}, + {"codec_type": "audio", "duration": "10", "sample_rate": "48000", "channels": 2}]}) + monkeypatch.setattr(media, "audio_continuity", lambda *_: { + "decoded_seconds": 10, "max_gap_seconds": 0}) + monkeypatch.setattr(media, "audio_metrics", lambda _: {"near_silence": True, "warnings": ["silent"]}) + result = media.verify("silent.mp4") + assert result["decoded_seconds"] == 10 + assert result["audio_quality"]["near_silence"] + assert "verdict" not in result + + +def test_gain_is_capped_and_peak_limited(): + quiet = media.pcm_metrics(pcm([.001, -.001] * 100 + [.01])) + gain = media.analysis_gain(quiet) + assert gain <= 20 + assert quiet["peak"] * 10 ** (gain / 20) <= 10 ** (-1 / 20) + 1e-6 + transient = media.pcm_metrics(pcm([.01] * 1000 + [.5])) + assert media.analysis_gain(transient) == pytest.approx(-1 - transient["peak_dbfs"]) + assert media.analysis_gain(media.pcm_metrics(pcm([1., .1]))) == 0 + assert media.analysis_gain(media.pcm_metrics(pcm([0.] * 100))) == 0 + + +def test_analysis_preserves_input_and_records_exact_gain(tmp_path, monkeypatch): + source = tmp_path / "original.wav" + source.write_bytes(b"original unchanged") + decoded = pcm([.02, -.03, .01]) + monkeypatch.setattr(media.subprocess, "check_output", lambda *a, **kw: decoded) + commands = [] + + def encode(command, **kwargs): + commands.append(command) + (tmp_path / "result.analysis.wav").write_bytes(kwargs["input"]) + + monkeypatch.setattr(media.subprocess, "run", encode) + derivative, record = media.prepare_analysis(source, tmp_path / "result.json") + assert source.read_bytes() == b"original unchanged" + assert derivative.name == "result.analysis.wav" + assert record["gain_db"] > 0 + assert record["analysis_quality"]["clipped_samples"] == 0 + assert record["source"] == str(source) + assert "-n" in commands[0] + with pytest.raises(FileExistsError): + media.prepare_analysis(source, tmp_path / "result.json") + + +def test_near_silence_never_invokes_model_or_becomes_pass(tmp_path, monkeypatch): + monkeypatch.setattr(media, "prepare_analysis", lambda *_: ( + tmp_path / "analysis.wav", {"original_quality": {"near_silence": True}})) + monkeypatch.setitem(sys.modules, "faster_whisper", SimpleNamespace()) + result = media.transcribe(tmp_path / "silence.wav", "unused", tmp_path / "out.json", normalize=True) + assert result["segments"] == [] + assert result["evidence_status"] == "INCONCLUSIVE" + assert "verdict" not in result + + +@pytest.mark.parametrize("threshold", [None, .3]) +def test_transcript_is_unverified_even_if_noise_hallucinates_text(tmp_path, monkeypatch, threshold): + calls = [] + + class FakeModel: + def __init__(self, *args, **kwargs): + assert kwargs["local_files_only"] is True + + def transcribe(self, path, **kwargs): + calls.append(kwargs) + return [SimpleNamespace(start=1, end=2, text="Thank you for watching", + no_speech_prob=.2, avg_logprob=-.4, words=[])], SimpleNamespace(language_probability=1) + + monkeypatch.setitem(sys.modules, "faster_whisper", SimpleNamespace(WhisperModel=FakeModel)) + options = {} if threshold is None else {"vad_threshold": threshold} + result = media.transcribe(tmp_path / "noise.wav", "local-model", tmp_path / "out.json", **options) + assert calls[0]["vad_filter"] is True + assert calls[0]["condition_on_previous_text"] is False + assert calls[0]["vad_parameters"] == {"threshold": .5 if threshold is None else threshold} + assert result["vad_parameters"] == calls[0]["vad_parameters"] + assert "initial_prompt" not in calls[0] + assert result["evidence_status"] == "UNVERIFIED_TRANSCRIPT" + assert "verdict" not in result + with pytest.raises(FileExistsError): + media.transcribe(tmp_path / "noise.wav", "local-model", tmp_path / "out.json") + + +@pytest.mark.parametrize("threshold", [-.1, 1.1, math.nan, math.inf, True, None, "invalid"]) +def test_invalid_vad_threshold_is_rejected_before_artifacts(tmp_path, threshold): + with pytest.raises(ValueError, match="VAD threshold"): + media.transcribe(tmp_path / "source.wav", "unused", tmp_path / "out.json", + normalize=True, vad_threshold=threshold) + assert list(tmp_path.iterdir()) == [] + + +@pytest.mark.parametrize("threshold", [0, .3, .5, 1]) +def test_silent_normalized_provenance_preserves_valid_threshold(tmp_path, monkeypatch, threshold): + monkeypatch.setattr(media, "prepare_analysis", lambda *_: ( + tmp_path / "analysis.wav", {"original_quality": {"near_silence": True}})) + result = media.transcribe(tmp_path / "source.wav", "unused", tmp_path / "out.json", + normalize=True, vad_threshold=threshold) + assert result["evidence_status"] == "INCONCLUSIVE" + assert result["segments"] == [] + assert result["vad_parameters"] == {"threshold": float(threshold)} diff --git a/tests/test_dotty_av_capture.py b/tests/test_dotty_av_capture.py new file mode 100644 index 00000000..80249ad0 --- /dev/null +++ b/tests/test_dotty_av_capture.py @@ -0,0 +1,149 @@ +"""Capture failure must never trigger speaker playback (Codex GPT-6 assisted).""" +import importlib.util +import os +from pathlib import Path +import subprocess +import shutil + +import pytest + +ROOT = Path(__file__).resolve().parents[1] +spec = importlib.util.spec_from_file_location("media", ROOT / "scripts/dotty_av_media.py") +media = importlib.util.module_from_spec(spec) +spec.loader.exec_module(media) + + +@pytest.mark.parametrize("streams", [[], [{"codec_type": "video", "duration": "1"}], + [{"codec_type": "video", "duration": "nan"}, {"codec_type": "audio", "duration": "1"}], + [{"codec_type": "video", "duration": "5"}, {"codec_type": "audio", "duration": "1"}]]) +def test_invalid_streams_rejected(monkeypatch, streams): + monkeypatch.setattr(media, "probe", lambda _: {"streams": streams}) + with pytest.raises(ValueError): + media.verify("unused") + + +def test_complete_but_truncated_capture_rejected(monkeypatch): + monkeypatch.setattr(media, "probe", lambda _: {"streams": [ + {"codec_type": "video", "duration": "5", "width": 1280, "height": 720}, + {"codec_type": "audio", "duration": "5", "sample_rate": "32000", "channels": 2}]}) + with pytest.raises(ValueError, match="truncated"): + media.verify("unused", expected=30) + + +def test_timestamp_stretched_audio_rejected(monkeypatch): + monkeypatch.setattr(media, "probe", lambda _: {"streams": [ + {"codec_type": "video", "duration": "36", "width": 1280, "height": 720}, + {"codec_type": "audio", "duration": "36", "sample_rate": "32000", "channels": 2}]}) + monkeypatch.setattr(media, "audio_continuity", lambda *_: { + "decoded_seconds": 2.14, "decoded_samples": 68544, "max_gap_seconds": 2.3}) + with pytest.raises(ValueError, match="audio sample loss"): + media.verify("unused") + + +def test_failed_camera_prevents_speaker_playback(tmp_path): + binaries = tmp_path / "bin" + binaries.mkdir() + programs = { + "pactl": "echo test-sink", + "espeak-ng": 'while [ "$1" != "-w" ]; do shift; done; shift; printf speech > "$1"', + "ffprobe": "echo 2", + "ffmpeg": "exit 42", + "pw-play": f'touch "{tmp_path / "PLAYED"}"', + } + for name, body in programs.items(): + path = binaries / name + path.write_text("#!/bin/sh\n" + body + "\n") + path.chmod(0o755) + result = subprocess.run(["bash", ROOT / "scripts/dotty-av-test.sh", "run", "hello", "5", tmp_path / "capture.mp4"], + env={**os.environ, "PATH": str(binaries) + ":" + os.environ["PATH"], + "XDG_RUNTIME_DIR": str(tmp_path), "DOTTY_AV_VOLUME": "20"}, + capture_output=True, text=True, timeout=10) + assert result.returncode != 0 + assert "capture failed before playback" in result.stderr + assert not (tmp_path / "PLAYED").exists() + + +def test_unbounded_recording_rejected(tmp_path): + result = subprocess.run(["bash", ROOT / "scripts/dotty-av-test.sh", "record", "181", tmp_path / "capture.mp4"], + capture_output=True, text=True, timeout=5) + assert result.returncode != 0 + assert not (tmp_path / "capture.mp4").exists() + + +def capture_function(name, *args, lossless="1", backend="pulse"): + # Load definitions only: no command dispatch, devices, or playback. + definitions = (ROOT / "scripts/dotty-av-test.sh").read_text().split('\ncmd=', 1)[0] + return subprocess.run(["bash", "-c", 'source /dev/stdin; "$@"', "fixture", name, + *map(str, args)], input=definitions, capture_output=True, text=True, timeout=5, + env={**os.environ, "DOTTY_AV_LOSSLESS_AUDIO": lossless, + "DOTTY_AV_AUDIO_BACKEND": backend}) + + +def test_lossless_sidecar_maps_same_audio_and_bounds_each_output(tmp_path): + output = tmp_path / "unique-case.mp4" + result = capture_function("capture_output_args", "7", output) + assert result.returncode == 0, result.stderr + argv = result.stdout.splitlines() + split = argv.index(str(output)) + mp4, wav = argv[:split], argv[split + 1:] + assert mp4[mp4.index("-t") + 1] == "7" + assert wav[wav.index("-t") + 1] == "7" + assert mp4[:4] == ["-map", "0:v:0", "-map", "1:a:0"] + assert wav[:2] == ["-map", "1:a:0"] + assert wav[wav.index("-c:a") + 1] == "pcm_f32le" + assert wav[wav.index("-ar") + 1] == "32000" + assert wav[wav.index("-ac") + 1] == "2" + assert wav[-1] == str(output.with_suffix(".capture.wav")) + + +def test_default_capture_has_no_lossless_output_or_float_input(tmp_path): + result = capture_function("capture_output_args", "7", tmp_path / "default.mp4", lossless="0") + assert result.returncode == 0, result.stderr + assert ".capture.wav" not in result.stdout and "pcm_f32le" not in result.stdout + inputs = capture_function("capture_args", lossless="0") + assert "pcm_f32le" not in inputs.stdout + + +def test_opt_in_pulse_requests_float_before_reading_input(): + result = capture_function("capture_args") + argv = result.stdout.splitlines() + pulse = argv.index("pulse") + assert argv.index("pcm_f32le") > pulse + assert argv.index("pcm_f32le") < len(argv) - 2 # before final -i source + assert "pcm_f32le" not in capture_function("capture_args", backend="alsa").stdout + + +@pytest.mark.parametrize("kind", ["record", "run"]) +def test_existing_sidecar_is_never_overwritten_or_played(tmp_path, kind): + sidecar = tmp_path / "capture.capture.wav" + sidecar.write_bytes(b"existing evidence") + argv = ["bash", ROOT / "scripts/dotty-av-test.sh", kind] + if kind == "run": + argv.append("fixture no playback") + argv.extend(["5", tmp_path / "capture.mp4"]) + result = subprocess.run(argv, capture_output=True, text=True, timeout=5, + env={**os.environ, "DOTTY_AV_LOSSLESS_AUDIO": "1"}) + assert result.returncode != 0 + assert "output already exists" in result.stderr + assert sidecar.read_bytes() == b"existing evidence" + assert not (tmp_path / "capture.mp4").exists() + + +def test_same_process_synthetic_float_sidecar_preserves_overshoots(tmp_path): + if not shutil.which("ffmpeg") or not shutil.which("ffprobe"): + pytest.skip("FFmpeg unavailable") + output = tmp_path / "synthetic.mp4" + rendered = capture_function("capture_output_args", "1", output) + assert rendered.returncode == 0, rendered.stderr + subprocess.run(["ffmpeg", "-nostdin", "-n", "-v", "error", + "-f", "lavfi", "-i", "color=size=160x120:rate=10", + "-f", "lavfi", "-i", "aevalsrc=1.05*sin(2*PI*440*t)|0.2*sin(2*PI*660*t):s=32000", + *rendered.stdout.splitlines()], check=True, capture_output=True, timeout=15) + sidecar = output.with_suffix(".capture.wav") + streams = media.probe(sidecar)["streams"] + assert len(streams) == 1 + assert streams[0]["codec_name"] == "pcm_f32le" + assert streams[0]["channels"] == 2 and streams[0]["sample_rate"] == "32000" + assert float(streams[0]["duration"]) == pytest.approx(1, abs=.01) + assert media.audio_metrics(sidecar)["peak"] > 1.04 + assert media.verify(output)["audio_seconds"] == pytest.approx(1, abs=.05) diff --git a/tests/test_dotty_av_clips.py b/tests/test_dotty_av_clips.py new file mode 100644 index 00000000..6dba059b --- /dev/null +++ b/tests/test_dotty_av_clips.py @@ -0,0 +1,142 @@ +"""Clip provenance regression checks. AI-assisted: OpenAI Codex (GPT-6).""" +import importlib.util +import json +from pathlib import Path +import shutil +from types import SimpleNamespace + +import pytest + + +spec = importlib.util.spec_from_file_location("dotty_av_clips", Path(__file__).parents[1] / "scripts/dotty_av_clips.py") +clips = importlib.util.module_from_spec(spec) +spec.loader.exec_module(clips) + + +@pytest.fixture +def case(tmp_path, monkeypatch): + folder = tmp_path / "cases" / "funny-1" + folder.mkdir(parents=True) + (folder / "raw.mp4").write_bytes(b"original evidence") + (folder / "result.json").write_text(json.dumps({ + "verdict": "PASS", "commit": "abc123", "response_offset": 8, + "started": "2026-09-12T23:59:59+10:00", "case": {"prompt": ""}})) + (folder / "response.json").write_text(json.dumps({"segments": [ + {"start": 2, "end": 4, "text": " Hello there! "}]})) + monkeypatch.setattr(clips, "media_duration", lambda _: 30) + commands = [] + monkeypatch.setattr(clips, "execute", lambda command: commands.append(command) or SimpleNamespace(stdout="")) + return tmp_path, folder, commands + + +def test_captions_align_response_then_trim(): + srt, cues = clips.captions({"segments": [{"start": 2, "end": 6, "text": "Test"}]}, 8, 11, 13) + assert "00:00:00,000 --> 00:00:02,000" in srt + assert cues[0]["source_start"] == 10 + assert cues[0]["source_end"] == 14 + + +@pytest.mark.parametrize("name", ["../other", "/tmp/video", "foo/bar", "", "a\\b"]) +def test_reject_path_identifiers(name): + with pytest.raises(ValueError): + clips.identifier(name) + + +def test_export_preserves_source_and_defaults_to_review(case): + session, folder, commands = case + destination = clips.export(session, "funny-1", name="take-one", start=9, end=14) + assert destination.parent.name == "needs-review" + assert (folder / "raw.mp4").read_bytes() == b"original evidence" + manifest = clips.load(destination / "manifest.json") + assert manifest["source"]["start_seconds"] == 9 + assert manifest["source"]["case_started"] == "2026-09-12T23:59:59+10:00" + assert manifest["commit"] == "abc123" + assert manifest["captions"]["state"] == "provisional" + assert "00:00:01,000 --> 00:00:03,000" in (destination / "captions.srt").read_text() + assert all("-n" in command for command in commands) + assert "<Hello>" in (destination / "index.html").read_text() + with pytest.raises(FileExistsError): + clips.export(session, "funny-1", name="take-one", start=9, end=14) + assert len(commands) == 2 + + +def test_case_symlink_cannot_escape_session(case, tmp_path): + session, _, _ = case + (session / "cases" / "escape").symlink_to(session.parent) + with pytest.raises(ValueError, match="escapes"): + clips.export(session, "escape") + + +@pytest.mark.parametrize("start,end", [(-1, 2), (5, 4), (0, 31), (float("nan"), 10)]) +def test_invalid_intervals_create_no_export(case, start, end): + session, _, _ = case + with pytest.raises(ValueError): + clips.export(session, "funny-1", start=start, end=end) + assert not (session / "clips").exists() + + +def test_failed_test_stays_labelled_even_with_review_gates(case): + session, folder, _ = case + result = clips.load(folder / "result.json") + result["verdict"] = "FAIL" + (folder / "result.json").write_text(json.dumps(result)) + destination = clips.export(session, "funny-1", reviewed=clips.GATES) + assert destination.parent.name == "failures" + + +def test_missing_response_offset_does_not_guess_caption_timing(case): + session, folder, _ = case + result = clips.load(folder / "result.json") + del result["response_offset"] + (folder / "result.json").write_text(json.dumps(result)) + destination = clips.export(session, "funny-1") + assert not (destination / "captions.srt").exists() + assert "missing response_offset" in clips.load(destination / "manifest.json")["captions"]["state"] + + +def test_index_is_new_snapshot_and_local_links(case): + session, _, _ = case + clips.export(session, "funny-1") + first, second = clips.index(session), clips.index(session) + assert first != second and first.exists() and second.exists() + assert "../cases/funny-1/raw.mp4" in first.read_text() + + +@pytest.mark.skipif(not shutil.which("ffmpeg") or not shutil.which("ffprobe"), reason="FFmpeg unavailable") +def test_synthetic_media_round_trip(tmp_path): + folder = tmp_path / "cases" / "synthetic" + folder.mkdir(parents=True) + # Generated patterns and tone only: no household media, microphone or speaker. + clips.execute(["ffmpeg", "-nostdin", "-n", "-v", "error", "-f", "lavfi", "-i", + "color=c=blue:s=320x240:r=10:d=2", "-f", "lavfi", "-i", + "sine=frequency=440:duration=2", "-c:v", "libx264", "-pix_fmt", "yuv420p", + "-c:a", "aac", "-shortest", str(folder / "raw.mp4")]) + (folder / "result.json").write_text(json.dumps({"verdict": "INCONCLUSIVE"})) + destination = clips.export(tmp_path, "synthetic") + assert clips.media_duration(destination / "clean.mp4") == pytest.approx(2, abs=.2) + streams = json.loads(clips.execute(["ffprobe", "-v", "error", "-show_streams", "-of", "json", + str(destination / "clean.mp4")]).stdout)["streams"] + video = next(s for s in streams if s["codec_type"] == "video") + assert (video["width"], video["height"]) == (1080, 1920) + assert (destination / "thumbnail.jpg").stat().st_size > 0 + + +# --- vertical crop for Shorts framing. AI-assisted: Claude. --- + +def test_crop_fills_the_vertical_frame_and_is_recorded(case): + session, _folder, commands = case + destination = clips.export(session, "funny-1", name="cropped", start=9, end=14, crop="405:720:368:0") + video_filter = commands[0][commands[0].index("-vf") + 1] + assert video_filter.startswith("crop=405:720:368:0,scale=1080:1920") + assert "pad=" not in video_filter + manifest = clips.load(destination / "manifest.json") + assert manifest["layout"] == "1080x1920 cropped from source region 405:720:368:0" + + +@pytest.mark.parametrize("crop", ["405x720", "405:720:368", "0:720:0:0", "720:405:0:0", "a:b:c:d", "405:720:-1:0"]) +def test_crop_must_be_a_portrait_source_region(case, crop): + session, _folder, commands = case + with pytest.raises(ValueError): + clips.export(session, "funny-1", name="bad-crop", start=9, end=14, crop=crop) + assert commands == [] + assert not list((session / "clips").glob("*/bad-crop")) diff --git a/tests/test_dotty_overnight.py b/tests/test_dotty_overnight.py new file mode 100644 index 00000000..3d554dea --- /dev/null +++ b/tests/test_dotty_overnight.py @@ -0,0 +1,745 @@ +"""Evidence and crash-safety regression tests. AI-assisted: OpenAI Codex (GPT-6).""" +import importlib.util +import hashlib +import json +from pathlib import Path +import sys + +import pytest + +SCRIPTS = Path(__file__).resolve().parents[1] / "scripts" +sys.path.insert(0, str(SCRIPTS)) +spec = importlib.util.spec_from_file_location("dotty_overnight", SCRIPTS / "dotty_overnight.py") +runner = importlib.util.module_from_spec(spec) +spec.loader.exec_module(runner) + + +def after(**state): + return {"errors": {}, "services_running": True, "perception": {"robot": { + "sensor_stale": False, "current_state": "idle", "listening": False, + "sensor_age_s": 0, "last_event_t": 1000, "last_chat_status_t": 1000, **state}}} + + +def transcript(text): + return {"segments": [{"text": text, "start": 0, "end": 2, + "no_speech_prob": .01, "avg_logprob": -.1}]} + + +def test_substring_is_not_correct_answer(): + assert not runner.matches("I know the answer", [["no"]]) + assert runner.matches("No, not necessarily.", [["no"]]) + + +@pytest.mark.parametrize("speech,logs,failure", [ + ({"segments": []}, "结果: name\nSentenceType.FIRST", "response_transcription_unavailable"), + (transcript("Dotty"), "", "no_asr_or_no_wake"), + (transcript("Dotty"), "结果: name", "no_tts"), + (transcript("other"), "结果: name\nSentenceType.FIRST", "response_mismatch"), +]) +def test_independent_evidence_required(speech, logs, failure): + result = runner.evaluate({"asr": [["name"]], "reply": [["Dotty"]]}, speech, + logs, 10, after(), "robot") + assert result["failure"] == failure + assert result["verdict"] == ("INCONCLUSIVE" if failure == "response_transcription_unavailable" else "FAIL") + + +def test_no_visual_evidence_means_no_overall_pass(): + result = runner.evaluate({"asr": [["name"]], "reply": [["Dotty"]]}, transcript("Dotty"), + "结果: name\nSentenceType.FIRST", 10, after(), "robot") + assert result["interaction"] == "PASS" + assert result["verdict"] == "INCONCLUSIVE" + + +@pytest.mark.parametrize("state", [after(current_state="talk", listening=True), after()]) +def test_ready_for_followup_accepts_ended_tts_and_fresh_resting_state(state): + result = runner.evaluate({"recovery_policy": "ready_for_followup", "reply": [["Dotty"]]}, + transcript("Dotty"), "结果: name\nSentenceType.FIRST\nSentenceType.LAST", + 10, state, "robot") + assert result["recovery"] == "PASS" + assert result["interaction"] == "PASS" + assert result["verdict"] == "INCONCLUSIVE" + + +@pytest.mark.parametrize("logs", ["SentenceType.FIRST", "SentenceType.LAST\nSentenceType.FIRST", + "SentenceType.LAST\n发送第一段语音:"]) +def test_ready_for_followup_rejects_missing_or_prior_turn_last(logs): + result = runner.evaluate({"recovery_policy": "ready_for_followup"}, transcript("Dotty"), + "结果: name\n" + logs, 10, + after(current_state="talk", listening=True), "robot") + assert result["recovery"] == "FAIL" + assert result["failure"] == "not_ready_for_followup" + + +@pytest.mark.parametrize("state", [ + after(current_state="talk", listening=False), + after(current_state="idle", listening=True), + after(current_state="sleep", listening=False), + after(sensor_stale=True), + {**after(), "services_running": False}, + {**after(), "errors": {"bridge": "unavailable"}}, +]) +def test_ready_for_followup_rejects_unhealthy_or_wrong_state(state): + result = runner.evaluate({"recovery_policy": "ready_for_followup"}, transcript("Dotty"), + "结果: name\nSentenceType.FIRST\nSentenceType.LAST", 10, state, "robot") + assert result["recovery"] == "FAIL" + + +def test_tool_requires_concrete_call_marker_not_spoken_name(): + result = runner.evaluate({"expected_tool": "think_hard"}, transcript("I used think_hard"), + "结果: please use think_hard\nSentenceType.FIRST", 10, after(), "robot") + assert result["expected_tool_call"] == "INCONCLUSIVE" + assert result["interaction"] == "INCONCLUSIVE" + assert result["verdict"] == "INCONCLUSIVE" + + +def test_tool_marker_is_exact_and_does_not_prove_execution(): + logs = "结果: think\nSentenceType.FIRST\nPiClient: tool call name=think_hard id=call_123" + result = runner.evaluate({"expected_tool": "think_hard"}, transcript("No"), logs, 10, after(), "robot") + assert result["expected_tool_call"] == "PASS" + assert result["tool_execution"] == "INCONCLUSIVE" + other = runner.evaluate({"expected_tool": "think"}, transcript("No"), logs, 10, after(), "robot") + assert other["expected_tool_call"] == "INCONCLUSIVE" + + +def test_missing_tool_marker_and_words_cannot_establish_silence(): + result = runner.evaluate({"expected_tool": "think_hard"}, {"segments": []}, + "结果: think\nSentenceType.FIRST", 10, after(), "robot") + assert result["failure"] == "response_transcription_unavailable" + assert result["interaction"] == "INCONCLUSIVE" + + +def test_response_word_limit_is_strict_without_asserting_creative_quality(): + logs = "结果: joke\nSentenceType.FIRST" + result = runner.evaluate({"max_response_words": 3}, transcript("I'm a robot"), logs, 10, after(), "robot") + assert result["response_word_count"] == 3 + assert result["response_length"] == "PASS" + assert result["semantic_quality"] == "INCONCLUSIVE" + assert result["verdict"] == "INCONCLUSIVE" + longer = runner.evaluate({"max_response_words": 3}, transcript("I'm a tiny robot"), logs, 10, after(), "robot") + assert longer["response_word_count"] == 4 + assert longer["failure"] == "response_too_long" + assert longer["verdict"] == "FAIL" + + +@pytest.mark.parametrize("limit", [0, -1, True, "3", 3.5]) +def test_invalid_word_limit_is_rejected(limit): + with pytest.raises(ValueError, match="positive integer"): + runner.evaluate({"max_response_words": limit}, transcript("hello"), + "结果: hello\nSentenceType.FIRST", 10, after(), "robot") + + +def test_emoji_stripped_tts_does_not_imply_bad_expression(): + result = runner.evaluate({}, transcript("Dotty"), "结果: name\n发送第一段语音: Dotty", + 10, after(), "robot") + assert result["expression"] == "INCONCLUSIVE" + assert result["visual"] == "INCONCLUSIVE" + + +def receive_frame(frame): + return ("260912 22:24:45[0.9.3_SiWhPiLononoCh][core.handle.textMessageProcessor]-INFO-收到" + + frame["type"] + "消息:" + json.dumps(frame)) + + +@pytest.mark.parametrize("frame", [ + {"session_id": "fixture", "type": "listen", "state": "detect", "text": "Hi ESP"}, + {"type": "event", "name": "wake_word_detected", "data": {"phrase": "Hi ESP"}}, +]) +def test_real_received_wake_frame_proves_wake_with_idle_precondition(frame): + before = {**after(), "devices": ["robot"]} + result = runner.evaluate({"require_wake_event": True}, transcript("Dotty"), + receive_frame(frame) + "\n结果: name\nSentenceType.FIRST", + 10, after(), "robot", before=before) + assert result["wake_event"] == "PASS" + assert result["wake_precondition"] == "PASS" + assert result["interaction"] == "PASS" + + +@pytest.mark.parametrize("false_evidence", [ + receive_frame({"type": "listen", "state": "start", "mode": "auto"}), + receive_frame({"type": "event", "name": "state_changed", "data": {"state": "talk"}}), + '结果: {"type":"listen","state":"detect"}', + 'wake_word_detected', + 'PiClient: text wake_word_detected', + '[core.handle.textMessageProcessor]-INFO-收到listen消息:{"type":"event","name":"wake_word_detected"}', + '[core.handle.textMessageProcessor]-INFO-收到listen消息:{bad json}', +]) +def test_existing_listening_or_keyword_mentions_do_not_prove_wake(false_evidence): + result = runner.evaluate({"require_wake_event": True}, transcript("Dotty"), + false_evidence + "\n结果: name\nSentenceType.FIRST", 10, after(), "robot", + before={**after(), "devices": ["robot"]}) + assert result["failure"] == "no_wake_event" + assert result["interaction"] == "FAIL" + + +@pytest.mark.parametrize("before", [ + None, + {**after(current_state="talk", listening=True), "devices": ["robot"]}, + {**after(sensor_stale=True), "devices": ["robot"]}, + {**after(), "devices": ["robot", "other"]}, +]) +def test_absent_wake_precondition_blocks_even_if_wake_frame_seen(before): + result = runner.evaluate({"require_wake_event": True}, transcript("Dotty"), + receive_frame({"type": "listen", "state": "detect"}) + + "\n结果: name\nSentenceType.FIRST", 10, after(), "robot", before=before) + assert result["verdict"] == "BLOCKED" + assert result["failure"] == "wake_precondition" + + +def test_warm_followup_does_not_require_cold_wake(): + result = runner.evaluate({"recovery_policy": "ready_for_followup"}, transcript("Dotty"), + "结果: name\nSentenceType.FIRST\nSentenceType.LAST", 10, + after(current_state="talk", listening=True), "robot", + before={**after(current_state="talk", listening=True), "devices": ["robot"]}) + assert result["interaction"] == "PASS" + assert result["wake_event"] == "NOT_REQUIRED" + + +def test_run_case_blocks_before_playback_without_changing_robot(tmp_path, monkeypatch): + commands = [] + def fake_command(argv, **kwargs): + commands.append(argv) + assert argv == ["git", "rev-parse", "HEAD"] + return "fixture-commit\n" + monkeypatch.setattr(runner, "command", fake_command) + monkeypatch.setattr(runner, "snapshot", lambda host: { + **after(current_state="talk", listening=True), "devices": ["robot"]}) + result = runner.run_case(tmp_path, {"host": "unused", "device": "robot"}, + {"id": "wake", "require_wake_event": True}) + assert result["verdict"] == "BLOCKED" + assert result["failure"] == "wake_precondition" + assert result["capture"] == "INCONCLUSIVE" + assert len(commands) == 1 + + +@pytest.fixture +def prompt_capture_boundary(monkeypatch): + """Stop at the subprocess boundary; never touch real playback or hardware.""" + captured = [] + def fake_command(argv, **kwargs): + if argv == ["git", "rev-parse", "HEAD"]: + return "fixture-commit\n" + return {("pactl", "get-default-sink"): "fixture-speaker\n", + ("pactl", "get-sink-volume", "fixture-speaker"): "Volume: 35%\n", + ("pactl", "get-sink-mute", "fixture-speaker"): "Mute: no\n"}[tuple(argv)] + def capture(argv, log, session, timeout, env, deadline): + captured.append({"argv": argv, "env": env, + "evidence_at_launch": json.loads((Path(log).parent / "result.json").read_text())}) + raise RuntimeError("fixture_capture_stop") + monkeypatch.setattr(runner, "command", fake_command) + monkeypatch.setattr(runner, "snapshot", lambda host: {**after(), "devices": ["robot"]}) + monkeypatch.setattr(runner, "host_clock", lambda host: {"offset_seconds": 0, "uncertainty_seconds": 0}) + monkeypatch.setattr(runner, "guarded_process", capture) + return {"host": "unused", "device": "robot", "sink": "fixture-speaker", "volume": 35, + "stop_prompts_at": "2099-01-01T00:00:00+00:00"}, captured + + +def test_explicit_warm_state_case_blocks_before_capture_when_idle(tmp_path, prompt_capture_boundary): + config, captured = prompt_capture_boundary + result = runner.run_case(tmp_path, config, {"id": "sleep", "prompt": "Go to sleep.", + "require_warm_listening": True}) + assert result["verdict"] == "BLOCKED" + assert result["failure"] == "warm_listening_precondition" + assert result["warm_listening_precondition"] == "BLOCKED" + assert result["capture"] == "INCONCLUSIVE" + assert captured == [] + + +def test_fresh_state_event_cannot_revalidate_old_listening_status( + tmp_path, monkeypatch, prompt_capture_boundary): + config, captured = prompt_capture_boundary + monkeypatch.setattr(runner, "snapshot", lambda host: { + **after(current_state="idle", listening=True, last_event_t=1500, + last_chat_status_t=1000, sensor_age_s=1), "devices": ["robot"]}) + result = runner.run_case(tmp_path, config, {"id": "sleep", "prompt": "Go to sleep.", + "require_warm_listening": True}) + assert result["verdict"] == "BLOCKED" + assert result["failure"] == "warm_listening_precondition" + assert result["warm_listening_status_age_seconds"] == 501 + assert captured == [] + + +@pytest.mark.parametrize("state", ["idle", "talk", "story_time"]) +def test_explicit_warm_precondition_requires_listening_not_talk_state( + tmp_path, monkeypatch, prompt_capture_boundary, state): + config, captured = prompt_capture_boundary + monkeypatch.setattr(runner, "snapshot", lambda host: { + **after(current_state=state, listening=True), "devices": ["robot"]}) + result = runner.run_case(tmp_path, config, {"id": "sleep", "prompt": "Go to sleep.", + "require_warm_listening": True}) + assert result["failure"] == "fixture_capture_stop" + assert result["warm_listening_precondition"] == "PASS" + assert len(captured) == 1 + + +def test_window_metrics_keep_prompt_and_response_separate_without_verdict(tmp_path, monkeypatch): + media = tmp_path / "raw.mp4" + sidecar = tmp_path / "raw.capture.wav" + sidecar.write_bytes(b"fixture") + calls = [] + def metrics(path, start, duration): + calls.append((Path(path).name, start, duration)) + return {"clipping_ratio": .006 if start == 0 else 0, + "channel_policy": "native; no downmix or resampling"} + monkeypatch.setattr(runner, "audio_metrics", metrics) + result = runner.capture_window_metrics(media, 53, 6.4, 6.8) + assert calls == [("raw.mp4", 0, 6.4), ("raw.mp4", 6.8, 46.2), + ("raw.capture.wav", 0, 6.4), ("raw.capture.wav", 6.8, 46.2)] + assert result["container_audio"]["prompt"]["clipping_ratio"] == .006 + assert result["container_audio"]["response"]["clipping_ratio"] == 0 + assert result["lossless_capture"]["source"] == str(sidecar) + assert "verdict" not in result and "interaction" not in result + assert "not sample-aligned" in result["timing_note"] + + +@pytest.mark.parametrize("mapping", [None, {}, {"other-case": {"path": "/unrelated/prompt.wav"}}]) +def test_unconfigured_case_cannot_inherit_another_prompt_wav(tmp_path, monkeypatch, prompt_capture_boundary, mapping): + config, captured = prompt_capture_boundary + if mapping is not None: + config["prompt_wavs"] = mapping + monkeypatch.setenv("DOTTY_AV_PROMPT_WAV", "/unrelated/prompt.wav") + result = runner.run_case(tmp_path, config, {"id": "arithmetic", "prompt": "Twelve plus seven?"}) + assert result["failure"] == "fixture_capture_stop" + assert "DOTTY_AV_PROMPT_WAV" not in captured[0]["env"] + assert result["prompt_provenance"] == {"source": "harness_default", "renderer": "espeak-ng", + "text": "Twelve plus seven?"} + + +@pytest.mark.parametrize("relative", [False, True]) +def test_configured_prompt_is_bound_to_case_and_persisted_before_capture(tmp_path, monkeypatch, prompt_capture_boundary, relative): + config, captured = prompt_capture_boundary + wav = tmp_path / "original.wav" + wav.write_bytes(b"fixture waveform bytes") + entry = {"path": wav.name if relative else str(wav), "text": "Twelve plus seven?", + "sha256": hashlib.sha256(wav.read_bytes()).hexdigest(), "renderer": "local Piper fixture"} + config["prompt_wavs"] = {"arithmetic": entry} + monkeypatch.setenv("DOTTY_AV_PROMPT_WAV", "/unrelated/prompt.wav") + result = runner.run_case(tmp_path, config, {"id": "arithmetic", "prompt": entry["text"]}) + assert result["failure"] == "fixture_capture_stop" + selected = Path(captured[0]["env"]["DOTTY_AV_PROMPT_WAV"]) + assert selected != wav # Freeze per-attempt evidence against later source edits. + assert selected.read_bytes() == wav.read_bytes() + assert result["prompt_provenance"] == {**entry, "source_path": str(wav), + "path": str(selected), "source": "configured_wav"} + assert captured[0]["evidence_at_launch"]["prompt_provenance"] == result["prompt_provenance"] + wav.write_bytes(b"later replacement waveform") + assert hashlib.sha256(selected.read_bytes()).hexdigest() == entry["sha256"] + + +@pytest.mark.parametrize("replacement, failure", [ + ({"text": "Twelve plus ten?"}, "prompt_wav_text_mismatch"), + ({"text": "Twelve plus seven? "}, "prompt_wav_text_mismatch"), + ({"sha256": "0" * 64}, "prompt_wav_sha256_mismatch"), +]) +def test_mismatched_prompt_cannot_reach_capture(tmp_path, prompt_capture_boundary, replacement, failure): + config, captured = prompt_capture_boundary + wav = tmp_path / "prompt.wav" + wav.write_bytes(b"fixture waveform bytes") + config["prompt_wavs"] = {"arithmetic": { + "path": str(wav), "text": "Twelve plus seven?", + "sha256": hashlib.sha256(wav.read_bytes()).hexdigest(), "renderer": "fixture", **replacement}} + result = runner.run_case(tmp_path, config, {"id": "arithmetic", "prompt": "Twelve plus seven?"}) + assert result["failure"] == failure + assert result["capture"] == "INCONCLUSIVE" + assert captured == [] + + +@pytest.mark.parametrize("defect", ["missing_file", "empty_file", "path", "text", "sha256", "renderer", + "blank_renderer", "invalid_hash", "entry_not_object", "map_not_object"]) +def test_invalid_prompt_manifest_rejected_before_capture(tmp_path, prompt_capture_boundary, defect): + config, captured = prompt_capture_boundary + wav = tmp_path / "prompt.wav" + if defect != "missing_file": + wav.write_bytes(b"" if defect == "empty_file" else b"fixture waveform bytes") + entry = {"path": str(wav), "text": "Twelve plus seven?", + "sha256": hashlib.sha256(b"fixture waveform bytes").hexdigest(), "renderer": "fixture"} + if defect in entry: + del entry[defect] + elif defect == "blank_renderer": + entry["renderer"] = " " + elif defect == "invalid_hash": + entry["sha256"] = "not-a-sha256" + config["prompt_wavs"] = {"arithmetic": None if defect == "entry_not_object" else entry} + if defect == "map_not_object": + config["prompt_wavs"] = [] + result = runner.run_case(tmp_path, config, {"id": "arithmetic", "prompt": "Twelve plus seven?"}) + assert result["verdict"] == "FAIL" + assert result["failure"].startswith("prompt_wav") + assert result["capture"] == "INCONCLUSIVE" + assert captured == [] + + +@pytest.mark.parametrize("threshold", [None, .25, 0, 1]) +def test_response_vad_threshold_does_not_change_prompt_transcription(tmp_path, monkeypatch, prompt_capture_boundary, threshold): + config, _ = prompt_capture_boundary + expected = .5 if threshold is None else float(threshold) + if threshold is not None: + config["response_vad_threshold"] = threshold + config.update(python="fixture-python", model="fixture-model", + finish_at="2099-01-01T01:00:00+00:00") + boundary_command = runner.command + monkeypatch.setattr(runner, "command", lambda argv, **kwargs: + "" if argv[0] == "ffmpeg" else boundary_command(argv, **kwargs)) + monkeypatch.setattr(runner, "verify", lambda *args, **kwargs: {"audio_seconds": 40, "video_seconds": 40}) + monkeypatch.setattr(runner, "audio_metrics", lambda *args, **kwargs: {"fixture": True}) + def logs(host, start, end, directory): + (directory / "xiaozhi-esp32-server.log").write_text("结果: Twelve plus seven\nSentenceType.FIRST") + monkeypatch.setattr(runner, "read_log_window", logs) + transcriptions = {} + def process(argv, log, session, timeout, env=None, deadline=None): + directory = Path(log).parent + if "transcribe" not in argv: + for name, second in (("recording-start", 0), ("playback-start", 2), ("playback-end", 5)): + (directory / f"raw.{name}.txt").write_text(f"2026-09-12T12:00:0{second}+00:00") + return + label = Path(argv[3]).stem + transcriptions[label] = argv + threshold = float(argv[argv.index("--vad-threshold") + 1]) if "--vad-threshold" in argv else .5 + output = Path(argv[argv.index("--output") + 1]) + runner.write_json(output, {**transcript("nineteen" if label == "response" else "Twelve plus seven"), + "vad_parameters": {"threshold": threshold}, "vad_filter": True, + "analysis": {"enabled": "--normalize" in argv}, "model": "fixture-model"}) + monkeypatch.setattr(runner, "guarded_process", process) + result = runner.run_case(tmp_path, config, {"id": "arithmetic", "prompt": "Twelve plus seven?", + "asr": [["twelve"], ["seven"]], "reply": [["nineteen"]]}) + assert result["interaction"] == "PASS", result + response = transcriptions["response"] + assert float(response[response.index("--vad-threshold") + 1]) == expected + assert "--vad-threshold" not in transcriptions["prompt-heard"] + assert result["response_vad_threshold"] == expected + assert result["transcription_analysis"]["response"]["vad_parameters"] == {"threshold": expected} + assert result["transcription_analysis"]["prompt-heard"]["vad_parameters"] == {"threshold": .5} + + +@pytest.mark.parametrize("threshold", [-.1, 1.1, float("nan"), float("inf"), -float("inf"), + True, False, None, ".25", [], {}]) +def test_invalid_response_vad_threshold_rejected_before_capture(tmp_path, prompt_capture_boundary, threshold): + config, captured = prompt_capture_boundary + config["response_vad_threshold"] = threshold + result = runner.run_case(tmp_path, config, {"id": "arithmetic", "prompt": "Twelve plus seven?"}) + assert result["failure"] == "response_vad_threshold_must_be_finite_number_between_0_and_1" + assert result["capture"] == "INCONCLUSIVE" + assert captured == [] + + +@pytest.mark.parametrize("arguments, expected", [([], .5), (["--response-vad-threshold", "0.25"], .25)]) +def test_init_persists_response_vad_setting(tmp_path, monkeypatch, arguments, expected): + monkeypatch.setattr(sys, "argv", ["dotty_overnight.py", "init", "--session", str(tmp_path), + "--host", "fixture@fixture-host", *arguments]) + monkeypatch.setattr(runner, "admin", lambda host, route: {"devices": ["fixture-robot"]}) + monkeypatch.setattr(runner, "command", lambda *args, **kwargs: "fixture-speaker\n") + runner.main() + assert runner.load(tmp_path / "config.json")["response_vad_threshold"] == expected + + +def test_init_requires_explicit_deployment_host(tmp_path, monkeypatch): + monkeypatch.delenv("DOTTY_TEST_HOST", raising=False) + session = tmp_path / "not-created" + monkeypatch.setattr(sys, "argv", ["dotty_overnight.py", "init", "--session", str(session)]) + monkeypatch.setattr(runner, "admin", lambda *args: pytest.fail("must not contact an assumed host")) + with pytest.raises(SystemExit) as error: + runner.main() + assert error.value.code == 2 + assert not session.exists() + + +def test_init_accepts_explicit_host_environment(tmp_path, monkeypatch): + monkeypatch.setenv("DOTTY_TEST_HOST", "fixture@fixture-host") + monkeypatch.setattr(sys, "argv", ["dotty_overnight.py", "init", "--session", str(tmp_path)]) + seen = [] + def admin(host, route): + seen.append((host, route)) + return {"devices": ["fixture-robot"]} + monkeypatch.setattr(runner, "admin", admin) + monkeypatch.setattr(runner, "command", lambda *args, **kwargs: "fixture-speaker\n") + runner.main() + assert seen == [("fixture@fixture-host", "devices")] + assert runner.load(tmp_path / "config.json")["host"] == "fixture@fixture-host" + + +def test_preflight_reports_response_vad_setting(tmp_path, monkeypatch): + monkeypatch.setattr(runner, "snapshot", lambda host: after()) + monkeypatch.setattr(runner, "space_ok", lambda session: True) + result = runner.preflight(tmp_path, {"host": "unused", "model": str(tmp_path), "response_vad_threshold": .25}) + assert result["response_vad_threshold"] == .25 + + +def test_preflight_rejects_invalid_vad_before_external_checks(tmp_path, monkeypatch): + def no_snapshot(*args): + pytest.fail("invalid config must not access the live host") + monkeypatch.setattr(runner, "snapshot", no_snapshot) + with pytest.raises(ValueError, match="response_vad_threshold"): + runner.preflight(tmp_path, {"host": "unused", "response_vad_threshold": float("nan")}) + + +@pytest.mark.parametrize("command, threshold", [("init", "nan"), ("init", "inf"), ("init", "-0.1"), + ("init", "1.1"), ("run", ".25"), ("resume", ".25")]) +def test_vad_cli_cannot_silently_override_existing_config(tmp_path, monkeypatch, command, threshold): + session = tmp_path / "not-created" + monkeypatch.setattr(sys, "argv", ["dotty_overnight.py", command, "--session", str(session), + "--response-vad-threshold", threshold]) + with pytest.raises(SystemExit) as error: + runner.main() + assert error.value.code == 2 + assert not session.exists() + + +def test_blocked_prerequisite_does_not_increment_failure_threshold(tmp_path): + checkpoint = runner.restore_checkpoint(tmp_path) + runner.record_attempt(checkpoint, {"id": "wake"}, + {"verdict": "BLOCKED", "failure": "wake_precondition"}) + assert checkpoint["blocked"] == {"wake": "wake_precondition"} + assert checkpoint["counts"]["wake"]["failures"] == 0 + assert checkpoint["consecutive_failures"] == 0 + assert checkpoint["completed"] == [] + + +def test_cannot_pool_unrelated_asr_requests(): + result = runner.evaluate({"asr": [["purple"], ["robot"]]}, transcript("yes"), + "结果: purple\n结果: robot\nSentenceType.FIRST", 10, after(), "robot") + assert result["failure"] == "asr_mismatch" + + +def test_missing_listening_state_is_not_recovery(): + state = after() + del state["perception"]["robot"]["listening"] + result = runner.evaluate({"asr": [["name"]]}, transcript("Dotty"), + "结果: name\nSentenceType.FIRST", 10, state, "robot") + assert result["recovery"] == "FAIL" + + +def test_low_confidence_hallucination_is_not_audible_response(): + speech = transcript("Dotty") + speech["segments"][0]["avg_logprob"] = -2 + result = runner.evaluate({}, speech, "结果: name\nSentenceType.FIRST", 10, after(), "robot") + assert result["failure"] == "response_transcription_unavailable" + assert result["interaction"] == "INCONCLUSIVE" + + +def test_playback_window_crosses_midnight(tmp_path): + for name, stamp in { + "recording-start": "2026-09-12T23:59:56+00:00", + "playback-start": "2026-09-12T23:59:59+00:00", + "playback-end": "2026-09-13T00:00:05+00:00", + }.items(): + (tmp_path / f"raw.{name}.txt").write_text(stamp) + assert runner.playback_window(tmp_path, 30) == (3, 9.4) + with pytest.raises(ValueError, match="truncated"): + runner.playback_window(tmp_path, 9) + + +def test_host_timestamp_preserves_date_and_offset(): + assert runner.host_timestamp("2026-09-12T23:59:59+00:00", {"offset_seconds": 2}) == "2026-09-13T00:00:01+00:00" + + +def test_interrupted_case_quarantined_before_active_cleared(tmp_path): + result_path = tmp_path / "cases" / "unique-attempt" / "result.json" + runner.write_json(result_path, {"case_id": "memory-write", "verdict": "INCONCLUSIVE"}) + runner.write_json(tmp_path / "active.json", {"case": "unique-attempt"}) + checkpoint = runner.restore_checkpoint(tmp_path) + assert checkpoint["completed"] == [] + assert checkpoint["quarantined"] == ["memory-write"] + assert runner.load(tmp_path / "active.json") == {} + assert runner.load(result_path)["failure"] == "interrupted_case_requires_review" + assert runner.restore_checkpoint(tmp_path)["quarantined"] == ["memory-write"] + + +def test_unknown_interrupted_identity_blocks_replay(tmp_path): + runner.write_json(tmp_path / "active.json", {"case": "missing-attempt"}) + with pytest.raises(RuntimeError, match="identity unknown"): + runner.restore_checkpoint(tmp_path) + assert runner.load(tmp_path / "active.json") + + +def test_checkpointed_attempt_not_quarantined_after_clear_crash(tmp_path): + runner.write_json(tmp_path / "checkpoint.json", {"last_attempt": "finished-attempt"}) + runner.write_json(tmp_path / "active.json", {"case": "finished-attempt", "case_id": "identity"}) + checkpoint = runner.restore_checkpoint(tmp_path) + assert checkpoint["quarantined"] == [] + assert runner.load(tmp_path / "active.json") == {} + + +def test_success_requires_streak_and_all_evidence(tmp_path): + checkpoint = runner.restore_checkpoint(tmp_path) + case = {"id": "identity"} + good = dict.fromkeys(("capture", "playback", "interaction", "recovery"), "PASS") + runner.record_attempt(checkpoint, case, good) + runner.record_attempt(checkpoint, case, {**good, "playback": "INCONCLUSIVE"}) + runner.record_attempt(checkpoint, case, good) + runner.record_attempt(checkpoint, case, good) + assert checkpoint["completed"] == [] + runner.record_attempt(checkpoint, case, good) + assert checkpoint["completed"] == ["identity"] + + +def test_failure_threshold_quarantines_and_pauses(tmp_path): + checkpoint = runner.restore_checkpoint(tmp_path) + for _ in range(3): + runner.record_attempt(checkpoint, {"id": "identity"}, {"verdict": "FAIL"}) + assert checkpoint["consecutive_failures"] == 3 + assert checkpoint["quarantined"] == ["identity"] + assert checkpoint["completed"] == [] + + +def test_operator_interruption_is_not_automatically_retried(tmp_path): + checkpoint = runner.restore_checkpoint(tmp_path) + runner.record_attempt(checkpoint, {"id": "memory-write"}, {"exception": "InterruptedError"}) + assert checkpoint["quarantined"] == ["memory-write"] + + +def test_device_lock_prevents_two_output_sessions(): + config = {"device": "pytest-lock-only-no-hardware"} + with runner.device_lock(config): + with pytest.raises(BlockingIOError): + with runner.device_lock(config): + pass + + +def test_guarded_process_does_not_leave_child_on_stop(tmp_path, monkeypatch): + monkeypatch.setattr(runner, "STOP", True) + with pytest.raises(InterruptedError): + runner.guarded_process([sys.executable, "-c", "import time; time.sleep(30)"], + tmp_path / "process.log", tmp_path, 60) + + +def test_deadline_prevents_launch(tmp_path): + with pytest.raises(TimeoutError, match="before process launch"): + runner.guarded_process(["this-command-must-never-start"], tmp_path / "log", + tmp_path, 60, deadline=0) + assert not (tmp_path / "log").exists() + + +# --- #182: unattended operation and evaluator false alarms. AI-assisted: Claude. --- + +def test_reference_mic_mishearing_is_inconclusive_when_service_tts_matches(): + logs = "结果: name\n发送音频消息: SentenceType.FIRST, My name is Dotty\nSentenceType.LAST" + result = runner.evaluate({"asr": [["name"]], "reply": [["Dotty"]]}, transcript("Your name is Dadi"), + logs, 2, after(), "robot") + assert result["failure"] == "response_transcript_disagrees_with_tts" + assert result["interaction"] == result["verdict"] == "INCONCLUSIVE" + assert result["service_tts_text"] == "My name is Dotty" + + +def test_wrong_reply_still_fails_when_service_tts_also_lacks_the_answer(): + logs = "结果: name\n发送音频消息: SentenceType.FIRST, I like trains\nSentenceType.LAST" + result = runner.evaluate({"asr": [["name"]], "reply": [["Dotty"]]}, transcript("I like trains"), + logs, 2, after(), "robot") + assert result["failure"] == "response_mismatch" + assert result["interaction"] == "FAIL" + + +def test_extra_recognised_speech_makes_the_case_inconclusive(): + logs = ("结果: what is your name\n识别文本: what is your name\n" + "发送音频消息: SentenceType.FIRST, My name is Dotty\nSentenceType.LAST\n" + "结果: I made a house\n识别文本: I made a house\n") + result = runner.evaluate({"asr": [["name"]], "reply": [["Dotty"]]}, transcript("My name is Dotty"), + logs, 2, after(), "robot") + assert result["extra_speech"] == ["I made a house"] + assert result["failure"] == "extra_speech_in_capture" + assert result["interaction"] == result["verdict"] == "INCONCLUSIVE" + + +def test_rejected_or_empty_recognitions_are_not_extra_speech(): + logs = ("结果: what is your name\n识别文本: what is your name\n" + "发送音频消息: SentenceType.FIRST, My name is Dotty\nSentenceType.LAST\n" + "结果: Thank you.\nASR-REJECT no_speech=0.811 > 0.60 | dropped 'Thank you.'\n结果: \n") + result = runner.evaluate({"asr": [["name"]], "reply": [["Dotty"]]}, transcript("My name is Dotty"), + logs, 2, after(), "robot") + assert result["extra_speech"] == [] + assert result["failure"] is None + + +@pytest.mark.parametrize("heard,asr_ok,expected", [ + ("repeat purple robot", True, ("PASS", "reference_mic")), + ("repeat purple brisket", True, ("PASS", "robot_asr")), + ("repeat purple brisket", False, ("INCONCLUSIVE", None)), + ("", False, ("INCONCLUSIVE", None)), +]) +def test_playback_is_proven_by_reference_mic_or_by_the_robot_hearing_it(heard, asr_ok, expected): + assert runner.playback_verdict(heard, asr_ok, [["robot"]]) == expected + + +def test_inconclusive_attempt_resets_streak_without_counting_as_failure(tmp_path): + checkpoint = runner.restore_checkpoint(tmp_path) + good = dict.fromkeys(("capture", "playback", "interaction", "recovery"), "PASS") + runner.record_attempt(checkpoint, {"id": "identity"}, good) + for _ in range(4): + runner.record_attempt(checkpoint, {"id": "identity"}, {**good, "interaction": "INCONCLUSIVE"}) + counts = checkpoint["counts"]["identity"] + assert (counts["streak"], counts["failures"]) == (0, 0) + assert checkpoint["consecutive_failures"] == 0 + assert checkpoint["quarantined"] == [] + + +def test_mic_opener_refreshes_a_stale_listening_status_before_capture( + tmp_path, monkeypatch, prompt_capture_boundary): + config, captured = prompt_capture_boundary + calls = [] + snapshots = iter([ + {**after(), "devices": ["robot"]}, # idle, mic closed + {**after(current_state="talk", listening=True, sensor_age_s=1), "devices": ["robot"]}, # too fresh + {**after(current_state="talk", listening=True, sensor_age_s=5), "devices": ["robot"]}, # settled + ]) + monkeypatch.setattr(runner, "snapshot", lambda host: next(snapshots)) + monkeypatch.setattr(runner, "admin", lambda host, route, body=None: calls.append((route, body)) or {"ok": True}) + monkeypatch.setattr(runner.time, "sleep", lambda seconds: None) + result = runner.run_case(tmp_path, {**config, "mic_opener": "admin_say"}, + {"id": "identity", "prompt": "What is your name?", + "require_warm_listening": True}) + assert calls == [("say", {"device_id": "robot", "text": "Ready."})] + assert result["mic_opened_by"] == "admin_say" + assert result["warm_listening_precondition"] == "PASS" + assert len(captured) == 1 # reached the capture boundary + assert result["failure"] == "fixture_capture_stop" + + +def test_mic_opener_that_never_opens_the_mic_still_blocks(tmp_path, monkeypatch, prompt_capture_boundary): + config, captured = prompt_capture_boundary + monkeypatch.setattr(runner, "admin", lambda host, route, body=None: {"ok": True}) + monkeypatch.setattr(runner.time, "sleep", lambda seconds: None) + monkeypatch.setattr(runner, "MIC_OPEN_POLLS", 3) + result = runner.run_case(tmp_path, {**config, "mic_opener": "admin_say"}, + {"id": "identity", "prompt": "What is your name?", + "require_warm_listening": True}) + assert result["verdict"] == "BLOCKED" + assert result["failure"] == "warm_listening_precondition" + assert captured == [] + + +def test_init_persists_mic_opener(tmp_path, monkeypatch): + monkeypatch.setattr(runner, "admin", lambda host, route, body=None: {"devices": ["robot"]}) + monkeypatch.setattr(runner, "command", lambda argv, **kwargs: "fixture-speaker\n") + monkeypatch.setattr(sys, "argv", ["dotty_overnight.py", "init", "--session", str(tmp_path), + "--host", "user@host", "--mic-opener", "admin_say"]) + runner.main() + assert json.loads((tmp_path / "config.json").read_text())["mic_opener"] == "admin_say" + + +def test_misheard_prompt_is_an_asr_mismatch_not_extra_speech(): + logs = ("结果: Repeat these words purple frispen\n识别文本: Repeat these words purple frispen\n" + "发送音频消息: SentenceType.FIRST, purple frispen\nSentenceType.LAST\n") + result = runner.evaluate({"asr": [["purple"], ["Brisbane"]], "reply": [["purple"]]}, + transcript("purple frispen"), logs, 2, after(), "robot") + assert result["extra_speech"] == [] + assert result["failure"] == "asr_mismatch" + assert result["interaction"] == "FAIL" + + +def test_repeatedly_inconclusive_case_is_set_aside_instead_of_retried_forever(tmp_path): + checkpoint = runner.restore_checkpoint(tmp_path) + murky = {**dict.fromkeys(("capture", "playback", "recovery"), "PASS"), "interaction": "INCONCLUSIVE"} + for _ in range(2): + runner.record_attempt(checkpoint, {"id": "repeat"}, murky) + assert checkpoint["unresolved"] == [] + runner.record_attempt(checkpoint, {"id": "repeat"}, murky) + assert checkpoint["unresolved"] == ["repeat"] + assert checkpoint["quarantined"] == [] and checkpoint["counts"]["repeat"]["failures"] == 0 + assert runner.eligible_cases([{"id": "repeat"}, {"id": "other"}], checkpoint) == [{"id": "other"}] + + +def test_a_pass_clears_the_inconclusive_count(tmp_path): + checkpoint = runner.restore_checkpoint(tmp_path) + good = dict.fromkeys(("capture", "playback", "interaction", "recovery"), "PASS") + murky = {**good, "interaction": "INCONCLUSIVE"} + for attempt in (murky, murky, good, murky, murky): + runner.record_attempt(checkpoint, {"id": "repeat"}, attempt) + assert checkpoint["unresolved"] == [] diff --git a/tests/test_state_phrase_routing.py b/tests/test_state_phrase_routing.py new file mode 100644 index 00000000..31bdbe73 --- /dev/null +++ b/tests/test_state_phrase_routing.py @@ -0,0 +1,110 @@ +"""Real ASR-to-state routing regressions; AI-assisted by OpenAI Codex (GPT-6). + +Container imports are isolated by the existing fixture module. Only external +intent/STT/MCP boundaries and the executor are faked; routing remains real. +""" + +import asyncio +import json +import types +from unittest.mock import AsyncMock, MagicMock, patch + +import pytest + +from tests.test_asr_name_corrections import _module + + +def route(text, state="idle"): + conn = types.SimpleNamespace( + logger=MagicMock(), need_bind=False, max_output_size=0, + client_is_speaking=False, current_state=state, session_id="fixture", + websocket=types.SimpleNamespace(send=AsyncMock()), + executor=MagicMock(), chat=MagicMock(), + _dotty_toggles_synced=True, + ) + commands = AsyncMock() + with patch.object(_module, "handle_user_intent", new=AsyncMock(return_value=False)), \ + patch.object(_module, "send_stt_message", new=AsyncMock()), \ + patch.object(_module, "_mcp_call_tool", new=commands): + asyncio.run(_module.startToChat(conn, json.dumps({"content": text}))) + states = [call.args[2]["state"] for call in commands.await_args_list + if call.args[1] == "self.robot.set_state"] + return states, conn.executor.submit.call_args.args[1] + + +def test_explicit_do_not_sleep_is_conversation_not_state_command(): + states, prompt = route("Do not go to sleep.") + assert states == [] + assert prompt == "Do not go to sleep." + + +@pytest.mark.parametrize("negation", ["Do not", "Don't", "Don’t", "Never"]) +@pytest.mark.parametrize("command", ["go to sleep", "tell me a story", "keep watch"]) +def test_explicit_negated_state_commands_do_not_dispatch(negation, command): + states, prompt = route(f"Please {negation} {command}.") + assert states == [] + assert negation in prompt + + +@pytest.mark.parametrize("quote", [('"', '"'), ("'", "'"), ('“', '”'), ('‘', '’')]) +def test_quoted_story_command_is_a_mention_not_a_state_change(quote): + text = f"What does {quote[0]}tell me a story{quote[1]} mean?" + states, prompt = route(text) + assert states == [] + assert prompt == text + + +@pytest.mark.parametrize("text, expected", [ + ("Goodnight Dotty.", "sleep"), + ("Good night, Dotty.", "sleep"), + ("Please go to sleep.", "sleep"), + ("Please keep watch.", "security"), + ("Security mode.", "security"), + ("Watch the room.", "security"), + ("Please tell me a story.", "story_time"), + ("Story time.", "story_time"), + ("Don't keep watch. Tell me a story.", "story_time"), + ("Don't keep watch, but tell me a story.", "story_time"), + ('The phrase "never" is interesting. Please go to sleep.', "sleep"), +]) +def test_affirmative_state_commands_still_dispatch(text, expected): + states, _ = route(text) + assert states == [expected] + + +@pytest.mark.parametrize("state", ["sleep", "security", "story_time"]) +@pytest.mark.parametrize("text", ["Wake up.", "Come back.", "Are you there?", + "Don't go to sleep, wake up."]) +def test_affirmative_wake_escape_still_dispatches_idle(state, text): + states, _ = route(text, state) + assert states == ["idle"] + + +@pytest.mark.parametrize("text", [ + 'What does "wake up" mean?', "Don't wake up.", "Never come back.", +]) +def test_quoted_or_negated_wake_is_not_an_escape_command(text): + states, _ = route(text, "story_time") + assert states == [] + + +@pytest.mark.parametrize("text", [ + "Please don't ever go to sleep.", + "Never, please tell me a story.", + "Don't you go to sleep.", + "Do not, tell me a story.", + 'Say "go to sleep" out loud.', + 'Explain “security mode” to me.', + "What does 'don't go to sleep' mean?", + "What's the meaning of 'tell me a story'?", + "Tell me. A story is something I dislike.", + "Explain the story timetable.", +]) +def test_punctuation_and_mentions_do_not_manufacture_state_commands(text): + states, _ = route(text) + assert states == [] + + +def test_unrelated_contraction_does_not_hide_an_affirmative_command(): + states, _ = route("I'm ready, please tell me a story.") + assert states == ["story_time"] diff --git a/tests/test_voice_camera_policy.py b/tests/test_voice_camera_policy.py new file mode 100644 index 00000000..2e34b080 --- /dev/null +++ b/tests/test_voice_camera_policy.py @@ -0,0 +1,87 @@ +"""Voice-only camera privacy regressions; AI-assisted by Codex (GPT-6).""" +import asyncio +import json +import types +from unittest.mock import AsyncMock, MagicMock, patch + +import pytest + +from tests.test_asr_name_corrections import _module + + +@pytest.fixture +def policy(tmp_path, monkeypatch): + path = tmp_path / "kid-mode" + monkeypatch.setattr(_module, "_KID_MODE_STATE_FILE", str(path)) + monkeypatch.setenv("DOTTY_KID_MODE", "false") + monkeypatch.setattr(_module, "VISION_BRIDGE_URL", "http://synthetic.invalid") + return path + + +@pytest.mark.parametrize("value, denied", [("true", True), ("1", True), ("yes", True), + ("false", False), ("0", False), ("no", False), (" FALSE\n", False), + ("", True), ("garbage", True)]) +def test_camera_policy_uses_explicit_current_state(policy, value, denied): + policy.write_text(value) + assert _module._camera_access_denied() is denied + + +def test_missing_unreadable_or_invalid_policy_denies_even_with_adult_env(policy): + assert _module._camera_access_denied() is True + policy.mkdir() + assert _module._camera_access_denied() is True + policy.rmdir() + policy.write_bytes(b"\xff") + assert _module._camera_access_denied() is True + + +def test_camera_policy_refreshes_without_process_restart(policy): + for value, denied in [("false", False), ("true", True), ("false", False)]: + policy.write_text(value) + assert _module._camera_access_denied() is denied + + +def test_physical_capture_boundary_denies_before_dispatch(policy): + policy.write_text("true") + conn = types.SimpleNamespace(logger=MagicMock(), headers={"device-id": "fixture"}) + dispatch = AsyncMock(side_effect=AssertionError("camera dispatch must not happen")) + with patch.object(_module, "_mcp_call_tool", dispatch): + with pytest.raises(PermissionError, match="Kid Mode"): + asyncio.run(_module._handle_vision(conn, "What do you see?")) + dispatch.assert_not_awaited() + + +def test_voice_denial_never_becomes_successful_photo_prompt(policy): + policy.write_text("true") + conn = types.SimpleNamespace(logger=MagicMock(), need_bind=False, max_output_size=0, + client_is_speaking=False, current_state="idle", session_id="fixture", + websocket=types.SimpleNamespace(send=AsyncMock()), executor=MagicMock(), + chat=MagicMock(), headers={"device-id": "fixture"}, _dotty_toggles_synced=True) + dispatch = AsyncMock(side_effect=AssertionError("camera dispatch must not happen")) + with patch.object(_module, "handle_user_intent", AsyncMock(return_value=False)), \ + patch.object(_module, "send_stt_message", AsyncMock()), \ + patch.object(_module, "_mcp_call_tool", dispatch): + asyncio.run(_module.startToChat(conn, json.dumps({"content": "What do you see?"}))) + dispatch.assert_not_awaited() + prompt = conn.executor.submit.call_args.args[1] + assert "Kid Mode" in prompt + assert "just used your camera" not in prompt + assert "photo shows" not in prompt + + +def test_adult_physical_capture_preserves_result(policy, monkeypatch): + policy.write_text("false") + conn = types.SimpleNamespace(logger=MagicMock(), headers={"device-id": "fixture"}) + response = types.SimpleNamespace(status_code=200, json=lambda: {"description": "Synthetic cube."}) + async def run(): + loop = asyncio.get_running_loop() + async def synthetic_executor(*args): + return response + dispatch = AsyncMock() + with patch.object(loop, "run_in_executor", synthetic_executor), \ + patch.object(_module, "_mcp_call_tool", dispatch): + assert await _module._handle_vision(conn, "What do you see?") == "Synthetic cube." + dispatch.assert_awaited_once_with(conn, "self.camera.take_photo", {"question": "What do you see?"}) + # requests is imported by the real function, but its client must not run. + monkeypatch.setitem(__import__("sys").modules, "requests", types.SimpleNamespace(get=None)) + asyncio.run(run()) diff --git a/tests/test_whisper_no_speech_gate.py b/tests/test_whisper_no_speech_gate.py new file mode 100644 index 00000000..76b2a89f --- /dev/null +++ b/tests/test_whisper_no_speech_gate.py @@ -0,0 +1,76 @@ +"""WhisperLocal must not hand near-silence hallucinations to the LLM. + +AI-assisted (Claude). Observed on the physical robot 2026-10-05: 1.3 s of +near-silence after Dotty finished speaking was transcribed as "Thank you." +with no_speech_prob=0.81, and Dotty replied to it. Every real utterance +captured that day scored <= 0.32. +""" +import asyncio +import importlib.util +import pathlib +import sys +import types +from contextlib import contextmanager +from unittest.mock import MagicMock + +import numpy as np + +_ROOT = pathlib.Path(__file__).parent.parent +_STUBS = ("faster_whisper", "config", "config.logger", "core", "core.providers", + "core.providers.asr", "core.providers.asr.base", "core.providers.asr.dto", + "core.providers.asr.dto.dto") + + +@contextmanager +def _stubs(): + missing = object() + previous = {name: sys.modules.get(name, missing) for name in _STUBS} + try: + for name in _STUBS: + sys.modules[name] = types.ModuleType(name) + sys.modules["faster_whisper"].WhisperModel = MagicMock() + sys.modules["config.logger"].setup_logging = lambda: MagicMock() + sys.modules["core.providers.asr.base"].ASRProviderBase = type( + "ASRProviderBase", (), {"__init__": lambda self: None}) + sys.modules["core.providers.asr.dto.dto"].InterfaceType = types.SimpleNamespace(LOCAL="local") + yield + finally: + for name, module in previous.items(): + if module is missing: + sys.modules.pop(name, None) + else: + sys.modules[name] = module + + +def _load(): + with _stubs(): + spec = importlib.util.spec_from_file_location( + "whisper_local_under_test", _ROOT / "custom-providers/asr/whisper_local.py") + module = importlib.util.module_from_spec(spec) + spec.loader.exec_module(module) + return module + + +def _transcribe(no_speech_probs, text, config=None): + module = _load() + provider = module.ASRProvider.__new__(module.ASRProvider) + provider.no_speech_threshold = module._no_speech_threshold(config or {}) + segments = [types.SimpleNamespace(text=text if i == 0 else "", no_speech_prob=p, avg_logprob=-0.5) + for i, p in enumerate(no_speech_probs)] + provider._transcribe_blocking = lambda audio: (segments, types.SimpleNamespace(language_probability=1.0)) + artifacts = types.SimpleNamespace(pcm_bytes=np.zeros(16000, dtype=np.int16).tobytes(), file_path="x.wav") + result, _path = asyncio.run(provider.speech_to_text([], "session", artifacts=artifacts)) + return result + + +def test_high_no_speech_transcript_is_dropped(): + assert _transcribe([0.811], " Thank you.") == {"content": ""} + + +def test_real_speech_passes_through(): + assert _transcribe([0.32], " See you soon!") == {"content": "See you soon!"} + + +def test_threshold_is_configurable_and_can_be_disabled(): + assert _transcribe([0.811], " Thank you.", {"no_speech_threshold": 1.0}) == {"content": "Thank you."} + assert _transcribe([0.5], " Hello.", {"no_speech_threshold": 0.4}) == {"content": ""}