Skip to content

Release candidate: server-v0.2.0-rc.1 - #185

Open
BrettKinny wants to merge 42 commits into
mainfrom
release/server-v0.2.0-rc.1
Open

BrettKinny wants to merge 42 commits into
mainfrom
release/server-v0.2.0-rc.1

Conversation

@BrettKinny

Copy link
Copy Markdown
Owner

Collects everything intended for server-v0.2.0-rc.1 on one branch so the Docker host can run a single commit instead of a mix of branches.

Contents

Merge notes

Verification

  • pytest tests/ custom-providers/pi_voice/tests/: 479 passed
  • pytest dotty-behaviour/tests/: 254 passed
  • ruff check .: clean; dotty-pi-ext npm test: 11/11 suites
  • Physical robot, before the three PR merges: six voice cases run unattended to three passes each (five completed, V02-repeat unresolved on a reference-microphone transcription limit), no silent turns or restarts.
  • Not yet verified on the robot: this exact merged commit; photo/song/memory/think-hard tools; sleep/wake states; cold wake word; dashboard changes from fix(dashboard): relay perception feed, drop dead waits, disable sound turner #184.

Not for unattended merge

Per AI_TRANSPARENCY.md, a human reviews and merges. This PR was assembled by an AI agent (Claude Opus 5.5) under the maintainer's direction.

🤖 Generated with Claude Code

BrettKinny and others added 30 commits July 12, 2026 18:52
Correct the pinned firmware source, first-boot provisioning path, effective OTA configuration, and stale compatibility metadata. Add regression guards, protect raw UAT evidence from accidental commits, and keep the calendar route test offline.

Co-Authored-By: OpenAI Codex <noreply@openai.com>
Correlate new_session acknowledgements by request ID and surface reset failures instead of accepting stale responses.

Co-Authored-By: OpenAI Codex <noreply@openai.com>
…ions

Replace the broken direct-ALSA capture default with the C920 PipeWire source and validate decoded sample coverage. Add bounded session coordination, honest evidence gates, checkpoint quarantine, and local clip review exports.

Co-Authored-By: GPT-6 <noreply@openai.com>
Add PCM clipping and silence metrics, optional pre-rendered prompts, truthful follow-up readiness, and a source-based feature inventory. Preserve unknown visual and tool execution outcomes.

Co-Authored-By: GPT-6 <noreply@openai.com>
…wups

Keep VAD enabled with peak-limited analysis gain and immutable provenance. Require actual wake frames for cold-wake claims; block warm tests unless the live microphone is listening. Use deployed follow-up readiness contracts.

Co-Authored-By: GPT-6 <noreply@openai.com>
Gate ambient turns on current motion ownership and refresh the post-chat quiet interval. Eight deterministic failures reproduced before the fix; 44 focused tests pass. Physical acceptance remains pending.

Co-Authored-By: GPT-6 <noreply@openai.com>
Microphone evidence proved VAD can miss genuine quiet robot replies. Keep original case evidence and do not infer silence from an empty transcript. 57 coordinator regressions pass.

Co-Authored-By: GPT-6 <noreply@openai.com>
Discover all test suites including cached take_photo, execute previously skipped memory integration checks against invented SQLite fixtures, and clear live service test overrides. Eleven suites and 124 assertions pass without household data.

Co-Authored-By: GPT-6 <noreply@openai.com>
Preserve per-attempt prompt provenance and reject stale or mismatched waveforms before playback. Clear inherited overrides so one case cannot accidentally play another prompt. 75 coordinator regressions pass.

Co-Authored-By: GPT-6 <noreply@openai.com>
Keep default threshold 0.5 and prompt transcription unchanged. Validate response-only prospective settings and persist actual transcription provenance; preserve inconclusive silence and unverified text semantics.

Co-Authored-By: GPT-6 <noreply@openai.com>
Content-only and speaker-without-language envelopes were treated as spoken JSON. Decode content independently of optional metadata and reject non-text payloads before noise filtering. Regression tests exercise startToChat with mocked side effects; seven pre-fix scenarios failed and now pass. Live deployment remains pending.

Co-Authored-By: GPT-6 <noreply@openai.com>
Document touch/admin wake, disabled physical face and voice producers, neutral parking, and historical release-pin differences. Correct the current voice-tool catalogue to seven without changing product behavior.

Co-Authored-By: GPT-6 <noreply@openai.com>
Physical V06 correctly recognized a small-robot joke, then the sliding edit-distance correction manufactured a story command. Limit known-phrase cleanup to exact word sequences with punctuation normalization, preserving sentence boundaries and explicit observed name aliases. Actual startToChat replay and negative phrase regressions pass. Negated/quoted substring routing remains a separate known limitation; deployment pending.

Co-Authored-By: GPT-6 <noreply@openai.com>
Protect literal state and wake shortcuts from immediately negated commands and paired quoted mentions. Preserve quote context during ASR punctuation normalization. Exercise the actual ASR-to-MCP boundary with affirmative and escape regressions; indirect and hypothetical language remains outside this bounded guard.

Co-Authored-By: GPT-6 <noreply@openai.com>
Local safety-review candidate only; not deployed. Deny physical voice capture and cached voice descriptions unless the shared guardian policy explicitly enables adult mode. Re-read per access and fail closed on policy errors. Keep ambient/admin camera behavior unchanged.

Regression evidence: red-first tests, 359 root tests plus 25 subtests and 250 behaviour tests passed; Compose validation passed. Human safety/red-team review and separate deployment approval remain required.

Co-Authored-By: GPT-6 <noreply@openai.com>
Remove workstation-specific host/path defaults from public harness code and documentation. Initialization fails before network access without --host or DOTTY_TEST_HOST; existing session config remains authoritative.

Co-Authored-By: GPT-6 <noreply@openai.com>
…eadiness

Map an opt-in lossless sidecar from the same bounded capture process, request float Pulse input, and reject existing outputs. Add explicit warm-listening prerequisites for state-command trials. Verify with synthetic dual-output capture and isolated runner tests; no hardware or volume changes.

Co-Authored-By: GPT-6 <noreply@openai.com>
Bound native-rate float decoding by evidence windows and preserve prompt clipping warnings independently of the quiet response tail. Include optional lossless sidecar measurements with explicit non-sample-aligned timing caveats; do not infer interaction verdicts. Replayed J02 metrics without changing historical evidence; all380 root tests and25 subtests pass offline.

Co-Authored-By: GPT-6 <noreply@openai.com>
Gate warm-case capture on the perception server's own clock: compute chat-listening status age from sensor_age_s, last_event_t and last_chat_status_t, so unrelated sensor events cannot make a stale listening flag look fresh. Missing, non-finite or inconsistent timestamps block the case as warm_listening_precondition rather than passing silently. Capture is a prerequisite for a new trial only; recovery after a long capture is separate. 108 overnight harness tests pass.

Co-Authored-By: Qwen3.8-Max <noreply@qwen.ai>
A device or admin abort set conn.client_abort and nothing on the nointent
chat path cleared it, so every later reply on that connection was cut off
until the robot reconnected. Reproduced on the physical robot 2026-10-05
(abort, then 'What is your name?' -> ASR + LLM ran, no TTS).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Near-silence after Dotty finished speaking was transcribed as 'Thank you.'
(no_speech_prob 0.81) and answered. Real utterances captured on the
physical robot 2026-10-05 all scored <= 0.32. Gate at Whisper's default
0.6, configurable via ASR.WhisperLocal.no_speech_threshold (1.0 disables).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The Docker host runs agent/post-release-automation's pi_client.py and
pi_voice.py (direct remember/recall/think_hard routing). Bring this branch
level with the deployment so further fixes do not revert it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
PiClient spawned pi with no --system-prompt, so pi's built-in coding-assistant
prompt was the only identity the model saw and it leaked into speech (#177).
Pass personas/pi_voice.md via --system-prompt (DOTTY_PI_SYSTEM_PROMPT_FILE
overrides; a missing file keeps the old behaviour).

Offline A/B in the dotty-pi container, 8 samples each of a self-review prompt:
before 1/8 'coding assistant' and 4/8 off-topic; after 0/8 and 0/8.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…uator slips as faults

Refs #182. Opt-in --mic-opener admin_say makes a warm case open the microphone
through /xiaozhi/admin/say instead of blocking, so a session can run
unattended. Playback is also proven by the robot's own recognition; a
reference-mic mishearing contradicted by service TTS text and any extra
speech in the capture are INCONCLUSIVE; inconclusive attempts no longer count
toward quarantine or the three-failure pause.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… turner

Dashboard pieces still pointed at producers that left the bridge in the
#36 / #111 cutovers:

- Activity feed opened /api/perception/feed on the bridge (404). The
  bridge now relays dotty-behaviour's SSE stream at /ui/perception/feed.
- Emoji / Say / Story / Dance waited 8 s on a turn stream nothing
  publishes to, then reported "no reply". They now return as soon as the
  inject is accepted.
- Voice tools card listed 5 of the 7 dotty-pi-ext tools; a test now
  pins the list to dotty-pi-ext/src/tools/.
- Bridge detail said "Raspberry Pi"; robot "Last seen" ignored chat
  activity now that turns no longer land in the bridge's convo log.

Sound localizer (#27, still stuck-left):

- Remove the sound-balance sparkline from the State card along with its
  getter and dotty-behaviour's /api/perception/sound-balance endpoint.
- Gate SoundTurner behind SOUND_TURN_ENABLED (default off) — every
  sound_event reports direction "left", so it snapped the head left on
  any noise.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The first pi_voice persona fixed the coding-assistant leak but made the 4B
chattier: offline, 'repeat these five words' was refused or hedged 3/8 and
'just say the answer' gained follow-up questions. With an explicit brevity
rule: repeat 8/8 exact, arithmetic 4/4 bare answers, self-review 8/8 about
itself with no coding-assistant mention.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The dashboard's Turns tab, error toast, Errors count and error report all
read conversation turns the bridge stopped seeing when voice moved to
dotty-pi (#111): /ui/events had subscribers and no producer, and the
convo log the Errors readers parsed was never written. The Content
filter card had the same gap for kid-mode hits on the voice path.

- PiVoiceLLM posts each completed turn to the bridge at
  POST /api/voice/turn and each kid-filter hit to
  POST /api/voice/filter-hit. Fire-and-forget on a daemon thread with a
  2 s timeout, so a missing dashboard never delays or breaks a turn.
  Target is DOTTY_DASHBOARD_URL, else the dotty-behaviour host on :8081.
- The bridge keeps the last 200 turns in memory only (nothing on disk),
  fans them out on /ui/events, and the Errors count / report / last-seen
  now read that ring instead of the dead convo log.
- Both endpoints require X-Admin-Token when DOTTY_ADMIN_TOKEN is set.

Closes #183

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…pt is not extra speech

Refs #182. The unattended trial retried one case 11 times because every
attempt was INCONCLUSIVE (the reference mic always mishears one word of the
reply), which neither completes nor fails a case. After max_inconclusive (3)
such attempts the case moves to checkpoint['unresolved'] and the run moves on.
Also: the first recognition in a capture is the prompt even when misheard.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
BrettKinny and others added 11 commits October 5, 2026 15:27
The default export fits the whole 16:9 frame into 1080x1920, leaving the robot
small between black bars. --crop W:H:X:Y takes a 9:16 source region and fills
the frame; the manifest records the region. Review gates are unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
# Conflicts:
#	README.md
#	docs/modes.md
# Conflicts:
#	custom-providers/pi_voice/pi_voice.py
# Conflicts:
#	custom-providers/pi_voice/pi_voice.py
…asses

The merge of #173 brought in a test that forbids the phrase used for the
35f701a pin. The statement is unchanged in meaning: that pin predates
StateManager.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The Docker host already runs the provider from bfe69b6 (PR #173 sync,
direct remember/recall/think_hard intents, Dotty system prompt), which
landed on the overnight branch after this branch was cut. Merging keeps
the dashboard turn reporting on top of what is deployed instead of
reverting it.

Conflict: pi_voice.py imports. response() is now a thin wrapper around
_respond() so the direct-tool reply paths are reported to the dashboard
too, and direct-tool failures carry their error.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… server-v0.2.0-rc.1

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…stored rows

Fixes #186. Every voice turn was logged to brain.db with the per-turn tool
routing and HARD CONSTRAINTS suffix attached (253 of 402 conversation rows in
a live db), and the direct recall route yielded memory_lookup's raw rows as
TTS text.

- turn_text.ts: stripTurnBoilerplate / cleanStoredTurn
- turn_logger stores only what the person said
- memory_lookup cleans rows written before this fix at read time
- explicit recall hands the search results to the model as context and lets
  it phrase the answer; no-match and error replies are unchanged

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@BrettKinny

Copy link
Copy Markdown
Owner Author

Status update, 2026-10-05 evening (AI-assisted, Claude Opus 5.5):

  • Branch head is now 39975e4. Added since the PR was opened: the latest fix(dashboard): relay perception feed, drop dead waits, disable sound turner #184 commit, and a fix for Memory: voice turns are logged with prompt scaffolding, and recall speaks stored rows verbatim #186 (prompt scaffolding stored in memory; recall speaking raw stored rows).
  • Deployed to the Docker host from this commit: voice server (22 mounted files), dotty-behaviour (85 files), bridge (69 files) and the dotty-pi extension (38 files). Checksums on the host match 39975e4 for all four; /health is 200 on the bridge and behaviour services. The two retired provider folders (zeroclaw, tier1_slim) and their mounts were removed from the voice server's compose on the host.
  • Tests at 39975e4: 482 + 254 Python tests, ruff clean, 12/12 extension suites.
  • Smoke test inside the deployed voice container (no robot): identity, arithmetic, explicit recall ("I remember that your favorite color is purple!"), a no-match recall, and a think-hard question (correct, ~30 s) all behaved; newly logged conversation rows contain only the spoken words.
  • Still not verified on the physical robot at this commit: the robot was powered off when the deploy finished. The six voice cases and the abort reproduction need to pass there before server-v0.2.0-rc.1 is tagged.

…iting

Fixes #187. VOICE_THINKER_TIMEOUT defaulted to 30 s while a cold
qwen3.6:27b-think load is 30-50 s, so the first think-hard after idle was
cancelled mid-load (llama-swap: 502, context canceled at 31 s). Default 90 s
(the value the pre-cutover bridge unit used), 105 s for the direct tool
runner, and the direct route yields a thinking-face preamble first.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@BrettKinny

Copy link
Copy Markdown
Owner Author

Update 2026-10-06 (AI-assisted, Claude Opus 5.5): head is now 7844b30, adding the think-hard cold-start fix (#187); deployed to the host and checksum-verified. The abort fix was re-verified on the robot at the RC via text injection. The six acoustic cases have not run at this commit (bench speakers were silent, then the room was in use), so the tag is still pending. Observed while testing: with the mic left open in a busy room the robot answered room audio continuously and turns queued behind the #175 turn lock; see #105.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants