Repository navigation
fix(native-chat): a chat cut off by a crash, quit or restart reads Interrupted until its state changes - #25670
brennanb2025 wants to merge 35 commits into
Conversation
…errupted until its state changes
A turn the host proved was cut short by a crash, quit, restart or eviction read red "Failed" on
every status surface, while the chat's turn bar read "Worked for N". It now reads Interrupted
everywhere: the sidebar dot and line, tab bar, Cmd+J, Activity, the notes menu, the phone's agent
list, notifications ("stopped"), and the turn bar ("Interrupted after N", with the one derived
notice that says why). A person's Stop is unchanged and gets no notice; a real failure stays red.
The status is derived from the chat's latest turn, so it lasts until the next turn replaces it.
The 30-minute freshness window no longer clears it on the sidebar card, tab bar or Cmd+J: a cut is
the chat's journal state, not a live report that goes quiet. A departed agent's retained cut reads
interrupted rather than done.
…blaming the agent After an Orca quit or restart, the cut turn's notice read "Codex stopped while this response was in progress" in error red: it blamed the agent for something Orca did, and its red contradicted the turn's Interrupted status. A reader cannot tell a quit, an eviction or a restart apart (the journal publishes no cause), so the notice now says "This response was interrupted. You can continue in this conversation." as a muted host status line, translated on desktop. The host writes the row a reopen leaves for an owner found dead the same way, and readers re-present such rows already stored in the old words. The agent's own unexpected-exit row is unchanged.
…t-reads-interrupted
… as its exit row does The agent exiting on its own mid-turn ends the turn interrupted with no verdict, as a quit does, so every status surface and the turn bar read it Interrupted under a red "<Agent> stopped" row. One shared reader now derives the verdict: a cut that the agent's own exit row explains reads `failure`. The host's status projection and the turn bar both use it, so the published row and the chat agree. Nothing is stored, so no completion event (and no notification) is sent for it.
…death still reads failed
…ntil the chat changes The card and tab bar drop an agent row 30 minutes after its evidence. That window is for hook reports, which can go quiet; a native chat's row is re-derived from its journal by its host. Only a cut turn was exempt, so a person's Stop, a failure and "Couldn't confirm" on a native chat vanished from the card and tab while the agent rows and the phone kept them. The status bridge now marks its rows as the chat's journal status (renderer-local), and the dot keeps any settled verdict of such a row. A clean done, hook rows and live work still age out, and the liveness check used for subagents, attention and cleanup is unchanged. Drops the Cmd+J exemption and its test: Cmd+J never shows a native chat's status.
… it stopped A settled turn folds every row but its answer. The interrupted notice is muted, so it was neither an answer nor exempt, and folded away with the turn's work: by default nothing said the turn was cut short. The notice now never folds, as a compaction report doesn't; the partial answer stays the answer and the notice sits under it. The phone hides only tool activity behind its turn caret, so the notice was already shown there.
The reader re-worded every reopen row by its id, so a newer host's own wording for that row (a recorded cause, for instance) would be replaced by the generic sentence on this client, and an early build's row lost the exit detail it quoted. Only the legacy shape (no presentation, error red or the exit fact) is re-worded now; anything else is kept as written. Comments no longer claim the owner that went away was Orca: the evidence proves only that the agent's process is gone.
No completion event is sent for a cut, so the formatter's interruption arm is unreachable; the wording change described nothing a user could see.
…stored beside the words Re-wording an older red row built a new body, dropping any field a host stored on it, such as why Orca stopped. It now changes only the words, presentation and tone. A quit's row in the same family (`stale-session:…:shutdown-…`) already counts as the cut's explanation by its writer prefix; a test now pins that, and that such a row with its own presentation is kept as written.
… words came from The status row type ties a failure fact to the sentence built from it, so the fact cannot stay beside the new words. Every other stored field (why Orca stopped) is kept.
… finish on the card
A native chat's settled verdict kept past the freshness window went through the card's fresh
flags, so a week-old Failed pinned the worktree card red over another agent's live work, and an
old Stop, cut or "Couldn't confirm" hid another agent's fresh Done. A mark kept only because it
never expires now lands in the card's retained tier, as a departed agent's does: live work and a
fresh finish show over it, and alone on the card it still shows.
The row's provenance is now the host's own `structuredHost` ('owned' | 'held'), stamped by the
status bridge as the host ingest stamps it and compared in the bridge's equality check, instead
of a second renderer-only field.
…terrupted line A host that records why Orca stopped under a turn writes its row with presentation `orca-stop` and its cause beside the words, keeping the red words and exit fact for older clients. This client mutes it in the interrupted words, keeps the presentation and cause for a client that words them, never folds it with the turn's work, and never takes its red tone as the turn's failure report.
… show it A client older than the interrupted presentation folds every row but an error one under a collapsed turn and derives no notice beside a reopen row, so the muted row this host wrote left a crash cut unexplained there until the turn was expanded. The host now stores the row red, with the neutral words and the `response-interrupted` presentation; the shared reader mutes any red row with that presentation, so this client, desktop and phone, draws it as before. Tests pin the phone path for older red rows, `orca-stop` rows and the stored red row, and an older client's fold.
…which is why it is stored red
…t-reads-interrupted
…ative chat's does The failed-start row #24904 writes carried no native-chat provenance, so its Failed faded from the tab and card after 30 minutes while every other native chat's Failed stays until the chat changes. It now carries the same marker ('held': no host runs its agent), and past the freshness window it sits in the card's retained tier like the rest.
… a kept chat mark A departed agent's done never expires, and it outranked a native chat's kept Interrupted or "Couldn't confirm", so the card flipped to Done 30 minutes after the cut with no change in the chat. Kept marks now rank above a retained done and below only a fresh finish.
…t-reads-interrupted # Conflicts: # src/main/native-chat/agent-session-wire/structured-agent-session-conversation-open.ts # src/renderer/src/components/native-chat/NativeChatNoticeRow.test.tsx # src/renderer/src/components/native-chat/use-structured-agent-session.ts # src/renderer/src/i18n/locales/en.json # src/renderer/src/i18n/locales/es.json # src/renderer/src/i18n/locales/fr.json # src/renderer/src/i18n/locales/ja.json # src/renderer/src/i18n/locales/ko.json # src/renderer/src/i18n/locales/zh.json # src/shared/structured-agent-session-projection.ts # src/shared/structured-agent-session-turn-timing.ts
…t-reads-interrupted
…t-reads-interrupted
…t-reads-interrupted
Review statusThe problem. When Orca crashes, quits, restarts or evicts a chat while the agent is mid-answer, the answer is cut off. On main the sidebar card, agent row, tab bar, Activity, the notes send menu and the phone's agent list all marked that chat red Failed, while the chat's own turn bar said "Worked for N". On the worktree card and tab bar the red mark also vanished about 30 minutes later, though nothing about the chat had changed. The line under the cut turn said "Codex stopped while this response was in progress…" in error red, blaming the agent for something Orca did. What changes for you.
Fixed during review
Deferred
Open question
Differences from the common pattern
VerifiedLive on a hidden, isolated test build on a second Mac at c4c8446, driven over the app's debugging protocol, with a scripted stand-in agent (no real login):
Crash, collapsed: muted line under the partial answer Automated: typechecks (node, web, mobile) and the named test set (12,453 desktop and 357 mobile tests) pass locally at c1962a0 (before the last main merge), and each fix was reverted with its tests kept to show they fail without it. CI at e7820df: typecheck, package, cross-version wire compatibility, Mobile Checks and test shards 3–5 pass. Shard 1 fails only on Not verified
|
… same at any age A native chat's Failed, Interrupted or Couldn't-confirm mark moved from the fresh tier to the retained tier 30 minutes after the turn ended, so the card could flip (Failed to Working, or Interrupted to Done) with nothing changed in any chat. The mark is the chat's state, so it now always sits in the retained tier: live work and another agent's fresh finish show over it, and alone it still shows.
…t-reads-interrupted
This branch's derived cut-turn notice names no agent, so it takes no options; the Grok tests from
main passed `{ agentName: 'Grok' }`. Grok's own exit row keeps "Grok stopped..." from the host.
…t-reads-interrupted
|
Live QA, round 2, at a56b6f1: all nine scenarios pass. Screenshots are in the PR description under Visual Proof. Setup: a macOS dev build with an isolated profile and home folder and stand-in agents, driven over the Chrome DevTools Protocol by a separate agent. Waits past 30 minutes ran on real time.
CI at this head: 21 checks pass. The one red on the first run, Noticed during QA, not changed by this PR: no resume offer after a crash, only after a graceful quit; the "1 chat to resume" pill stays after Close; an "Enjoying Orca?" card can cover the composer's Send/Stop corner. |
# Conflicts: # src/main/native-chat/agent-session-wire/structured-agent-session-dead-generation-settlement.ts # src/renderer/src/components/native-chat/native-chat-transcript-slots.ts
|
Main merges Round 3: merged The main change that overlaps this PR is #25675 (Orca says why it stopped a reply, and offers Continue). The two now combine like this:
Four restart tests failed after that merge: #25675's tests expected the stored quit row to match what the chat shows. With this PR the stored row is red and the chat shows it muted. cdcc646 updates those tests; production code is unchanged. CI on cdcc646 passed (21 passed, 16 skipped, 0 failed). Round 4: merged
These main commits in round 4 touch nearby code:
The earlier QA still holds: none of these change the cut-turn notice, the row Orca writes when it stops a reply, or Continue. The only one in the same file adds a separate sign-in button beside them. Checks on 7fa999c: running; I'll update this comment when they finish. |
Main's #25675 now writes the quit's own row about a cut turn, and these restart tests read it twice: as stored and as the transcript shows it. The host stores it red so an older client keeps it on screen, and this branch's reader shows it muted, so the two views differ only in that tone. The stored-row assertions now expect the red row; the reader assertions keep the muted notice.
# Conflicts: # src/renderer/src/components/native-chat/NativeChatNoticeRow.tsx # src/renderer/src/i18n/locales/en.json # src/renderer/src/i18n/locales/es.json # src/renderer/src/i18n/locales/fr.json # src/renderer/src/i18n/locales/ja.json # src/renderer/src/i18n/locales/ko.json # src/renderer/src/i18n/locales/zh.json
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configuration
📒 Files selected for processing (82)
💤 Files with no reviewable changes (2)
Included review availability: This review used your included allowance. Your plan provides up to 10 included reviews per hour; 8 remain after this review. 📝 WalkthroughWalkthroughThe change distinguishes interruption verdicts from failures across turn status, activity labels, and worktree summaries. Journal explanations identify agent exits that classify a cut turn as a failure. Reader-facing cut-turn notices use generic interruption wording, and host-status handling normalizes qualifying older rows. Settled native-chat verdicts remain visible beyond the freshness window when they have structured-host provenance. Priority: ➖ Normal Merge Risk: ⚪ Minimal · up to Native chats cut off by a crash, quit or restart now read as Interrupted, while agent failures remain Failed. No merge-blocking risk was found in the supplied changes. 🚥 Pre-merge checks | ✅ 4 | ❓ 1❌ Failed checks (1 inconclusive)
✅ Passed checks (4 passed)
Full details: Docstring CoverageExplanation Docstring coverage is 37.84% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 37 functions across 50 files. (30 skipped: 7 unsupported, 23 over the file limit.) ✨ Finishing Touches 💡 1📝 Generate docstrings 💡
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |










ELI5
When Orca crashes, quits, restarts, or evicts a chat while the agent is mid-answer, that answer is cut off. Orca marked it with a red "Failed" in the sidebar and everywhere else, but nothing failed: the chat was interrupted. Now it reads "Interrupted" everywhere, and it keeps saying so until the chat actually does something else. If the agent itself crashes, that is still a red "Failed".
What Changed
The problem: a turn cut off by an Orca crash, quit, restart, or eviction (Orca proved the agent process gone and nobody pressed Stop) showed a red Failed dot and "Failed" line on the sidebar card and agent row, the tab bar, Activity, the notes send menu, and the phone's agent list. The chat's own turn bar said "Worked for N", which disagreed with all of them.
The notice under the cut turn also blamed the wrong party: after an Orca quit or restart it read "Codex stopped while this response was in progress. You can continue in this conversation." in error red, though nothing shows the agent did anything wrong.
Separately, on desktop a native chat's settled status (a Stop's Interrupted, a Failed, a "Couldn't confirm") vanished from the worktree card and tab bar about 30 minutes after the chat's last activity, and right away after a restart for an older chat, while the agent row under the card and the phone kept showing it.
What you see now: a cut turn shows the existing Interrupted state on every status surface (the muted dot and the word "Interrupted", the same one a Stop uses, without "by user"). The turn bar reads "Interrupted after N", and the row under the turn reads "This response was interrupted. You can continue in this conversation." as a muted line, translated into the app's language. That line stays visible when the settled turn is collapsed, under the turn's partial answer.
The agent's own crash is unchanged in substance: when the agent process exits on its own while Orca runs, its row still reads " stopped while this response was in progress…" in red, and the status surfaces and the turn bar read red Failed / "Failed after N" to match it. A person's Stop, a real failure and "Couldn't confirm" read as before. A cut sends no notification, as before.
How long it lasts: a native chat's settled status is read from its latest turn, so it stays until the chat's state changes, for example when a new message starts a turn and it reads Working, or until the chat is closed. That includes after Orca restarts: a chat whose last turn was stopped, failed, cut off or could not be confirmed shows that on its tab and worktree card however long ago it happened. Opening or viewing the chat changes nothing, and no timer clears it. A clean "done" still fades from the card after 30 minutes as before. Only the latest turn's mark is shown: once a newer turn runs, an earlier turn's Failed, Stop or cut lives only in the transcript, under that turn.
On the worktree card, which sums up every chat and terminal agent in a workspace, a native chat's mark now ranks the same however old it is. Live work in another chat or terminal always shows over it. Another agent's fresh finish shows over an Interrupted or Couldn't confirm, but not over a Failed, which stays until that chat starts a new turn or is closed. Alone, the mark shows. Before, a native chat's mark ranked higher for its first 30 minutes and then dropped, so the card could change with nothing new happening. For example, chat A failed 10 minutes ago and chat B is working: the card read Failed, then switched to Working once A's failure turned 30 minutes old. Now it reads Working from the start, while chat A's tab and agent row read Failed. Likewise, a chat that was just stopped no longer hides another agent's fresh Done on the card.
Mechanism, the mark:
agentVerdictDisplayMark(desktop) and its phone mirroragentRowVerdictMarkmap the host'sinterruptionverdict tointerrupted, as a Stop or a replaced turn, instead offailed. The turn bar'ssettledTurnStatusKeymaps it tointerruptedAfter. Activity titles follow the same grouping.Mechanism, the agent's own crash: the journal records the same "interrupted, no verdict" turn for the agent's own exit as for a quit, so the cause is read from the row that explains the cut. One shared reader (
structuredAgentTurnVerdictReaderinnative-chat-cut-turn-explanation.ts) gives a cut turn that the agent's own exit row explains the verdictfailure. The host's status projection and the chat's turn bar both use it, so the published row, older clients and the turn bar all read Failed. Nothing is stored, so no turn-completion event (and no notification) is sent for it, as before. If the exit's own bookkeeping failed and a later reopen wrote the row instead, the cause cannot be told apart from a quit, and the turn reads Interrupted.Mechanism, lifetime: the worktree card and tab bar ignore an agent row once its evidence is 30 minutes old, because a hook stream that goes quiet proves nothing. A native chat's row is not a live report; its host re-derives it from the chat's journal whenever that changes. The status bridge now carries the host's own marker for a native chat's row (
structuredHost, which the host already stamps on that row), andisSettledNativeChatVerdictkeeps a done row with a settled verdict from such a row on those two surfaces past the window. On the worktree card such a mark, at any age, joins the tier a departed agent's mark already uses (applyRetainedAgentMarkinworktree-agent-activity-summary.ts): live work always ranks above it, a fresh finish ranks above an Interrupted or Couldn't confirm but not above a Failed (the order the card already uses for a departed agent's failure), and alone it still shows. It no longer moves between tiers when its evidence turns 30 minutes old. Hook and terminal rows, a clean done, and live work decay exactly as before, and the shared freshness check used for subagent rows, attention sorting and cleanup is untouched.Mechanism, notice: the chat cannot tell a quit, an eviction or a restart apart, because the host publishes no cause for them, so the notice uses one wording true for every cause. It is a host status line (
response-interruptedinagent-session-host-status-rows.ts), like the existing "Part of this chat's history couldn't be loaded" line: desktop words it in the reader's language, and itsnoticetone keeps it muted on a client that does not know the line. The collapsed-turn fold never hides it, as it never hides a compaction report. The host writes its reopen row (left when a reopen finds the previous agent process gone) with the same words and presentation, stored red for older clients, and an older host's red reopen row is shown in the new words; any other reopen row, such as one a newer host words itself, is shown as written.Replaces #25293, which cleared the red mark once the chat was viewed; that approach was rejected because viewing a chat does not change whether it was interrupted. Builds on #25043, which introduced the one derived notice row for a cut turn.
Why
The status should describe the chat's state. An Orca crash, quit or restart interrupts a turn; it is not the agent failing, while the agent's own crash is. The turn bar and the sidebar should agree, and the mark should hold for exactly as long as the chat stays in that state, whoever ended the turn.
Alternatives considered: keeping red Failed until the chat is viewed (#25293) was rejected because it makes a status depend on whether you looked. Recording a
failureoutcome in the journal for the agent's own exit was rejected because a recorded outcome sends a turn-completion event, which would add a notification that is not sent today. Exempting native-chat rows inside the shared freshness check (isExplicitAgentStatusFresh) was rejected because that check also decides whether subagent rows under a parent may claim live work; a dead parent's leftover children would have read Working forever. Exempting only the cut from the 30-minute window was rejected because a Stop shows the same Interrupted mark, and the two would have vanished at different times. Showing old marks only in the transcript was rejected because the latest turn's failure is the chat's current state, and the chat's row shows it. On the card, keeping a fresh native-chat Failed above live work (main's order for a fresh failure) was rejected: with no expiry it would pin the card red over live work for as long as the chat sits failed.Differences from the common pattern
Linked Issue
Fixes #
Visual Proof
Live QA on macOS: an isolated dev build with stand-in agents, driven over the Chrome DevTools Protocol by a separate agent. Waits past 30 minutes ran on real time; the clock was not faked.
After main's Continue feature was merged in (round 3, all scenarios pass):
Crash mid-reply, then relaunch. The turn reads "Interrupted after 31s", with the cause line and Continue. The card and tab dots are grey.
After pressing Continue. The agent carries on and finishes.
Graceful quit mid-reply, then relaunch. The quit wording of the cause line, with Continue.
32 minutes later, after viewing the chat. Still Interrupted.
Neighbouring states and the workspace card (round 2, all nine scenarios pass):
Stop pressed (unchanged). "Interrupted by user", with no notice line.
The agent process dies while Orca runs (unchanged). Red Failed.
One chat failed, another working, in one workspace. The card shows Working; the failed chat's tab stays red.
After the working chat finishes. The card shows the failure, and doesn't change on its own.
Noticed during QA, not changed by this PR: no resume offer after a crash, only after a graceful quit; after a quit, the resume dialog, Continue and the "1 chat to resume" pill are all offered at once; Continue adds a message bubble styled like the user's own.
Testing
New tests: a Stop, a cut, a failure and "Couldn't confirm" on a native chat stay on the worktree card and tab bar more than 30 minutes after the chat's last activity, while a clean done, a terminal hook row and a stale working row still fade, and the subagent liveness check still reads the row stale; on the card, another agent's live work reads Working over a native chat's Failed and a fresh Done reads Done over its Stop, both for a mark minutes old and one past 30 minutes, and a cut alone reads Interrupted; the status bridge carries the host's marker on its rows; a cut that the agent's own exit explains reads failed on the status row and "Failed after N" on the turn bar and sends no completion event, while a quit, a reopen and an exit row answering a later message keep it Interrupted; a collapsed cut turn shows its partial answer and the notice (derived, the host's reopen row, and an older host's re-worded red row), and the notice alone for a turn with no answer; a reopen row with its own presentation, and an early build's untoned row, are kept as written; Activity titles a cut "Agent interrupted"; the phone's row reader shows Interrupted for a cut (also checked through the cross-version phone test, which loads the phone's real row reader).
Updated tests, each pinning the old rule (red Failed or "Worked for N" for a cut): the shared verdict table, the turn-bar labels (desktop and phone), the tab bar, the worktree card summary and agent rows, Activity grouping, the dashboard row, the paired-client mirror, and the host tests for a close, quit, restart and crash (
structured-agent-session-close-verdict,-crash-turn-end,-recovered-turn-clock,-stop-event-restart).Each rule was ablated (the change reverted with the tests kept) and its tests went red: the verdict mark, the turn-bar label, the agent-exit failure reading and its precedence over a later reopen row, the settled-verdict lifetime and its done-only guard, the retained tier for old marks on the card, the bridge's host marker, the never-folded notice, the phone mirror, the Activity title, the notice words, the host's reopen row, the re-wording of only legacy stored rows, and the notice tone.
AI Disclosure
Review
Agent skill upstream boundary
docs/reference/agent-skill-sharing-upstream-boundary.mdand copies or mechanically translates no upstream skill-installer source, tests, fixtures, registry entries, path tables, comments, or documentation.Notes
Remote and SSH: the execution host still owns the verdict and publishes it on the same row as before. There is no wire change:
failureandinterruptionare existing verdicts, and the row marker is the host's existingstructuredHostfield. Losing contact with a remote host keeps the last settled status, as for every settled row.Mixed versions: a phone or desktop client older than this change still draws a cut chat as Failed, because the drawing happens on the client. A newer client paired with an older host draws a cut as Interrupted; for an agent's own exit it shows "Failed after N" on the turn bar but Interrupted on the sidebar, until the host updates. An older client reading a newer host shows the new reopen row in red with its neutral English words, because the host stores it red so a client that folds every other row under a collapsed turn still shows it; a newer client draws that row muted, and shows an older host's red reopen row in the new words.
Unchanged on purpose: attention sorting still ranks a cut like a completion (news you have not seen), unlike a Stop, which it demotes. Cmd+J shows no status for native chats, before or after this change.
Checklist
N/Awith reasonpnpm lint,pnpm typecheck,pnpm test, andpnpm buildpass (or CI will cover; local preferred)