You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Environment maintenance hold with observable provider drain
#16359
T3's server can continue opening provider sessions, probes, background tasks and terminals during host maintenance. Closing a conversation or seeing an empty session map does not establish that provider and MCP cleanup finished. Session release can be persisted while scope closure is still running.
I propose an opt-in maintenance control for one server environment. Normal behavior and defaults would stay unchanged. The control would keep status and recovery available while it prevents new execution and observes the cleanup of previously admitted work.
Proposed behavior
Establish a durable maintenance generation before acknowledging the hold. Enroll openings under the same admission boundary before they prepare MCP or initialize adapters.
Cover cached and retained runtime handles, queued and delegated work, provider probes, standalone generation, server commands and terminal opens, restarts and writes.
Keep queued work intact. Maintenance is a pause, not a provider failure or a retry. Existing user queue holds stay in effect after reopening.
Track acquisitions and actual closing obligations separately from the live session map. Report incomplete or failed-held status when cleanup fails or descendant completion cannot be established.
Reload the hold before startup recovery can open providers. A lost runner or server restart must not reopen admission automatically.
Reopen only through an explicit request for the current generation and controller boot, after verified drain and operator-supplied maintenance readiness. Preserve interrupted conversation state rather than silently restarting turns.
This would be an environment boundary. Independent CLIs, another server environment and externally hosted providers remain separate operator obligations. T3 must not claim machine-wide exclusion.
Why existing release and retry paths are insufficient
ProviderSessionManager.releaseEntry removes a live entry before joining its detached scope closure. Its ordinary timeout and release record do not attest completed drain. Cached handles can also resume through the adapter without opening a new manager entry.
EffectOutbox.claimNext increments the attempt count. Routing maintenance through ordinary retry handling can consume the five-attempt budget. Startup reconciliation also cancels pending process-bound effects. A maintenance pause must preserve never-executed work without replaying genuinely process-bound work.
ProviderContinuationService consumes a request before dispatch. Its failure path reoffers delegated completion only, so an ordinary maintenance rejection can discard other buffered wakes.
Provider completion needs an explicit decision. Some SDK cleanup paths suppress failures, some descendant checks are no-ops, and external OpenCode connections have no owned exit handle. Unsupported completion must block a successful drain receipt.
Implementation and verification scope
Reuse one application-owned Effect service, existing atomic persistence, typed RPC authorization and worker receipt patterns. Avoid a parallel transport, fake maintenance thread or new selective-routing framework.
Focused tests would force these orderings with Deferreds and worker drains:
Opening before versus after hold, including an acquisition paused before map insertion and a retained cached handle.
Hold during detached cleanup, timeout or finalizer failure. A release record must not become a drain receipt.
Hold around effect claim and external execution. Pending identities and attempt counts remain correct.
Hold around queued promotion and a consumed continuation request. Work remains available after reopen without clearing user holds.
Server loss, unreadable state and failed persistence. Admission stays closed and no stale boot claims completed drain.
Repeated hold and reopen requests, stale generations and unsupported provider containment.
The proposal does not include installation, live activation, provider shutdown on an operator's machine or a coordinated client upgrade. I would keep the initial implementation unqualified for live cutover wherever an opening path or strict completion contract remains unsupported.
Is this maintenance direction and scope acceptable? If so, which provider completion guarantees should the initial supported cut require? I will link explicit direction and scope approval in any implementation PR.
reacted with thumbs up emoji reacted with thumbs down emoji reacted with laugh emoji reacted with hooray emoji reacted with confused emoji reacted with heart emoji reacted with rocket emoji reacted with eyes emoji
Uh oh!
There was an error while loading. Please reload this page.
T3's server can continue opening provider sessions, probes, background tasks and terminals during host maintenance. Closing a conversation or seeing an empty session map does not establish that provider and MCP cleanup finished. Session release can be persisted while scope closure is still running.
I propose an opt-in maintenance control for one server environment. Normal behavior and defaults would stay unchanged. The control would keep status and recovery available while it prevents new execution and observes the cleanup of previously admitted work.
Proposed behavior
This would be an environment boundary. Independent CLIs, another server environment and externally hosted providers remain separate operator obligations. T3 must not claim machine-wide exclusion.
Why existing release and retry paths are insufficient
ProviderSessionManager.releaseEntryremoves a live entry before joining its detached scope closure. Its ordinary timeout and release record do not attest completed drain. Cached handles can also resume through the adapter without opening a new manager entry.EffectOutbox.claimNextincrements the attempt count. Routing maintenance through ordinary retry handling can consume the five-attempt budget. Startup reconciliation also cancels pending process-bound effects. A maintenance pause must preserve never-executed work without replaying genuinely process-bound work.ProviderContinuationServiceconsumes a request before dispatch. Its failure path reoffers delegated completion only, so an ordinary maintenance rejection can discard other buffered wakes.Provider completion needs an explicit decision. Some SDK cleanup paths suppress failures, some descendant checks are no-ops, and external OpenCode connections have no owned exit handle. Unsupported completion must block a successful drain receipt.
Implementation and verification scope
Reuse one application-owned Effect service, existing atomic persistence, typed RPC authorization and worker receipt patterns. Avoid a parallel transport, fake maintenance thread or new selective-routing framework.
Focused tests would force these orderings with Deferreds and worker drains:
The proposal does not include installation, live activation, provider shutdown on an operator's machine or a coordinated client upgrade. I would keep the initial implementation unqualified for live cutover wherever an opening path or strict completion contract remains unsupported.
Is this maintenance direction and scope acceptable? If so, which provider completion guarantees should the initial supported cut require? I will link explicit direction and scope approval in any implementation PR.
All reactions