You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Opt-in bounded auto-resume after transient HTTP 429 without a reset time
#16289
Provide an optional, bounded recovery workflow for a coding turn that stops after transient HTTP 429 retries are exhausted, including providers that do not supply an allowance reset time. The user should be able to leave the task running and have T3 Code continue the same conversation when requests become available again, with visible waiting state and cancellation controls.
Problem to solve
An observed Codex run stopped with exceeded retry limit, last status: 429 Too Many Requests. The retained error code was responseTooManyFailedAttempts, with retryable: null and no reset information in the retained failure. At the initial inspection, the run was failed and no continuation was queued. A later user-authored continuation message started another run in the same native session and further tool activity appeared. Requiring that manual intervention leaves unattended tasks stopped after a recoverable interruption.
T3 Code already documents automatic resume for limited threads when a provider reports a reset time. Providers without one offer manual retry. Recognizing the original stop as Limited and exposing manual recovery can be fixed independently; this proposal adds automatic recovery for a transient stop without a known reset time.
Proposed behavior
Let users enable automatic recovery for supported transient provider limits. Preserve the failed turn's context and completed work, display that the thread is waiting to retry, and show the next attempt or remaining wait. Use available provider retry guidance; when no reset or retry time exists, provide a bounded cooldown policy rather than requiring a new user message.
Respect a user cancellation, pause, new message, archive, or settlement before any scheduled continuation starts. Coordinate retries across affected work so a burst of parent and child failures does not immediately repeat the same overload. Resume eligible work in the existing conversation without blindly replaying completed tool calls or successful child tasks.
Distinguish transient throttling from permanent quota, billing, authentication, or configuration failures using available evidence. An ambiguous error should keep a clear manual recovery path rather than trigger unlimited attempts. Stop at a visible retry or elapsed-time bound, explain what prevented recovery, and let the user retry manually. A public provider status page may add context, but a status incident or its resolution should not by itself determine that a specific request can safely resume.
Acceptance criteria
With automatic recovery enabled, an isolated running turn that exhausts transient HTTP 429 retries without a reported reset time enters a visible waiting state and schedules a bounded continuation without a new user message.
If the provider begins accepting requests again, the continuation uses the existing conversation, makes further task progress, and preserves earlier work without duplicate successful tool actions or child results.
If throttling persists, attempts remain bounded and separated by cooldowns; the UI shows the waiting state, next attempt, cancellation, and eventual exhaustion reason.
Cancelling or pausing the recovery, sending a new message, archiving, or settling the thread prevents the pending automatic continuation. Disabling the setting leaves recovery manual.
Known permanent quota, billing, authentication, and configuration failures do not loop automatically. A non-429 retry-exhaustion error does not qualify solely because it shares the same broad error code.
Existing reset-time auto-resume and manual recovery remain available. Restarting the environment must not create duplicate continuations or lose the user's cancellation decision.
Affected area
Provider failure interpretation, thread recovery settings, continuation scheduling, and the web, desktop, and mobile presentation of recovery state.
Non-goals
Increasing provider quotas, bypassing limits, replacing the user's provider or model without consent, or guaranteeing the completion of an arbitrary agent task. This proposal does not require replaying all failed subagents automatically or assigning the observed 429 to a particular upstream service.
Alternatives considered
Manual continuation is available and worked in the observed incident, but it requires the user to notice and intervene. Reset-time scheduling covers providers that supply a reset time, while the observed terminal error supplied none. Bounded automatic continuation closes the unattended-work gap and keeps a user-controlled stop condition.
Supporting context
The observed desktop installation reports T3 Code Nightly 0.0.46-nightly.20261005.2702. The retained native transcript reports Codex CLI 0.160.1, and the session selected gpt-6.1-sol with ultra reasoning and the default service tier through a custom provider route. No fresh rate-limit reproduction was run. The provider-side cause, response headers, and time at which requests became available again are unknown.
The installed build's thread recovery documentation describes reset-time auto-resume and manual retry for providers without a reset time.
The user also pointed to OpenAI's public status page. At inspection it listed Codex Cloud elevated errors and API Platform login and administration API errors in monitoring. Neither notice establishes the cause of this custom-route HTTP 429 or proves that it affected the same request path. They provide context for the recovery use case, not a causal diagnosis.
The supplied screenshot and read-only session inspection support the stopped-turn and later manual-continuation observations. Private task text, session identifiers, and machine paths are omitted from this report.
Related proposals have different completion criteria. Connection-loss recovery requires a preceding connection failure and limits its initial scope to the built-in OpenAI route. Capacity recovery explicitly excludes rate and usage limits. This request covers eligible transient HTTP 429 stops without a reported reset time, including supported custom provider routes, while preserving the same selected provider and conversation.
Triage assessment
Impact level: not-applicable
Assessment status: not-applicable
Impact basis: Automatic continuation without a reported reset time is a new capability beyond the documented manual-retry workflow. The independently supported classification and manual-recovery defect has a separate closure condition.
Workaround status: not-applicable
Workaround basis: This is a capability request. Manual continuation worked in the motivating incident, but does not provide unattended recovery.
reacted with thumbs up emoji reacted with thumbs down emoji reacted with laugh emoji reacted with hooray emoji reacted with confused emoji reacted with heart emoji reacted with rocket emoji reacted with eyes emoji
Uh oh!
There was an error while loading. Please reload this page.
Summary
Provide an optional, bounded recovery workflow for a coding turn that stops after transient HTTP 429 retries are exhausted, including providers that do not supply an allowance reset time. The user should be able to leave the task running and have T3 Code continue the same conversation when requests become available again, with visible waiting state and cancellation controls.
Problem to solve
An observed Codex run stopped with
exceeded retry limit, last status: 429 Too Many Requests. The retained error code wasresponseTooManyFailedAttempts, withretryable: nulland no reset information in the retained failure. At the initial inspection, the run was failed and no continuation was queued. A later user-authored continuation message started another run in the same native session and further tool activity appeared. Requiring that manual intervention leaves unattended tasks stopped after a recoverable interruption.T3 Code already documents automatic resume for limited threads when a provider reports a reset time. Providers without one offer manual retry. Recognizing the original stop as
Limitedand exposing manual recovery can be fixed independently; this proposal adds automatic recovery for a transient stop without a known reset time.Proposed behavior
Let users enable automatic recovery for supported transient provider limits. Preserve the failed turn's context and completed work, display that the thread is waiting to retry, and show the next attempt or remaining wait. Use available provider retry guidance; when no reset or retry time exists, provide a bounded cooldown policy rather than requiring a new user message.
Respect a user cancellation, pause, new message, archive, or settlement before any scheduled continuation starts. Coordinate retries across affected work so a burst of parent and child failures does not immediately repeat the same overload. Resume eligible work in the existing conversation without blindly replaying completed tool calls or successful child tasks.
Distinguish transient throttling from permanent quota, billing, authentication, or configuration failures using available evidence. An ambiguous error should keep a clear manual recovery path rather than trigger unlimited attempts. Stop at a visible retry or elapsed-time bound, explain what prevented recovery, and let the user retry manually. A public provider status page may add context, but a status incident or its resolution should not by itself determine that a specific request can safely resume.
Acceptance criteria
Affected area
Provider failure interpretation, thread recovery settings, continuation scheduling, and the web, desktop, and mobile presentation of recovery state.
Non-goals
Increasing provider quotas, bypassing limits, replacing the user's provider or model without consent, or guaranteeing the completion of an arbitrary agent task. This proposal does not require replaying all failed subagents automatically or assigning the observed 429 to a particular upstream service.
Alternatives considered
Manual continuation is available and worked in the observed incident, but it requires the user to notice and intervene. Reset-time scheduling covers providers that supply a reset time, while the observed terminal error supplied none. Bounded automatic continuation closes the unattended-work gap and keeps a user-controlled stop condition.
Supporting context
The observed desktop installation reports T3 Code Nightly
0.0.46-nightly.20261005.2702. The retained native transcript reports Codex CLI0.160.1, and the session selectedgpt-6.1-solwithultrareasoning and the default service tier through a custom provider route. No fresh rate-limit reproduction was run. The provider-side cause, response headers, and time at which requests became available again are unknown.The installed build's thread recovery documentation describes reset-time auto-resume and manual retry for providers without a reset time.
The user also pointed to OpenAI's public status page. At inspection it listed Codex Cloud elevated errors and API Platform login and administration API errors in monitoring. Neither notice establishes the cause of this custom-route HTTP 429 or proves that it affected the same request path. They provide context for the recovery use case, not a causal diagnosis.
The supplied screenshot and read-only session inspection support the stopped-turn and later manual-continuation observations. Private task text, session identifiers, and machine paths are omitted from this report.
Related proposals have different completion criteria. Connection-loss recovery requires a preceding connection failure and limits its initial scope to the built-in OpenAI route. Capacity recovery explicitly excludes rate and usage limits. This request covers eligible transient HTTP 429 stops without a reported reset time, including supported custom provider routes, while preserving the same selected provider and conversation.
Triage assessment
All reactions