Skip to content

Make job rescue paginated to protect against full batch of no-rescue jobs#1318

Open
brandur wants to merge 1 commit into
brandur-pilot-job-get-stuckfrom
brandur-paginated-rescue
Open

Make job rescue paginated to protect against full batch of no-rescue jobs#1318
brandur wants to merge 1 commit into
brandur-pilot-job-get-stuckfrom
brandur-paginated-rescue

Conversation

@brandur

@brandur brandur commented Jul 21, 2026

Copy link
Copy Markdown
Contributor

This one's aimed at fixing a long-standing bug accidentally detected
while working on another rescue-related feature.

The JobRescuer works by fetching a full batch of jobs to rescue then
going through each one to determine what it should be doing about it.
The default batch size is 10k, so this generally works perfectly fine.

Some stuck jobs are potentially not rescued if their timeout is
configured to be -1 (no timeout).

Codex identified a tail bug possibility in which if you had an entire
batch worth of jobs with -1 timeouts, they'd block any jobs after them
from being rescued. So the rescuer would rescue 10k jobs, determine none
of them needed rescue, then go back to sleep, stranding any jobs after
that.

This is such a tail possibility that it's probably fine, but just since
we're making changes to the rescue driver boundary right now anyway,
it's not a bad time to just get this one fixed up.

…jobs

This one's aimed at fixing a long-standing bug accidentally detected
while working on another rescue-related feature.

The `JobRescuer` works by fetching a full batch of jobs to rescue then
going through each one to determine what it should be doing about it.
The default batch size is 10k, so this generally works perfectly fine.

Some stuck jobs are potentially not rescued if their timeout is
configured to be -1 (no timeout).

Codex identified a tail bug possibility in which if you had an entire
batch worth of jobs with -1 timeouts, they'd block any jobs after them
from being rescued. So the rescuer would rescue 10k jobs, determine none
of them needed rescue, then go back to sleep, stranding any jobs after
that.

This is such a tail possibility that it's probably fine, but just since
we're making changes to the rescue driver boundary right now anyway,
it's not a bad time to just get this one fixed up.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant