Make job rescue paginated to protect against full batch of no-rescue jobs#1318
Open
brandur wants to merge 1 commit into
Open
Make job rescue paginated to protect against full batch of no-rescue jobs#1318brandur wants to merge 1 commit into
brandur wants to merge 1 commit into
Conversation
…jobs This one's aimed at fixing a long-standing bug accidentally detected while working on another rescue-related feature. The `JobRescuer` works by fetching a full batch of jobs to rescue then going through each one to determine what it should be doing about it. The default batch size is 10k, so this generally works perfectly fine. Some stuck jobs are potentially not rescued if their timeout is configured to be -1 (no timeout). Codex identified a tail bug possibility in which if you had an entire batch worth of jobs with -1 timeouts, they'd block any jobs after them from being rescued. So the rescuer would rescue 10k jobs, determine none of them needed rescue, then go back to sleep, stranding any jobs after that. This is such a tail possibility that it's probably fine, but just since we're making changes to the rescue driver boundary right now anyway, it's not a bad time to just get this one fixed up.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This one's aimed at fixing a long-standing bug accidentally detected
while working on another rescue-related feature.
The
JobRescuerworks by fetching a full batch of jobs to rescue thengoing through each one to determine what it should be doing about it.
The default batch size is 10k, so this generally works perfectly fine.
Some stuck jobs are potentially not rescued if their timeout is
configured to be -1 (no timeout).
Codex identified a tail bug possibility in which if you had an entire
batch worth of jobs with -1 timeouts, they'd block any jobs after them
from being rescued. So the rescuer would rescue 10k jobs, determine none
of them needed rescue, then go back to sleep, stranding any jobs after
that.
This is such a tail possibility that it's probably fine, but just since
we're making changes to the rescue driver boundary right now anyway,
it's not a bad time to just get this one fixed up.