Skip to content

fix(e2e): wait up to ~16 min to find the dispatched E2E run - #101

Closed
rotemamsa wants to merge 1 commit into
mainfrom
fix/e2e-find-run-window
Closed

rotemamsa wants to merge 1 commit into
mainfrom
fix/e2e-find-run-window

Conversation

@rotemamsa

@rotemamsa rotemamsa commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

Problem

run-tests in regolibrary has failed on every run since Oct 5 with Could not find workflow run, even though the E2E tests it dispatches pass.

The job dispatches to armosec/shared-workflows, then polls the run list for a run whose name has the correlation ID. It gives up after 15 attempts, about 6 minutes.

Evidence (2026-10-06)

  • Receiver run regolibrary-37487278173 was created at 15:42:21. With a user token it showed up in the same event=repository_dispatch listing within 38 seconds.
  • The App token poller in the same run only found it on attempt 14 of 15, at 15:47:54, about 5.5 minutes later.
  • Runs regolibrary-37482457489 and regolibrary-37485687683 were created 4 seconds after dispatch and passed, but the poller never found them in 6 minutes.
  • On Sep 29 the same lookup found the run in about 30 seconds.

Fix

Raise the attempts from 15 to 35, about 16 minutes in total, in both reusable workflows that use this lookup:

  • kubescape-cli-e2e-tests.yaml (regolibrary)
  • incluster-comp-pr-merged.yaml (node-agent and other in-cluster components)

Follow-up

  • regolibrary pins this workflow by SHA, so it needs a pin bump after merge.
  • kubescape/kubescape 00-pr-scanner.yaml and both helm-charts 02-e2e-test.yaml have their own copy of the same loop.

AI-skills: armosec-shared-rules:docs_factcheck,high-confidence-review,superpowers:systematic-debugging

The App token sees a new repository_dispatch run several minutes after it
is created. The ~6 min lookup window was too short, so run-tests failed with
"Could not find workflow run" even when the E2E tests passed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Rotem Refael <rotem@armosec.io>
@rotemamsa rotemamsa added the ai-assisted Created through Armosec AI tooling (armosec-shared-rules plugin) label Oct 6, 2026
@coderabbitai

coderabbitai Bot commented Oct 6, 2026

Copy link
Copy Markdown

Important

  • 🔍 Trigger review

This repository does not receive automatic reviews because it has fewer than 10 stars.

⚙️ Run configuration
  • Configuration used: defaults
  • Review profile: CHILL
  • Plan: Advanced
  • Run ID: e2fda2fa-e512-46a6-aa40-74fe95c82605
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actions

github-actions Bot commented Oct 6, 2026

Copy link
Copy Markdown

Summary:

  • License scan: failure
  • Credentials scan: failure
  • Vulnerabilities scan: failure
  • Unit test: success
  • Go linting: failure

1 similar comment
@github-actions

github-actions Bot commented Oct 6, 2026

Copy link
Copy Markdown

Summary:

  • License scan: failure
  • Credentials scan: failure
  • Vulnerabilities scan: failure
  • Unit test: success
  • Go linting: failure

@matthyx matthyx left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed 9cf76cdab032d5311374300e4b9d8f13626e68dd against main at e697027077be46d848d4a519179cef1f5bdb3672.

No confirmed introduced code defects. The two 15→35 retry changes preserve early success, correlation matching, and explicit failure. Exhaustion adds 20 API requests and 600 seconds of sleep (370→970 seconds, plus API latency). This is a plausible mitigation, but readiness needs one clarification: can you provide an App-token lookup trace showing a previously failing correlation ID becoming discoverable after the old window and within the new window, or a controlled run demonstrating that recovery? The existing PR checks do not exercise either changed E2E discovery job.

Necessity is supported: regolibrary run 37482457489 dispatched at 15:04:03 and exhausted 15 attempts at 15:10:41, while its receiver was created at 15:04:07 and passed. Run 37487278173 found its receiver on attempt 14 at 15:47:54. These establish failed discovery and variable latency; they do not independently establish App-token-specific causation or recovery inside sixteen minutes.

Validation: extracted base/head shell blocks passed syntax checks and 24 mocked visibility scenarios in a credential-free, network-isolated bubblewrap environment. At simulated 400-second visibility, both base workflows fail and both changed workflows succeed; immediate/330-second discovery, the 940/941-second boundary, and permanent absence behave as expected. Actionlint reports only the same pre-existing ubuntu-large label diagnostic on base and head. All current head checks pass, including pr-checks; the bot's failure-summary comments are not failed check conclusions. No live dispatch was triggered.

History: searched all returned repository PRs/issues (200/100 limits), plus dispatch, correlation, poll, Could not find, and the affected CLI filename. #75, #76, and #80 established the approved orchestration; none supersedes this window change. Closed #85 concerns payload fields, with no recorded maintainer rejection of polling. Closed #43 requested testing evidence for a different generic-workflow change; closure does not establish rejection. Open #102 overlaps files only through action pinning. No duplicate window fix found within these search limits; searches are not proof of absence.

Architecture watch, nonblocking: the existing one-hour App token is reused for up to an hour of monitoring, and discovery still searches only the latest 30 runs. Extra retries do not fix token expiry or pagination. These are existing limitations, not confirmed new blockers. Regolibrary currently pins 547a98f361e0713832308e3f34165bb89fd9b08c, so adoption requires a subsequent pin bump.

Verdict: needs clarification. No request-changes defect identified; approval withheld pending evidence that the longer window recovers the reported failure. Independent code and architecture reviews agree that the implementation is bounded but recovery remains unverified.

@rotemamsa

Copy link
Copy Markdown
Contributor Author

Closing in favor of a redesign: dispatch with workflow_dispatch + return_run_details so the caller gets the run ID directly, instead of polling the run list.

@rotemamsa rotemamsa closed this Oct 7, 2026
@rotemamsa
rotemamsa deleted the fix/e2e-find-run-window branch October 7, 2026 06:38
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ai-assisted Created through Armosec AI tooling (armosec-shared-rules plugin)

Projects

Status: To Archive

Development

Successfully merging this pull request may close these issues.

2 participants