You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Our classroom is one org (m323-ix24) with 1 classroom, 74 students and 70 assignments: about 2,000 accepted repos, 2,300 submit/* releases stored in scores.json, 2,100 repos in the org. We run collect-scores.yaml nightly and after lessons.
Every run downloads every release's result.json again, whether or not scores.json already holds it. In the run above, 2,297 of the 4,592 requests were downloads of releases that were already stored; the run changed 28 entries.
Because a fine-grained PAT has a primary limit of 5,000 requests per hour, one run uses about 92% of the service token's hourly budget. Two runs in the same hour (a cancelled run followed by a restart, or "Sync now" twice) fail:
##[error]m323-ix24: collection was throttled by GitHub (HTTP 403, X-RateLimit-Remaining: 0, resets at 2026-09-11T10:03:56Z) and did not recover after retrying. The service token is fine, do NOT rotate it; re-run once the limit resets.
This is the follow-up to #825: that fix removed the org-listing cost, but the walk still scales with accepted repos × stored releases, and the cost is paid again on every run.
What did you expect to happen?
A collect run over a classroom this size should take a few minutes, re-collecting an unchanged submission should not cost a download, and a single run should never need most of the token's hourly budget.
Steps to reproduce
A classroom with ~70 assignments and ~70 students where most assignments are accepted and submitted (~2,000 repos with submit/* releases).
gh workflow run collect-scores.yaml --repo <org>/classroom50 (all classrooms), then read the final collect: N GitHub API request(s) in X.Xs line: ~4,600 requests, ~30 minutes, on every run.
Run it again within the same hour: the run fails with the throttle error above.
Version and environment
Skeleton scripts from main at 969b6bc (CLI/web 1.48.2); the org's .github/scripts/collect_scores.py was identical to cli/gh-teacher/skeleton/dotgithub/scripts/collect_scores.py at that commit.
GitHub-hosted runner, ubuntu-latest, Python 3.14.
Relevant logs or output
collecting all classrooms
m323-ix24: 2089 repo(s) visible to the service token (19.8s)
m323-ix24/m323-lu01-a03-funktionaler-bubblesort: 60/74 submitted, 8 with pushes but no graded submission
m323-ix24/m323-lu01-a04-funktionaler-ggt: 64/74 submitted, 1 with pushes but no graded submission
...
m323-ix24: 28 updated submission(s) (1810.5s)
collect: 28 total submission(s) updated across 1 classroom(s)
collect: 4592 GitHub API request(s) in 1810.5s
Request breakdown, from the log and scores.json: ~1,996 release listings (one per accepted repo), ~2,297 result.json downloads (all already stored), ~270 for repos without a release (marker history + commits), ~25 listing and team reads.
For comparison, with the fix in the linked PR on the same classroom: first run 4614 GitHub API request(s) in 348.0s (the walk in parallel; every entry stamped), second run 2301 GitHub API request(s) in 175.4s (stored releases reused). Same entries and counts in both.
Before submitting
I searched existing issues and this is not a duplicate.
I removed any tokens, secrets, or private student data from the logs above.
Which part is affected?
Autograder / grading
What happened?
Our classroom is one org (
m323-ix24) with 1 classroom, 74 students and 70 assignments: about 2,000 accepted repos, 2,300submit/*releases stored inscores.json, 2,100 repos in the org. We runcollect-scores.yamlnightly and after lessons.Every run takes about 30 minutes:
Two things add up to that:
result.jsonagain, whether or notscores.jsonalready holds it. In the run above, 2,297 of the 4,592 requests were downloads of releases that were already stored; the run changed 28 entries.Because a fine-grained PAT has a primary limit of 5,000 requests per hour, one run uses about 92% of the service token's hourly budget. Two runs in the same hour (a cancelled run followed by a restart, or "Sync now" twice) fail:
This is the follow-up to #825: that fix removed the org-listing cost, but the walk still scales with accepted repos × stored releases, and the cost is paid again on every run.
What did you expect to happen?
A collect run over a classroom this size should take a few minutes, re-collecting an unchanged submission should not cost a download, and a single run should never need most of the token's hourly budget.
Steps to reproduce
submit/*releases).gh workflow run collect-scores.yaml --repo <org>/classroom50(all classrooms), then read the finalcollect: N GitHub API request(s) in X.Xsline: ~4,600 requests, ~30 minutes, on every run.Version and environment
mainat 969b6bc (CLI/web 1.48.2); the org's.github/scripts/collect_scores.pywas identical tocli/gh-teacher/skeleton/dotgithub/scripts/collect_scores.pyat that commit.ubuntu-latest, Python 3.14.Relevant logs or output
collecting all classrooms m323-ix24: 2089 repo(s) visible to the service token (19.8s) m323-ix24/m323-lu01-a03-funktionaler-bubblesort: 60/74 submitted, 8 with pushes but no graded submission m323-ix24/m323-lu01-a04-funktionaler-ggt: 64/74 submitted, 1 with pushes but no graded submission ... m323-ix24: 28 updated submission(s) (1810.5s) collect: 28 total submission(s) updated across 1 classroom(s) collect: 4592 GitHub API request(s) in 1810.5sRequest breakdown, from the log and
scores.json: ~1,996 release listings (one per accepted repo), ~2,297result.jsondownloads (all already stored), ~270 for repos without a release (marker history + commits), ~25 listing and team reads.For comparison, with the fix in the linked PR on the same classroom: first run
4614 GitHub API request(s) in 348.0s(the walk in parallel; every entry stamped), second run2301 GitHub API request(s) in 175.4s(stored releases reused). Same entries and counts in both.Before submitting