Repository navigation
Conversation
…de from Encina.UnitTests (#1441) Bump dotnet-stryker to 5.0.0, set test-runner mtp and project Encina.csproj in stryker-config.json (no solution, test-projects or test-case-filter keys), run Stryker from tests/Encina.UnitTests, and stop passing --log-to-file by default because the MTP runner's trace JSON-RPC logs grow by about 40 MB per test run.
…d missing and zero-kill shards (#1441, #1440, #1682) Replace FOLDERS/FILTERS with SHARDS of ';'-separated mutate globs (20 shards of at most about 150 tested mutants, large folders split by file with an exclusion-based remainder shard), set provisional TIMEOUTS from the #1395 formula, drop the inert test-case-filter patching, build only Encina.UnitTests, log runner memory and disk, upload reports even when a step fails, fail a shard and the aggregate when a report tested mutants but killed none, and warn in the summary when a shard report is missing.
…hards, guards)
…l only in matrix mode, forward --configuration only when given (#1441)
…er-resources step log and optional --configuration (#1441)
|
Important Draft PR not reviewedDraft PRs are not automatically reviewed by default.
To automatically review draft PRs, update your CodeRabbit configuration: reviews:
auto_review:
drafts: true
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
Review of PR #1713 (phase 1): merge after fixes and the validation runMulti-dimension review with adversarial verification: each finding was checked by two independent skeptics. CONFIRMED = both found it real; PLAUSIBLE = one did. Refuted findings are listed at the end. Confirmed (5)
Plausible (3)
Refuted (11)
|
Stryker 5.0.0 under-reports kills with the MTP runner above concurrency 1 (stryker-net#3832), and each slot keeps a multi-GB test host alive. The value lives in the config because run-stryker.cs reads -c as --configuration.
…tant merge for span-split files (#1441) The Run mutation tests step caps the managed heap at 4 GiB (DOTNET_GCHeapHardLimit), raises its own OOM score and runs Stryker under nice so a memory or CPU squeeze takes Stryker or a test host instead of the runner agent, and tees Stryker's console to stryker-console.log with pipefail. A monitor publishes runner readings every minute (every PUBLISH_EVERY minutes on the GitHub API, sized to the shard count) to a 'mutation telemetry (shard N)' check run that survives a lost runner; a final step and the aggregate job complete it. SHARDS can split one file with Stryker's character-offset span suffix; Build matrix rejects a span glob without its twin, and the aggregate merges reports mutant by mutant, drops empty and 'Removed by mutate filter' placeholders, and fails when two reports hold the same mutant. SHARDS and TIMEOUTS stay marked pending the phase-2 measurement.
… splits and per-mutant merge of the mutation workflow (#1441)
…, telemetry, span splits) (#1441)
…nd slower telemetry updates (#1441) A shard whose Stryker console reports an OutOfMemoryException fails and does not upload its report, since a test failing on the heap cap may count as a kill; the visible-failure effect of the cap is marked as a hypothesis for the calibration runs. The check-run token no longer reaches Stryker or the test hosts, every gh call times out, check-run updates stay near 360 per hour for the whole matrix, a failed oom_score_adj write only warns, and the readings also count processes naming Encina.UnitTests.
…ration; document the OOM guard, publish rate, token handling and process counts (#1441)
…lemetry rate, Debug build) (#1441)
… for custom scopes, restore the weekly schedule (#1441)
…erver (#1441 phase 2f) Phase 2e measured that glibc's dynamic mmap threshold kept the native memory freed after each run of the ABAC EEL tests, so the reused test server grew ~2.5 GB per run. Fixed 128 KiB thresholds cut the first-run peak from 7.6 GB to 4.9 GB and the growth to ~230 MB per run with no slowdown.
…during Stryker runs (#1441 phase 2f) Adds an MTP TestingPlatformBuilderHook to Encina.UnitTests that registers an ITestSessionLifetimeHandler only when STRYKER_MUTANT_FILE is set. At the start of a test session after the first one in the process, it kills the test server when its private memory is above ENCINA_MTP_RECYCLE_MB (default 6144 MB) and writes one [encina-mtp-recycle] line to stderr. Stryker 5.0.0 treats the lost connection as a crashed host, discards the server and reruns the same mutant on a fresh one (RunAssemblyTestsInternalAsync, two attempts), so the verdict comes from the second attempt. The first session of a process never recycles, so the retry cannot fail because of the hook.
#1441 phase 2f) Review fixes for the MTP test server recycler: - Stryker 5.0.0 discards the test server's stderr unless --log-to-file is set, so the recycle line is also appended to the file named by ENCINA_MTP_RECYCLE_LOG when that variable is set. - ENCINA_MTP_RECYCLE_MB values too large to express in bytes are capped instead of overflowing into a negative threshold that recycles every session. - The builder hook moves to its own file and takes the environment reader as a parameter, so tests prove it registers nothing without STRYKER_MUTANT_FILE.
…cle count in the shard summary (#1441 phase 2f)
|
#1441 phase 2f, ready for verification. Branch head Changes:
Local Windows runs with and without a recycle gave identical verdicts. Verification run 37371127806 uses
Related: #1858 / PR #1860 (EEL tests share a static compiler) and #1859 (collectible ALC in |
Codecov Report✅ All modified and coverable lines are covered by tests. 📢 Thoughts on this report? Let us know! |
…ance timeout margin (#1441)
|
Follow-up for the rewritten |
…arallel strategies directly (#1441)
Part of #1441 (phase 1), #1440 and #1682. The closing keywords, and the knowledge records for #1440 and #1682, are added in phase 2, once the CI validation run proves that mutants are killed and that a missing shard is reported.
Summary
Under xUnit v3 the VsTest runner of Stryker.NET 4.x killed 0 mutants (#1440), per-folder
test-case-filterwas the only thing keeping shards inside their timeouts, and shard 14 could lose its runner and upload nothing while the run stayed green (#1682). The Mutation Tests workflow now runs Stryker.NET 5.0.0 with the Microsoft Testing Platform (MTP) runner and fails loudly when a shard proves nothing..config/dotnet-tools.json,.github/stryker-config.json):test-runner: mtp,coverage-analysis: off, project mode (project: Encina.csproj) fromtests/Encina.UnitTests. Thesolution,test-projectsandtest-case-filterkeys are gone (MTP ignores the filter, stryker-net#3757). Absorbs Dependabot PR deps: Bump dotnet-stryker from 4.16.0 to 5.0.0 #1527.run-stryker.cs: runs Stryker fromtests/Encina.UnitTests;--log-to-fileis off by default (under MTP it writes about 40 MB per test run per server, gigabytes for a 150-mutant shard; still available with-- --log-to-file);-c/--configurationis forwarded only when given (it was parsed and never passed before).mutation-tests.yml): aSHARDSarray of;-separated mutate globs (with!exclusions) replacesFOLDERS/FILTERS. 20 shards of about 150 tested mutants or fewer, from run 36959365309 (Modules/Isolation from run 36304862790). A large folder is split by naming its biggest files; its last shard takes the folder glob minus those files, so a new file always lands in a shard. Modules.Isolation is four shards; the two permission-script generators (201 and 185 mutants) get one shard each.TIMEOUTSfrom the [INFRA] Mutation Tests: 8 of 17 shards hit the 60-minute job limit, so the weekly run never completes or publishes #1395 formula (20 min setup + 35.6 s per mutant, measured locally): 81 to 209 minutes. Phase 2 re-sizes them from the validation run.Encina.UnitTests;test-baselineruns onlyEncina.UnitTests.::warning, a "Missing shard reports" summary block and a row in the table ([INFRA] Mutation shard 14 (Modules.Isolation) loses its hosted runner and uploads nothing, and the run stays green #1682); report and summary uploads run withif: always(); memory and disk go torunner-resources.logevery minute and every fifth reading to the step log ([INFRA] Mutation shard 14 (Modules.Isolation) loses its hosted runner and uploads nothing, and the run stays green #1682). The fix(ci): mutation aggregate keeps every shard's mutants and fails on overlap or a count mismatch (#1681) #1684 overlap and count guards are kept. Permissions are declared per job (AGENTS.md §10); the workflow-level block is gone.docs/testing/mutation-measurement-methodology.md,docs/en/guides/MUTATION_TESTING.md; knowledge recorddocs/knowledge/issues/1441.md(outcomepartial). TheAGENTS.md§9 mutation rule now describesSHARDS/TIMEOUTSunder the MTP runner. Stale pages outside this diff: [DEBT] Mutation docs and rules outside the methodology page still describe the VsTest runner, per-folder test filters and the 17-shard matrix #1711.Verification
SequentialDispatchStrategy.cs): 22,069 tests found, 4 mutants tested, 4 killed, score 100 %; 4 mutants in 92 s at concurrency 2.**/Dispatchers/Strategies/*.csplus!entries selected only the remaining file's 4 mutants (the remainder-shard design works).mutation-tests.yml;run-stryker.csbuilds with no warnings.src/Encina: no gap or overlap against the old 17 folders; counts and timeouts match. Four minors fixed (step-log readings, zero-kill guard only in matrix mode, configuration message, a docs claim).custom_scope=**/Dispatchers/Strategies/*.cs) measures the CI time per mutant (local figure is from 32 cores; runners have 4 vCPU). Then the full 20-shard validation run, re-sizedTIMEOUTS(max(60, ceil(minutes × 1.5))), and the record and docs updated with the observed figures.Cross-cutting checklist (ADR-018)
Not applicable: CI and test tooling only; no entity, store, behavior, service or external integration in
src/.