Repository navigation
fix: deflake //rs/tests/nns:delete_subnet_test_local - #11797
Open
github-actions[bot] wants to merge 1 commit into
Open
github-actions[bot] wants to merge 1 commit into
github-actions[bot] wants to merge 1 commit into
Conversation
Contributor
There was a problem hiding this comment.
🟢 Approval recommended
The focused change correctly addresses the identified race while preserving immediate failure for state-removal errors.
0 open findings
What changed in this PR
Updates the shared node-unassignment helper to retry while filesystem trimming is still in progress.
Changes:
- Replaces a premature assertion with a retryable
ensure!. - Documents the orchestrator timing race.
| File | Description |
|---|---|
rs/tests/consensus/utils/src/node.rs |
Makes the fstrim metric check retryable. |
🧠 Review effort: Balanced
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This PR was created by https://github.com/dfinity/ic/actions/runs/37827535031 following
.claude/skills/fix-flaky-tests/SKILL.mdto deflake://rs/tests/nns:delete_subnet_test_localDeflakes //rs/tests/nns:delete_subnet_test_local.
Flaky labels:
Root cause
The test flaked once in the last week, on master at commit 8d8880f (https://github.com/dfinity/ic/actions/runs/37752534307). Attempt 1 failed and attempt 2 passed. The
testtask failed on theassert!(metrics[&fs_trim_duration][0] > 0)inassert_node_is_unassigned_with_ssh_session(rs/tests/consensus/utils/src/node.rs). The test calls this helper at the end, once for each of the 6 nodes of the 3 deleted subnets.The helper works in two steps:
orchestrator_state_removal_failed_totalandorchestrator_fstrim_duration_millisecondsin aretry_with_msg!loop and asserts that the fstrim duration is greater than 0.The orchestrator does things in a different order (
remove_statein rs/orchestrator/src/upgrade.rs). It first deletes the state and the CUP. Then it runs/opt/ic/bin/sync_fstrim.shand waits for it. Only then does it set the fstrim duration gauge. In between, the node already looks unassigned but the gauge is still 0. In the failed attempt this window lasted 27 to 213 ms, depending on the node.The failed attempt hit that window. On the failing node, the orchestrator removed the subnet state and started sync_fstrim.sh at 09:10:10.249. The test's next SSH poll saw the node as unassigned right after that. The test then fetched the metrics and found the gauge still at 0. The orchestrator set the gauge to 46 ms a moment later.
The
assert!panics inside the retry closure instead of returning anErr, soretrynever retried. To hit the race, a poll has to land in a window of tens to hundreds of milliseconds, and the helper polls every 10 s. That is why this failure is rare.Fix
Replace the
assert!withanyhow::ensure!. The retry loop then retries every 10 s, for up to 120 s, until the node has finished trimming its filesystem.The
assert_eq!onorchestrator_state_removal_failed_totalstays. That counter never goes back to 0, so failing right away is correct.node_reassignment_testand the subnet recovery tests use the same helper, so they get the same fix. None of them flaked or failed on this assertion in the last month.Verification
cargo check --all-targets --all-features -p ic_consensus_system_test_utils,cargo fmtand./ci/scripts/rust-lint.shall pass (exit 0). No dependencies changed.bazel build //rs/tests/consensus/utils:utilssucceeds. This is the only direct reverse dependency of the changed file. The depth-2 query for affected tests returns no tests.bazel test --test_output=errors --runs_per_test=3 --local_test_jobs=2 //rs/tests/nns:delete_subnet_test_local: 3 of 3 runs passed (88.8 to 101.7 s). None of them happened to hit the race window.--runs_per_test=2:This PR was created following the steps in .claude/skills/fix-flaky-tests/SKILL.md.
🤖 Generated with Claude Code (https://claude.com/claude-code)