Skip to content

fix(#8177): sample the matching bootstrap baseline only after the setup commit reached the disk - #8844

Merged
robfrank merged 4 commits into
mainfrom
fix/8177-flaky-bootstrap-window-gate-matching-peer
Oct 1, 2026
Merged

robfrank merged 4 commits into
mainfrom
fix/8177-flaky-bootstrap-window-gate-matching-peer

Conversation

@robfrank

@robfrank robfrank commented Oct 1, 2026 •

Copy link
Copy Markdown
Collaborator

Closes #8177

Issue7519BootstrapWindowGateTest.aPeerThatMatchesTheBaselineIsNeverHeldOutOfTheService sampled the baseline with BootstrapFingerprint.compute(dir) right after setUp committed a transaction, and the state machine recomputed it from the same directory a moment later. A commit hands its pages to the asynchronous flush thread, so the two reads could land on either side of the flush, hash different bytes, and send a peer whose copy IS the baseline down the mismatch arm. The fix is test-side: a small helper (SettledBootstrapFingerprint.of(db)) drains the flush queue (waitAllPagesOfDatabaseAreFlushed, asserted true) before fingerprinting, and the two tests with this shape use it. A new regression class makes the race deterministic by holding the flush with PageManager.suspendFlushAndExecute around the commit.

Root cause evidence (scratch probe, not committed)

Probe Digests differ
create + 1 commit, fingerprint, waitAllPagesOfDatabaseAreFlushed, fingerprint 5 / 200
create only (no commit), same sequence 0 / 200
commit inside suspendFlushAndExecute, fingerprint inside, then after resume + drain 20 / 20
after the drain, fingerprint, sleep 20 ms, fingerprint again 0 / 200 (stable once settled)

Test plan

  • Issue8177BootstrapFingerprintInFlightPagesTest (new): a baseline sampled while the commit is in flight takes the mismatch arm (pins the mechanism); its control, sampled from the settled copy after the same held commit, matches, installs nothing, marks nothing, reason null
  • Issue7519BootstrapWindowGateTest (7/7), Issue8368BootstrapPassWindowTest (18/18)
  • Siblings that fingerprint an open database: Issue7011BootstrapSourceRaceTest, ArcadeStateMachineBootstrap{Mismatch,Divergence,BaselinePersistence}Test, ArcadeStateMachineAppliedIndexPerDatabaseTest, ArcadeStateMachineDeferredDropTest, Issue8651StaleEntryAfterRaftInstallBoundaryTest, Issue8368FollowerHeldDuringBootstrapPassIT - 83 tests, 0 failures

Completeness

Invariant: a test that hands the state machine "this peer's own fingerprint" samples the same on-disk bytes the state machine recomputes.

Sweep - grep -rln "BootstrapFingerprint.compute" ha-raft/src/test server/src/test engine/src/test:

Test (fingerprint of an open db) Commit before sampling? Outcome depends on a match? Disposition
Issue7519BootstrapWindowGateTest (reported) yes (setUp) yes fixed here
Issue8368BootstrapPassWindowTest.matchingBaseline() yes (setUp) yes (bootstrapWindowReason() is null) fixed here
Issue7011BootstrapSourceRaceTest:186 yes no: originatedLocally=true, the source arm accepts any copy with localLastTxId >= baseline argued, untouched
ArcadeStateMachineBootstrapDivergenceTest, ArcadeStateMachineBootstrapMismatchTest, ArcadeStateMachineAppliedIndexPerDatabaseTest, ArcadeStateMachineBootstrapBaselinePersistenceTest, Issue8651StaleEntryAfterRaftInstallBoundaryTest no (create only) some argued: the create-only probe shows 0/200, untouched
ArcadeStateMachineDeferredDropTest:190 yes no: asserts only that the baseline is recorded, which happens before the decision argued, untouched
RaftBootstrapFromLocalDatabaseIT, RaftBootstrapFingerprintMismatchSameLsnIT, engine/.../BootstrapFingerprintTest n/a not a state-machine round trip on an open db with a fresh commit argued, untouched

Production callers (ArcadeStateMachine.readLocalBootstrapState, BootstrapElection.computeLocalStates, PostBootstrapStateHandler) have the same shape and are not changed here: filed as #8843. Waiting for the flush there is not free (waitAllPagesOfDatabaseAreFlushed keeps waiting while progress is observed, so on a database under sustained writes it blocks as long as the writes do), and it would run on the Raft apply thread and an HTTP handler.

Modified existing tests: two lines (the sample) in Issue7519BootstrapWindowGateTest and Issue8368BootstrapPassWindowTest. The flaky line IS the defect, so it cannot be fixed by adding tests only. No assertion was changed or loosened.

Known gaps

Residual risk

The fix removes the race from the two tests that sample after a commit. It does not change production behaviour (see #8843).

Adversarial pass

Not run: this session had no subagent tool available to spawn the independent reviewer. Noted per the workflow's error table; not a gate.

Review cycles

Cycle Head Changes claude review
1 11d6fde1e7 initial fix LGTM, optional polish; applied the rename of the settled-baseline test (it does not exercise the race)
2 108d8b92a9 rename no blockers; applied: assertion message says what a red in-flight test means, and a commit-only helper for the control test
3 e366c3388d message + commitWhileTheFlushIsHeld() nothing blocking; applied: comment on why the in-flight test drains before the apply, control-test note
4 987868d71d comments only nothing blocking, no actionable items: clean approval

Deferred items

  • "Using SettledBootstrapFingerprint.of in those five tests too is a one-line change each and would remove the need for the argument entirely." (cycle 1) - skipped: those tests do not commit before sampling (0/200 in the probe), and the workflow forbids modifying existing tests beyond the defect itself.
  • "If HA bootstrap fingerprint is computed over an open database without draining in-flight page flushes #8843 adds a production-side drain, consider folding or deleting SettledBootstrapFingerprint" (cycles 2, 3) - skipped: a production-side drain would not make the helper redundant, because the test's own sample is taken before the state machine runs, and a naive sample there still reads the pre-flush bytes.
  • "Trim the Javadocs" (cycle 2) and "shared fixture for the stubbed-server harness" (cycles 3, 4) - skipped as optional style; a third self-contained copy kept the change minimal.
  • "Consider a precondition that the in-flight sample equals a pre-commit fingerprint" (cycle 1) - skipped: the reviewer noted the existing isNotEqualTo already fails loudly on a vacuous pass, and the cycle-2 message now names the cause.

Final state: clean-approval (cycle 4 of 4).

🤖 Generated with Claude Code

Summary by CodeRabbit

  • Tests
    • Expanded high-availability bootstrap coverage to verify how database fingerprints behave while page writes are still being flushed.
    • Confirmed that a fingerprint sampled during an in-flight flush triggers a bootstrap install, while one sampled after the database settles avoids an unnecessary install and reconciliation.
    • Updated baseline test setups to sample fingerprints from settled database copies.

…up commit reached the disk

The matching-peer tests fingerprinted an open database right after committing,
while the async flush thread could still be writing that commit's pages. The
state machine's recomputation then hashed different bytes and took the
mismatch arm. Drain the flush first, and pin the mechanism deterministically
with a held flush.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
@robfrank robfrank added this to the 26.10.1 milestone Oct 1, 2026
@mergify

mergify Bot commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

This pull request does not currently match the merge queue conditions, so it cannot be queued from here. The box comes back if it matches again.

@coderabbitai

coderabbitai Bot commented Oct 1, 2026 •

Copy link
Copy Markdown

Review in Change Stack →

Navigate logical layers of code changes, visualize relationships, and explore their blast radius.

🧰 Additional context used
📚 Code guidelines (1)
ha-raft/CLAUDE.md — auto-discovered

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Advanced

Run ID: 766e6b1a-f13c-4ad5-ae2b-38ca0942382a

📥 Commits

Reviewing files that changed from the base of the PR and between 6abf36d and 987868d.

📒 Files selected for processing (4)
  • ha-raft/src/test/java/com/arcadedb/server/ha/raft/Issue7519BootstrapWindowGateTest.java
  • ha-raft/src/test/java/com/arcadedb/server/ha/raft/Issue8177BootstrapFingerprintInFlightPagesTest.java
  • ha-raft/src/test/java/com/arcadedb/server/ha/raft/Issue8368BootstrapPassWindowTest.java
  • ha-raft/src/test/java/com/arcadedb/server/ha/raft/SettledBootstrapFingerprint.java

Included review availability: This review used your included allowance. Your plan provides up to 4 included reviews per hour; 1 remain after this review.


📝 Walkthrough

Walkthrough

The changes add a helper that waits for database pages to flush before computing a bootstrap fingerprint. Two existing matching-baseline tests use it. New tests compare fingerprints sampled during and after a held page flush and check the resulting bootstrap state.

Changes

Bootstrap fingerprint tests

Layer / File(s) Summary
Settled fingerprint helper and matching-baseline tests
ha-raft/src/test/java/com/arcadedb/server/ha/raft/SettledBootstrapFingerprint.java, ha-raft/src/test/java/com/arcadedb/server/ha/raft/Issue7519BootstrapWindowGateTest.java, ha-raft/src/test/java/com/arcadedb/server/ha/raft/Issue8368BootstrapPassWindowTest.java
The helper waits for database pages to flush before computing a fingerprint. Both matching-baseline tests use the helper.
In-flight page regression tests
ha-raft/src/test/java/com/arcadedb/server/ha/raft/Issue8177BootstrapFingerprintInFlightPagesTest.java
New tests compare a fingerprint sampled during a held flush with a settled fingerprint. They assert the bootstrap state after applying each baseline.

Priority: ⬇️ Low

Estimated code review effort: 3 (Moderate) | ~20 minutes

Change: Bug fix · Severity of issue fixed: Low

Merge Risk: ⚪ Minimal · up to 98786

This PR only stabilizes flaky bootstrap fingerprint tests and does not change production behavior. No actionable merge-blocking risk was found.

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 46.15% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 13 functions across 4 files. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Title check ✅ Passed The title clearly and concisely describes the primary change: sampling the matching bootstrap baseline after the setup commit reaches disk.
Description check ✅ Passed The description is detailed and covers the change, motivation, related issue, test results, known gaps, and review notes. It does not use the template headings exactly and does not confirm that `mvn c…
Linked Issues check ✅ Passed Issue #8177 requires the matching-baseline test to keep a peer reconciled and to install nothing. SettledBootstrapFingerprint.of(LocalDatabase) drains pending page flushes before it computes the dir…
Out of Scope Changes check ✅ Passed The changes are limited to HA raft test support and regression coverage for issue #8177. The helper, the updates to the affected matching-baseline tests, and the in-flight page tests all support the l…
  • Fix all pre-merge checks with AI
✨ Finishing Touches 💡 1
📝 Generate docstrings 💡
  • Commit to this branch
  • Create a new PR
🧪 Generate unit tests (beta)
  • Commit to this branch
  • Create a new PR
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Autopilot is currently an internal CodeRabbit preview.


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@codacy-production

codacy-production Bot commented Oct 1, 2026 •

Copy link
Copy Markdown

Up to standards ✅

🟢 Issues 0 issues

Results:
0 new issues

View in Codacy

🟢 Coverage ∅ diff coverage · -5.91% coverage variation

Metric Results
Coverage variation ✅ -5.91% coverage variation
Diff coverage ✅ ∅ diff coverage

View coverage diff in Codacy

Coverage variation details
Coverable lines Covered lines Coverage
Common ancestor commit (e01f4f1) 205671 174005 84.60%
Head commit (987868d) 238445 (+32774) 187637 (+13632) 78.69% (-5.91%)

Coverage variation is the difference between the coverage for the head and common ancestor commits of the pull request branch: <coverage of head commit> - <coverage of common ancestor commit>

Diff coverage details
Coverable lines Covered lines Diff coverage
Pull request (#8844) 0 0 ∅ (not applicable)

Diff coverage is the percentage of lines that are covered by tests out of the coverable lines that the pull request added or modified: <covered lines added or modified>/<coverable lines added or modified> * 100%

NEW Get contextual insights on your PRs based on Codacy's metrics, along with PR and Jira context, without leaving GitHub. Enable AI reviewer
TIP This summary will be updated as you push new changes.

@claude

claude Bot commented Oct 1, 2026

Copy link
Copy Markdown

Review

Overall: a clean, well-scoped test-side fix. The root-cause analysis (probe table, 5/200 differing digests) is convincing and the scope split with #8843 for production callers is sensible. No blockers from my side.

What I checked

  • SettledBootstrapFingerprint.of uses PageManager.waitAllPagesOfDatabaseAreFlushed(Database) (public, PageManager.java:542) and asserts the boolean result. That is the right call, since it returns false on a flush timeout and silently ignoring it would reintroduce the flake.
  • Imports in Issue7519/Issue8368 are consistent after the change (File is still used for the cleanup code, the BootstrapFingerprint import is correctly removed). No assertion was loosened.
  • Style matches CLAUDE.md: AssertJ assertThat(...).isTrue() style, final on params, no stray System.out, license header present.

Suggestions (non-blocking)

  1. Pin test 1 against future changes to the suspend window. aBaselineSampledWhileTheLastCommitIsStillInFlight... relies on commit pages inside suspendFlushAndExecute being deferred and then flushed asynchronously after the callback returns. The PageManager comments around suspendFlushAndExecute (lines ~558-600, Issue7586BackupArchiveEntriesSitAtArchiveRootIT is intermittently red on main (restored database has 0 records, expected 50) #8111/Fix backup flakiness, importer tx leak, HA latch race, redis retry duplication, hardcoded test ports #8128) show this path has been reworked several times. If a later change made the resume write synchronously, or drained before returning, the isNotEqualTo(inFlight) assertion would still hold (the on-disk state differs from the in-window sample), so the test is robust here. Good. But the assertion message says the pages "reached the disk after the sample was taken", which is only true if the commit actually deferred. Consider also asserting inFlight differs from a fingerprint taken before the commit, or at least that the commit really did change the on-disk digest, so a vacuous pass (commit produced no on-disk change at all, e.g. page already full/unchanged) is impossible. Today that would show up as isNotEqualTo failing, so it fails loudly, which is fine.
  2. Test 2 does not exercise the race. aBaselineSampledFromTheSettledCopyMatches... is deterministic by construction (the drain is explicit), so it is a regression guard for the helper rather than for the race. That is fine, but the name "WhateverTheFlushTiming" slightly oversells it; the in-flight commit is already flushed by the time settled is computed. Consider a shorter name, e.g. aBaselineSampledFromTheSettledCopyMatchesThePeer.
  3. Helper placement. SettledBootstrapFingerprint is a package-private class in ha-raft test sources. If the siblings listed as "argued, untouched" ever need it (e.g. Issue7011, DeferredDrop), it is reusable as is. If other modules (server tests) later need the same helper it may be worth promoting it to the engine test-jar, but not needed now.
  4. Sweep residual risk. The "create only, 0/200" argument for the untouched siblings rests on the empirical probe. createDocumentType and DatabaseFactory.create() also go through transactions/schema writes, so that guarantee is empirical, not structural. Using SettledBootstrapFingerprint.of in those five tests too is a one-line change each and would remove the need for the argument entirely. Your call whether that churn is worth it.
  5. Cleanup in tearDown. localDb.close() while the test-1 commit may still be flushing is fine (close drains), just noting no leak of the suspension window: the callback throws nothing that would leave the window open, since suspendFlushAndExecute releases in a finally.

Test coverage

Good: new regression class plus the two modified tests, and the 83-test sibling run listed in the PR body. Since these are test-only changes, no new @Tag seems needed (both new tests are short).

Security / performance

N/A for test-only changes; production behaviour is untouched and tracked in #8843.

LGTM from a correctness standpoint; items above are optional polish.

🤖 Generated with Claude Code

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
@claude

claude Bot commented Oct 1, 2026

Copy link
Copy Markdown

Review

Overall: a good, well-scoped test-side fix. The root-cause evidence is convincing, the helper is small, and the regression class pins the mechanism deterministically with suspendFlushAndExecute instead of racing the flush thread. I found nothing blocking.

What I checked

  • waitAllPagesOfDatabaseAreFlushed(Database) returns boolean and delegates to the flush thread (PageManager.java:542), so asserting true in SettledBootstrapFingerprint.of is sound. A silent timeout would otherwise turn back into the original flake.
  • Removing import com.arcadedb.database.BootstrapFingerprint is safe in both modified tests, and java.io.File is still used in each (new File(DB_PATH) cleanup).
  • The new test follows the harness of Issue7519BootstrapWindowGateTest (stubbed server, zero install retries, awaitLifecycleTasksForTesting before asserting getBootstrapInstallsInFlight()), so the in-flight assertion matches the existing pattern for a failed install.
  • CLAUDE.md style: assertThat(...).isTrue() form, final params, no System.out, no new dependencies, no wall-clock assertions, no hardcoded ports. Nothing committed to production code.

Suggestions (all non-blocking)

  1. Test 1 depends on the flush being deferred, not just on timing. aBaselineSampledWhileTheLastCommitIsStillInFlight... relies on suspendFlushAndExecute deferring the commit's pages until the callback returns, and on PageManager's resume re-enqueuing them asynchronously. That is deterministic today, but it is a coupling to flush internals. If a future change makes the commit write synchronously inside a suspension, isNotEqualTo(inFlight) fails with a message that suggests a fingerprint bug. Consider adding a precondition inside the callback or in the as(...) text saying "if this fails, suspendFlushAndExecute no longer defers commit pages", so the next reader looks in the right place.
  2. Test 2 discards the in-flight sample on purpose, but that is easy to misread. commitAndSampleWhileItsPagesAreInFlight() is called only for its side effect (the commit). A tiny rename or a one-line comment, or splitting out a commitWhileFlushIsHeld() helper, would make that explicit. The same helper also computes a fingerprint (an extra directory hash) that test 2 does not need.
  3. Helper could be shared. SettledBootstrapFingerprint lives in ha-raft test sources only. That is fine now, but HA bootstrap fingerprint is computed over an open database without draining in-flight page flushes #8843 fixes the same shape in production; once it lands this helper may become redundant, so a // remove once #8843 lands style pointer in the Javadoc would help avoid it lingering. (Keep it to one short line per the CLAUDE.md comment rules.)
  4. Naming in Issue8177...Test. The class and method Javadocs are long and mostly restate the PR description. CLAUDE.md prefers short comments that explain only the non-obvious WHY. The first paragraph of each is valuable; the rest could be trimmed.
  5. Sibling tests. The sweep table argues the create-only tests are unaffected based on a 0/200 probe. That is reasonable evidence; if any of them later commits in setUp, they will need the helper. A short note in SettledBootstrapFingerprint's Javadoc ("use this whenever an open database has committed before sampling") already says roughly this, so I think that is covered.

Production follow-up

Agreed that touching readLocalBootstrapState / BootstrapElection.computeLocalStates / PostBootstrapStateHandler is out of scope here, given that waitAllPagesOfDatabaseAreFlushed can block for as long as writes continue and those callers run on the Raft apply thread and an HTTP handler. Tracking it in #8843 is the right call. When that is picked up, a snapshot-style read (openSnapshot) rather than a full drain may be worth considering, since it does not stall on sustained writes.

Test coverage

The new class covers both sides of the race (in-flight sample takes the mismatch arm; settled sample matches, installs nothing, no unreconciled mark, reason null). I did not re-run the suites locally; the PR description reports 83 tests green across the siblings, and CI should confirm.

Looks good to merge once CI is green.

…ut sampling where no sample is wanted

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
@claude

claude Bot commented Oct 1, 2026

Copy link
Copy Markdown

Review of #8844

Test-only fix for the flaky Issue7519BootstrapWindowGateTest (#8177). No production code changes, so no performance or security impact. The root-cause analysis is solid and the probe table in the description makes it convincing.

What looks good

  • The fix targets the actual defect: the baseline sample races the async flush. Draining via waitAllPagesOfDatabaseAreFlushed (asserted true, so a timed-out drain fails loudly instead of silently re-introducing the flake) removes the race.
  • Issue8177BootstrapFingerprintInFlightPagesTest makes the race deterministic with suspendFlushAndExecute rather than relying on timing, and the first test's .as(...) message tells the next reader what a failure means. This is the right shape for a regression test.
  • No assertions were loosened, no wall-clock assertions, no hardcoded ports, and final usage and import style follow CLAUDE.md. The two now-unused BootstrapFingerprint imports were removed, and java.io.File is still used in both modified tests.
  • The sweep of other BootstrapFingerprint.compute callers and the split-off follow-up for production callers (HA bootstrap fingerprint is computed over an open database without draining in-flight page flushes #8843) are well scoped.

Minor suggestions (non-blocking)

  1. Name the second test's role. aBaselineSampledFromTheSettledCopyMatchesThePeer is effectively the negative control for the first test (same in-flight commit, helper drains first). A one-line comment saying so would make the intent explicit.
  2. First test relies on the helper's drain as a side effect. In aBaselineSampledWhileTheLastCommitIsStillInFlightIsNotTheCopyOnDisk, the state machine's own recomputation sees settled bytes only because SettledBootstrapFingerprint.of(localDb) ran just before it inside the assertThat. That is correct, but a short comment would stop someone from removing the "unrelated" call.
  3. Helper lifetime. If HA bootstrap fingerprint is computed over an open database without draining in-flight page flushes #8843 adds a production-side drain, consider folding or deleting SettledBootstrapFingerprint so the two do not drift.
  4. Duplicated fixture. stateMachine() repeats stubbedServer from Issue7519BootstrapWindowGateTest. Acceptable for a self-contained regression class, but worth a shared fixture if a third test needs it.

Test coverage

Adequate for a flake fix: the mechanism is pinned and the fixed path is checked. I did not run the suite locally; the PR description lists the sibling tests run (83 tests, 0 failures).

Nothing here blocks merge.

…and which test is its control

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
@claude

claude Bot commented Oct 1, 2026

Copy link
Copy Markdown

Review

Overall this is a solid, well-scoped test-side fix. I found nothing that blocks the merge.

What I checked

  • The root cause is credible. BootstrapFingerprint.compute hashes on-disk bytes, and a commit hands its pages to the async flush thread, so two reads can land on either side of the flush. The probe table in the description backs this up.
  • PageManager.waitAllPagesOfDatabaseAreFlushed(Database) returns boolean, so asserting true in SettledBootstrapFingerprint.of is correct. A drain timeout becomes a clear failure instead of a quietly wrong sample.
  • suspendFlushAndExecute drains first, then defers the pages of commits made inside its window and re-enqueues them asynchronously on resume. That is what makes the regression test deterministic. The SettledBootstrapFingerprint.of(localDb) assertion in the first test both pins the mechanism and forces the drain before the state machine reads the directory.
  • Both modified tests still use java.io.File (DB_PATH cleanup), so removing only the BootstrapFingerprint import is right. No assertions were changed or loosened.
  • CLAUDE.md conventions: AssertJ style, final on params, no System.out, no new dependency, no wall-clock assertions, no hardcoded ports.

Minor suggestions (non-blocking)

  1. The new test's setUp, tearDown, stateMachine() and stubbed server are close copies of the Issue7519BootstrapWindowGateTest / Issue8368BootstrapPassWindowTest harness. This is the third copy. A shared fixture could be a follow-up, but it is fine to leave here given the minimal-change goal.
  2. In aBaselineSampledFromTheSettledCopyMatchesThePeer, the Javadoc says "deterministic by construction". That is true for the drain. Still, this test would also pass with a raw compute most of the time, because the race is probabilistic after resume. It is a guard for the helper, not a proof that the helper is needed. The first test is what proves the mechanism, and the Javadoc already says so, so I would leave it.
  3. The sm.getBootstrapInstallsInFlight() assertion in the first test after awaitLifecycleTasksForTesting(30_000) relies on the failed download (retries = 0) leaving the DB listed. This mirrors the Issue7519 harness, so it should be stable. If it ever flakes, that is the first place to look.
  4. Production callers (readLocalBootstrapState, BootstrapElection.computeLocalStates, PostBootstrapStateHandler) have the same race shape. Deferring them to HA bootstrap fingerprint is computed over an open database without draining in-flight page flushes #8843 is reasonable, given the drain cost on the Raft apply thread and the HTTP handler that the description points out.

Test coverage: good. There is a deterministic reproduction of the race plus a control, and the two flaky tests now sample from the settled copy. I did not run the tests locally. Please make sure CI is green on the ha-raft module before merging.

@codecov

codecov Bot commented Oct 1, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 76.35%. Comparing base (e01f4f1) to head (987868d).
⚠️ Report is 27 commits behind head on main.

Additional details and impacted files
@@             Coverage Diff              @@
##               main    #8844      +/-   ##
============================================
- Coverage     76.35%   76.35%   -0.01%     
- Complexity     3780     3781       +1     
============================================
  Files          2101     2102       +1     
  Lines        205671   205851     +180     
  Branches      43315    43375      +60     
============================================
+ Hits         157043   157169     +126     
- Misses        31708    31741      +33     
- Partials      16920    16941      +21     

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

@robfrank
robfrank merged commit 700e982 into main Oct 1, 2026
31 of 37 checks passed
@robfrank
robfrank deleted the fix/8177-flaky-bootstrap-window-gate-matching-peer branch October 1, 2026 20:53
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Flaky: Issue7519BootstrapWindowGateTest.aPeerThatMatchesTheBaselineIsNeverHeldOutOfTheService marks a matching peer unreconciled

1 participant