Conversation
Concurrent openInteractiveContextOffloadStoreForWrite calls each awaited the on-disk lease check before claiming or joining the per-lease single flight. Those checks finish in no fixed order, so a later caller with different limits could become the writer and reject the earlier callers. Check the lease synchronously and claim or join the slot before the first await; the root identity is still verified on disk before a writer is returned. The single-flight test now awaits the conflicting open with the others, and a new test replays the race to guard the ordering. Fixes apache#5801 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
hqhq1025
left a comment
There was a problem hiding this comment.
This change claims the per-lease context-offload writer slot before the first await, so the first valid caller's snapshotted limits win even if concurrent on-disk root-identity checks complete in a different order. The opening path still validates the lease and root identity before constructing and returning the SQLite-backed writer; an existing writer is revalidated and rechecked after the await. I found no actionable issue in these paths.
I reviewed the open/join/close paths and the new concurrent-limits regression. On Node 24, the core and storage builds and all five focused context-offload-store tests pass, including the 500-round concurrent-open test. The change has no schema migration; a fresh-main merge-tree and diff check are clean. Only the label hosted check is visible on this head, so I cannot treat the full CI suite or packaged Host/Desktop behavior as verified. This is not merge approval.
Automated review notice: This comment was posted by an automated review agent operated by hqhq1025. It is not an independent human review and does not replace one.
Summary
openInteractiveContextOffloadStoreForWriteawaited the on-disk lease check before it claimed or joined the per-lease single flight. Concurrent callers finish that check in no fixed order, so a later caller with different limits could become the writer and reject the earlier callers with "different limits". This madesingle-flights one limit-bound writer and snapshots admitted inputsflaky.The opener now validates the lease synchronously (
assertStorageRootLeaseActive) and claims or joins the slot before its firstawait, so the earliest caller always binds the limits. The root identity is still verified on disk before a writer is returned:runWithStorageRootLease, which checks the root identity, and is re-checked afterwards (unchanged);assertStorageRootLeaseexplicitly, and re-enters if the writer was closed during that check.The other storage authorities use the same single-flight pattern but take no limits, so every caller gets the same writer whichever call wins; only this store needed the change.
Fixes #5801
Verification
binds the first caller limits whatever order concurrent lease checks finish inreplays the race 500 times (~0.4 s). Without the fix it failed 10/10 runs in isolation and 9/10 runs as part of the file; with the fix the file passed 30/30 runs.@maka/storagebuild (tsc) and fulldistsuite: 1548 tests, 0 failures (Node 24.15.0, macOS arm64).biome checkon the changed files: clean.npm test.Behavior change
Concurrent openers are now admitted in call order. One minor difference in error precedence: when a writer already exists, a limits mismatch is now reported before the on-disk root identity check, so a caller with both a relocated root and mismatched limits sees "different limits" rather than the lease error.
AI use
Tool(s) and scope: Claude Code analysed the race, wrote the fix and tests, and ran the verification; I reviewed and directed the change. The commit carries a
Generated-by: Claude Codetrailer.Checklist
Does this PR entail a change in behavior?