Repository navigation
fix: lock newly allocated pages before publishing mappings - #42
Konstantin V Shvachko (shvachko) wants to merge 1 commit into
Conversation
Acquire the base/mini page write guard under the mapping insertion mutex, before snapshot iterators can discover the new ID. Reproduce both publication races with bounded Shuttle regressions and document the initialization invariant. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Problem definition and reproductionThe allocation panic is caused by publishing the page-table entry before acquiring its exclusive initialization lock. A valid interleaving on the original code is:
The same sequence applies to both The required invariant is: a new mapping is already exclusively locked when its ID becomes visible, and the allocator keeps that guard until initialization is finished. Acquiring a blocking lock after publication is insufficient because the sweep could already be reading the new, not-yet-initialized page. The two regression tests run a real page-table allocation against a snapshot-style reader under bounded Shuttle DFS scheduling. The reader probes instead of spinning under an unfair scheduler; if it obtains a guard, it holds it while joining the allocator. On the original code, both tests reproduce: To run: For the negative control, keep the final regression tests but restore the two allocation call sites to Separate remaining issue: combined downstream validation with the snapshot fix still reproduced the recovery |
Problem
Both
PageTable::alloc_base_page_mappingandinsert_mini_page_mappingpublish an unlocked mapping before callingtry_write().unwrap(). Snapshot sweep discovers pages directly through the page-table iterator and can take a reader lock in that interval. Allocation then panics on real reader contention. The publication window also permits reading a page before initialization is finished.This is a demonstrated publication race, not merely a conjecture about spurious weak-CAS failure. The regression tests reproduce the original panics at
src/storage.rs:150:44andsrc/storage.rs:183:44with Shuttle.Fix
MappingTable::insert_with, which executes an initialization callback while the existing insertion mutex still prevents iterators from discovering the new ID.doc/snapshot-recovery.md.The callback only locks the fresh, unpublished entry and must not reenter the table. Changing to a blocking write lock after publication would not protect the initialization window; changing weak CAS to strong CAS would not remove genuine reader contention.
Regression tests
Two tests in
storage::tests::allocation_publicationcover base-page allocation and mini-page insertion. Each explores up to 10,000 DFS schedules using the existing Shuttle dependency. A snapshot-style iterator probes for a reader lock on the new entry; if acquired, it retains the guard while waiting for allocation to finish. Allocation must complete without panic/deadlock, and the reader must observe completed initialization.The nonblocking read probe avoids an unfair DFS schedule spinning indefinitely in the blocking lock. Shuttle's ordering/model limitations still apply; these tests do not claim exhaustive weak-memory verification.
Validation
try_write().unwrap()sites.Related downstream complete-state suite with both fixes combined:
src/snapshot.rs:1288:83, returningLockedThe failing integration run is retained, not retried until green. This PR fixes allocation publication; it does not fix or claim full downstream acceptance of the independent recovery lock-upgrade path.
Scope
Independent PR against upstream
main(ca6130715eeba872682747852b5da2cde32838aa). It does not include the cache-conversion fix from #38 / #41 and can be reviewed separately. No downstream source or dependency pin is changed.