Repository navigation
Fix: Stack overflow in ScanIter::next when scanning across many delete… - #40
Open
nless (u-less) wants to merge 7 commits into
Open
nless (u-less) wants to merge 7 commits into
nless (u-less) wants to merge 7 commits into
Conversation
根因:bf-tree 合并页大小使用 u16/i16,32 KiB leaf page 合并时超过 65,535 字节发生溢出,导致错误分裂并触发 335, 105, 2364 panic。
- Fix WAL durability for the first appended entry - Flush pending WAL data before stopping the background worker - Replace recursive WAL segment rotation with an iterative wait loop - Write only modified, block-aligned WAL regions during flush - Safely handle invalid operation types during WAL recovery - Replay delete operations correctly when restoring snapshots - Change the default WAL segment size from 1 GiB to 1 MiB - Avoid temporary key allocations during range scans - Handle short boundary keys without panicking - Stop bounded scans promptly when deleted records exceed the bound - Add regression tests for WAL durability, segment rotation, and key comparison - Register the intentionally dormant SPDK configuration for lint checks
- replay WAL records using the actual on-disk WriteOp format - preserve WAL entries and monotonically increasing LSNs across restarts - handle short positioned I/O and malformed WAL segments safely - use lazy per-thread io_uring instances with a shared polling work queue - remove cache-line flushes from leaf search and avoid reference-bit pollution - reduce promotion RNG overhead and use native cross-platform atomic waits - harden aligned allocations and snapshot deserialization - reduce Shuttle snapshot memory usage and add recovery regression tests
Use allocation-bounded leaf views, preserve allocator provenance, and copy independent snapshot images before releasing leaf locks. Reduce read and scan overhead and record complete performance evidence. Validated with 130 unit tests, integration and doctests, 27 default Miri checks, and native interop coverage. Document remaining inner-node risks and the blocked file-I/O Miri retest.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fix: #39
BF-Tree Optimization Report
1. Overview
The optimization was performed on the local code.
39d9dab2. Correctness Improvements
Write-Ahead Log
WriteOpformat.BfTree.Snapshot Recovery
Tree and Leaf Operations
Configuration
3. Performance Improvements
Leaf Search
The production linear-search path previously executed
_mm_clflushfor every inspected metadata entry. This explicitly evicted useful cache lines and significantly increased lookup latency.The cache flush was removed, producing the largest measured improvement.
io_uring
RefCellaccess caused by thread-ID hash collisions.SendandSyncimplementations.IORING_SETUP_ATTACH_WQ.This follows the Linux per-thread ring and shared worker-pool model documented for high-performance io_uring workloads.
Promotion Decisions
0%and100%.Rc<UnsafeCell<SmallRng>>wrapper with direct thread-local RNG access.Atomic Waiting
atomic-waitimplementation.WaitOnAddress.WAL Startup
WAL startup now searches segments backward to find the latest LSN. Normally, only the final segment needs to be parsed instead of scanning the complete WAL.
4. Benchmark Methodology
Environment:
The baseline and optimized versions were built in separate target directories and executed in alternating order.
5. Benchmark Results
Latency changes:
The binary-search control changed by only 1.7%, confirming that the large linear-search gain primarily comes from removing
_mm_clflush, rather than environmental variation.The promotion-disabled fast path showed no material end-to-end improvement because page lookup and copying dominate that workload.
6. Validation
The optimized code passed:
cargo test --all-targets: 82/82 tests7. Limitations
8. Conclusion
The optimization delivers a verified 2.73–3.53× improvement on the affected linear-search paths without materially regressing the default binary-search path.
It also resolves several high-impact correctness issues involving WAL recovery, restart durability, short I/O, cache-only logging, snapshot alignment, prefix comparison, unsafe allocation, and io_uring concurrency.