Skip to content

Fix: Stack overflow in ScanIter::next when scanning across many delete… - #40

Open
nless (u-less) wants to merge 7 commits into
microsoft:mainfrom
u-less:fix#39
Open

nless (u-less) wants to merge 7 commits into
microsoft:mainfrom
u-less:fix#39

Conversation

@u-less

@u-less nless (u-less) commented Sep 3, 2026 •

Copy link
Copy Markdown

Fix: #39

BF-Tree Optimization Report

1. Overview

The optimization was performed on the local code.

  • Baseline: commit 39d9dab
  • Version base: BF-Tree v0.5.6
  • Optimized version: current uncommitted working tree
  • Primary objectives:
    • Fix correctness and recovery issues
    • Eliminate unsafe behavior
    • Improve read-path and concurrent I/O performance
    • Strengthen cross-platform support

2. Correctness Improvements

Write-Ahead Log

  • Fixed recovery parsing to use the actual on-disk WriteOp format.
  • Changed recovery to return the reconstructed BfTree.
  • Prevented recovered operations from being written to the WAL again.
  • Preserved existing WAL data when reopening a database.
  • Continued LSN allocation monotonically after restart.
  • Added bounds checking for malformed and partially written WAL entries.
  • Made WAL shutdown wait until the background worker has flushed and exited.
  • Fixed WAL usage in cache-only trees.
  • Added regression coverage for WAL reopening, corruption, recovery, and cache-only operation.

Snapshot Recovery

  • Fixed short-read handling in snapshot metadata and mapping data.
  • Preserved the required alignment for snapshot buffers.
  • Removed an intermediate byte-vector allocation during mapping recovery.
  • Fixed potentially misaligned conversion from byte buffers to typed vectors.
  • Added allocation-failure handling.
  • Reduced Shuttle snapshot memory from 1 GB per iteration to 4 MB.

Tree and Leaf Operations

  • Fixed prefix-compressed leaf lookups when the search key is shorter than the stored prefix.
  • Prevented negative lookups from marking the next greater record as referenced.
  • Fixed an out-of-bounds pointer calculation for inner-node keys shorter than the look-ahead size.
  • Added null checks for circular-buffer, inner-node, leaf-node, and aligned-buffer allocations.

Configuration

  • Removed panics for non-UTF-8 file paths.
  • Fixed default WAL path construction for paths without a parent directory.

3. Performance Improvements

Leaf Search

The production linear-search path previously executed _mm_clflush for every inspected metadata entry. This explicitly evicted useful cache lines and significantly increased lookup latency.

The cache flush was removed, producing the largest measured improvement.

io_uring

  • Replaced eagerly allocated, hash-selected rings with lazily created thread-local rings.
  • Eliminated possible concurrent RefCell access caused by thread-ID hash collisions.
  • Removed unsafe manual Send and Sync implementations.
  • Polling rings now share a kernel work queue through IORING_SETUP_ATTACH_WQ.

This follows the Linux per-thread ring and shared worker-pool model documented for high-performance io_uring workloads.

Promotion Decisions

  • Added fast paths for promotion rates of 0% and 100%.
  • Avoided random-number generation for those common configurations.
  • Replaced the cloneable Rc<UnsafeCell<SmallRng>> wrapper with direct thread-local RNG access.

Atomic Waiting

  • Replaced duplicated platform-specific wait implementations with the maintained atomic-wait implementation.
  • Linux uses futexes.
  • Windows uses WaitOnAddress.
  • macOS now uses a real blocking wait instead of busy polling.

WAL Startup

WAL startup now searches segments backward to find the latest LSN. Normally, only the final segment needs to be parsed instead of scanning the complete WAL.

4. Benchmark Methodology

Environment:

  • CPU: Intel Core i7-11700K, 8 cores / 16 logical processors
  • Operating system: Windows
  • Rust: 1.98.0
  • Build: Release, optimization level 3, Thin LTO, one codegen unit
  • Records: 65,536
  • Operations per sample: 524,288
  • Samples per version and workload: 15
  • Query sequence: deterministic
  • Result validation: all before/after checksums matched

The baseline and optimized versions were built in separate target directories and executed in alternating order.

5. Benchmark Results

Workload Before After Throughput Gain Speedup
Linear-search hit, 1 thread 0.889 Mops/s 3.075 Mops/s +246.1% 3.46×
Linear-search miss, 1 thread 0.869 Mops/s 3.067 Mops/s +252.8% 3.53×
Linear-search hit, 8 threads 5.718 Mops/s 15.605 Mops/s +172.9% 2.73×
Binary-search hit, 1 thread—control 4.278 Mops/s 4.350 Mops/s +1.7% 1.02×
Base-page hit, promotion disabled 3.734 Mops/s 3.749 Mops/s +0.4% 1.00×

Latency changes:

Workload Before After Reduction
Linear-search hit, 1 thread 1,125 ns/op 325 ns/op 71.1%
Linear-search miss, 1 thread 1,150 ns/op 326 ns/op 71.7%
Linear-search hit, 8 threads 175 ns/op 64 ns/op 63.4%

The binary-search control changed by only 1.7%, confirming that the large linear-search gain primarily comes from removing _mm_clflush, rather than environmental variation.

The promotion-disabled fast path showed no material end-to-end improvement because page lookup and copying dominate that workload.

6. Validation

The optimized code passed:

  • cargo test --all-targets: 82/82 tests
  • Documentation tests: 10 passed, 2 ignored
  • Strict library Clippy with all features
  • Release build
  • Formatting and diff validation
  • Linux x86-64 cross-compilation
  • macOS x86-64 cross-compilation
  • macOS ARM64 cross-compilation
  • Shuttle core-tree scheduling: 4 × 4,000 executions
  • Shuttle cache-only snapshot scheduling: 4 × 1,000 executions
  • Shuttle disk snapshot scheduling: 4 × 100 executions

7. Limitations

  • io_uring performance was not measured because the current host is Windows. Reliable results require Linux, a supported kernel, and preferably a dedicated NVMe device.
  • WAL changes primarily improve correctness, durability, restart behavior, and startup scalability; they are not represented by the in-memory lookup benchmark.
  • The repository’s full 100-million-record benchmark was not run because it requires a dedicated Linux environment with substantially more memory and storage resources.

8. Conclusion

The optimization delivers a verified 2.73–3.53× improvement on the affected linear-search paths without materially regressing the default binary-search path.

It also resolves several high-impact correctness issues involving WAL recovery, restart durability, short I/O, cache-only logging, snapshot alignment, prefix comparison, unsafe allocation, and io_uring concurrency.

根因:bf-tree 合并页大小使用 u16/i16,32 KiB leaf page 合并时超过 65,535 字节发生溢出,导致错误分裂并触发 335, 105, 2364 panic。
- Fix WAL durability for the first appended entry
- Flush pending WAL data before stopping the background worker
- Replace recursive WAL segment rotation with an iterative wait loop
- Write only modified, block-aligned WAL regions during flush
- Safely handle invalid operation types during WAL recovery
- Replay delete operations correctly when restoring snapshots
- Change the default WAL segment size from 1 GiB to 1 MiB
- Avoid temporary key allocations during range scans
- Handle short boundary keys without panicking
- Stop bounded scans promptly when deleted records exceed the bound
- Add regression tests for WAL durability, segment rotation, and key comparison
- Register the intentionally dormant SPDK configuration for lint checks
- replay WAL records using the actual on-disk WriteOp format
- preserve WAL entries and monotonically increasing LSNs across restarts
- handle short positioned I/O and malformed WAL segments safely
- use lazy per-thread io_uring instances with a shared polling work queue
- remove cache-line flushes from leaf search and avoid reference-bit pollution
- reduce promotion RNG overhead and use native cross-platform atomic waits
- harden aligned allocations and snapshot deserialization
- reduce Shuttle snapshot memory usage and add recovery regression tests
@u-less nless (u-less) changed the title Fix Stack overflow in ScanIter::next when scanning across many delete… Fix: Stack overflow in ScanIter::next when scanning across many delete… Sep 6, 2026
Use allocation-bounded leaf views, preserve allocator provenance, and copy independent snapshot images before releasing leaf locks. Reduce read and scan overhead and record complete performance evidence.

Validated with 130 unit tests, integration and doctests, 27 default Miri checks, and native interop coverage. Document remaining inner-node risks and the blocked file-I/O Miri retest.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Stack overflow in ScanIter::next when scanning across many deleted records

1 participant