Skip to content

perf(scan): reduce system column pipeline overhead - #8747

Merged
Xuanwo merged 7 commits into
lance-format:mainfrom
jiaoew1991:perf/system-column-scan-pipeline
Sep 4, 2026
Merged

perf(scan): reduce system column pipeline overhead#8747
Xuanwo merged 7 commits into
lance-format:mainfrom
jiaoew1991:perf/system-column-scan-pipeline

Conversation

@jiaoew1991

@jiaoew1991 jiaoew1991 commented Aug 25, 2026

Copy link
Copy Markdown
Contributor

Summary

Reduce fixed per-batch work in the system-column scan pipeline after the row-ID and version cursors have removed decoder prefix scans.

This supersedes #8717, which was merged into its temporary integration base while the dependency stack was being refreshed. No version of #8717 was merged into main.

This change:

  • reads stable row IDs ahead in batch-aligned chunks and serves uniform adjacent batches as zero-copy Arrow slices;
  • validates the complete prefetched selection before exposing or caching it;
  • drains the current row-ID chunk and falls back to exact per-task decoding when non-final task sizes vary, avoiding repeated boundary copies or cursor rewinds;
  • assembles all requested system columns into one RecordBatch for uniform task streams, while variable task streams retain the prior incremental assembly;
  • caches structurally equivalent schemas for uniform task streams, not only pointer-identical Arc<Schema> values, and disables that cache after task sizes vary;
  • reuses a full-size zero-column batch template;
  • skips no-op projections for system-only reads only when schemas are strictly equal;
  • bounds read-ahead by the actual range selection, including empty and tail reads.

For batches up to 64K rows, each cached row-ID chunk is at most 64K u64 values (512 KiB). A larger batch is decoded as one batch without additional read-ahead. Zero-copy slices can keep a chunk alive until adjacent batches are released; an irregular stream can make one boundary copy while draining the existing chunk.

Performance

Benchmarks compare main at a8bec27d5 (including #8716) with this branch, use release-with-debug, pin both revisions to the same CPU, and run each case in a fresh process.

The 100K-row uniform-task matrix improved all 16 combinations of batch size (64/1024), bitmap density (holes every 2/17 values), payload (absent/present), and system columns (_rowid/all):

  • batch size 64: _rowid improved 30.3%-33.7%; all system columns improved 51.4%-56.8%;
  • batch size 1024: _rowid improved 13.1%-26.5%; all system columns improved 29.5%-36.7%.

A standard Criterion run for the representative batch-size-64, holes-17, zero-payload, all-system-columns case measured 3.7000 ms on main and 1.7455 ms here (-52.8%). CPU profiles attribute the change: RecordBatchExt::try_with_column (7.47% self time on main) and SchemaExt::try_with_column (4.59%) both fell below the 0.1% reporting threshold.

The review reproducer uses 10M rows, holes-17, one payload column, _rowid, and alternating 32,768/32,769-row tasks. With isolated base/head target directories, standard Criterion measured 10.387-10.468 ms on main and 10.449-10.499 ms here (+0.51% by point estimate, within run noise). A second pair under perf record measured 10.699-10.746 ms and 10.652-10.724 ms, respectively. Profiles show the same dominant workload (SegmentCursorState::extend_dense_range, 42.17% / 42.22% self time); RecordBatchExt::try_with_column was only 0.16% / 0.10%, and schema reconstruction and boundary copying were not hotspots.

Validation

  • cargo test -p lance-table --lib (384 passed)
  • cargo test -p lance --lib dataset::fragment (108 passed)
  • cargo check -p lance-table --tests --benches
  • cargo clippy --all --tests --benches -- -D warnings
  • cargo fmt --all -- --check
  • git diff --check
  • Python integration: test_to_batches_with_partial_last_batch, test_to_batches, test_scan_no_columns, and test_roundtrip_reader (4 passed)
  • read-only production-scale matrix: 32/32 cases passed against an independent metadata oracle, with the dataset version asserted unchanged before and after

Targeted regressions cover mixed payload + all-system projection schema/order, uniform structurally equal schemas with different Arcs, variable-task fallback, stable-row-ID unsorted indices and metadata underfill, read-ahead tail/chunk boundaries, empty tasks, and system-only read_all / read_ranges.

Dependencies #8713, #8715, and #8716 are merged into main. This PR's diff is limited to the system-column benchmark and the two implementation files.

@jiaoew1991
jiaoew1991 changed the base branch from jiaoew/stack-system-column-base-v5 to main August 25, 2026 12:28
lance-gatekeeper[bot]

This comment was marked as outdated.

@lance-gatekeeper lance-gatekeeper Bot added the K-approved Latest Gatekeeper recommendation permits acceptance. label Aug 25, 2026
BubbleCal pushed a commit that referenced this pull request Aug 26, 2026
## Summary

Reuse a stable row-ID cursor across ordered record-batch tasks and
bulk-decode range-backed segments.

Stable row IDs represented as `RangeWithBitmap` / `RangeWithHoles`
previously rebuilt selection state for every batch. Sequential scans
therefore rescanned an ever-growing prefix, approaching quadratic work
as batch count increased.

This change:

- persists `RowIdSequenceCursor` across ordered tasks and caches segment
lengths;
- adds an exact-capacity contiguous-range path;
- adds `SegmentCursorState::extend_range` for bulk expansion of range
and bitmap segments;
- preserves the direct/random selection fallback and rejects unsorted
indices explicitly;
- reports truncated stable row-ID metadata as `CorruptFile` instead of
panicking or returning a short batch.

## Performance

100K-row synthetic sequential scan, identical Criterion harness:

| Batch size | `main` | This PR | Speedup |
|---:|---:|---:|---:|
| 64 | 17.177 ms | 1.174 ms | 14.6x |
| 1024 | 1.868 ms | 236.16 us | 7.9x |

CPU profiling on `main` attributed 51.73% to `RowIdSequence::select` and
33.45% to `U64Segment::len`, matching repeated prefix traversal. After
this change those hotspots are replaced by
`SegmentCursorState::extend_range`; remaining time is fixed
allocation/schema work.

## Validation

- `cargo test -p lance-table --lib` (336 passed)
- `cargo check -p lance-table --tests --benches`
- `cargo clippy -p lance-table --all-targets --no-deps -- -D warnings`
- `cargo fmt --all -- --check`
- `git diff --check`

This is the root of the row-ID optimization stack and has no dependency
on the bitmap or version-cursor follow-ups.

Follow-up PRs:

- #8715: dense / near-dense bitmap decode paths
- #8716: dataset-version RLE cursor
- #8747: system-column scan pipeline fixed-cost reduction
@jiaoew1991
jiaoew1991 force-pushed the perf/system-column-scan-pipeline branch from 53925c6 to 8229ef7 Compare September 3, 2026 03:21
@lance-gatekeeper lance-gatekeeper Bot removed the K-approved Latest Gatekeeper recommendation permits acceptance. label Sep 3, 2026
lance-gatekeeper[bot]

This comment was marked as outdated.

@lance-gatekeeper lance-gatekeeper Bot added the K-changes Latest Gatekeeper recommendation requests changes. label Sep 3, 2026
@lance-gatekeeper lance-gatekeeper Bot removed the K-changes Latest Gatekeeper recommendation requests changes. label Sep 3, 2026
lance-gatekeeper[bot]

This comment was marked as outdated.

@lance-gatekeeper lance-gatekeeper Bot added the K-changes Latest Gatekeeper recommendation requests changes. label Sep 3, 2026
@lance-gatekeeper lance-gatekeeper Bot removed the K-changes Latest Gatekeeper recommendation requests changes. label Sep 4, 2026
lance-gatekeeper[bot]

This comment was marked as outdated.

@lance-gatekeeper lance-gatekeeper Bot added the K-changes Latest Gatekeeper recommendation requests changes. label Sep 4, 2026
@lance-gatekeeper lance-gatekeeper Bot removed the K-changes Latest Gatekeeper recommendation requests changes. label Sep 4, 2026
lance-gatekeeper[bot]

This comment was marked as outdated.

@lance-gatekeeper lance-gatekeeper Bot added K-approved Latest Gatekeeper recommendation permits acceptance. and removed K-approved Latest Gatekeeper recommendation permits acceptance. labels Sep 4, 2026
@lance-gatekeeper lance-gatekeeper Bot removed the K-approved Latest Gatekeeper recommendation permits acceptance. label Sep 4, 2026

@lance-gatekeeper lance-gatekeeper Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Gate recommendation: approve.

The rebase preserves all five substantive patches exactly. Current-head stream and fragment regressions remain clean, and the intervening row-ID optimization is isolated to index/cache lookup rather than this stream-selection path, so the prior correctness and base-equivalent variable-task performance evidence remains valid.

@lance-gatekeeper lance-gatekeeper Bot added the K-approved Latest Gatekeeper recommendation permits acceptance. label Sep 4, 2026
@Xuanwo
Xuanwo merged commit fb46239 into lance-format:main Sep 4, 2026
38 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

K-approved Latest Gatekeeper recommendation permits acceptance. performance

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants