Skip to content

25862: fix: apply pushed-down fetch to filter statistics - #378

Open
martin-augment wants to merge 1 commit into
mainfrom
pr-25862-2026-09-29-06-32-10
Open

martin-augment wants to merge 1 commit into
mainfrom
pr-25862-2026-09-29-06-32-10

Conversation

@martin-augment

Copy link
Copy Markdown
Owner

25862: To review by AI

@coderabbitai

coderabbitai Bot commented Sep 29, 2026 •

Copy link
Copy Markdown

Review in Change Stack →

Navigate logical layers of code changes, visualize relationships, and explore their blast radius.

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Advanced

Run ID: b641cf36-7446-4617-91f0-9816e9cc0a8f

📥 Commits

Reviewing files that changed from the base of the PR and between d2ef8e5 and 0f34623.

📒 Files selected for processing (2)
  • datafusion/physical-plan/src/filter.rs
  • datafusion/physical-plan/src/operator_statistics/mod.rs

Included review availability: This review used your included allowance. Your plan provides up to 1 included review per hour; 0 remain after this review.


Summary

  • FilterExec now recognizes bare IS NULL and IS NOT NULL checks on columns and uses them to refine statistics.
  • Fetch limits now cap row counts in filter statistics. The calculation handles overall and per-partition statistics, updates byte and column counts, and preserves supported exact facts.
  • FilterStatisticsProvider applies the fetch limit before projecting statistics.
  • Tests cover fetch-limited singleton statistics, execution and statistics agreement across one or two partitions, null-column statistics, and fetch-limited provider statistics.

Walkthrough

FilterExec now applies fetch limits to overall and per-partition statistics. It also handles bare IS NULL predicates and constructs exact zero counts for contradictory predicates. FilterStatisticsProvider applies fetch limits before projecting statistics. Tests cover fetch-limited counts, partition behavior, null statistics, and singleton column statistics.

Priority: ⬇️ Low

Merge Risk: ⚪ Minimal · up to 0f346

No material fetch-statistics issue remains identified; this change is mergeable after normal checks.

✨ Finishing Touches
📝 Generate docstrings
  • Commit to this branch
  • Create a new PR
🧪 Generate unit tests (beta)
  • Commit to this branch
  • Create a new PR

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review: FilterExec Fetch Statistics

Summary

The change adds fetch-aware statistics to FilterExec, refactors empty-column-statistics construction, and introduces a null_check_column helper. Overall the logic is reasonable, but there are several issues worth addressing.


Issues & Recommendations

1. null_check_column uses TypeId + downcast_ref instead of downcast_ref alone (Minor / Style)

let type_id = (predicate.as_ref() as &dyn Any).type_id();
let (arg, is_null) = if type_id == TypeId::of::<IsNullExpr>() {
    (predicate.downcast_ref::<IsNullExpr>()?.arg(), true)
} else if type_id == TypeId::of::<IsNotNullExpr>() {
    ...

The TypeId check is redundant: downcast_ref::<T>() already returns None when the concrete type doesn't match. The comment claims "One type check keeps every other predicate cheap," but downcast_ref performs the same TypeId comparison internally. This adds an extra import (std::any::{Any, TypeId}) and complexity for no measurable benefit.

Recommendation: Simplify to:

if let Some(expr) = predicate.downcast_ref::<IsNullExpr>() {
    Some((expr.arg().downcast_ref::<Column>()?.index(), true))
} else if let Some(expr) = predicate.downcast_ref::<IsNotNullExpr>() {
    Some((expr.arg().downcast_ref::<Column>()?.index(), false))
} else {
    None
}

2. statistics_under_fetch early-return condition may be too conservative (Correctness / Logic)

if stats.num_rows.get_value().is_some_and(|rows| *rows <= fetch) {
    // No partition reaches the fetch, so no rows are dropped.
    return Ok(stats);
}

This is only valid when partition.is_none() (overall stats). When partition = Some(p) and the overall num_rows <= fetch, that's fine per-partition too. But consider the inverse: when partition.is_none() and num_rows > fetch, the code proceeds to compute bound = fetch * partitions. That's correct.

However, when partition = Some(p) and num_rows > fetch, single = true and bound = fetch, which is correct. The logic holds, but the comment "No partition reaches the fetch" is misleading when partition.is_some() — it's the given partition that doesn't reach the fetch. Consider clarifying.


3. empty_column_statistics() drops min_value/max_value/sum_value semantics (Potential regression)

The old code explicitly set:

col_stat.min_value = Precision::Absent;
col_stat.max_value = Precision::Absent;
col_stat.sum_value = Precision::Absent;

The new helper uses ..ColumnStatistics::new_unknown(). You should verify that new_unknown() sets min_value, max_value, and sum_value to Precision::Absent (not Inexact or some other default). If new_unknown() leaves them as Absent, this is equivalent and fine. If not, this is a silent behavior change. Please confirm.

Also note the old code preserved input_stats.to_inexact() as the base (so byte_size etc. were Inexact), while the new code builds fresh stats. Since all fields are overwritten to Exact(0)/Absent, this should be equivalent — but worth a test asserting sum_value/min_value/max_value are Absent.


4. statistics_under_fetch mutates null_count to stats.num_rows without checking precision (Correctness)

if let Some((column, true)) = null_check_column(self.predicate())
    && let Some(column_stats) = stats.column_statistics.get_mut(column)
{
    // Every surviving row is still null after the fetch.
    column_stats.null_count = stats.num_rows;
}

stats.num_rows may be Inexact or Absent. Assigning an Inexact/Absent value to null_count is fine type-wise, but the comment "Every surviving row is still null" asserts an exact fact. If num_rows is Absent, null_count becomes Absent — losing the exactness that the predicate guarantees. Consider:

if let Precision::Exact(rows) = stats.num_rows {
    column_stats.null_count = Precision::Exact(rows);
}

or at least document that null_count inherits num_rows' precision.


5. cap_at_rows behavior when num_rows is Absent (Verify)

column_stats.null_count = cap_at_rows(column_stats.null_count, stats.num_rows);

Ensure cap_at_rows handles Precision::Absent for the row count gracefully (returning the original or Absent). Not shown in the diff — worth a quick check.


6. Test test_filter_fetch_statistics_match_execution assumes exact fetch semantics (Test robustness)

assert_eq!(stats.num_rows.get_value(), Some(&(fetch * partitions)));

This asserts Exact(fetch * partitions). For fetch = 0 and partitions = 2, that's Exact(0) — fine. But the test relies on TestMemoryExec reporting exact row counts and the fetch being applied uniformly. If statistics_with_fetch ever returns Inexact for the multi-partition case (which the code path suggests it might via to_inexact()), this assertion would fail. The test passes today, but it's tightly coupled to the implementation's precision choices. Consider asserting get_value() equality rather than exact precision, or add a comment explaining why Exact is expected.


7. Duplicate fetch application risk (Correctness)

statistics_with_fetch is now called in two places:

  • FilterExec::statistics_from_inputs (line ~755)
  • FilterStatisticsProvider::compute (operator_statistics/mod.rs)

If both paths can be invoked for the same plan (e.g., provider delegates to statistics_from_inputs), the fetch could be applied twice. Verify these are mutually exclusive code paths. If FilterStatisticsProvider calls into FilterExec::statistics_from_inputs internally, this is a double-application bug.


8. fetch.saturating_mul(partitions) overflow semantics (Minor)

fetch.saturating_mul(partitions) saturates at usize::MAX. That's safe, but the resulting bound passed to with_fetch may be nonsensical (larger than any real row count). Since with_fetch presumably caps at num_rows, this is likely harmless — but a comment noting the saturation intent would help.


9. statistics_with_fetch visibility (API surface)

pub(crate) is appropriate, but the doc comment says "for partition, or for all partitions when partition is None." The parameter is Option<usize> — consider documenting what Some(p) means when p >= partition_count (out-of-range). Currently it's silently treated as a single partition.


Positive Notes

  • Good test coverage, including an execution-vs-statistics consistency test.
  • The empty_column_statistics helper reduces duplication.
  • The null_check_column helper is a clean abstraction (modulo issue #1).
  • Comments explain the "fetch only drops rows" invariant well.

Action Items

Priority Item
High Verify #3 (new_unknown() field defaults) and #7 (double fetch application)
Medium Fix #4 (null_count precision) and #1 (redundant TypeId)
Low Clarify comments (#2, #8, #9), review test precision assumptions (#6)

@claude

claude Bot commented Sep 29, 2026

Copy link
Copy Markdown

Review

Overall this is a sound fix. FilterExec statistics now respect a pushed-down fetch, and both the statistics_from_inputs path and FilterStatisticsProvider use the same helper. I did not build or run the tests.

What I checked and found correct

  • The early return when rows <= fetch is safe for multi-partition plans. If the total is at most fetch, no partition can exceed it.
  • For multiple partitions, to_inexact() combined with a fetch * partitions bound is a reasonable upper-bound estimate. Saturating the multiplication avoids overflow.
  • null_check_column runs on the pre-projection stats, so the column index matches the input schema.
  • Carrying stats.num_rows into null_count for IS NULL keeps the precision consistent.

Suggestions

  1. Provider ignores partitions. FilterStatisticsProvider passes None for the partition. A partition-specific query through the registry therefore gets fetch * partitions rather than fetch. That is still a valid upper bound, but it is looser than the statistics_from_inputs path. A short comment saying so would help, or the provider could take the partition from its context if one is available.
  2. Behaviour change in the infeasible-predicate path. It now builds fresh empty_column_statistics() instead of starting from input_stats.to_inexact(). That is arguably more correct, but it is a semantic change that is not tied to the fetch fix. It also drops the previous "value stats undefined" handling on sum_value, which is now covered by new_unknown(). Please mention it in the PR description, and consider adding a small test for the infeasible case if none exists.
  3. Skewed partitions. fetch * partitions assumes every partition can reach fetch. Under skew, the true total can be smaller than min(rows, fetch * partitions). That is fine because the result is Inexact, but the doc comment could say the value is only an upper-bound estimate.
  4. Singleton preservation. This relies on null_count == Exact(0) plus is_singleton(), and it only keeps distinct_count when it is Exact(1). That is correct, but the logic is dense. Extracting it into a small helper would make the loop easier to read.
  5. Nit. #[inline] on a pub(crate) method that only dispatches is unnecessary. The PR description is also just a placeholder, so please fill it in per the PR template.

Tests
Coverage is good. It includes a stats-versus-execution consistency check across 1 and 2 partitions with fetch of 0 and 3, the singleton cases, the IS NULL byte-size cases, and a provider test. Two additions would help:

  • a case with fetch larger than the row count on multiple partitions;
  • an IS NOT NULL predicate combined with a fetch.

Please run cargo fmt --all and cargo clippy --all-targets --all-features -- -D warnings before merging, as required by CLAUDE.md.

@augmentcode

augmentcode Bot commented Sep 29, 2026

Copy link
Copy Markdown
🤖 Augment PR Summary

Summary: Applies FilterExec's pushed-down per-partition fetch limit to reported output statistics.

Changes:

  • Adds a shared helper that applies filter fetch bounds for overall and partition-specific statistics.
  • Uses the helper from the standard `ExecutionPlan` statistics path.
  • Uses the same helper from the enhanced `FilterStatisticsProvider` path.
  • Scales row and byte estimates through the existing `Statistics::with_fetch` implementation.
  • Downgrades aggregate multi-partition statistics when partition distribution is unknown.
  • Preserves exact empty, null-free, singleton, and all-null column invariants where derivable.
  • Centralizes exact-empty column statistics construction.
  • Adds coverage for singleton columns, null columns, execution/statistics agreement, and the provider path.

Technical Notes: Fetch is enforced independently by each FilterExec partition, so aggregate estimates use fetch × partition_count as their bound.

🤖 Was this summary useful? React with 👍 or 👎

@augmentcode augmentcode Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review completed. 1 suggestion posted.

Fix All in Augment

Comment augment review to trigger a new review at any time.

stats.to_inexact()
};

let mut stats = stats.with_fetch(Some(bound), 0, 1)?;

@augmentcode augmentcode Bot Sep 29, 2026 •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

datafusion/physical-plan/src/filter.rs:496: When fetch == 0 and the filtered input row count is Inexact or Absent, Statistics::with_fetch returns Inexact(0), so this exact-zero branch is skipped. Execution emits no rows for a zero fetch, but the output retains inexact zero/null/byte statistics and potentially value bounds, preventing consumers from recognizing the provably empty result.

Severity: medium

Fix This in Augment

🤖 Was this useful? React with 👍 or 👎, or 🚀 if it prevented an incident/outage.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants