Skip to content

fix(postprocessing): preserve NULL grouping index values in pivot() - #43694

Open
Archita-kale wants to merge 5 commits into
apache:masterfrom
Archita-kale:fix/pivot-null-grouping-keys
Open

Archita-kale wants to merge 5 commits into
apache:masterfrom
Archita-kale:fix/pivot-null-grouping-keys

Conversation

@Archita-kale

@Archita-kale Archita-kale commented Aug 30, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Follow-up fix for #43547 to preserve NULL/NaN grouping values through the pandas post-processing pivot() operator.

Problem

NULL values in the columns parameter were already handled using Superset's existing NULL_STRING before calling pivot_table(). However, index columns did not have equivalent handling.

As a result, when a NULL grouping value survived the aggregate step, pandas pivot_table() could drop that row during pivoting because the grouping key contained NaN.

This caused valid NULL groups to disappear from the pivot output.

Solution

  • Fill NULL/NaN values in index columns with the existing NULL_STRING before calling pivot_table().
  • Handle categorical index columns by adding NULL_STRING to their categories before filling.
  • Handle categorical columns values when the configured fill value is not already present in the categories.
  • Preserve the existing drop_missing_columns / dropna behavior.
  • Reuse the existing NULL_STRING constant.
  • Add regression tests covering NULL index values in flat and MultiIndex pivots, including categorical dtypes.

Tests

Added regression tests for:

  • Flat pivot with NULL index values
  • Pivot with NULL index values and columns grouping
  • Categorical index with NULL values
  • Categorical index and columns with NULL values

Tested with:

pytest tests/unit_tests/pandas_postprocessing/test_pivot.py -v

@kokhlo I’ve submitted this follow-up PR for the NULL index handling issue identified in the discussion. The change preserves NULL grouping values through the pivot post-processing path and includes regression tests. Would appreciate your review!

Fixes #43547

@dosubot dosubot Bot added the change:backend Requires changing the backend label Aug 30, 2026
@bito-code-review

bito-code-review Bot commented Aug 30, 2026 •

Copy link
Copy Markdown
Contributor

Code Review Agent Run #523df3

Actionable Suggestions - 0
Filtered by Review Rules

Bito filtered these suggestions based on rules created automatically for your feedback. Manage rules.

  • superset/utils/pandas_postprocessing/pivot.py - 1
Review Details
  • Files reviewed - 2 · Commit Range: 464adc9..464adc9
    • superset/utils/pandas_postprocessing/pivot.py
    • tests/unit_tests/pandas_postprocessing/test_pivot.py
  • Files skipped - 0
  • Tools
    • MyPy (Static Code Analysis) - ✔︎ Successful
    • Astral Ruff (Static Code Analysis) - ✔︎ Successful
    • Whispers (Secret Scanner) - ✔︎ Successful
    • Detect-secrets (Secret Scanner) - ✔︎ Successful

Bito Usage Guide

Commands

Type the following command in the pull request comment and save the comment.

  • /review - Manually triggers an incremental AI Review.

  • /review full - Manually triggers a full AI Review.

  • /pause - Pauses automatic reviews on this pull request.

  • /resume - Resumes automatic reviews.

  • /resolve - Marks all Bito-posted review comments as resolved.

  • /abort - Cancels all in-progress reviews.

Refer to the documentation for additional commands.

Configuration

This repository uses Superset You can customize the agent settings here or contact your Bito workspace admin at evan@preset.io.

Documentation & Help

AI Code Review powered by Bito Logo

Comment thread superset/utils/pandas_postprocessing/pivot.py Outdated
@bito-code-review

Copy link
Copy Markdown
Contributor

The flagged issue is correct. Filling a datetime index with a string like <NULL> will cause a TypeError because pandas datetime dtypes cannot hold string values. To resolve this, you should check the dtype of the index column before filling it. If the column is a datetime type, you should either leave the NaT values as they are or convert the column to an object type before filling it with the string.

superset/utils/pandas_postprocessing/pivot.py

for col in index:
        if (
            isinstance(df[col].dtype, pd.CategoricalDtype)
            and NULL_STRING not in df[col].cat.categories
        ):
            df[col] = df[col].cat.add_categories([NULL_STRING])
        elif pd.api.types.is_datetime64_any_dtype(df[col].dtype):
            continue
        df[col] = df[col].fillna(value=NULL_STRING)

@codecov

codecov Bot commented Aug 30, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 81.52%. Comparing base (ca3d356) to head (f2c9974).
⚠️ Report is 102 commits behind head on master.

Additional details and impacted files
@@           Coverage Diff           @@
##           master   #43694   +/-   ##
=======================================
  Coverage   81.51%   81.52%           
=======================================
  Files        2973     2973           
  Lines      180273   180319   +46     
  Branches    41734    41739    +5     
=======================================
+ Hits       146949   146998   +49     
+ Misses      30601    30599    -2     
+ Partials     2723     2722    -1     
Flag Coverage Δ
hive 36.69% <6.66%> (-0.01%) ⬇️
mysql 55.98% <46.66%> (-0.02%) ⬇️
postgres 55.98% <46.66%> (-0.02%) ⬇️
presto 38.61% <6.66%> (-0.02%) ⬇️
python 85.66% <100.00%> (+0.01%) ⬆️
sqlite 55.71% <46.66%> (-0.01%) ⬇️
unit 78.40% <100.00%> (+0.01%) ⬆️

Flags with carried forward coverage won't be shown. Click here to find out more.

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

@Archita-kale

Copy link
Copy Markdown
Contributor Author

I’ve implemented the follow-up fix for the NULL grouping issue in the pivot() post-processing path.

The fix preserves NULL/NaN/NaT grouping values in index columns while handling categorical and datetime dtypes appropriately. I’ve also added regression coverage for the affected pivot scenarios.

The changes are now included in this PR and ready for review. Thanks for identifying the remaining aggregate → pivot issue!

@Archita-kale

Copy link
Copy Markdown
Contributor Author

Thanks for the coverage feedback! I’ve added the missing regression test coverage for the newly introduced pivot handling, including the relevant categorical and NULL-value paths. The updated tests cover the previously uncovered lines, and the existing pivot test suite has been run successfully. The changes have been pushed to this PR for review.

@bito-code-review

bito-code-review Bot commented Aug 30, 2026 •

Copy link
Copy Markdown
Contributor

Code Review Agent Run #6f5c70

Actionable Suggestions - 0
Filtered by Review Rules

Bito filtered these suggestions based on rules created automatically for your feedback. Manage rules.

  • superset/utils/pandas_postprocessing/pivot.py - 1
Review Details
  • Files reviewed - 2 · Commit Range: 464adc9..8d13196
    • superset/utils/pandas_postprocessing/pivot.py
    • tests/unit_tests/pandas_postprocessing/test_pivot.py
  • Files skipped - 0
  • Tools
    • MyPy (Static Code Analysis) - ✔︎ Successful
    • Astral Ruff (Static Code Analysis) - ✔︎ Successful
    • Whispers (Secret Scanner) - ✔︎ Successful
    • Detect-secrets (Secret Scanner) - ✔︎ Successful

Bito Usage Guide

Commands

Type the following command in the pull request comment and save the comment.

  • /review - Manually triggers an incremental AI Review.

  • /review full - Manually triggers a full AI Review.

  • /pause - Pauses automatic reviews on this pull request.

  • /resume - Resumes automatic reviews.

  • /resolve - Marks all Bito-posted review comments as resolved.

  • /abort - Cancels all in-progress reviews.

Refer to the documentation for additional commands.

Configuration

This repository uses Superset You can customize the agent settings here or contact your Bito workspace admin at evan@preset.io.

Documentation & Help

AI Code Review powered by Bito Logo

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Preserves NULL grouping keys during pandas pivot post-processing.

Changes:

  • Adds dtype-aware dimension NULL filling.
  • Adds regression coverage for categorical, datetime, numeric, and MultiIndex pivots.

Reviewed changes

Copilot reviewed 2 out of 2 changed files in this pull request and generated 2 comments.

File Description
superset/utils/pandas_postprocessing/pivot.py Fills missing pivot dimensions before pivoting.
tests/unit_tests/pandas_postprocessing/test_pivot.py Adds NULL-preservation regression tests.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread superset/utils/pandas_postprocessing/pivot.py Outdated
Comment thread superset/utils/pandas_postprocessing/pivot.py Outdated
@rusackas

rusackas commented Sep 4, 2026

Copy link
Copy Markdown
Member

@Archita-kale nice catch on the root cause, but _fill_dimension_column in pivot.py:17-21 adds the <NULL> category to every categorical dimension even when there's no null present. Per @sadpandajoe's comment, this creates a fake zero-value group in pivots that never had one. Should be gated on s.isna().any() like the datetime branch.

Also still open: the datetime path stringifies the whole column instead of just the missing values, which per copilot's thread breaks the epoch serializer downstream in charts/data/api.py, and timedelta64 columns aren't caught by that check so they'll still hit fillna() and raise. Those threads are still unresolved, want to take another pass before this is mergeable?

@rusackas

Copy link
Copy Markdown
Member

No update since my last comment naming the three open threads (the datetime-NaT TypeError risk, the spurious <NULL> category on zero-value dims, and the datetime-stringify/timedelta64 crash risk). What has changed is CI — most required checks are now red, though the failure pattern looks more like stale/infra breakage than new logic failures. I'll see about a rebase, which might clear CI...

@rusackas
rusackas force-pushed the fix/pivot-null-grouping-keys branch from 8d13196 to 4c58a06 Compare September 10, 2026 04:00
@github-actions

Copy link
Copy Markdown
Contributor

⚠️ Translation Regression Detected

A source change in this PR renamed or reworded strings, invalidating existing translations (they are now #, fuzzy) in zh. Please resolve the affected .po files before merging.

Note: neither intentionally deleting a translatable string nor filling a previously-untranslated entry with a fuzzy guess (e.g. an AI backfill) is a regression — only a confirmed translation that a renamed/reworded source string turned fuzzy is flagged here.

Language Invalidated translations
zh 7

How to fix

1. Install dependencies (if not already set up):

pip install -r superset/translations/requirements.txt
sudo apt-get install gettext   # or: brew install gettext

2. Re-extract strings and sync .po files:

./scripts/translations/babel_update.sh

This rewrites superset/translations/messages.pot from the current source files and merges the changes into every .po file. Strings whose msgid changed will be marked #, fuzzy.

3. Resolve the fuzzy entries in the affected language files (zh):

grep -n '#, fuzzy' superset/translations/<lang>/LC_MESSAGES/messages.po

For each fuzzy entry, either rewrite the msgstr to match the new string and remove the #, fuzzy line, or clear the msgstr to "" if you cannot provide a translation.

4. Commit your changes to the .po files.

@bito-code-review

bito-code-review Bot commented Sep 10, 2026 •

Copy link
Copy Markdown
Contributor

Code Review Agent Run #ec8993

Actionable Suggestions - 0
Review Details
  • Files reviewed - 2 · Commit Range: 8d13196..4c58a06
    • superset/utils/pandas_postprocessing/pivot.py
    • tests/unit_tests/pandas_postprocessing/test_pivot.py
  • Files skipped - 0
  • Tools
    • MyPy (Static Code Analysis) - ✔︎ Successful
    • Astral Ruff (Static Code Analysis) - ✔︎ Successful
    • Whispers (Secret Scanner) - ✔︎ Successful
    • Detect-secrets (Secret Scanner) - ✔︎ Successful

Bito Usage Guide

Commands

Type the following command in the pull request comment and save the comment.

  • /review - Manually triggers an incremental AI Review.

  • /review full - Manually triggers a full AI Review.

  • /pause - Pauses automatic reviews on this pull request.

  • /resume - Resumes automatic reviews.

  • /resolve - Marks all Bito-posted review comments as resolved.

  • /abort - Cancels all in-progress reviews.

Refer to the documentation for additional commands.

Configuration

This repository uses Superset You can customize the agent settings here or contact your Bito workspace admin at evan@preset.io.

Documentation & Help

AI Code Review powered by Bito Logo

Comment thread superset/utils/pandas_postprocessing/pivot.py
@rusackas

Copy link
Copy Markdown
Member

@Archita-kale just echoing that @sadpandajoe caught one more issue on pivot.py:217. Nullable extension dtypes (Int64, Float64, boolean) still crash on fillna(<NULL>) the same way the datetime/timedelta and categorical cases from the earlier threads do. That's four separate dtype paths that can raise or silently misreport, still open.

Might be worth handling this by upcasting to object dtype once in _fill_dimension_column, rather than adding a branch every time a new dtype turns up.

Archita-kale and others added 4 commits September 28, 2026 22:41
…categorical dtypes in pivot() (apache#43547)

- Fill NULL/NaN values in index columns with NULL_STRING (<NULL>) prior to calling pivot_table()
- For categorical index and column dtypes, add the fill value to cat.categories if not already present before calling fillna()
- Preserve existing drop_missing_columns behavior without altering dropna=drop_missing_columns
- Add regression tests for flat index, MultiIndex with columns, and categorical index/column dimensions with NULL values
…ors in pivot() (apache#43547)

- Convert datetime dimensions containing NaT to string representation with NULL_STRING (<NULL>)
- Prevent TypeError on datetime64 arrays and mixed-type MultiIndex sorting
- Add regression tests for datetime index with NaT (flat, MultiIndex, and timezone-aware)
…ic pivot dimensions (apache#43547)

- test_pivot_preserves_null_index_value_categorical_already_in_categories
- test_pivot_categorical_column_with_null
- test_pivot_categorical_column_already_in_categories
- test_pivot_preserves_null_numeric_index_value
Co-Authored-By: Evan Rusackas <evan@preset.io>
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
@rusackas
rusackas force-pushed the fix/pivot-null-grouping-keys branch from 4c58a06 to bcafb71 Compare September 29, 2026 05:48
@bito-code-review

bito-code-review Bot commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

Code Review Agent Run #723801

Actionable Suggestions - 0
Additional Suggestions - 1
  • tests/unit_tests/pandas_postprocessing/test_pivot.py - 1
    • Duplicated test boilerplate · Line 885-1142
      The 11 added tests repeat the same pattern (build DataFrame, call `pivot()`, assert `NULL_STRING in result.index`/columns, assert value) with only the dtype/scenario differing. Consider `pytest.mark.parametrize` over the cases to cut ~250 duplicated lines and make the dtype matrix explicit.
Review Details
  • Files reviewed - 2 · Commit Range: 783e99b..bcafb71
    • superset/utils/pandas_postprocessing/pivot.py
    • tests/unit_tests/pandas_postprocessing/test_pivot.py
  • Files skipped - 0
  • Tools
    • MyPy (Static Code Analysis) - ✔︎ Successful
    • Astral Ruff (Static Code Analysis) - ✔︎ Successful
    • Whispers (Secret Scanner) - ✔︎ Successful
    • Detect-secrets (Secret Scanner) - ✔︎ Successful

Bito Usage Guide

Commands

Type the following command in the pull request comment and save the comment.

  • /review - Manually triggers an incremental AI Review.

  • /review full - Manually triggers a full AI Review.

  • /pause - Pauses automatic reviews on this pull request.

  • /resume - Resumes automatic reviews.

  • /resolve - Marks all Bito-posted review comments as resolved.

  • /abort - Cancels all in-progress reviews.

Refer to the documentation for additional commands.

Configuration

This repository uses Superset You can customize the agent settings here or contact your Bito workspace admin at evan@preset.io.

Documentation & Help

AI Code Review powered by Bito Logo

…pivot() NULL fill

Addresses the three review threads still open on this PR: the categorical
branch added NULL_STRING as a category unconditionally, creating a spurious
all-zero '<NULL>' group in pivots that never had a missing value; the
datetime special case stringified every value (not just the missing ones)
to work around NaT, rewriting valid labels and breaking the epoch
serializer downstream; and timedelta64 (kind "m") and pandas' nullable
extension dtypes (Int64, Float64, boolean) weren't handled at all, so a
NULL group on one of those raised TypeError instead of being preserved.

Rewrites _fill_dimension_column with one rule: skip columns with no missing
values entirely (fixes the categorical case and avoids unnecessary casts),
and for dtypes that reject a string sentinel outright -- datetime64,
timedelta64, and any pandas ExtensionDtype -- cast to `object` first so
valid values keep their real type and only the missing slots become the
sentinel, instead of stringifying the whole column.

Adds five regression tests: categorical with no nulls (asserts no spurious
group), datetime with a null (asserts valid entries stay real Timestamp
objects), timedelta64 with a null, and Int64/boolean nullable dtypes with a
null. Confirmed each fails on the prior code before this fix and passes
after.

Rebased onto current master to clear the stale-branch babel-extract/pot
failures; full pandas_postprocessing suite (180 tests) and ruff both pass.
pylint verified manually against the project's own venv (10/10) -- the
pre-commit hook itself failed here on a broken local environment (system
`pylint` on PATH resolves to an unrelated pyenv shim, not this project's
venv), unrelated to this change.

Co-Authored-By: Archita Kale <Archita-kale@users.noreply.github.com>
Co-Authored-By: Evan Rusackas <evan@preset.io>
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
@netlify

netlify Bot commented Sep 29, 2026

Copy link
Copy Markdown

✅ Deploy Preview for superset-docs-preview ready!

Name Link
🔨 Latest commit f2c9974
🔍 Latest deploy log https://app.netlify.com/projects/superset-docs-preview/deploys/6abbd78b98f9770008c1f96a
😎 Deploy Preview https://deploy-preview-43694--superset-docs-preview.netlify.app
📱 Preview on mobile
Toggle QR Code...

QR Code

Use your smartphone camera to open QR code link.
🤖 Make changes Run an agent on this branch

To edit notification comments on pull requests, go to your Netlify project configuration.

@bito-code-review

bito-code-review Bot commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

Code Review Agent Run #990cac

Actionable Suggestions - 0
Review Details
  • Files reviewed - 2 · Commit Range: bcafb71..f2c9974
    • superset/utils/pandas_postprocessing/pivot.py
    • tests/unit_tests/pandas_postprocessing/test_pivot.py
  • Files skipped - 0
  • Tools
    • MyPy (Static Code Analysis) - ✔︎ Successful
    • Astral Ruff (Static Code Analysis) - ✔︎ Successful
    • Whispers (Secret Scanner) - ✔︎ Successful
    • Detect-secrets (Secret Scanner) - ✔︎ Successful

Bito Usage Guide

Commands

Type the following command in the pull request comment and save the comment.

  • /review - Manually triggers an incremental AI Review.

  • /review full - Manually triggers a full AI Review.

  • /pause - Pauses automatic reviews on this pull request.

  • /resume - Resumes automatic reviews.

  • /resolve - Marks all Bito-posted review comments as resolved.

  • /abort - Cancels all in-progress reviews.

Refer to the documentation for additional commands.

Configuration

This repository uses Superset You can customize the agent settings here or contact your Bito workspace admin at evan@preset.io.

Documentation & Help

AI Code Review powered by Bito Logo

# Mirrors the column fill above; pivot_table() drops NaN index rows
# regardless of the dropna= setting (dropna only governs the column axis).
for col in index:
_fill_dimension_column(df, col, NULL_STRING)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

A null timestamp now makes the pivot index a mixed object Index, so Timeseries charts with resampling enabled fail with Resample operation requires DatetimeIndex instead of resampling the remaining dates. Could the null-preservation policy account for the temporal index required by the following resample operation?

@github-actions github-actions Bot added the requires:rebase Requires rebasing on top of current master label Oct 2, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

change:backend Requires changing the backend requires:rebase Requires rebasing on top of current master size/L

Projects

None yet

Development

Successfully merging this pull request may close these issues.

NULL values are visualized inconsistently across different plugins

4 participants