Skip to content

fix(chart-data): validate select()'s exclude option instead of raising 500s - #42410

Merged
rusackas merged 6 commits into
apache:masterfrom
SEPURI-SAI-KRISHNA:fix/post-processing-validation-errors
Aug 27, 2026
Merged

rusackas merged 6 commits into
apache:masterfrom
SEPURI-SAI-KRISHNA:fix/post-processing-validation-errors

Conversation

@SEPURI-SAI-KRISHNA

@SEPURI-SAI-KRISHNA SEPURI-SAI-KRISHNA commented Jul 25, 2026 •

Copy link
Copy Markdown
Contributor

SUMMARY

SqlaTable.exec_query converts only InvalidPostProcessingError into a
QueryObjectValidationError (→ 400). Anything else raised during
post-processing escapes as an unhandled 500. select's exclude option did
exactly that for malformed post_processing input on
POST /api/v1/chart/data.

select's exclude option was never validated.

@validate_column_args("columns", "drop", "rename")   # "drop" is not a parameter
def select(df, columns=None, exclude=None, rename=None):

validate_column_args only inspects argnames that appear in the options, so
naming a non-existent "drop" meant exclude was unchecked and
df.drop(exclude, axis=1) raised a raw KeyError:

{"operation": "select", "options": {"exclude": ["does_not_exist"]}}

→ KeyError: "['does_not_exist'] not found in axis", surfacing as a 500.

I ran an AST scan over every @validate_column_args usage in the tree; this
is the only decorator/signature mismatch.

Fixing the name alone left two sibling 500s in the same function:

  • exclude is applied after the columns projection, so
    columns=["y"], exclude=["label"] passes validation against the incoming
    DataFrame and then KeyErrors on the narrowed one. That path is now a
    validation error naming the offending column.
  • validate_column_args normalises a scalar through scalar_to_sequence to
    validate, then hands the original value on, so a bare exclude="label"
    was iterated character by character. It is normalised before use.

The exclude docstring claimed post-rename names should be referenced, which
is backwards — it is applied before rename — so it is corrected.

What changed since the first review

This PR originally carried a second, unrelated fix: a malformed i18n
placeholder in exec_post_processing, where

_("Unsupported post processing operation: %(operation)s", type=operation)

raised KeyError: 'operation' from flask_babel while constructing
InvalidPostProcessingError, so the intended 400 never surfaced.

That fix has since landed on master independently, as part of #43337's
rewrite of exec_post_processing for the EXTRA_PANDAS_POSTPROCESSING_OPS
extension point. On merging master I took its version of query_object.py
wholesale, so that file no longer appears in this diff.

What is retained is the regression guard. #43337 added
test_exec_post_processing_unknown_op_raises, which asserts the error type
but not the message, so the placeholder could regress unnoticed. Rather than
ship a near-duplicate test, this PR adds the one distinguishing assertion —
that the message names the offending operation — to that existing test.

The PR is therefore now scoped to select.py plus tests.

BEFORE/AFTER SCREENSHOTS OR ANIMATED GIF

Not applicable — API error handling, no UI change.

TESTING INSTRUCTIONS

pytest tests/unit_tests/pandas_postprocessing/test_select.py \
       tests/unit_tests/queries/query_object_test.py

Adds test_select_invalid_exclude and test_select_exclude_accepts_scalar,
plus test_exec_post_processing_missing_operation.

To confirm they are genuine regression tests, revert the source change and
re-run — select(df, exclude=["abc"]) goes back to
KeyError: "['abc'] not found in axis".

Behaviour after the change — every previously-valid call is unchanged:

call result
exclude=["abc"] InvalidPostProcessingError (was KeyError)
exclude=["label"] ["y"]
exclude="label" (scalar) ["y"] (was iterated per character)
rename={"y":"y1"}, exclude=["label"] ["y1"]
columns=["y"], exclude=["label"] InvalidPostProcessingError (was KeyError)
columns=["y","label"], exclude=["label"] ["y"]
columns=["y","label"] ["y", "label"]

Wider sweep, no regressions:

pytest tests/unit_tests/common/ tests/unit_tests/queries/ tests/unit_tests/pandas_postprocessing/

ADDITIONAL INFORMATION

  • Has associated issue:
  • Required feature flags:
  • Changes UI
  • Includes DB Migration (follow approval process in SIP-59)
    • Migration is atomic, supports rollback & is backwards-compatible
    • Confirm DB migration upgrade and downgrade tested
    • Runtime estimates and downtime expectations provided
  • Introduces new feature or API
  • Removes existing feature or API

@dosubot dosubot Bot added the api Related to the REST API label Jul 25, 2026
@bito-code-review

bito-code-review Bot commented Jul 25, 2026 •

Copy link
Copy Markdown
Contributor

Code Review Agent Run #05e3ff

Actionable Suggestions - 0
Review Details
  • Files reviewed - 4 · Commit Range: ddf9090..ddf9090
    • superset/common/query_object.py
    • superset/utils/pandas_postprocessing/select.py
    • tests/unit_tests/pandas_postprocessing/test_select.py
    • tests/unit_tests/queries/query_object_test.py
  • Files skipped - 0
  • Tools
    • MyPy (Static Code Analysis) - ✔︎ Successful
    • Astral Ruff (Static Code Analysis) - ✔︎ Successful
    • Whispers (Secret Scanner) - ✔︎ Successful
    • Detect-secrets (Secret Scanner) - ✔︎ Successful

Bito Usage Guide

Commands

Type the following command in the pull request comment and save the comment.

  • /review - Manually triggers a full AI review.

  • /pause - Pauses automatic reviews on this pull request.

  • /resume - Resumes automatic reviews.

  • /resolve - Marks all Bito-posted review comments as resolved.

  • /abort - Cancels all in-progress reviews.

Refer to the documentation for additional commands.

Configuration

This repository uses Superset You can customize the agent settings here or contact your Bito workspace admin at evan@preset.io.

Documentation & Help

AI Code Review powered by Bito Logo

@github-actions github-actions Bot removed the api Related to the REST API label Jul 25, 2026
@codecov

codecov Bot commented Jul 25, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 78.98%. Comparing base (b4f7114) to head (d19fb06).
⚠️ Report is 28 commits behind head on master.

Additional details and impacted files
@@            Coverage Diff             @@
##           master   #42410      +/-   ##
==========================================
- Coverage   79.04%   78.98%   -0.07%     
==========================================
  Files        2877     2877              
  Lines      165371   165393      +22     
  Branches    38215    38211       -4     
==========================================
- Hits       130723   130634      -89     
- Misses      32186    32279      +93     
- Partials     2462     2480      +18     
Flag Coverage Δ
hive 37.99% <57.14%> (-0.03%) ⬇️
mysql 57.77% <57.14%> (-0.04%) ⬇️
postgres 57.80% <57.14%> (-0.04%) ⬇️
presto 39.91% <57.14%> (-0.03%) ⬇️
python 83.65% <100.00%> (+0.06%) ⬆️
sqlite 57.49% <57.14%> (-0.04%) ⬇️
unit 73.74% <100.00%> (+0.06%) ⬆️

Flags with carried forward coverage won't be shown. Click here to find out more.

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Note

Copilot couldn't run its full agentic review because it didn't start before the timeout. Make sure your repository has a runner available, or add a copilot-code-review.yml file specifying one with the runs-on attribute. See the docs for more details.

This PR hardens chart data post-processing error handling so malformed post_processing inputs surface as validation errors (400) instead of unhandled exceptions (500).

Changes:

  • Fixes select post-processing validation for exclude, aligning decorator args and preventing pandas KeyErrors.
  • Fixes i18n placeholder mismatch so unsupported operations correctly raise InvalidPostProcessingError.
  • Adds regression tests covering unsupported/missing operation and invalid exclude scenarios.

Reviewed changes

Copilot reviewed 4 out of 4 changed files in this pull request and generated 2 comments.

File Description
tests/unit_tests/queries/query_object_test.py Adds regression tests for exec_post_processing invalid operation inputs.
tests/unit_tests/pandas_postprocessing/test_select.py Adds regression tests ensuring invalid exclude becomes a validation error.
superset/utils/pandas_postprocessing/select.py Fixes validation decorator arg mismatch and adds runtime validation for exclude after columns projection.
superset/common/query_object.py Fixes gettext placeholder kwarg so unsupported operations raise the intended exception.

Comment thread superset/utils/pandas_postprocessing/select.py Outdated
Comment thread tests/unit_tests/queries/query_object_test.py Outdated

@rusackas rusackas left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Verified both root causes (the "drop" vs exclude decorator mismatch, and type=operation vs %(operation)s) and the fix plus tests look right. Copilot's two open nits (select.py:60 wanting the specific missing column names in the message, query_object_test.py:392 preferring str(excinfo.value) over .message) are fair polish but not worth blocking on. Let me know if you intend to tackle those.

@SEPURI-SAI-KRISHNA

Copy link
Copy Markdown
Contributor Author

Thanks for the review @rusackas — yes, tackled both. Pushed a follow-up.

Copilot's select.py nit — name the missing columns

Done, and it uses the interpolation style already present in the package (cum.py
does _("Invalid cumulative operator: %(operator)s", operator=operator)):

if missing := [column for column in exclude if column not in df_select.columns]:
    raise InvalidPostProcessingError(
        _(
            "Referenced columns not available in DataFrame: %(columns)s",
            columns=", ".join(missing),
        )
    )

One wrinkle worth flagging, since it means the nit is only partly addressable
without widening scope. The generic wording I originally used wasn't arbitrary —
it's the same string validate_column_args raises at utils.py:136. So the
decorator still intercepts the "column isn't in the frame at all" case before my
guard runs, and that path keeps the unnamed message:

exclude=["abc"]                        -> Referenced columns not available in DataFrame.
columns=["y"], exclude=["label"]       -> Referenced columns not available in DataFrame: label

I left the decorator's message alone deliberately: it's shared by every operation
that uses validate_column_args, so renaming it is a broader change (and a new
msgid for translators). The upside is the two messages now carry different
information — unnamed means "not in the query result", named means "your
columns selection removed it". Happy to make the decorator name its columns too
if you'd rather they were uniform; it's a small change, just wider blast radius
than I wanted to take unilaterally on an approved PR.

Copilot's query_object_test.py nit — .message coupling

Done, now str(excinfo.value). Equivalent today —
SupersetException.__init__ ends with super().__init__(self.message) — but not
coupled to that staying true.

One thing I noticed while verifying

The two fixes overlap more than the PR description implies. Because my guard
checks the projected frame, and the projected frame is always a subset of the
input, the guard catches everything the decorator's exclude validation would.
Reverting the decorator argname on its own now leaves the tests green; reverting
the guard on its own fails with the original KeyError.

I'd still keep the decorator change — "drop" is not a parameter of select(),
so leaving it is simply wrong and would silently mislead the next person — but
the guard is the load-bearing half, and the description overstates the split.
Say the word if you'd like me to reword it.

Verified against the pinned toolchain: ruff 0.9.7, mypy 1.15.0 with the hook's
stub set, pylint 3.3.7 with superset.extensions.pylint (10.00/10). Tests:
pandas_postprocessing, queries and common all pass (the two test_prophet
failures are the optional prophet extra not being installed locally).

@netlify

netlify Bot commented Jul 29, 2026

Copy link
Copy Markdown

✅ Deploy Preview for superset-docs-preview ready!

Name Link
🔨 Latest commit 0233d1c
🔍 Latest deploy log https://app.netlify.com/projects/superset-docs-preview/deploys/6a6988c5fd152300086c4d64
😎 Deploy Preview https://deploy-preview-42410--superset-docs-preview.netlify.app
📱 Preview on mobile
Toggle QR Code...

QR Code

Use your smartphone camera to open QR code link.
🤖 Make changes Run an agent on this branch

To edit notification comments on pull requests, go to your Netlify project configuration.

Comment thread superset/utils/pandas_postprocessing/select.py
@SEPURI-SAI-KRISHNA

Copy link
Copy Markdown
Contributor Author

Good catch, this one is real — fixed and pushed.

@rusackas heads-up, since you'd already approved: this is a third commit fixing an
actual regression, not polish.

exclude="label" was being iterated character by character, so the guard reported
l, a, b, e, l as missing and rejected a valid call:

Referenced columns not available in DataFrame: l, a, b, e, l

The scalar form is supported — validate_column_args normalises every column
argument through scalar_to_sequence to validate it, then passes the original
value on, so the decorator accepts the string and hands the raw string through.
Fixed by normalising with the same helper the decorator uses, so validation and
execution agree on the shape:

exclude = list(scalar_to_sequence(exclude))

Two corrections to the report, in both directions:

It's older than the comment suggests. The regression came in with the first
commit on this branch, not the follow-up that added the column names. The original
any(column not in df_select.columns for column in exclude) had exactly the same
defect — the follow-up only made it visible by printing the characters. I verified
by reverting to each form in turn; the new test fails against both:

commit 1 form -> Referenced columns not available in DataFrame.
commit 2 form -> Referenced columns not available in DataFrame: l, a, b, e, l

Before this branch, df.drop("label", axis=1) handled the scalar natively, which
is what makes it a regression rather than a pre-existing limitation.

But "Major" overstates the reach. It isn't reachable through the chart-data
API: superset/charts/schemas.py:658 declares exclude = fields.List(fields.String()),
so the payload is always a list. Only direct callers of the utility are affected.
Worth fixing regardless — it's a behaviour regression against master and the
decorator's own contract.

test_select_exclude_accepts_scalar covers both the valid scalar and a scalar
naming a column that doesn't exist. Verified with ruff 0.9.7, mypy 1.15.0, pylint
3.3.7 (10.00/10); 290 passed across pandas_postprocessing, queries and
common (the two test_prophet failures are the optional prophet extra missing
locally).

@bito-code-review

bito-code-review Bot commented Jul 29, 2026 •

Copy link
Copy Markdown
Contributor

Code Review Agent Run #b9b797

Actionable Suggestions - 0
Review Details
  • Files reviewed - 3 · Commit Range: ddf9090..b82d8db
    • superset/utils/pandas_postprocessing/select.py
    • tests/unit_tests/pandas_postprocessing/test_select.py
    • tests/unit_tests/queries/query_object_test.py
  • Files skipped - 0
  • Tools
    • MyPy (Static Code Analysis) - ✔︎ Successful
    • Astral Ruff (Static Code Analysis) - ✔︎ Successful
    • Whispers (Secret Scanner) - ✔︎ Successful
    • Detect-secrets (Secret Scanner) - ✔︎ Successful

Bito Usage Guide

Commands

Type the following command in the pull request comment and save the comment.

  • /review - Manually triggers a full AI review.

  • /pause - Pauses automatic reviews on this pull request.

  • /resume - Resumes automatic reviews.

  • /resolve - Marks all Bito-posted review comments as resolved.

  • /abort - Cancels all in-progress reviews.

Refer to the documentation for additional commands.

Configuration

This repository uses Superset You can customize the agent settings here or contact your Bito workspace admin at evan@preset.io.

Documentation & Help

AI Code Review powered by Bito Logo

rusackas and others added 2 commits August 14, 2026 17:45
@SEPURI-SAI-KRISHNA

Copy link
Copy Markdown
Contributor Author

Rebased onto master to clear the conflict with #42927.

The collision was purely textual: both that PR and this one append tests to the end of tests/unit_tests/queries/query_object_test.py, off the same anchor. Both sets are kept; the imports merged cleanly.

The two changes are complementary rather than overlapping. #42927 drops options an operation no longer accepts, and deliberately leaves a missing or unknown operation untouched:

# A missing or unknown operation is left untouched, so that exec_post_processing reports it as InvalidPostProcessingError.

That is the error this PR fixes, it was interpolating type=operation into a %(operation)s placeholder, so the message raised at render time instead of naming the operation. So #42927 routes those cases to exec_post_processing, and this PR makes the message it produces correct. Both of this PR's tests still pass unchanged against the merged tree.

Verified after the merge: 369 passed across pandas_postprocessing, queries and common; 9 of those are the post-processing tests, 2 from this PR and 7 from #42927. ruff 0.9.7 clean, formatting unchanged, pylint 9.71/10 with the
four remaining findings all pre-existing on master (identical codes, same functions), the score is marginally above master's 9.69.

@rusackas thanks for merging master in on the 15th, flagging this since your approval predates the conflict.

@rusackas

Copy link
Copy Markdown
Member

@SEPURI-SAI-KRISHNA looks like there are fresh conflicts... would you mind tackling those, or would you like a hand?

@SEPURI-SAI-KRISHNA

Copy link
Copy Markdown
Contributor Author

Thanks @rusackas, on it, no hand needed.

Worth flagging what the conflict turned out to be, since it changes what's left of this PR.

The conflict is with #43337, which rewrote exec_post_processing to add the EXTRA_PANDAS_POSTPROCESSING_OPS extension point. That rewrite already carries this PR's query_object.py fix, master now reads operation=operation where it used to read type=operation, so the "Unsupported post processing operation" message interpolates properly on master today. I've resolved that file by taking master's version wholesale; nothing of mine needs to survive there.

What remains, and is still entirely unfixed on master, is superset/utils/pandas_postprocessing/select.py:

  • the @validate_column_args decorator still names "drop", a parameter select() does not have, so exclude was never validated at all;
  • exclude naming a column already removed by a preceding columns selection still reaches df.drop() and surfaces as a bare pandas KeyError, a 500, not a 400;
  • a scalar exclude="label" is iterated character by character, because validate_column_args normalises through scalar_to_sequence to validate but then passes the original value through.

So the PR is now scoped to select.py plus its tests. I'll retitle it accordingly if you'd prefer that over leaving the current title.

One test note: my test_exec_post_processing_unsupported_operation overlapped with #43337's new test_exec_post_processing_unknown_op_raises almost exactly, so rather than ship a near-duplicate I dropped mine and added its one distinguishing assertion, that the message actually names the offending operation, to theirs. That keeps the
regression guard on the %(operation)s placeholder without a redundant test.

@rusackas

Copy link
Copy Markdown
Member

Retitling the PR and updating/amending the description on the PR definitely helps for reviews and for historical record.

@SEPURI-SAI-KRISHNA SEPURI-SAI-KRISHNA changed the title fix(chart-data): raise validation errors instead of 500s for invalid post-processing fix(chart-data): validate select()'s exclude option instead of raising 500s Aug 26, 2026
@SEPURI-SAI-KRISHNA

Copy link
Copy Markdown
Contributor Author

retitled and rewrote the description, the summary now covers only the select() fix, with a section recording what #43337 absorbed and how the regression guard was preserved.

@rusackas

Copy link
Copy Markdown
Member

Rechecked this against the rebased head rather than trusting the old approval. Root cause still checks out against validate_column_args, and all three nits (Copilot's two, plus Codeant's scalar_to_sequence catch) are genuinely fixed with tests. Still LGTM, merging.

@rusackas
rusackas merged commit 926e0e6 into apache:master Aug 27, 2026
75 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants