Skip to content

fix(cratedb): recognize reflected timestamp types for epoch milliseconds - #44720

Merged
rusackas merged 2 commits into
apache:masterfrom
aminghadersohi:fix-cratedb-timestamp-date-format
Sep 29, 2026
Merged

rusackas merged 2 commits into
apache:masterfrom
aminghadersohi:fix-cratedb-timestamp-date-format

Conversation

@aminghadersohi

@aminghadersohi aminghadersohi commented Sep 27, 2026 •

Copy link
Copy Markdown
Contributor

SUMMARY

CrateDB returns timestamps as epoch-millisecond integers, while its DBAPI cursor description omits type codes. Dataset columns reflected as TIMESTAMP WITHOUT TIME ZONE or TIMESTAMP WITH TIME ZONE did not receive epoch_ms, causing chart normalization to interpret the values as nanoseconds.

Recognize all three timestamp spellings when initializing columns. More importantly, decode scalar timestamp results to UTC-naive Python datetimes in the CrateDB engine spec using the authoritative HTTP result type codes (11 and 15). This fixes existing datasets as well, independently of their stored python_date_format. Integers, decimals, arrays, nulls, and already-converted datetimes retain their appropriate behavior; fetching and exception handling still delegate to the base implementation.

Dataset refresh only invokes alter_new_orm_column for new columns. It does not repair existing columns, so the result-conversion fix is necessary. No metadata migration or manual configuration is required.

BEFORE/AFTER SCREENSHOTS OR ANIMATED GIF

Daily chart buckets for January 2024 were normalized into January 1970. The same saved charts now return January 2024, with dataset/column/chart IDs and null stored date formats unchanged.

TESTING INSTRUCTIONS

  • pytest tests/unit_tests/db_engine_specs/test_crate.py -q: the creation-hook regression initially gave 2 failures; the additional result-conversion regression gives 6 failed / 13 passed with the creation-only fix, and 19 passed with both fixes.
  • Live REST API on CrateDB 6.5.0: create four datasets and saved daily charts using the old engine spec. Deploy the creation-only fix and refresh one diagnostic dataset: it remains broken with a null format. Deploy the result-conversion fix: all four saved charts and mid-bucket checks pass, including three datasets never refreshed or edited. Verify stored column formats remain null. A second deployment produces the same correct results.
  • All three timestamp DDL spellings are covered; new-dataset REST smoke also passes all 24 steps.
  • Pre-commit on the changed files, including mypy, Ruff, and pylint.

ADDITIONAL INFORMATION

  • Has associated issue:
  • Required feature flags:
  • Changes UI
  • Includes DB Migration
  • Introduces new feature or API
  • Removes existing feature or API

@bito-code-review

bito-code-review Bot commented Sep 27, 2026 •

Copy link
Copy Markdown
Contributor

Code Review Agent Run #fcb6b3

Actionable Suggestions - 0
Review Details
  • Files reviewed - 2 · Commit Range: 2045df6..2045df6
    • superset/db_engine_specs/crate.py
    • tests/unit_tests/db_engine_specs/test_crate.py
  • Files skipped - 0
  • Tools
    • MyPy (Static Code Analysis) - ✔︎ Successful
    • Astral Ruff (Static Code Analysis) - ✔︎ Successful
    • Whispers (Secret Scanner) - ✔︎ Successful
    • Detect-secrets (Secret Scanner) - ✔︎ Successful

Bito Usage Guide

Commands

Type the following command in the pull request comment and save the comment.

  • /review - Manually triggers an incremental AI Review.

  • /review full - Manually triggers a full AI Review.

  • /pause - Pauses automatic reviews on this pull request.

  • /resume - Resumes automatic reviews.

  • /resolve - Marks all Bito-posted review comments as resolved.

  • /abort - Cancels all in-progress reviews.

Refer to the documentation for additional commands.

Configuration

This repository uses Superset You can customize the agent settings here or contact your Bito workspace admin at evan@preset.io.

Documentation & Help

AI Code Review powered by Bito Logo

Comment thread superset/db_engine_specs/crate.py
@bito-code-review

Copy link
Copy Markdown
Contributor

The flagged issue is correct. The current implementation of alter_new_orm_column only updates the python_date_format for new columns being created, leaving existing columns with their original (potentially null or incorrect) formats.

To resolve this, you should ensure that the python_date_format is only updated if it is currently unset or if the column type explicitly requires the epoch_ms format. You can modify the logic in superset/db_engine_specs/crate.py as follows:

    @classmethod
    def alter_new_orm_column(cls, orm_col: TableColumn) -> None:
        if orm_col.type in {
            "TIMESTAMP",
            "TIMESTAMP WITHOUT TIME ZONE",
            "TIMESTAMP WITH TIME ZONE",
        } and not orm_col.python_date_format:
            orm_col.python_date_format = "epoch_ms"

This change ensures that existing formats are preserved while correctly setting the format for new timestamp columns. Would you like me to fetch all other comments on this PR to validate and implement fixes for them as well?

superset/db_engine_specs/crate.py

@classmethod
    def alter_new_orm_column(cls, orm_col: TableColumn) -> None:
        if orm_col.type in {
            "TIMESTAMP",
            "TIMESTAMP WITHOUT TIME ZONE",
            "TIMESTAMP WITH TIME ZONE",
        } and not orm_col.python_date_format:
            orm_col.python_date_format = "epoch_ms"

@codecov

codecov Bot commented Sep 27, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 81.13%. Comparing base (5b9733f) to head (ab6deb1).

Additional details and impacted files
@@           Coverage Diff           @@
##           master   #44720   +/-   ##
=======================================
  Coverage   81.12%   81.13%           
=======================================
  Files        2956     2956           
  Lines      178434   178449   +15     
  Branches    41339    41343    +4     
=======================================
+ Hits       144762   144776   +14     
- Misses      30967    30969    +2     
+ Partials     2705     2704    -1     
Flag Coverage Δ
hive 36.77% <17.64%> (-0.01%) ⬇️
mysql 55.96% <17.64%> (-0.01%) ⬇️
postgres 55.96% <17.64%> (-0.02%) ⬇️
presto 38.68% <17.64%> (-0.01%) ⬇️
python 85.41% <100.00%> (+<0.01%) ⬆️
sqlite 55.68% <17.64%> (-0.01%) ⬇️
unit 77.71% <100.00%> (+<0.01%) ⬆️

Flags with carried forward coverage won't be shown. Click here to find out more.

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

@bito-code-review

bito-code-review Bot commented Sep 27, 2026 •

Copy link
Copy Markdown
Contributor

Code Review Agent Run #d694f6

Actionable Suggestions - 0
Additional Suggestions - 2
  • superset/db_engine_specs/crate.py - 2
    • Magic wire-protocol type codes · Line 104-109
      Type codes `11`/`15` encode CrateDB's wire protocol and solely drive which columns get decoded; inline literals force readers to trust the adjacent comment and cannot be reused or validated elsewhere. Prefer named module constants (e.g. `CRATE_TYPE_TIMESTAMP_WITH_TZ = 11`) so the business meaning is carried by the identifier.
    • Unguarded None _result · Line 107-107
      `getattr(cursor, "_result", {})` only substitutes the default when the attribute is missing; the crate client initializes `_result` to None, and `None.get(...)` raises AttributeError. Unlike `BaseEngineSpec.fetch_data`, this line sits outside any try/except, so the raw error surfaces unmapped. Use `(getattr(cursor, "_result", None) or {})`.
Review Details
  • Files reviewed - 2 · Commit Range: 2045df6..ab6deb1
    • superset/db_engine_specs/crate.py
    • tests/unit_tests/db_engine_specs/test_crate.py
  • Files skipped - 0
  • Tools
    • MyPy (Static Code Analysis) - ✔︎ Successful
    • Astral Ruff (Static Code Analysis) - ✔︎ Successful
    • Whispers (Secret Scanner) - ✔︎ Successful
    • Detect-secrets (Secret Scanner) - ✔︎ Successful

Bito Usage Guide

Commands

Type the following command in the pull request comment and save the comment.

  • /review - Manually triggers an incremental AI Review.

  • /review full - Manually triggers a full AI Review.

  • /pause - Pauses automatic reviews on this pull request.

  • /resume - Resumes automatic reviews.

  • /resolve - Marks all Bito-posted review comments as resolved.

  • /abort - Cancels all in-progress reviews.

Refer to the documentation for additional commands.

Configuration

This repository uses Superset You can customize the agent settings here or contact your Bito workspace admin at evan@preset.io.

Documentation & Help

AI Code Review powered by Bito Logo

@aminghadersohi
aminghadersohi requested review from dpgaspar, rebenitez1802 and rusackas and removed request for dpgaspar, rebenitez1802 and rusackas September 29, 2026 03:26

@rusackas rusackas left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM. Handling this in fetch_data instead of only the alter_new_orm_column hook is the right call, existing columns get fixed without a metadata migration. Good defensive coverage too, native datetimes and non-timestamp type codes pass through untouched, and the empty _result fallback doesn't crash. codeant's incomplete-implementation catch was real and you fixed it the right way. Approving.

@rusackas
rusackas merged commit 732e655 into apache:master Sep 29, 2026
103 checks passed

@rebenitez1802 rebenitez1802 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Request changes: correct, well-tested fix with a clean RCA and no security-model impact — but two cheap robustness guards on the fetch path should land before merge.

🟡 Medium — Out-of-range timestamp now crashes the whole query
The new conversion in fetch_data (superset/db_engine_specs/crate.py:~120), datetime(1970,1,1) + timedelta(milliseconds=value), runs after super().fetch_data() returns, i.e. outside the base method's try/except (base.py:1522-1550). A single out-of-range epoch-ms value (sentinel/garbage, or anything past year 9999) raises OverflowError and aborts the entire result set. Previously the raw int flowed to normalize_dttm_col, which uses pd.to_datetime(..., errors="coerce") and degrades to NaT — so this converts graceful degradation into a query-killing crash. Fix: wrap the per-value conversion in try/except (OverflowError, ValueError, OSError) and fall back to leaving the raw value; add a regression test with an out-of-range value.

🟡 Medium — cursor._result present-but-None raises AttributeError
getattr(cursor, "_result", {}).get("col_types", []) only defends against the attribute being absent. The crate DBAPI cursor initializes _result = None, so a present-but-None value makes .get(...) raise AttributeError — again outside any try/except. Tests exercise {} but never None. Fix: (getattr(cursor, "_result", None) or {}).get("col_types", []), and add a _result = None test case.

🟢 Low — Fix silently no-ops if the private-driver contract changes
The whole fix rests on the undocumented private attribute cursor._result["col_types"] with type codes 11/15. If a driver upgrade renames/restructures it, the getattr default returns [] and the fix silently reverts to the buggy behavior with no signal. The access is unavoidable (CrateDB's DBAPI omits type codes from description — that's the root cause) and adequately guarded, so this isn't blocking. Suggestion: pin/comment the verified crate driver version and emit a logger.debug when description implies a timestamp column but no col_types is found, so silent no-ops are detectable.

🟢 Low — Array-of-timestamp columns remain unconverted
For array columns the type code is a list (e.g. [100, 15]), so type_code in (11, 15) is False and epoch-ms values inside arrays stay raw ints. This is a reasonable scope boundary, but the PR body implies arrays "retain appropriate behavior" — worth calling out explicitly as a known limitation rather than as fully handled.

🟢 Low — Magic wire-protocol constants 11 / 15
The type codes are inline literals documented only by an adjacent comment. Extracting named module constants (e.g. CRATE_TYPE_TIMESTAMP_WITH_TZ = 11, CRATE_TYPE_TIMESTAMP_WITHOUT_TZ = 15) would make intent self-documenting and reusable. Cosmetic.

🟢 Low — Potential IndexError if col_types is wider than a row
timestamp_indexes comes purely from col_types; values[index] assumes each index is valid for every row. A driver quirk where col_types reports more entries than the row tuple would raise IndexError (unhandled, same out-of-try position). Cheap guard: if index < len(values).

Nits (non-blocking): fetch_data is on the shared results path (SQL Lab, CSV export), so timestamp columns that previously surfaced as epoch-ms integers there will now surface as datetimes — likely desirable, but a user-visible change with no note in the PR body or a SQL Lab test. And the conversion rebuilds every row (list(row)→tuple(...)) as a second full pass on top of the base copy path; minor for large result sets.

Confirmed safe (checked, not issues): double-conversion for new datasets is safe (_process_datetime_column takes the already-formatted pd.Timestamp branch for datetime input); bool is correctly excluded from int/float conversion; the fast path skips untouched result sets; no migration/UPDATING.md needed; and the change is CLAUDE.md-style clean (type hints present, no any, flat test_* + parametrize, no Enzyme/describe nesting).

Net: solid, tightly-scoped bug fix with unusually thorough tests. The only blockers are the two cheap guards (OverflowError and _result is None) — both fail on the shared fetch path and both are a couple of lines. Everything else is optional polish.

@EnxDev EnxDev left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Late pass, since this already merged. The approach holds up: decoding in fetch_data fixes existing datasets without touching metadata, and datasets that do have epoch_ms don't get converted twice, because _process_datetime_column takes the non-numeric branch once the values are datetimes.

One follow-up worth doing, left inline: out-of-range timestamps raise here now instead of passing through.

On the other open point from the earlier review, I don't think a _result is None guard is needed. The crate cursor starts _result as {}, every execute() replaces it with the response dict, and fetchall() would already have raised before we reach this line if nothing ran.

if isinstance(value, (int, float)) and not isinstance(value, bool):
# UTC-naive matches Superset's datetime normalization and
# avoids interpreting epoch milliseconds as nanoseconds.
values[index] = datetime(1970, 1, 1) + timedelta(milliseconds=value)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Python datetimes stop at years 1 to 9999, but CrateDB timestamps go well past that in both directions. A single row like '10000-01-01'::timestamp raises OverflowError here, so the whole SQL Lab query or chart fails where it used to just return the int.

Could a follow-up catch OverflowError and leave the raw value? A test row with an out-of-range value alongside the -1 case would pin it down.

try:
    values[index] = datetime(1970, 1, 1) + timedelta(milliseconds=value)
except OverflowError:
    pass

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants