test: characterization suite for timestamps with time zone - #25175
Conversation
|
Self-review (QA pass) of our own PR, at head The main hazard of this file: a pinned bug without a comment reads as an intended contract. So I checked each expectation that records a bug or a divergence from PostgreSQL. For each one I asked two questions: does a comment name the problem, and does it link the specific issue? Result. The file passes at this head, alone and merged with current PostgreSQL answers below come from Findings (most important first)1. CI: "Check License Header" fails at this head
Fix: put 2. Section 8 links the wrong issue, and several of its pins have no commentThe Section 8 header says: "This is the area corrupted by unwrap_cast", and links #25095. But no query in Section 8 has the shape of 25095. That issue is about a naive column compared with an aware literal under a non-UTC session time zone. Section 8 pins a different behaviour: a naive literal compared with an aware column is read in the time zone of the column, not in the session time zone. These expectations record that behaviour:
No issue in our list covers this exactly. The closest open issue is #13212 (" 3. Two
|
| Engine | Bin for 2024-03-11T12:30:00Z |
|---|---|
| DataFusion (this file) | 2024-03-11T01:00:00-06:00 |
PostgreSQL 15, date_bin |
2024-03-11 01:00:00-06 |
DuckDB 1.5.2, time_bucket |
2024-03-11 01:00:00-06 |
The pin is correct, and #25168 is the right link. But the comment must say that PostgreSQL and DuckDB give the same answer. Otherwise a fix can claim PostgreSQL parity that does not exist.
8. Low-priority notes
- The header "Related issues" list does not include Document that
date_binwith an explicit origin drifts an hour across a DST transition #25168. - Line 1181 says: "the naive side is treated as UTC". That is not true in general. I joined
cmp_denverwithnaive_col, with the session time zone unset. The naive2024-07-01 12:00:00matches2024-07-01T12:00:00-06:00, which is 18:00 UTC. So the naive side is read in the zone of the column. PostgreSQL 15 underUTCmatches the 12:00 UTC row. - The PR body says that the 40
SET datafusion.execution.time_zonestatements onmainare "nearly all" in two files. Onmainthey are in 9 files.set_variable.slthas 1 of them. - docs: document the
TIMESTAMP WITH TIME ZONEtype mapping and fix a stale comment #25171 adds docs that say DataFusion "discards theZ" in'...Z'::timestamptz. Lines 52-56 of this file say that the offset "is still honoured". Lines 90-95 confirm this:+05:30becomes06:30:00. The two texts must agree. The offset is applied, and only the time zone label is lost.
Merge-order collisions, checked against the file
I merged each PR into this branch at bafa661263 and ran this file, except where the diff shows no possible effect.
| PR (head) | Change | Result of the merged run | Expectations that change |
|---|---|---|---|
#25165 (3e42f8b102) |
AT TIME ZONE on an aware value returns a naive value |
2 failures | lines 323 and 331 (5b): Timestamp(ns, "Europe/Brussels") 2024-01-15T13:00:00+01:00 becomes Timestamp(ns) 2024-01-15T13:00:00, and the Denver query changes the same way |
#25163 (bdbfe1f8cd) |
date_part accepts timezone, timezone_hour and timezone_minute |
2 failures | lines 800 and 803: the two query error blocks now succeed |
#25161 (fb72121aef) |
from_unixtime uses the session time zone |
pass | none. The comment at lines 835-838 becomes stale |
#25173 (aeb4122371) |
generate_series precision |
not run | none: this file has no generate_series, and the PR changes only generate_series.rs and table_functions.slt |
#25171 (095e6d3a31) |
a code comment and docs | not run | none: no behaviour change. See the Z note in finding 8 |
current main |
includes #24920 and #25182 | pass | none. At a32121369b, #24920 broke at least 10 Section 16 error expectations. bafa661263 fixes them |
The updated #25165 adds a CASE limitation test in timestamps.slt. This file applies AT TIME ZONE to no CASE expression, so that test adds no collision.
What I checked and found correct
- The file passes at
bafa661263(cargo test --profile ci -p datafusion-sqllogictest --test sqllogictests -- timestamps_timezone), and also with currentmainmerged in. - These comments name the problem and link the right issue: Sections 0-4 (TIMESTAMP WITH TIME ZONE can resolve to a timezone-naive type, and casting to it discards an existing timezone #25166), the
'+05:30'literal in 5a (AT TIME ZONE '+05:30'uses the opposite sign convention from PostgreSQL #25170), 5b, 6b and the first round trip in 7 (Casting existing timestamp to timestamp again strips timezone information #12218), the Denver and Kolkatadate_bin/date_trunccases (date_binanddate_truncdisagree on timezone-aware timestamps #25167), the origin drift (Document thatdate_binwith an explicit origin drifts an hour across a DST transition #25168),from_unixtime(Make from_unixtime aware of execution timezone #12892), and all of Section 16 (Cast fromTimestamp(_, None)to a named timezone errors on DST boundaries #25084). - I measured the PostgreSQL answers that the comments quote. They are correct for 5a, 5b, the composed idiom, 6b, the round trips, the Section 8 row counts, the
168 dayscase,to_timestamp,current_time, theUNIONtype, and all four Section 16 DST values. - The new Section 16 expectations match the stable
Arrow error: ...tail, so they still separate the column path (Cannot cast timezone to different timezone) from the literal path (error computing timezone offset).
🤖 Generated with Claude Code
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## main #25175 +/- ##
==========================================
- Coverage 82.48% 82.48% -0.01%
==========================================
Files 1140 1140
Lines 437614 437614
Branches 437614 437614
==========================================
- Hits 360986 360975 -11
- Misses 54834 54843 +9
- Partials 21794 21796 +2 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
…pache#25161) ## Which issue does this PR close? - Closes apache#12892 ## Rationale for this change ### The bug, in SQL Set a session time zone. Every other date and time function honours it. `from_unixtime` does not. ```sql SET datafusion.execution.time_zone = 'America/Denver'; -- now() and to_timestamp() honour the session time zone SELECT arrow_typeof(now()), arrow_typeof(to_timestamp(1704110400)); -- Timestamp(ns, "America/Denver") | Timestamp(ns, "America/Denver") -- from_unixtime() does not SELECT arrow_typeof(from_unixtime(1704110400)), from_unixtime(1704110400); -- Timestamp(s) | 2024-01-01T12:00:00 ``` The last row is the defect. The user asks for a Denver session, but the value comes back as a bare wall clock with no zone. `2024-01-01T12:00:00` is the UTC clock, not the Denver clock. After this PR the same query returns the session time zone: ```sql SET datafusion.execution.time_zone = 'America/Denver'; SELECT arrow_typeof(from_unixtime(1704110400)), from_unixtime(1704110400); -- Timestamp(s, "America/Denver") | 2024-01-01T05:00:00-07:00 ``` The instant is the same. Only the declared type and the rendered offset change. ### What changes for a user - `from_unixtime(expr)` returns a value in `datafusion.execution.time_zone`. - `from_unixtime(expr, 'tz')` is unchanged. An explicit argument still wins. - With the session time zone unset, which is the default, nothing changes at all. - A value that no time zone can render is now an error, not a process panic. ### Field research: PostgreSQL and DuckDB Neither engine has a function named `from_unixtime`. The closest equivalent in both is `to_timestamp(<epoch>)`, so that is the comparison I ran. Both engines are timezone-aware there, and both render in the session time zone. **PostgreSQL 17.11** ```sql SHOW TimeZone; -- Etc/UTC SELECT to_timestamp(1704110400), pg_typeof(to_timestamp(1704110400)); -- 2024-01-01 12:00:00+00 | timestamp with time zone SET TimeZone = 'America/Denver'; SELECT to_timestamp(1704110400); -- 2024-01-01 05:00:00-07 SELECT extract(epoch from to_timestamp(1704110400)); -- 1704110400.000000 ``` PostgreSQL has no timezone-naive form of `to_timestamp`. A user must ask for one: ```sql SELECT to_timestamp(1704110400) AT TIME ZONE 'UTC'; -- 2024-01-01 12:00:00 | timestamp without time zone ``` **DuckDB v1.5.2** ```sql SELECT current_setting('TimeZone'), to_timestamp(1704110400), typeof(to_timestamp(1704110400)); -- America/Chicago | 2024-01-01 06:00:00-06 | TIMESTAMP WITH TIME ZONE SET TimeZone = 'America/Denver'; SELECT to_timestamp(1704110400); -- 2024-01-01 05:00:00-07 SELECT epoch(to_timestamp(1704110400)); -- 1704110400.0 ``` DuckDB defaults its session time zone to the operating system zone, so it is never naive. The naive form is `make_timestamp(<microseconds>)`, which returns `TIMESTAMP`. **What the comparison supports, and what it does not** - The type and the render rule of this PR agree with both engines. A one-argument epoch conversion is timezone-aware and follows the session time zone. - The default differs, and this PR keeps DataFusion's default. DataFusion leaves the session time zone unset, so `from_unixtime` stays naive out of the box. PostgreSQL and DuckDB always have a session zone. - The fixed-offset string form is not a meaningful comparison against PostgreSQL. `SET TimeZone = '+08:00'` in PostgreSQL means UTC-08:00, because PostgreSQL reads the string POSIX-style. DataFusion reads `'+08:00'` as UTC+08:00. That divergence is separate and is tracked in apache#25170. - The out-of-range comparison is not meaningful either. The three engines have three different timestamp ranges. PostgreSQL raises `ERROR: timestamp out of range` at `to_timestamp(-8334601211039)`. DuckDB renders the same value as `262145-12-31 (BC) 23:59:59.000448-04:56`. DataFusion sits between the two, and this PR makes it raise an error rather than panic. ## What changes are included in this PR? ### The fix - `FromUnixtimeFunc` carries an `Option<Arc<str>> timezone` taken from `config.execution.time_zone`, and implements `ScalarUDFImpl::with_updated_config`. `NowFunc` and the `to_timestamp*` functions already work this way. The single-argument form reports and produces `Timestamp(Second, <session tz>)` from both `return_field_from_args` and `invoke_with_args`. - `from_unixtime` moves to `make_udf_function_with_config!`, and its `expr_fn` gains the `@config` marker. The session `ConfigOptions` now reach the function at registration time, and again on every `SET` or `RESET`. - `FromUnixtimeFunc::new()` is deprecated since `56.0.0` in favour of `FromUnixtimeFunc::new_with_config()`. This mirrors the `to_timestamp*` constructors. `Default` still yields the previous timezone-naive behaviour. - The `from_unixtime` documentation description was wrong before this PR. It claimed an RFC3339 string with nanosecond precision and a `Z` suffix, but the function returns an Arrow `Timestamp(Second, ...)`, and the example directly below it shows a `-04:00` offset. The description now states what the function returns, and `docs/source/user-guide/sql/scalar_functions.md` is regenerated with `dev/update_function_docs.sh`. ### The guard Arrow renders a `Timestamp(Second, Some(tz))` through `chrono`'s `DateTime::naive_local`, which panics when the shifted value leaves `NaiveDateTime`'s range. The panic fires at render time, far from `from_unixtime`. The two-argument form can always reach it. Once the one-argument form applies a session zone, it can reach it too. So `invoke_with_args` now checks the bound and returns a `DataFusionError`: ``` Execution error: Cannot convert -8334601211039 to a timestamp in timezone "America/New_York" for function from_unixtime: the local date and time is outside the supported range ``` The check has three parts: 1. A timezone-naive result is never shifted, so the check returns at once. The full `NaiveDateTime` range stays usable, exactly as on `main`. 2. A fast path reads only the array's `min` and `max`. Inside `[NaiveDateTime::MIN + 86400, NaiveDateTime::MAX - 86400]` no zone offset can push a value out of range, so no per-value work happens. 3. Outside that window the check resolves the zone offset per value. `86400` is a deliberate over-bound. The largest offset in the whole IANA database is `Asia/Manila` at `-15:56:08` local mean time, and the largest positive one is `America/Metlakatla` at `+15:13:42`. The largest fixed offset Arrow's `Tz` parser accepts is `+23:59`, which is `86340` seconds. `+24:00` is rejected by the parser. ## What is the testing strategy for this PR? **New sqllogictest file** `datafusion/sqllogictest/test_files/from_unixtime_timezone.slt`, modelled on `to_timestamp_timezone.slt`. It asserts both `arrow_typeof` and the value for: - the session time zone unset, - a fixed offset (`+08:00`), - a named IANA zone (`America/Denver`), - the explicit two-argument form under a set session time zone, which must ignore it, - array (not constant folded) input, - `NULL` input, - equality of one instant across two time zones, - `RESET`, which must restore the timezone-naive behaviour. The `to_unixtime` / `from_unixtime` round trip from the issue is in the file. So are the out-of-range bounds, in both directions, for a named zone and a fixed offset zone, for the one-argument and the two-argument forms, and for array input. **New unit tests** in `datafusion/functions/src/datetime/from_unixtime.rs` cover `return_field_from_args` and `invoke_with_args` with a session time zone set, an explicit time zone that overrides it, and the out-of-range bound. Every bounds test renders its result. The unit tests go through `arrow::util::display::ArrayFormatter`, and the sqllogictest cases use `query P`. This is deliberate: the panic they guard against fires when the value is formatted, not when it is produced, so a test that only computes the value would pass either way. **Verified locally** - `cargo test -p datafusion-sqllogictest --test sqllogictests` — 506 files pass - `cargo test -p datafusion-functions --lib datetime::from_unixtime` — 10 tests pass - `cargo clippy -p datafusion-functions -p datafusion-sql --all-targets -- -D warnings` — clean - `cargo fmt --all --check` and `./ci/scripts/doc_prettier_check.sh` — clean **Verified against a `main` build.** I built `main` at the merge base and this branch as two `datafusion-cli` binaries and diffed their output. With the session time zone unset, eleven probes are byte-identical, including the error paths. The two-argument form is byte-identical across 140 pairs of value and zone. The full detail is in a separate comment on this PR. **Cost of the guard.** `SELECT count(from_unixtime(v)) FROM generate_series(1, 100000000) t(v)`, three runs each, `ci` profile: 3.95 s / 3.65 s / 3.07 s with the zone unset, and 3.09 s / 3.78 s / 3.33 s with `America/Denver`. The two extra kernel passes sit inside the noise. ## Are there any user-facing changes? Yes, four. **1. The behaviour change, which is the bug fix.** When `datafusion.execution.time_zone` is set, `from_unixtime(expr)` returns `Timestamp(Second, <session tz>)` instead of `Timestamp(Second, None)`. The epoch value is unchanged. Only the declared type and the rendered offset differ. With the default unset session time zone nothing changes. **2. An API deprecation.** `FromUnixtimeFunc::new()` is deprecated since `56.0.0` in favour of `FromUnixtimeFunc::new_with_config()`. `Default::default()` stays available and behaves like the old `new()`. **3. A public factory signature change.** `datafusion::functions::datetime::from_unixtime()` becomes `from_unixtime(&ConfigOptions)`, because it now uses `make_udf_function_with_config!`. This breaks callers of that factory. It is the established convention for config-aware functions in this module: `now` and all five `to_timestamp*` functions already have this signature on `main`, with no zero-argument shim. The `expr_fn` wrapper `datafusion::functions::expr_fn::from_unixtime(expr)` keeps its signature. **4. A new error in place of a panic.** See below. ### Out of range values Arrow renders a `Timestamp(Second, Some(tz))` value by a shift of the UTC instant by the zone's offset. It does that with `chrono`'s `DateTime::naive_local`, which **panics** when the shifted value leaves `NaiveDateTime`'s range. The panic fires when the value is formatted, not when it is cast, so it surfaces far from where it starts: ```sql SELECT from_unixtime(-8334601211039, 'America/New_York'); -- thread 'main' panicked at chrono-0.4.45/src/datetime/mod.rs:579:14: -- Local time out of range for `NaiveDateTime` ``` This PR replaces that panic with an error. The check applies to both forms. Its deliberate limits: - **Timezone-naive results are unaffected.** They are never shifted, so the full `NaiveDateTime` range stays usable, exactly as on `main`. - **A value that is out of range in UTC as well is left to the cast**, which already reports it (`Cast error: Failed to convert 8210266876800 to datetime for ...`). No error message changes. - **A query that consumed an out-of-range value without a render now returns an error.** On `main` such a value is only fatal at render time, so `SELECT to_unixtime(from_unixtime(-8334601211039, 'America/New_York'))` returns `-8334601211039` today and errors after this PR. Every query that produced a rendered result keeps its behaviour. ### Known interaction with issue 16594 This removes the panic reproducer in apache#16594, which I verified directly. `from_unixtime(-8334601211038 - 1, 'America/New_York')` errors instead of a panic, and `from_unixtime(8210266876799 + 1, 'America/New_York')` keeps its existing cast error. It does **not** close that issue. The issue asks that *all* `Int64` values convert, and that is not reachable while `chrono::NaiveDateTime` renders the value. The underlying defect is an unguarded `naive_local()` in the Arrow timestamp-with-timezone formatter. It stays reachable with no `from_unixtime` at all: ```sql SELECT arrow_cast(-8334601211039, 'Timestamp(Second, Some("America/New_York"))'); -- same panic SET datafusion.execution.time_zone = 'America/New_York'; SELECT to_timestamp_seconds(-8334601211039); -- same panic, on main and on this branch ``` A proper guard belongs in Arrow. The check here only stops `from_unixtime` from handing the formatter a value the formatter cannot render. ### Merge order with PR 25175 apache#25175 pins today's `from_unixtime` behaviour in a characterization test file. **No assertion in it breaks.** Its SECTION 10 cases run with the session time zone unset, and this PR keeps `Timestamp(Second, None)` in that case. I replayed both assertions on this branch and they still pass: ``` Timestamp(s) 2024-07-01T00:00:00 Timestamp(s, "America/Denver") 2024-06-30T18:00:00-06:00 ``` The comment above them says `from_unixtime` "ignores `datafusion.execution.time_zone`". That comment goes stale when this PR lands, so whichever merges second must reword it. 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
6df7863 to
1254960
Compare
…ate_part` (apache#25163) ## Which issue does this PR close? - Part of apache#10368. That issue lists several timezone gaps. This PR covers one bullet: "Extracting the time offset using `date_part` might be a nice to have". The other bullets stay open, so this PR does not close the issue. ## Rationale for this change ### What a user sees today A timezone-aware timestamp carries a UTC offset, but nothing in DataFusion can read it back: ```sql > SELECT date_part('timezone', TIMESTAMP '2024-07-01T12:00:00' AT TIME ZONE 'Europe/Brussels') AS utc_offset_seconds; Execution error: Date part 'timezone' not supported ``` `timezone_hour` and `timezone_minute` fail the same way. There is no workaround. The offset is not a property of the type either, so a user cannot look it up once and reuse it: for a named zone it moves with daylight saving time. ### What a user sees after this PR ```sql > SELECT date_part('timezone', TIMESTAMP '2024-07-01T12:00:00' AT TIME ZONE 'Europe/Brussels') AS utc_offset_seconds; +--------------------+ | utc_offset_seconds | +--------------------+ | 7200 | +--------------------+ ``` `7200` is +02:00, the summer offset of `Europe/Brussels`. ### The new capability, in plain terms Three new `date_part` fields report the UTC offset of a timezone-aware timestamp: | part | meaning | `Europe/Brussels`, July | | --- | --- | --- | | `timezone` | the whole UTC offset, in seconds | `7200` | | `timezone_hour` | whole hours of the offset | `2` | | `timezone_minute` | whole minutes of the offset, without the hours | `0` | Three things follow from the definition: - The value depends on the instant, not only on the type. `Europe/Brussels` gives `3600` in January and `7200` in July. - Offsets that are not whole hours work. `Asia/Kolkata` gives `19800 / 5 / 30`. - A negative offset carries its sign on the minute as well as the hour. `America/St_Johns` in January gives `-12600 / -3 / -30`. `EXTRACT(TIMEZONE_HOUR FROM ...)` is the alternative syntax for the same thing. ## What changes are included in this PR? Two commits. **1. `refactor: move parse_tz from date_trunc into datetime::common`.** No behaviour change. `parse_tz` turns the optional timezone string of `DataType::Timestamp` into an `arrow::array::timezone::Tz`. It was private to `date_trunc.rs`. It now sits next to the other shared datetime helpers, so `date_part` reuses it instead of a second parser. **2. `feat: support timezone, timezone_hour and timezone_minute in date_part`.** The three fields, with these design points: - **Per row, not per type.** Each value goes through `as_datetime_with_timezone::<T>`, and the code reads its offset back. One pass of `PrimitiveArray::unary_opt` computes the field, so there is no intermediate offset array. Input nulls survive. - **Timezone-naive input is an error, not an implicit UTC.** A `Timestamp(_, None)` carries no offset. The error text is `Date part 'timezone' is not supported for timezone-naive timestamps, got Timestamp(ns)`. Dates, times, intervals and durations get a similar error. PostgreSQL rejects these inputs too. DuckDB instead returns `0`, which this PR treats as the weaker choice: a plain `0` is indistinguishable from a real UTC value. - **The return type is `Int32`.** DataFusion already returns `Int32` for every integral `date_part` field, and the two exceptions are `epoch` (`Float64`) and `nanosecond` (`Int64`). PostgreSQL returns `double precision` for all fields, but DataFusion diverges from that today for `hour`, `minute` and the rest. DuckDB returns `BIGINT`. So an integer type is both consistent here and precedented elsewhere. - **`EXTRACT` needs no planner change.** `SQLExpr::Extract` forwards the `DateTimeField` to `date_part` as its `Display` string. sqlparser already carries the `Timezone`, `TimezoneHour` and `TimezoneMinute` variants, so `EXTRACT(TIMEZONE_HOUR FROM ...)` arrives as `"TIMEZONE_HOUR"` and matches without regard to case. No existing field changes behaviour. The new parts sit in the `Err(_)` arm that used to return `Date part '{part}' not supported`. ### Field research The sign of `timezone_minute` on a negative offset is easy to get backwards, so I measured it rather than assumed it. **PostgreSQL 17.11**, in Docker, with `SET TimeZone TO '<zone>'` and a `timestamptz` literal: | zone | offset kind | instant | `timezone` | `timezone_hour` | `timezone_minute` | | --- | --- | --- | --- | --- | --- | | `Asia/Kolkata` | positive, 30 minutes | 2024-07-01T12:00:00Z | 19800 | 5 | 30 | | `Asia/Kathmandu` | positive, 45 minutes | 2024-07-01T12:00:00Z | 20700 | 5 | 45 | | `America/Denver` | negative, whole hour | 2024-01-01T12:00:00Z | -25200 | -7 | 0 | | `America/St_Johns` | negative, 30 minutes | 2024-01-01T12:00:00Z | **-12600** | **-3** | **-30** | | `Pacific/Marquesas` | negative, 30 minutes | 2024-07-01T12:00:00Z | -34200 | -9 | -30 | | `Europe/Brussels` | DST zone, winter | 2024-01-01T12:00:00Z | 3600 | 1 | 0 | | `Europe/Brussels` | DST zone, summer | 2024-07-01T12:00:00Z | 7200 | 2 | 0 | | `Pacific/Chatham` | DST zone, 45 minutes | 2024-01-01T12:00:00Z | 49500 | 13 | 45 | | `Pacific/Chatham` | DST zone, 45 minutes | 2024-07-01T12:00:00Z | 45900 | 12 | 45 | | `UTC` | zero | 2024-07-01T12:00:00Z | 0 | 0 | 0 | So the minute follows the sign of the offset. `Pacific/Chatham` is in the southern hemisphere, so its January value is the DST one. Also measured on the same server: - `NULL` in gives `NULL` out. - The part name ignores case. - `tz` is not an accepted spelling. - `date`, `time` and `interval` input are all rejected. **DuckDB 1.5.2 supports all three fields.** This is worth stating plainly, because the opposite would also be a useful finding. DuckDB agrees with PostgreSQL on every zone and instant in the table above, sign included. Two differences from PostgreSQL: - DuckDB returns `BIGINT`, where PostgreSQL returns `double precision`. - DuckDB returns `0` for a timezone-naive `TIMESTAMP`, where PostgreSQL raises `unit "timezone" not supported for type timestamp without time zone`. DuckDB rejects `DATE` and `INTERVAL` input, as this PR does. This PR matches both engines on every value, and it follows PostgreSQL on the timezone-naive error. ## What is the testing strategy for this PR? **Unit tests** in `datafusion/functions/src/datetime/date_part.rs`, 9 new tests. `timezone_parts_match_postgres` is a table of 15 zone and instant combinations, asserted for all three fields. It covers every row of the PostgreSQL table above. The arrays are built by a relabel of UTC instants, so the cases line up exactly. The other 8 tests cover: - case insensitivity, including the spelling that `EXTRACT` produces; - an array across a DST transition, with a null row; - all four `TimeUnit` variants; - scalar input; - `NULL` input; - the timezone-naive rejection; - the non-timestamp rejection; - the `Int32` return field. **sqllogictest** cases appended to `datafusion/sqllogictest/test_files/datetime/date_part.slt`: - fixed offsets `+05:30` and `-03:30`; - the exact query from the issue; - named zones in standard time and in daylight saving time; - the 45-minute `Pacific/Chatham` cases; - negative offsets that are not whole hours; - `EXTRACT(...)` syntax; - `arrow_typeof` of all three results; - `NULL` input; - a column across the exact Brussels spring-forward boundary, 01:59:59 against 03:00:00 local; - four error cases. Verified locally, on each commit separately: - commit 1: `cargo check --workspace --all-targets` clean, `cargo clippy -p datafusion-functions --all-targets -- -D warnings` clean, 344 lib tests pass. - commit 2: `cargo clippy --all-targets -- -D warnings` clean across the workspace, 353 lib tests pass, the full sqllogictest suite passes. - `cargo fmt --all`, `./ci/scripts/doc_prettier_check.sh` and `typos` are all clean. ## Are there any user-facing changes? Yes, and they are additive. Three new `date_part` and `EXTRACT` fields. No public API changes, and no existing field changes behaviour, so this PR needs no `api change` label. `docs/source/user-guide/sql/scalar_functions.md` comes from the `user_doc!` block through `dev/update_function_docs.sh`. It gains the three parts, a note on the DST and sign behaviour, the timezone-only restriction, and a worked example. One known limit: `datafusion.execution.time_zone` defaults to `None`, so `now()` is timezone-naive and `date_part('timezone', now())` errors. `SET datafusion.execution.time_zone` makes it work. That is apache#25166 and it is out of scope here. ### Merge order apache#25175 is a characterization suite. Its SECTION 9 pins the current rejection: ``` query error DataFusion error: Execution error: Date part 'timezone_hour' not supported SELECT date_part('timezone_hour', ts) FROM day_denver ``` Once this PR lands, `timezone_hour` and `timezone_minute` return `-6` and `0` for every row of `day_denver`. So whichever of the two merges second must update the other, and its CI goes red until it does. That PR records the same collision in its own description. 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
kosiew
left a comment
There was a problem hiding this comment.
Thanks for putting this together. This is a thorough characterization suite for the current timestamp-with-time-zone behavior, especially around session zones, casts, comparisons, DST, and mixed-zone coercion. I did not find anything blocking.
I left one small cleanup suggestion below.
| # `date_trunc` is included as a control: it gets both rows right. | ||
| # See https://github.com/apache/datafusion/issues/25168 | ||
| statement ok | ||
| CREATE TABLE dst_origin AS |
There was a problem hiding this comment.
Small cleanup suggestion: dst_origin is created here but isn't dropped in the cleanup block. The per-file context makes this harmless today, but could we add DROP TABLE dst_origin with the other cleanup statements? That would keep the suite self-contained if the context lifetime ever changes.
There was a problem hiding this comment.
Good catch, thanks. dst_origin was the only table in the file created without being dropped. Added DROP TABLE dst_origin to the cleanup section, in creation order after day_denver, in 5f74a75.
Adds `test_files/datetime/timestamps_timezone.slt`, ~1500 lines in 18 labelled sections that record what DataFusion actually does today with timezone-aware timestamps. DataFusion has a long tail of timezone correctness bugs, and the reason they keep recurring is structural: almost nothing in the test suite pins the behaviour down, so a change to timezone semantics can land without producing a single test diff. `SET datafusion.execution.time_zone` appears roughly 40 times in the whole sqllogictest corpus, nearly all of it in two files, and DST-boundary dates are essentially absent. This file does not change any behaviour. It writes the current behaviour down, including the parts that are known wrong or known to disagree with PostgreSQL -- those carry a comment saying so and a link to the issue, rather than being "fixed" here. When a fix lands, the diff is the specification of what changed. Every case is exercised as a literal and, where it matters, as a real CTAS column, since several of these behaviours only reproduce on columns. Sections cover literal typing under five session zones, `AT TIME ZONE`, casts in all four directions, round trips, tz-aware/tz-naive comparison (with `EXPLAIN` of the rewritten filter), `date_bin`/`date_trunc`/`date_part`, `to_char` / `to_local_time` / `to_timestamp*` / `from_unixtime`, `now`/`current_date`, interval arithmetic across both DST transitions, aggregates, joins and coercion, the US and EU DST boundaries by all four routes into a named zone, and a no-DST zone plus a half-hour zone. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Section 9 already had a `date_bin(..., origin)` case, but all of its data sits inside a single July day, so it never crosses a transition and the drift is invisible. Adds a case whose two rows are the same local time of day on either side of the America/Denver spring-forward transition, with the origin chosen so bins land on local midnight. The first row does; the second lands on 01:00 local. That is the bug in apache#25168 -- `date_bin` steps a fixed number of nanoseconds from a fixed instant, so it cannot track a local day that is 23 or 25 hours long. `date_trunc` is asserted alongside as a control: it gets both rows right, so the drift is a property of the origin arithmetic rather than of the data. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Upstream apache#24920 (merged 2026-09-06) adds field context to cast errors. The DST-boundary section pinned the full error chain, so 16 expectations started to fail once CI merged this PR onto current main -- 10 for America/Denver and 6 for Europe/Brussels: expected: DataFusion error: Arrow error: Cast error: Cannot cast timezone ... got: DataFusion error: Failed to cast field 'column1' from Timestamp(ns) to Timestamp(ns, "America/Denver") caused by Arrow error: Cast error: Cannot cast timezone to different timezone The wrapper chain is not what these cases pin. They pin which error class each path raises: the column path raises a Cast error, and the literal path raises a Parser error. So each expectation now matches only the `Arrow error: ...` tail. This uses the single-line `query error <regex>` form. In sqllogictest 0.29.1 an inline pattern goes through `Regex::is_match`, which is an unanchored search, while the multiline `----` form requires an exact match. The tail appears verbatim in both the old and the new message, so these expectations hold on either side of apache#24920 and will not break on the next wrapper change. Expectations and SQL are otherwise unchanged. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…ization suite Upstream apache#25182 now checks license headers in `.slt` files, and main uses `#` on header lines 8 and 10. This file had empty lines there, so the license check fails. A self-review of this PR found annotations that misdescribe what is pinned: - Section 8 and two later comments linked 25095 (the unwrap_cast bug), but no query here has that shape. What they pin is 13212, the zone a tz-naive side is read in, and 25166, a tz-naive `'...Z'::timestamptz` literal. Links re-pointed; the header now says 25095 is not exercised. - The UNION comment said "the instants are preserved", but its second row is the 12:00Z row where 06:00Z was requested -- the Section 8 problem again. - The explicit-origin `date_bin` case was labelled KNOWN BUG. PostgreSQL 15.19 (`date_bin`) and DuckDB 1.5.2 (`time_bucket`) return the same `01:00-06`, so it is inherent to instant-based binning, not a divergence. Reframed. - Added missing links or divergence notes for `::timestamp` giving the UTC wall clock (12218), the `'+05:30'` sign convention (25170), `date_bin` versus `date_trunc` (25167), and the `timezone_hour` gap that PR 25163 fills. Comments and header only. No query or expected output changes. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Addresses the rest of the self-review of this PR.
New expectations for issues the header listed but no query exercised:
- 25095: under an America/Denver session zone, a tz-naive column compared to
`'...Z'::timestamptz` matches no rows because unwrap_cast moves the cast
onto the literal. Pinned with the count and the rewritten logical plan.
- 12892: `from_unixtime` under a session zone. PR 25161 is already in this
branch's base, so this pins the fixed behaviour and the stale "ignores
the session zone" comment is corrected.
- 25166 (second part): casting a Europe/Brussels value to `::timestamptz`
under a UTC session replaces its zone, which changes `date_part('hour')`.
Annotation fixes:
- The tz-naive join comment said the naive side is treated as UTC; it is
read in the aware side's zone (13212).
- Linked 25166 on the naive-result notes for `to_timestamp*`, `now()`,
`from_unixtime`, naive interval arithmetic, and the CASE type oddity.
- The `date_part` and `168 days` notes now say this follows from the zone
living on the value, and that no issue tracks it as a bug.
- The `'...Z'::timestamptz` comparison comment now says the offset is
applied and the zone-less result is re-read in the column's zone, rather
than that the `Z` is discarded.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LZLtz9CYVo7MQedhVpdSh7
…erence apache#25163 is now on main, so `date_part('timezone_hour' | 'timezone_minute')` succeed and the two `query error` pins fail. Replace them with value queries for `timezone`, `timezone_hour` and `timezone_minute` on the Denver column and on a half-hour Asia/Kolkata value. Also replace the "no issue tracks it as a bug" wording on the date_part and `168 days` notes with one sentence on the cause: PostgreSQL has only the session zone, while Arrow stores the zone in each column's type. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LZLtz9CYVo7MQedhVpdSh7
1254960 to
080fea0
Compare
`dst_origin` was the only table the file created without dropping it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LZLtz9CYVo7MQedhVpdSh7
|
Sending to merge since this is just tests. |
) ## Which issue does this PR close? None. This PR adds tests only. It records current behaviour, bugs included, for these issues: - apache#12218: `::timestamp` on a tz-aware value gives the UTC wall clock, and `tstz AT TIME ZONE` stays tz-aware - apache#12892: `from_unixtime` and the session time zone (fixed by apache#25161, pinned as fixed) - apache#13212: a tz-naive side of a comparison is read in the other operand's zone, not the session zone - apache#25084: casting a tz-naive value into a named zone errors at DST boundaries - apache#25095: unwrap_cast rewrites a tz-naive column compared with a tz-aware literal under a non-UTC session zone - apache#25166: `TIMESTAMP WITH TIME ZONE` / `::timestamptz` can resolve to a tz-naive type, and a cast to `::timestamptz` replaces an existing zone - apache#25167: `date_bin` and `date_trunc` disagree on tz-aware input - apache#25168: `date_bin` with an explicit origin drifts across a DST transition (PostgreSQL and DuckDB behave the same way) - apache#25170: `AT TIME ZONE '+05:30'` uses the opposite sign convention from PostgreSQL - apache#10368: `date_part('timezone_hour' | 'timezone_minute')` (implemented by apache#25163) ## Rationale for this change Timezone bugs keep coming back in DataFusion because the result depends on several things at once: the session time zone (set or unset), the value's own zone, DST, and optimizer rewrites such as unwrap_cast and constant folding. Few tests cover these combinations. Outside this file, the sqllogictest suite has 47 `SET datafusion.execution.time_zone` statements across 10 files. The PostgreSQL-compatibility tests had no timezone coverage until apache#25164. This file records what DataFusion does today for timestamps with a time zone, so any change to those semantics shows up as a diff in review instead of shipping silently. Where the recorded behaviour is wrong, or disagrees with PostgreSQL, a comment says so and links the issue. When a fix lands, the expected output changes and the comment goes away. apache#25164 is the companion PR. It checks against a real PostgreSQL the cases where the two engines agree. This file covers the cases where they don't. ## What changes are included in this PR? One new file (1674 lines), `datafusion/sqllogictest/test_files/datetime/timestamps_timezone.slt`, in 18 sections: | Section | What it pins | |---|---| | 0–4 | Type and value of `::timestamptz`, `TIMESTAMP [WITH TIME ZONE]` and table columns with the session zone unset, `+00:00`, `+05:30`, `America/Denver`, and `Europe/Brussels` | | 5 | `AT TIME ZONE` on tz-naive and tz-aware values, the composed `(tstz AT TIME ZONE z)::timestamp` idiom, and the `'+05:30'` sign convention | | 6 | Casts in every direction (naive/named/fixed offset), including `::timestamptz` on an already-zoned value | | 7 | Naive → aware → naive round trips, and the reverse | | 8 | Comparisons between tz-aware and tz-naive values, including the unwrap_cast case, with EXPLAIN of the rewritten filters | | 9 | `date_bin`, `date_trunc`, `date_part` (including `timezone*` fields) and `extract` on tz-aware input, including the explicit-origin DST case | | 10 | `to_char`, `from_unixtime`, `to_unixtime`, `to_timestamp*`, `to_local_time` | | 11 | `now`, `current_date`, `current_time`, `make_date` under a session zone | | 12 | `+`/`-` interval across DST, `'1 day'` vs `'24 hours'` | | 13–15 | Aggregates, GROUP BY, ORDER BY, DISTINCT, joins with mixed zones, and UNION / CASE / COALESCE / `greatest` type coercion | | 16 | Ambiguous and non-existent local times (US and EU transitions) | | 17 | A zone with no DST (America/Phoenix) and a half-hour zone (Asia/Kolkata) | No production code changes. ### Interaction with other open PRs Some fix PRs change expectations in this file. Whichever lands second updates the pins. I'd prefer the fixes land first so they stay small, and this file absorbs the change: - apache#25161 (`from_unixtime` follows the session zone) is already on main. The file pins the fixed behaviour. - apache#25163 (`date_part` timezone fields) is already on main. SECTION 9 pins `timezone`, `timezone_hour` and `timezone_minute` as values. - apache#25165 (`AT TIME ZONE` on a tz-aware value returns naive) changes the two SECTION 5b queries. - apache#25171 (docs) doesn't touch this file. Its wording that `'...Z'::timestamptz` "discards the `Z`" should match this file: the offset is applied, but the result carries no zone. - apache#25173 (`generate_series` precisions) doesn't touch this file. ## What is the testing strategy for this PR? The PR is the test. I generated expected output with sqllogictest `--complete`. Then I read every case and rewrote the ones that weren't testing what they meant to test, and added comments with issue links where the behaviour is wrong or differs from PostgreSQL. The file runs in DataFusion only. The PostgreSQL 15 answers quoted in comments come from separate runs, with the PostgreSQL `TimeZone` given in each comment. DuckDB 1.5.2 answers are quoted where they matter (the `date_bin` origin case, and the `'+05:30'` offset string). DST error expectations match only the stable `Arrow error: ...` tail, so changes to the wrapper error chain (for example apache#24920) don't break them. Run with: ``` cargo test -p datafusion-sqllogictest --test sqllogictests -- datetime/timestamps_timezone ``` ## Are there any user-facing changes? No. Tests only. --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…stale comment (apache#25171) ## Which issue does this PR close? This PR closes no issue. It changes documentation and one comment, and it changes no behaviour. - Related to apache#25166 ## Rationale for this change ### What a user hits There is no bug to reproduce. There is a mapping a user cannot look up. To find out what `TIMESTAMP WITH TIME ZONE` produces, a user must run it: ```sql -- default configuration: datafusion.execution.time_zone is unset SELECT arrow_typeof('2000-01-01T00:00:00'::TIMESTAMPTZ); +---------------+ | Timestamp(ns) | +---------------+ ``` The type has no timezone. Set the session timezone and the same expression gains one: ```sql SET datafusion.execution.time_zone = 'America/New_York'; SELECT arrow_typeof('2000-01-01T00:00:00'::TIMESTAMPTZ); +-----------------------------------+ | Timestamp(ns, "America/New_York") | +-----------------------------------+ ``` `docs/source/user-guide/sql/data_types.md` holds the SQL-to-Arrow type table. Three things were absent from it: - any row for `TIMESTAMPTZ` or `TIMESTAMP WITH TIME ZONE` - any note that the result depends on `datafusion.execution.time_zone` - any note that `TIMESTAMP(p)` is accepted, and which values of `p` are valid A user who reads the page learns none of the above. The comment in `datafusion/sql/src/planner.rs` was worse than absent. It stated the opposite of the code: ```rust // OUTPUT: [ArrowDataType] Timestamp<TimeUnit, Some(Time Zone)> self.context_provider.options().execution.time_zone.clone() ``` The expression is an `Option<String>` that is `None` by default, so the arm produces `Timestamp(unit, None)` on a default install. ### What changes for a user The type table gains the two absent spellings and the precision rules, so a user can look the mapping up instead of a run of `arrow_typeof`. The page also states that the timezone component comes from a session setting that is unset by default, and it links the open issue on that default. No behaviour changes. ### The technical detail The comment went stale in apache#18359, which changed `datafusion.execution.time_zone` from `String` to `Option<String>` with a `None` default. That PR dropped the `Some(...)` wrapper from the expression as a mechanical consequence of the type change, and left the two comment lines above it untouched. Whether the timezone-naive default is right is under discussion in apache#25166. This PR only makes the documentation match the current code. It deliberately changes no behaviour. ### Field research The table is exactly where a user who moves from another engine looks, so I measured both reference engines. PostgreSQL 17.11 (`postgres:17`): ``` => CREATE TABLE t(b TIMESTAMPTZ, c TIMESTAMP WITH TIME ZONE); => SELECT column_name, data_type FROM information_schema.columns WHERE table_name='t'; b | timestamp with time zone c | timestamp with time zone => SELECT pg_typeof('2024-01-01'::timestamptz); timestamp with time zone ``` DuckDB 1.5.2: ``` D SELECT typeof('2024-01-01 00:00:00'::TIMESTAMPTZ); TIMESTAMP WITH TIME ZONE D SELECT typeof('2024-01-01'::TIMESTAMP WITH TIME ZONE); TIMESTAMP WITH TIME ZONE ``` Both engines always resolve the type to a zone-aware type. Neither has an unset session timezone, so neither can produce a naive result from this SQL type. DataFusion produces a naive result on a default install. That divergence is the subject of apache#25166. The precision rules diverge too: - PostgreSQL accepts `p` from 0 to 6. `TIMESTAMP(1)` and `TIMESTAMP(4)` are valid. `TIMESTAMP(9)` raises a warning and reduces `p` to 6. - DuckDB accepts `p` from 0 to 9 and rounds up to the next storage unit. `TIMESTAMP(1)` gives `timestamp_ms`, `TIMESTAMP(4)` gives `timestamp` and `TIMESTAMP(7)` gives `timestamp_ns`. - DataFusion accepts 0, 3, 6 and 9 alone, and rejects every other value. ## What changes are included in this PR? - `datafusion/sql/src/planner.rs`: correct the `SQLDataType::Timestamp` comment to say `Timestamp<TimeUnit, Time Zone>`, note that the configured time zone is an `Option` that is unset by default, and point at apache#25166. There is no code change. - `docs/source/user-guide/sql/data_types.md`: add the `TIMESTAMP WITH TIME ZONE` and `TIMESTAMPTZ` row to the Date/Time Types table, document the optional `(p)` precision and the units it selects, and state that the timezone component comes from `datafusion.execution.time_zone`, which is unset by default, with a link to the open issue. ## What is the testing strategy for this PR? There are no tests. This is a comment fix plus a documentation fix, with no behaviour change. Every mapping in the docs was verified against a `cargo build --bin datafusion-cli` build of this branch rather than read off the source: ``` > select arrow_typeof('2000-01-01T00:00:00'::TIMESTAMP); -- Timestamp(ns) > select arrow_typeof('2000-01-01T00:00:00'::TIMESTAMP(0)); -- Timestamp(s) > select arrow_typeof('2000-01-01T00:00:00'::TIMESTAMP(3)); -- Timestamp(ms) > select arrow_typeof('2000-01-01T00:00:00'::TIMESTAMP(6)); -- Timestamp(µs) > select arrow_typeof('2000-01-01T00:00:00'::TIMESTAMP(9)); -- Timestamp(ns) > select arrow_typeof('2000-01-01T00:00:00'::TIMESTAMP(1)); -- Error: Unsupported SQL type TIMESTAMP(1) > select arrow_typeof('2000-01-01T00:00:00'::TIMESTAMP(2)); -- Error: Unsupported SQL type TIMESTAMP(2) -- default configuration (datafusion.execution.time_zone unset) > select arrow_typeof('2000-01-01T00:00:00'::TIMESTAMPTZ); -- Timestamp(ns) > select arrow_typeof(CAST('2000-01-01T00:00:00' AS TIMESTAMP WITH TIME ZONE)); -- Timestamp(ns) > select arrow_typeof(CAST('2000-01-01T00:00:00' AS TIMESTAMP(6) WITH TIME ZONE)); -- Timestamp(µs) > select arrow_typeof('2000-01-01T00:00:00'::TIMESTAMPTZ(3)); -- Timestamp(ms) > SET datafusion.execution.time_zone = 'America/New_York'; > select arrow_typeof('2000-01-01T00:00:00'::TIMESTAMPTZ); -- Timestamp(ns, "America/New_York") > select arrow_typeof(CAST('2000-01-01T00:00:00' AS TIMESTAMP(3) WITH TIME ZONE)); -- Timestamp(ms, "America/New_York") > select arrow_typeof('2000-01-01T00:00:00'::TIMESTAMP); -- Timestamp(ns) > select arrow_typeof('2000-01-01T00:00:00'::TIMESTAMP(0)); -- Timestamp(s) ``` The DDL path gives the same answers: ``` > CREATE TABLE t(a TIMESTAMP, b TIMESTAMPTZ, c TIMESTAMP(6), d TIMESTAMP(0) WITH TIME ZONE); > describe t; a | Timestamp(ns) b | Timestamp(ns) c | Timestamp(µs) d | Timestamp(s) ``` `./ci/scripts/doc_prettier_check.sh`, `cargo fmt --all` and `cargo clippy --all-targets -- -D warnings` all pass. The Sphinx docs build cleanly. Its one warning is the pre-existing absent generated `_static/data/deps.svg`, which is unrelated to this change. ## Are there any user-facing changes? Documentation only. There are no behaviour changes and no API changes. ## Merge order There is no file-level conflict with any PR in flight. apache#25175 adds only `datafusion/sqllogictest/test_files/datetime/timestamps_timezone.slt`, and this PR touches `planner.rs` and `data_types.md`. The two do share a dependency. Sections 0 to 4 of apache#25175 pin the same timezone-naive result that the new paragraph describes. A fix for apache#25166 must update this page and that file in one change. 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…pache#25164) ## Which issue does this PR close? This PR closes no issue. It adds tests only. Related to these issues. Each one is a divergence from PostgreSQL, so this PR keeps its queries out of the new file: - apache#12218 - apache#12892 - apache#13212 - apache#25084 - apache#25095 - apache#25166 - apache#25170 Companion PR: apache#25175 pins the current DataFusion behaviour, bugs included. ## Rationale for this change Set the session time zone to `America/Denver` in both engines. Then this query gives the same answer in DataFusion and in PostgreSQL: ```sql -- ts is 2024-03-09 12:00 MST, one day before the spring-forward transition SELECT date_part('hour', ts + INTERVAL '1 day')::bigint AS plus_1_day, date_part('hour', ts + INTERVAL '24 hours')::bigint AS plus_24_hours FROM tstz_den WHERE id = 1; ``` | Engine | `plus_1_day` | `plus_24_hours` | | --- | --- | --- | | DataFusion | `12` | `13` | | PostgreSQL 15 | `12` | `13` | `INTERVAL '1 day'` keeps the local wall clock. `INTERVAL '24 hours'` adds 24 hours of elapsed time, so the local hour moves by one. **What this PR guarantees.** CI runs each query of the new file on DataFusion and on PostgreSQL 15. If the answers differ, the build fails. So a change cannot move DataFusion away from PostgreSQL on these queries without a red build. ### Why time zone bugs recur - Time zone rules are hard to check by reading code. Two engines can each look correct and still give different answers. - On `main` before this PR, the sqllogictest corpus has 40 `SET datafusion.execution.time_zone` statements, in 9 files. - The PostgreSQL differential harness (`test_files/pg_compat/`) has no time zone coverage. No file in it refers to a time zone or to `timestamptz`. - So no test compares the time zone behaviour of DataFusion with a reference engine. A change can alter an answer, and no test fails. ## What changes are included in this PR? This PR adds one test file and changes no production code: `datafusion/sqllogictest/test_files/pg_compat/pg_compat_timestamptz.slt` (about 700 lines). ### Sections Each block sets the session time zone in both engines: `SET TimeZone` for PostgreSQL, and `SET datafusion.execution.time_zone` for DataFusion. Each block builds its table in the same zone. | Block | Session time zone | Coverage | | --- | --- | --- | | A | `UTC` (DataFusion `+00:00`) | `timestamptz` literals with and without an offset; `AT TIME ZONE` on naive values; comparison by instant; `DISTINCT`, `count(DISTINCT)`, `min`, `max`, `GROUP BY`, `ORDER BY`; filters with aware literals; a self join; `UNION`, `COALESCE`, `CASE`, `greatest`, `least`; `date_part` and `extract`; `to_timestamp`; `date_bin` with an explicit origin; interval arithmetic | | B | `America/Denver` | local `date_part` readings; `'1 day'` against `'24 hours'` across both DST transitions; the repeated hour at fall back; `date_trunc` at day, month and hour; `date_bin`; week, month and year arithmetic; `min`, `max` and `ORDER BY` by instant | | C | `Asia/Kolkata`, then `America/Phoenix` | a half-hour offset, and a zone with no DST | ### How we made the expected output 1. We ran the sqllogictest `--complete` mode against a real PostgreSQL (`PG_COMPAT=true PG_URI=... --complete`). So the answer of PostgreSQL is the expected output, and DataFusion must match it. 2. We read each result. 3. We removed each query where the two engines disagree. The list is in "Field research" below. ### Harness constraints 1. **No aware value in a result.** The PostgreSQL runner cannot render `timestamptz`: `postgres_engine/mod.rs` calls `unimplemented!` for that type. So each result is a naive `timestamp`, a `bigint`, a `boolean` or `text`. 2. **Session zone against value zone.** PostgreSQL uses the session time zone to read a naive literal, to render a `timestamptz`, to apply a day interval and to run `date_trunc`. DataFusion uses the time zone of the value. The two agree only when both zones are the same. So each block sets them to the same zone. 3. **A naive projection must not hide a divergence.** Block A projects with `::timestamp`. That projection gives the UTC wall clock in both engines only because the session time zone is UTC. Blocks B and C read each value through `date_part`. The self-review found one query that breaks this rule (lines 112-118). A follow-up commit must fix it. ## What is the testing strategy for this PR? This PR adds tests only. Checks at head `2b0a37fb91`: - `PG_COMPAT=true PG_URI=... cargo test --profile ci -p datafusion-sqllogictest --features postgres --test sqllogictests -- pg_compat` against `postgres:15`: 7 of 7 files pass. - With `log_statement = 'all'` on the server, the PostgreSQL log shows the harness execute each statement of this file, with zero errors. So the run really compares the two engines. - The same command without `PG_COMPAT` (DataFusion only): the file passes. - CI on 2026-09-10: all 38 checks pass, the "Run sqllogictest with Postgres runner" job included. - On 2026-09-11, apache#25182 added `*.slt` to the license header check. The new file now has `#` on header lines 8 and 10, as `main` expects (fixed in `aafa2e7df4`). ### Field research #### PostgreSQL version - CI runs the harness against `postgres:15`: see the `sqllogictest-postgres` job in `.github/workflows/rust.yml`. - The harness compares text. If a value renders differently in another PostgreSQL version, CI fails, even when the instant is correct. So generate and check the expected output against `postgres:15`, not against a newer local server. - An earlier revision of this PR failed in CI on `extract(epoch ...)`. The expected output had `1719792000`, and CI produced `1719792000.000000`. The file now casts each `extract` and `date_part` result to `bigint`. - We measured that value in `psql` on PostgreSQL 15.19 and 17.11. Both give `numeric` `1719792000.000000`. So `psql` shows no version difference for this value, and the cause of the earlier mismatch is not confirmed. #### Divergences that this file leaves out We measured each answer. The DataFusion answers come from the pinned expectations in apache#25175, or from `datafusion-cli` at the same base. PostgreSQL answers come from `postgres:15`. | # | Query | Session time zone | DataFusion | PostgreSQL 15 | Issue | | --- | --- | --- | --- | --- | --- | | 1 | `arrow_typeof('2024-07-01 12:00:00Z'::timestamptz)` (PostgreSQL: `pg_typeof`) | DataFusion unset, PostgreSQL `UTC` | `Timestamp(ns)` (naive) | `timestamp with time zone` | apache#25166 | | 2 | `'2024-07-01 12:00:00'::timestamp::timestamptz::timestamp` | `America/Denver` | `2024-07-01T18:00:00` | `2024-07-01 12:00:00` | apache#12218 | | 3 | `ts AT TIME ZONE 'Europe/Brussels'`, where `ts` is the aware value `2024-07-01 12:00:00Z` | DataFusion `+00:00`, PostgreSQL `UTC` | `Timestamp(ns, "Europe/Brussels")` `2024-07-01T14:00:00+02:00` | `timestamp without time zone` `2024-07-01 14:00:00` | apache#12218 | | 4 | `TIMESTAMP '2024-07-01 12:00:00' AT TIME ZONE '+05:30'` | DataFusion unset, PostgreSQL `UTC` | `2024-07-01T12:00:00+05:30` (06:30 UTC) | `2024-07-01 17:30:00+00` | apache#25170 | | 5 | `WHERE ts > '2024-07-01 06:00:00'` on rows at 00:00, 06:00, 12:00 and 18:00 UTC; DataFusion column zone `America/Denver` | DataFusion unset, PostgreSQL `UTC` | 1 row (18:00 UTC) | 2 rows (12:00 and 18:00 UTC) | closest: apache#13212 | | 6 | `WHERE ts = '2024-07-01T06:00:00Z'::timestamptz` on the same rows | DataFusion unset, PostgreSQL `UTC` | the 12:00 UTC row | the 06:00 UTC row | apache#25166 | | 7 | naive column `= '2024-07-01T18:00:00Z'::timestamptz`, for the naive row `2024-07-01 12:00:00` | `America/Denver` | 0 rows | 1 row | apache#25095 | | 8 | `date_trunc('day', ts)` for `2024-07-01T00:00:00Z`; DataFusion column zone `America/Denver` | DataFusion unset, PostgreSQL `UTC` | `2024-06-30T00:00:00-06:00` | `2024-07-01 00:00:00+00` | no issue | | 9 | `(TIMESTAMP '2024-01-15 12:00:00' AT TIME ZONE 'America/Denver') = (TIMESTAMP '2024-07-01 12:00:00' AT TIME ZONE 'America/Denver') - INTERVAL '168 days'` | DataFusion `+00:00`, PostgreSQL `UTC` | `true` | `f` | no issue | | 10 | `'2024-11-03 01:30:00'::timestamptz` (repeated hour) | `America/Denver` | error: `error computing timezone offset` | `2024-11-03 01:30:00-07` | apache#25084 | | 11 | `'2024-03-10 02:30:00'::timestamptz` (hour that does not exist) | `America/Denver` | error: `error computing timezone offset` | `2024-03-10 03:30:00-06` | apache#25084 | | 12 | `date_part('timezone_hour', ts)` in July | `America/Denver` | error: `Date part 'timezone_hour' not supported` | `-6` | fix in apache#25163 | | 13 | `date_bin` with a `1 month` stride | DataFusion unset, PostgreSQL `UTC` | returns a bin | error: `timestamps cannot be binned into intervals containing months or years` | no issue | | 14 | `from_unixtime(1719792000)` (PostgreSQL: `to_timestamp(1719792000)`) | `America/Denver` (PostgreSQL `UTC`) | `Timestamp(s)` `2024-07-01T00:00:00` (naive) | `timestamp with time zone` | apache#12892 | Rows 5 to 9 have one cause in common: DataFusion uses the time zone of the value, and PostgreSQL uses the session time zone. That is why each block of this file sets the two to the same zone. ## Are there any user-facing changes? No. This PR adds one test file. It changes no production code, no public API and no behaviour. 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> Co-authored-by: Andrew Lamb <andrew@nerdnetworks.org>
The characterization suite from apache#25175 pinned the old result, a tz-aware value relabelled to the target zone. It now returns the naive wall clock, as PostgreSQL does, so SECTION 5b changes and loses its divergence note. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The characterization suite from apache#25175 pinned the old result, a tz-aware value relabelled to the target zone. It now returns the naive wall clock, as PostgreSQL does, so SECTION 5b changes and loses its divergence note. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Which issue does this PR close?
None. This PR adds tests only. It records current behaviour, bugs included, for these issues:
::timestampon a tz-aware value gives the UTC wall clock, andtstz AT TIME ZONEstays tz-awarefrom_unixtimeand the session time zone (fixed by fix:from_unixtimeshould respectdatafusion.execution.time_zone#25161, pinned as fixed)datafusion.execution.time_zoneis not used for basic time zone inference #13212: a tz-naive side of a comparison is read in the other operand's zone, not the session zoneTimestamp(_, None)to a named timezone errors on DST boundaries #25084: casting a tz-naive value into a named zone errors at DST boundariesTIMESTAMP WITH TIME ZONE/::timestamptzcan resolve to a tz-naive type, and a cast to::timestamptzreplaces an existing zonedate_binanddate_truncdisagree on timezone-aware timestamps #25167:date_binanddate_truncdisagree on tz-aware inputdate_binwith an explicit origin drifts an hour across a DST transition #25168:date_binwith an explicit origin drifts across a DST transition (PostgreSQL and DuckDB behave the same way)AT TIME ZONE '+05:30'uses the opposite sign convention from PostgreSQL #25170:AT TIME ZONE '+05:30'uses the opposite sign convention from PostgreSQLdate_part('timezone_hour' | 'timezone_minute')(implemented by feat: supporttimezone,timezone_hourandtimezone_minuteindate_part#25163)Rationale for this change
Timezone bugs keep coming back in DataFusion because the result depends on several things at once: the session time zone (set or unset), the value's own zone, DST, and optimizer rewrites such as unwrap_cast and constant folding. Few tests cover these combinations. Outside this file, the sqllogictest suite has 47
SET datafusion.execution.time_zonestatements across 10 files. The PostgreSQL-compatibility tests had no timezone coverage until #25164.This file records what DataFusion does today for timestamps with a time zone, so any change to those semantics shows up as a diff in review instead of shipping silently. Where the recorded behaviour is wrong, or disagrees with PostgreSQL, a comment says so and links the issue. When a fix lands, the expected output changes and the comment goes away.
#25164 is the companion PR. It checks against a real PostgreSQL the cases where the two engines agree. This file covers the cases where they don't.
What changes are included in this PR?
One new file (1674 lines),
datafusion/sqllogictest/test_files/datetime/timestamps_timezone.slt, in 18 sections:::timestamptz,TIMESTAMP [WITH TIME ZONE]and table columns with the session zone unset,+00:00,+05:30,America/Denver, andEurope/BrusselsAT TIME ZONEon tz-naive and tz-aware values, the composed(tstz AT TIME ZONE z)::timestampidiom, and the'+05:30'sign convention::timestamptzon an already-zoned valuedate_bin,date_trunc,date_part(includingtimezone*fields) andextracton tz-aware input, including the explicit-origin DST caseto_char,from_unixtime,to_unixtime,to_timestamp*,to_local_timenow,current_date,current_time,make_dateunder a session zone+/-interval across DST,'1 day'vs'24 hours'greatesttype coercionNo production code changes.
Interaction with other open PRs
Some fix PRs change expectations in this file. Whichever lands second updates the pins. I'd prefer the fixes land first so they stay small, and this file absorbs the change:
from_unixtimeshould respectdatafusion.execution.time_zone#25161 (from_unixtimefollows the session zone) is already on main. The file pins the fixed behaviour.timezone,timezone_hourandtimezone_minuteindate_part#25163 (date_parttimezone fields) is already on main. SECTION 9 pinstimezone,timezone_hourandtimezone_minuteas values.AT TIME ZONEon a timezone-aware timestamp returns a naive timestamp #25165 (AT TIME ZONEon a tz-aware value returns naive) changes the two SECTION 5b queries.TIMESTAMP WITH TIME ZONEtype mapping and fix a stale comment #25171 (docs) doesn't touch this file. Its wording that'...Z'::timestamptz"discards theZ" should match this file: the offset is applied, but the result carries no zone.generate_series/range#25173 (generate_seriesprecisions) doesn't touch this file.What is the testing strategy for this PR?
The PR is the test. I generated expected output with sqllogictest
--complete. Then I read every case and rewrote the ones that weren't testing what they meant to test, and added comments with issue links where the behaviour is wrong or differs from PostgreSQL.The file runs in DataFusion only. The PostgreSQL 15 answers quoted in comments come from separate runs, with the PostgreSQL
TimeZonegiven in each comment. DuckDB 1.5.2 answers are quoted where they matter (thedate_binorigin case, and the'+05:30'offset string).DST error expectations match only the stable
Arrow error: ...tail, so changes to the wrapper error chain (for example #24920) don't break them.Run with:
Are there any user-facing changes?
No. Tests only.