Skip to content

SNOW-3790287: Defer incompatible pyarrow warning until first use - #3047

Draft
cristiangirlea wants to merge 1 commit into
snowflakedb:mainfrom
cristiangirlea:SNOW-3790287-defer-pyarrow-version-warning
Draft

cristiangirlea wants to merge 1 commit into
snowflakedb:mainfrom
cristiangirlea:SNOW-3790287-defer-pyarrow-version-warning

Conversation

@cristiangirlea

Copy link
Copy Markdown

Please answer these questions before submitting your pull requests. Thanks!

  1. What GitHub issue is this PR addressing? Make sure that there is an accompanying issue to your PR.

    Fixes SNOW-3790287: Pyarrow version warning fires at import time even when pandas tools are never used #2950

  2. Fill out the following pre-review checklist:

    • I am adding a new automated test(s) to verify correctness of my new code
    • I am adding new logging messages
    • I am adding a new telemetry message
    • I am modifying authorization mechanisms
    • I am adding new credentials
    • I am modifying OCSP code
    • I am adding a new dependency
  3. Please describe how your code solves the related issue.

    The pyarrow version check lived in _import_or_missing_pandas_option(), which runs at module level in options.py. Any environment where pyarrow is outside the [pandas] extra's range therefore warned on import snowflake.connector, even when pyarrow came from an unrelated package (awswrangler in the report) and no pandas/Arrow API was used.

    This PR moves the check, unchanged, into warn_if_incompatible_pyarrow(). It is memoized with functools.lru_cache, so it runs at most once per process and doesn't re-read package metadata on every fetch. It is called from the places that hand pyarrow or pandas objects to the user:

    • ArrowResultBatch._create_iter when iter_unit == TABLE_UNIT, sync and aio. Every table path funnels through it: fetch_arrow_all/fetch_arrow_batches, fetch_pandas_all/fetch_pandas_batches, and ResultBatch.to_arrow/to_pandas. Row iteration (fetchone etc.) never triggers it.
    • write_pandas, sync and aio.

    Importing pandas/pyarrow and setting ARROW_DEFAULT_MEMORY_POOL still happen at import time; only the version check moved. Users of the pandas/Arrow APIs get the same warning text as before, on first use. One small hardening: if the connector's metadata has no pyarrow requirement for the pandas extra, the check now skips instead of raising AttributeError on None.

    Tests (in test/integ/pandas_it/test_unit_options.py, alongside the existing pyarrow-check test):

    • test_import_does_not_warn_about_incompatible_pyarrow: with an incompatible pyarrow faked through importlib.metadata.distribution, the import-time function emits no warning, even with warnings turned into errors.
    • test_incompatible_pyarrow_warns_once_on_use: the deferred check warns on its first call and is silent on the second.
    • test_arrow_result_batch_checks_pyarrow_only_for_tables: _create_iter calls the check for TABLE_UNIT and not for ROW_UNIT.
    • test_write_pandas_checks_pyarrow: write_pandas calls the check.
    • test_pandas_option_reporting: updated to call the new function for the "Cannot determine…" log.

    End-to-end check in a fresh interpreter, faking pyarrow 0.0.1 with -W error::UserWarning:

    • before: import snowflake.connector raises the "incompatible version of 'pyarrow'" warning;
    • after: the import is clean, and the same warning is raised on the first write_pandas call.

    Locally, the full sync unit suite and the aio unit tests for result batches, pandas tools and cursors have identical results with and without this change. pre-commit passes on the changed files.

The pyarrow version check ran in options.py at import time, so any
environment with a pyarrow outside the [pandas] extra's range warned on
`import snowflake.connector`, even when pyarrow was installed by an
unrelated package and no pandas/Arrow API was used.

Move the check into warn_if_incompatible_pyarrow(), memoized to run once
per process, and call it from the APIs that hand pyarrow or pandas
objects to the user: ArrowResultBatch._create_iter for table iteration
(covers fetch_arrow_*, fetch_pandas_*, ResultBatch.to_arrow/to_pandas)
and write_pandas, in both the sync and aio code paths.

Fixes snowflakedb#2950

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

SNOW-3790287: Pyarrow version warning fires at import time even when pandas tools are never used

1 participant