Skip to content

Implement __arrow_array__ on PandasArrayExtensionArray so ArrayXD columns can go back from pandas to Arrow - #8701

Open
JoeyTan21 wants to merge 1 commit into
huggingface:mainfrom
JoeyTan21:fix/pandas-array-arrow-array
Open

JoeyTan21 wants to merge 1 commit into
huggingface:mainfrom
JoeyTan21:fix/pandas-array-arrow-array

Conversation

@JoeyTan21

Copy link
Copy Markdown

Fixes #8699

Dataset.to_pandas() returns ArrayXD columns as PandasArrayExtensionArray, but that extension array never implemented pyarrow's __arrow_array__ protocol, so pa.Table.from_pandas raised ArrowTypeError: Did not pass numpy.dtype object. This broke Dataset.from_pandas(ds.to_pandas()) and ds.with_format("pandas").map(fn, batched=True) for any dataset with an Array2D..Array5D column.

This PR adds PandasArrayExtensionArray.__arrow_array__:

  • builds the nested list storage with the existing to_pyarrow_listarray helper (same code path as ArrowWriter), so fixed-shape and dynamic-first-dimension arrays and null rows are handled the same way as when writing;
  • wraps it in the matching ArrayNDExtensionType, taken from the requested type when pyarrow passes one (e.g. schema=), otherwise inferred from the numpy data (shape/dtype), so the column keeps its ArrayXD feature after the round trip;
  • if a non-extension Arrow type is requested, returns the storage cast to it.

Tests: test_array_xd_from_pandas (fixed + dynamic shape; covers pa.array(series.array), Dataset.from_pandas with/without features, and a pandas-formatted batched map) and test_array_xd_from_pandas_with_none (null row preserved as an Arrow null). tests/features/test_array_xd.py, test_features.py, test_formatting.py and the pandas/map subset of test_arrow_dataset.py pass; ruff check / ruff format --check clean on the changed files.

Disclosure: found, reproduced and the draft fix tested locally with the help of an AI coding agent (Claude Code); reviewed before filing.

馃 Generated with Claude Code

`Dataset.to_pandas()` returns ArrayXD columns as `PandasArrayExtensionArray`,
but that pandas extension array did not implement pyarrow's `__arrow_array__`
protocol, so converting the DataFrame back to Arrow failed with
`ArrowTypeError: Did not pass numpy.dtype object`. This broke
`Dataset.from_pandas(ds.to_pandas())` and `ds.with_format("pandas").map(fn,
batched=True)` for any dataset with an Array2D..Array5D column.

The new `__arrow_array__` builds the list storage with the existing
`to_pyarrow_listarray` helper and wraps it in the matching ArrayXD extension
type (inferred from the numpy data, or taken from the requested type), so the
column keeps its Array2D/3D/4D/5D feature after the round trip.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@HuggingFaceDocBuilderDev

Copy link
Copy Markdown

The docs for this PR live here. All of your documentation changes will be reflected on that endpoint. The docs are available until 30 days after the last update.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Dataset.from_pandas and pandas-formatted map fail on ArrayXD columns: PandasArrayExtensionArray has no __arrow_array__

2 participants