Describe the bug
Dataset.to_pandas() returns Array2D/Array3D/... columns as datasets.features.features.PandasArrayExtensionArray. That pandas extension array does not implement pyarrow's __arrow_array__ protocol, so converting such a DataFrame back to Arrow raises ArrowTypeError: Did not pass numpy.dtype object. This breaks:
Dataset.from_pandas(ds.to_pandas()), with or without features=
ds.with_format("pandas").map(fn, batched=True): batches are DataFrames and the returned DataFrame still contains the ArrayXD column, so the writer's pa.Table.from_pandas fails even when fn never touches that column.
Both fixed-shape and dynamic-first-dimension (shape=(None, ...)) arrays are affected. PandasArrayExtensionDtype.__from_arrow__ covers the Arrow -> pandas direction, but there is no pandas -> Arrow direction.
Steps to reproduce the bug
import numpy as np
from datasets import Array2D, Dataset, Features, Value
features = Features({"image": Array2D(shape=(2, 3), dtype="int32"), "label": Value("int64")})
ds = Dataset.from_dict({"image": [np.arange(6).reshape(2, 3)] * 3, "label": [0, 1, 2]}, features=features)
Dataset.from_pandas(ds.to_pandas())
# pyarrow.lib.ArrowTypeError: ('Did not pass numpy.dtype object', 'Conversion failed for column image with type array[int32]')
ds.with_format("pandas").map(lambda df: df, batched=True)
# same ArrowTypeError, raised from ArrowWriter.write_table -> pa.Table.from_pandas
Expected behavior
Both calls succeed and the column keeps its Array2D(shape=(2, 3), dtype='int32') feature, i.e. Dataset.from_pandas(ds.to_pandas())[:] == ds[:].
Implementing __arrow_array__ on PandasArrayExtensionArray (build the list storage with the existing to_pyarrow_listarray helper and wrap it in the matching ArrayNDExtensionType) fixes all paths above. I have a patch with tests ready and will open a PR referencing this issue.
Disclosure: found, reproduced and the draft fix tested locally with the help of an AI coding agent (Claude Code); reviewed before filing.
Environment info
- `datasets` version: 5.0.2.dev0 (main, a4ee9cf1)
- Platform: macOS-15.3.1-arm64-arm-64bit
- Python version: 3.12.13
- `huggingface_hub` version: 1.33.0
- PyArrow version: 25.0.1
- Pandas version: 3.0.6
- `fsspec` version: 2026.7.0
Describe the bug
Dataset.to_pandas()returnsArray2D/Array3D/... columns asdatasets.features.features.PandasArrayExtensionArray. That pandas extension array does not implement pyarrow's__arrow_array__protocol, so converting such a DataFrame back to Arrow raisesArrowTypeError: Did not pass numpy.dtype object. This breaks:Dataset.from_pandas(ds.to_pandas()), with or withoutfeatures=ds.with_format("pandas").map(fn, batched=True): batches are DataFrames and the returned DataFrame still contains the ArrayXD column, so the writer'spa.Table.from_pandasfails even whenfnnever touches that column.Both fixed-shape and dynamic-first-dimension (
shape=(None, ...)) arrays are affected.PandasArrayExtensionDtype.__from_arrow__covers the Arrow -> pandas direction, but there is no pandas -> Arrow direction.Steps to reproduce the bug
Expected behavior
Both calls succeed and the column keeps its
Array2D(shape=(2, 3), dtype='int32')feature, i.e.Dataset.from_pandas(ds.to_pandas())[:] == ds[:].Implementing
__arrow_array__onPandasArrayExtensionArray(build the list storage with the existingto_pyarrow_listarrayhelper and wrap it in the matchingArrayNDExtensionType) fixes all paths above. I have a patch with tests ready and will open a PR referencing this issue.Disclosure: found, reproduced and the draft fix tested locally with the help of an AI coding agent (Claude Code); reviewed before filing.
Environment info