Describe the bug
push_to_hub() and to_parquet() silently write nested (list/struct) columns UNCOMPRESSED, while flat columns get snappy. Measured impact: ~30x size inflation for nested-column datasets (a 114 MB JSONL with list/struct columns pushes as a 215 MB shard; the same data with compression applied is ~7 MB). The output remains valid, readable parquet — only the size is wrong.
Root cause: ArrowWriter._build_writer (src/datasets/arrow_writer.py, ~line 817) and its duplicate in src/datasets/io/parquet.py (~line 128) pass pyarrow a compression dict keyed by TOP-LEVEL feature names:
compression={col: "none" if require_storage_embed(feature) else "snappy" for col, feature in self._features.items()}
but pyarrow matches compression-dict keys against the full dotted leaf path (train.list.element...), so nested columns never match a key and fall back to UNCOMPRESSED. Introduced by #7971 first released in 4.6.0; present through 5.0.1 and on main.
Suggested fix: key the dict by the parquet leaf paths (derivable from self._schema), or pass a plain compression="snappy" string when no require_storage_embed column is present. Note #8576 fixes the related to_parquet(compression=...) TypeError but not this keying.
The same pattern reproduces through push_to_hub (verified on a 5.0.1 push of nested-column data: flat columns SNAPPY, nested UNCOMPRESSED, ~30x inflation). A fix must cover both dict sites — they are literal duplicates.
Steps to reproduce the bug
from datasets import Dataset
ds = Dataset.from_dict({
"question": ["hello world", "foo bar"],
"train": [[{"input": [[1, 2]], "output": [[3]]}]] * 2,
})
ds.to_parquet("/tmp/out.parquet")
import pyarrow.parquet as pq
rg = pq.ParquetFile("/tmp/out.parquet").metadata.row_group(0)
for i in range(rg.num_columns):
print(rg.column(i).path_in_schema, "->", rg.column(i).compression)
Observed:
question -> SNAPPY
train.list.element.input.list.element.list.element -> UNCOMPRESSED
train.list.element.output.list.element.list.element -> UNCOMPRESSED
Expected behavior
All columns should be SNAPPY — the dict's own intent is "none" only for require_storage_embed media columns, and this schema has none.
Environment info
datasets version: 5.0.1
pyarrow version: 25.0.1
- Python version: 3.12.13
- OS: macOS
Describe the bug
push_to_hub()andto_parquet()silently write nested (list/struct) columns UNCOMPRESSED, while flat columns get snappy. Measured impact: ~30x size inflation for nested-column datasets (a 114 MB JSONL with list/struct columns pushes as a 215 MB shard; the same data with compression applied is ~7 MB). The output remains valid, readable parquet — only the size is wrong.Root cause:
ArrowWriter._build_writer(src/datasets/arrow_writer.py, ~line 817) and its duplicate insrc/datasets/io/parquet.py(~line 128) pass pyarrow a compression dict keyed by TOP-LEVEL feature names:but pyarrow matches compression-dict keys against the full dotted leaf path (
train.list.element...), so nested columns never match a key and fall back to UNCOMPRESSED. Introduced by #7971 first released in 4.6.0; present through 5.0.1 and onmain.Suggested fix: key the dict by the parquet leaf paths (derivable from
self._schema), or pass a plaincompression="snappy"string when norequire_storage_embedcolumn is present. Note #8576 fixes the relatedto_parquet(compression=...)TypeError but not this keying.The same pattern reproduces through push_to_hub (verified on a 5.0.1 push of nested-column data: flat columns SNAPPY, nested UNCOMPRESSED, ~30x inflation). A fix must cover both dict sites — they are literal duplicates.
Steps to reproduce the bug
Observed:
Expected behavior
All columns should be SNAPPY — the dict's own intent is
"none"only forrequire_storage_embedmedia columns, and this schema has none.Environment info
datasetsversion: 5.0.1pyarrowversion: 25.0.1