Skip to content

push_to_hub() / to_parquet() leave nested list/struct columns uncompressed #8671

Description

@aksakalmustafa

Describe the bug

push_to_hub() and to_parquet() silently write nested (list/struct) columns UNCOMPRESSED, while flat columns get snappy. Measured impact: ~30x size inflation for nested-column datasets (a 114 MB JSONL with list/struct columns pushes as a 215 MB shard; the same data with compression applied is ~7 MB). The output remains valid, readable parquet — only the size is wrong.

Root cause: ArrowWriter._build_writer (src/datasets/arrow_writer.py, ~line 817) and its duplicate in src/datasets/io/parquet.py (~line 128) pass pyarrow a compression dict keyed by TOP-LEVEL feature names:

compression={col: "none" if require_storage_embed(feature) else "snappy" for col, feature in self._features.items()}

but pyarrow matches compression-dict keys against the full dotted leaf path (train.list.element...), so nested columns never match a key and fall back to UNCOMPRESSED. Introduced by #7971 first released in 4.6.0; present through 5.0.1 and on main.

Suggested fix: key the dict by the parquet leaf paths (derivable from self._schema), or pass a plain compression="snappy" string when no require_storage_embed column is present. Note #8576 fixes the related to_parquet(compression=...) TypeError but not this keying.

The same pattern reproduces through push_to_hub (verified on a 5.0.1 push of nested-column data: flat columns SNAPPY, nested UNCOMPRESSED, ~30x inflation). A fix must cover both dict sites — they are literal duplicates.

Steps to reproduce the bug

from datasets import Dataset

ds = Dataset.from_dict({
    "question": ["hello world", "foo bar"],
    "train": [[{"input": [[1, 2]], "output": [[3]]}]] * 2,
})
ds.to_parquet("/tmp/out.parquet")

import pyarrow.parquet as pq
rg = pq.ParquetFile("/tmp/out.parquet").metadata.row_group(0)
for i in range(rg.num_columns):
    print(rg.column(i).path_in_schema, "->", rg.column(i).compression)

Observed:

question                                                -> SNAPPY
train.list.element.input.list.element.list.element     -> UNCOMPRESSED
train.list.element.output.list.element.list.element    -> UNCOMPRESSED

Expected behavior

All columns should be SNAPPY — the dict's own intent is "none" only for require_storage_embed media columns, and this schema has none.

Environment info

  • datasets version: 5.0.1
  • pyarrow version: 25.0.1
  • Python version: 3.12.13
  • OS: macOS

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions