Skip to content

Honor explicit Parquet compression and encoding options - #8576

Draft
Mr-J-Forest wants to merge 1 commit into
huggingface:mainfrom
Mr-J-Forest:fix/parquet-writer-options
Draft

Mr-J-Forest wants to merge 1 commit into
huggingface:mainfrom
Mr-J-Forest:fix/parquet-writer-options

Conversation

@Mr-J-Forest

Copy link
Copy Markdown

Dataset.to_parquet(..., compression="gzip") raises TypeError: ParquetWriter() got multiple values for keyword argument 'compression' on current main. Explicit use_dictionary and column_encoding options fail for the same reason: the writer supplies defaults as named arguments while leaving the caller's values in **parquet_writer_kwargs.

Pop these three options from the forwarded kwargs, using the existing feature-dependent defaults only when an option is absent. This lets callers choose compression and encoding, including explicit None and False values.

The regression tests write to both paths and binary buffers, read the data back, and inspect Parquet metadata for the requested codecs and encodings. They cover global/per-column compression, disabling compression, dictionary selection, and explicit column encodings, along with unchanged defaults.

Validation on Windows, Python 3.12.13, PyArrow 25.0.1:

  • Before the fix: 14 new regression cases failed with the duplicate-keyword error; the two default-option controls passed.
  • After the fix: python -m pytest tests/io/test_parquet.py -q — 72 passed.
  • Ruff lint and format checks passed for both changed files; git diff --check passed.

Reproducer:

from io import BytesIO
from datasets import Dataset

dataset = Dataset.from_dict({"value": [1, 2, 3]})
dataset.to_parquet(BytesIO(), compression="gzip")

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant