Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Buffered
ArrowWriter.write_rowcalls concatenate the input tables before the usual schema cast. Compatible rows therefore fail when their columns arrive in a different order or a later row omits a column. This also breaksload_dataset("json", ...)for one-record JSON files, even though the same records work with immediate flushing or larger input tables.For heterogeneous buffered schemas, align each row to the effective writer schema using the existing
table_cast, then concatenate and write one batch. Homogeneous buffers retain their current path. The selected schema still governs missing-column null filling, casts and feature metadata; extra columns remain errors.The regressions cover buffering, reordered/missing columns, declared and established schemas, semantic feature metadata, invalid casts, and JSON arrays/selected fields with inferred and explicit features.
Validation on macOS arm64, Python 3.11.15 and PyArrow 25.0.1:
pytest tests/test_arrow_writer.py tests/packaged_modules/test_json.py -q: 131 passed, 11 skipped (optional agent-trace dependency).tests/test_arrow_dataset.pymap checks covering Arrow/pandas outputs, features, batching, column removal and multiprocessing: 12 passed.make qualityandgit diff --check: passed.AI assistance: OpenAI Codex generated the implementation, tests and this description, with GPT-6 Astra used for investigation. The listed checks were executed locally.