Repository navigation
Tell the model how many rows a Parquet or Avro file holds - #73925
Merged
Merged
Conversation
vatsrahul1001
approved these changes
Sep 30, 2026
ObjectStorageToolset's read_file returned a Parquet or Avro file as its schema and first 20 rows, with no row count, and silently ignored offset and limit. Asked how many rows a file held, a model could only retry with an offset and get the same 20 rows back. The result now opens with the file's row count and says that offset and limit do not apply, ahead of the sample so that cutting a long result never drops it. Parquet's count comes from the file footer. An Avro file is read in one pass over its blocks: the sample comes from the first blocks and the count from every block header, so only the blocks the sample reaches have records decoded. A file whose schema is not a record now shows its values in the sample too. A corrupt deflate block anywhere in an Avro file, and corrupt .xz text, raise zlib.error and lzma.LZMAError, which are not OSError or ValueError. Both are now refused like any other unreadable file instead of failing the run.
kaxil
force-pushed
the
commonai-object-storage-row-count
branch
from
September 30, 2026 06:01
e7747df to
ba6c64e
Compare
Lee-W
reviewed
Sep 30, 2026
Comment on lines
+329
to
+331
| f"offset and limit do not apply to {columnar.capitalize()} files.\n" | ||
| ) | ||
| return _cut(header + sample.text, self._max_output_bytes) |
Member
There was a problem hiding this comment.
Suggested change
| f"offset and limit do not apply to {columnar.capitalize()} files.\n" | |
| ) | |
| return _cut(header + sample.text, self._max_output_bytes) | |
| f"offset and limit do not apply to {columnar.capitalize()} files." | |
| ) | |
| return _cut(f"{header}\n{sample.text}", self._max_output_bytes) |
non-blocking, but i prefer doing something like this. although the difference is not noticeable
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Follow-up to #73899.
ObjectStorageToolset'sread_filereturned a Parquet or Avro file as its schema and first 20 rows, with no row count, and silently ignoredoffsetandlimit. In a run against a real model, asked how many rows a Parquet file held, the agent calledread_fileagain withoffset=21, got the same 20 rows back, and answered that it could not tell.The result now opens with the row count and says that
offsetandlimitdo not apply:The header comes before the sample, so cutting a long result never drops it. Parquet's count comes from the file footer. An Avro file is read in one pass over its blocks: the sample comes from the first blocks and the count from every block header, so only the blocks the sample reaches have their records decoded, and the file is opened once, under the existing
max_read_bytescheck.LLMFileAnalysisOperatorkeeps its current Avro path. A file whose Avro schema is not a record now shows its values in the sample instead of an empty list.Counting reaches every block, which exposed a gap: a corrupt deflate block raises
zlib.error, and corrupt.xztext raiseslzma.LZMAError, neither of which is anOSErrororValueError, so either one failed the run. Both are now refused like any other unreadable file, and the model is told why.{pr_number}.significant.rst, in airflow-core/newsfragments. You can add this file in a follow-up commit after the PR is created so you know the PR number.