Skip to content

Tell the model how many rows a Parquet or Avro file holds - #73925

Merged
kaxil merged 1 commit into
commonai-native-docsfrom
commonai-object-storage-row-count
Sep 30, 2026
Merged

kaxil merged 1 commit into
commonai-native-docsfrom
commonai-object-storage-row-count

Conversation

@kaxil

@kaxil kaxil commented Sep 29, 2026

Copy link
Copy Markdown
Member

Follow-up to #73899. ObjectStorageToolset's read_file returned a Parquet or Avro file as its schema and first 20 rows, with no row count, and silently ignored offset and limit. In a run against a real model, asked how many rows a Parquet file held, the agent called read_file again with offset=21, got the same 20 rows back, and answered that it could not tell.

The result now opens with the row count and says that offset and limit do not apply:

Rows: 250. The schema and the first 20 rows follow; offset and limit do not apply to Parquet files.
Schema: order_id: int64, region: string, total: double
Sample rows:
...

The header comes before the sample, so cutting a long result never drops it. Parquet's count comes from the file footer. An Avro file is read in one pass over its blocks: the sample comes from the first blocks and the count from every block header, so only the blocks the sample reaches have their records decoded, and the file is opened once, under the existing max_read_bytes check. LLMFileAnalysisOperator keeps its current Avro path. A file whose Avro schema is not a record now shows its values in the sample instead of an empty list.

Counting reaches every block, which exposed a gap: a corrupt deflate block raises zlib.error, and corrupt .xz text raises lzma.LZMAError, neither of which is an OSError or ValueError, so either one failed the run. Both are now refused like any other unreadable file, and the model is told why.


  • Read the Pull Request Guidelines for more information. Note: commit author/co-author name and email in commits become permanently public when merged.
  • For fundamental code changes, an Airflow Improvement Proposal (AIP) is needed.
  • When adding dependency, check compliance with the ASF 3rd Party License Policy.
  • For significant user-facing changes create newsfragment: {pr_number}.significant.rst, in airflow-core/newsfragments. You can add this file in a follow-up commit after the PR is created so you know the PR number.

ObjectStorageToolset's read_file returned a Parquet or Avro file as its
schema and first 20 rows, with no row count, and silently ignored offset
and limit. Asked how many rows a file held, a model could only retry with
an offset and get the same 20 rows back.

The result now opens with the file's row count and says that offset and
limit do not apply, ahead of the sample so that cutting a long result
never drops it. Parquet's count comes from the file footer. An Avro file
is read in one pass over its blocks: the sample comes from the first
blocks and the count from every block header, so only the blocks the
sample reaches have records decoded. A file whose schema is not a record
now shows its values in the sample too.

A corrupt deflate block anywhere in an Avro file, and corrupt .xz text,
raise zlib.error and lzma.LZMAError, which are not OSError or ValueError.
Both are now refused like any other unreadable file instead of failing the
run.
@kaxil
kaxil force-pushed the commonai-object-storage-row-count branch from e7747df to ba6c64e Compare September 30, 2026 06:01
@kaxil
kaxil merged commit d4ba567 into main Sep 30, 2026
63 of 75 checks passed
@kaxil
kaxil deleted the commonai-object-storage-row-count branch September 30, 2026 06:04
Comment on lines +329 to +331
f"offset and limit do not apply to {columnar.capitalize()} files.\n"
)
return _cut(header + sample.text, self._max_output_bytes)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
f"offset and limit do not apply to {columnar.capitalize()} files.\n"
)
return _cut(header + sample.text, self._max_output_bytes)
f"offset and limit do not apply to {columnar.capitalize()} files."
)
return _cut(f"{header}\n{sample.text}", self._max_output_bytes)

non-blocking, but i prefer doing something like this. although the difference is not noticeable

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants