Skip to content

fix(csv): sniff dialect after leading empty lines - #4467

Merged
PeterStaar-IBM merged 1 commit into
docling-project:mainfrom
sergioperezcheco:fix/csv-leading-blank-dialect
Oct 1, 2026
Merged

PeterStaar-IBM merged 1 commit into
docling-project:mainfrom
sergioperezcheco:fix/csv-leading-blank-dialect

Conversation

@sergioperezcheco

Copy link
Copy Markdown
Contributor

Fixes #4466.

Delimiter detection

An empty first physical line sends detection straight to the larger sample. For ragged records, csv.Sniffer cannot find a consistent delimiter count in that sample and the backend falls back to comma. The same file without the leading empty line is correctly detected from its header.

Skip genuinely empty physical lines before sniffing, and start the lazy 4096-character fallback window at the first nonempty line. Parsing still starts at offset zero. Empty-field records, quoted empty fields and whitespace-only fields are not skipped, and existing blank-record filtering remains unchanged. Starting the fallback window at the header also prevents a long prefix of empty lines from exhausting it before a multiline quoted field is reached.

This complements #4316, which filters empty records after parsing, and retains the first-line-first detection strategy introduced in #3985.

Regression coverage

35 added public DocumentConverter cases cover Path and stream inputs, all five supported delimiters, LF/CRLF prefixes, multiline quoted headers after 4096 empty lines, and preservation of nonempty first records. On the original implementation, these cases produce 25 failures and 10 passes. The complete CSV suite now passes all 52 cases; existing reference outputs are unchanged.

make validate passes all applicable project hooks, including Ruff, ty, Tach and module coverage. No full non-CSV test suite was run.

Implementation, investigation and local validation used AI assistance through Hermes with OpenAI GPT-6.1-Sol.

Start the bounded fallback sample at the first nonempty physical line without changing the input passed to csv.reader.

Assisted-by: OpenAI GPT-6.1-Sol <noreply@openai.com>
Signed-off-by: sergioperezcheco <checo520@outlook.com>
@github-actions

github-actions Bot commented Oct 1, 2026

Copy link
Copy Markdown
Contributor

✅ DCO Check Passed

Thanks @sergioperezcheco, all your commits are properly signed off. 🎉

@mergify

mergify Bot commented Oct 1, 2026

Copy link
Copy Markdown
Contributor

Merge Protections

🟢 Merge protection satisfied — ready to merge.

Show 1 satisfied protection

🟢 Enforce conventional commit

Make sure that we follow https://www.conventionalcommits.org/en/v1.0.0/

  • title ~= ^(fix|feat|docs|style|refactor|perf|test|build|ci|chore|revert)(?:\(.+\))?(!)?:

@codecov

codecov Bot commented Oct 1, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

@PeterStaar-IBM PeterStaar-IBM left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

lgtm!

@PeterStaar-IBM
PeterStaar-IBM merged commit 2dd8c53 into docling-project:main Oct 1, 2026
26 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

CSV leading empty lines can cause incorrect delimiter detection

2 participants