Skip to content

fix(pdf): keep failed-page sizes on the conversion, not on the pipeline - #4486

Open
RizgarOzan wants to merge 1 commit into
docling-project:mainfrom
RizgarOzan:fix/page-sizes-per-conversion
Open

RizgarOzan wants to merge 1 commit into
docling-project:mainfrom
RizgarOzan:fix/page-sizes-per-conversion

Conversation

@RizgarOzan

Copy link
Copy Markdown
Contributor

Issue resolved by this Pull Request:
Resolves #4478

StandardPdfPipeline kept _page_sizes_by_no on the pipeline instance and reset it at the start and end of every _build_document(). DocumentConverter reuses one pipeline instance per (pipeline_cls, options), so two concurrent convert() calls on the same converter shared that dict: the second conversion wiped the first one's sizes, and a failed page could get another document's size or Size(0, 0).

The dict now lives on the ConversionResult as a private attribute, next to _pdf_outline, which already follows the same build → assemble lifecycle. The pipeline writes to conv_res._page_sizes_by_no while producing pages and reads it back in _add_failed_pages_to_document; the instance attribute and its two resets are gone. No behaviour change for a single conversion.

How I tested

  • tests/test_pdf_streaming_foundation.py::test_failed_page_sizes_are_kept_per_conversion builds two conversions on one bare pipeline (the file's existing synthetic backend, no models), then adds the first one's missing pages. On main pages 2–4 come back as Size(0, 0); with this change all four keep 100×200.
  • pytest tests/test_pdf_streaming_foundation.py tests/test_failed_pages.py: 8 passed, 4 skipped (the skips are the existing fixture-less tests).
  • ruff check / ruff format --check clean, tach check OK, ty check on the two touched modules reports the same 18 pre-existing warnings as main and no errors.

Checklist:

  • Documentation has been updated, if necessary.
  • Examples have been added, if necessary.
  • Tests have been added, if necessary.

StandardPdfPipeline stored the page sizes it records for failed pages in an
instance dict and reset it at the start and end of every _build_document()
call. DocumentConverter reuses one pipeline instance for every conversion, so
two concurrent convert() calls on the same converter shared and wiped each
other's sizes, and a failed page could end up with another document's size or
with Size(0, 0). Keep the dict on the ConversionResult as a private attribute,
next to _pdf_outline which already follows that lifecycle.

Resolves docling-project#4478

Signed-off-by: Rızgar Ozan <rizgarozan7@gmail.com>
@github-actions

github-actions Bot commented Oct 1, 2026

Copy link
Copy Markdown
Contributor

✅ DCO Check Passed

Thanks @RizgarOzan, all your commits are properly signed off. 🎉

@mergify

mergify Bot commented Oct 1, 2026

Copy link
Copy Markdown
Contributor

Merge Protections

🟢 Merge protection satisfied — ready to merge.

Show 1 satisfied protection

🟢 Enforce conventional commit

Make sure that we follow https://www.conventionalcommits.org/en/v1.0.0/

  • title ~= ^(fix|feat|docs|style|refactor|perf|test|build|ci|chore|revert)(?:\(.+\))?(!)?:

@codecov

codecov Bot commented Oct 2, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

@dolfim-ibm dolfim-ibm left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

lgtm

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

StandardPdfPipeline: unsynchronized self._page_sizes_by_no can race when the same cached pipeline instance converts documents concurrently

2 participants