feat(pdf): infer heading levels from PDF bookmarks/ToC - #3688
Conversation
Follow-up to docling-project#3633, which inferred PDF heading levels from numbering and font style but left out the document's own bookmarks/table-of-contents - usually the most reliable hierarchy signal. This wires bookmarks in as the authoritative source, with precedence bookmarks > numbering > style. - Extract the outline via a shared pypdfium2 helper on PdfDocumentBackend (active under both the pypdfium2 and the default docling-parse backends); surfaced on ConversionResult.pdf_outline. - Match bookmarks to detected headings by fuzzy title (difflib) + page, comparing with and without leading numbering markers; a confident match is authoritative, and a confidently matched list-item is promoted to a heading (layout models often mis-classify headings as list-items). - Partial/noisy outlines never degrade the numbering result: unmatched entries fall back to numbering/style. New options use_bookmarks (default true) and bookmark_match_threshold; no chunker changes needed. Signed-off-by: Daniel Nguyen <danielnguyenh07@gmail.com>
|
✅ DCO Check Passed Thanks @DanielNg0729, all your commits are properly signed off. 🎉 |
Merge Protections🟢 All 2 merge protections satisfied — ready to merge. Show 2 satisfied protections🟢 Enforce conventional commitMake sure that we follow https://www.conventionalcommits.org/en/v1.0.0/
🟢 Require two reviewer for test updatesWhen test data is updated, we require two reviewers
|
Codecov Report❌ Patch coverage is 📢 Thoughts on this report? Let us know! |
|
@DanielNg0729 Love the work: I would suggest that we add at least 1 example pdf with an existing outline in the tests. |
Signed-off-by: Daniel Nguyen <danielnguyenh07@gmail.com>
|
Thanks, Dr. @PeterStaar-IBM! I've added a committed sample PDF ( |
|
@DanielNg0729 great to see this continuation PR. I have not reviewed deeply yet, however one feedback already below:
|
…contents Per review: DoclingParseDocumentBackend (+v2/v4) and ThreadedDoclingParseDocumentBackend now read the outline via docling-parse's native get_table_of_contents() instead of the internal pypdfium handle, so the signal works without a pypdfium dependency (incl. the threaded backend, which has none). The pypdfium2 backend keeps its richer extraction (title + page + position). The native ToC carries title + hierarchy only, so page_no/y_top are left None and the matcher falls back to title-only matching for those entries. Signed-off-by: Daniel Nguyen <danielnguyenh07@gmail.com>
Signed-off-by: Nguyen Hoang Duong <danielnguyenh07@gmail.com>
|
Thank you very much for the feedback Dr @cau-git I've implemented it as requested: |
get_table_of_contents() returns None for PDFs with no embedded outline, so outline_from_docling_parse crashed on None.children and failed heading-hierarchy conversion (use_bookmarks defaults on). Return [] in that case and fall back to numbering/style. Adds a regression test. Signed-off-by: Daniel Nguyen <danielnguyenh07@gmail.com>
|
Hi Dr. @cau-git, Dr. @PeterStaar-IBM Quick update on the failing Could you kindly re-trigger the CI when you have a moment? Thank you very much! |
of course! |
…dule Per review: PdfOutlineItem is an internal backend->heading-stage data structure (already excluded from serialization), not public datamodel. Move it out of base_models into docling/utils/pdf_outline.py, next to the extractors that produce it. Kept as our own type rather than reusing docling-parse's PdfTocEntry, which has no vertical-position (y_top) field and would couple the pypdfium2 backend to docling-parse. Signed-off-by: Daniel Nguyen <danielnguyenh07@gmail.com>
|
@cau-git Let's review this one internally! |
cau-git
left a comment
There was a problem hiding this comment.
@DanielNg0729 Thanks for the updates! I went through the code and posted a few remarks and questions below.
…ternal Signed-off-by: Daniel Nguyen <danielnguyenh07@gmail.com>
…rom_docling_parse for naming consistency Signed-off-by: Daniel Nguyen <danielnguyenh07@gmail.com>
…line), reset after the heading stage consumes it Signed-off-by: Daniel Nguyen <danielnguyenh07@gmail.com>
|
Thanks @DanielNg0729 and reviewing team. |
|
This is amazing work @DanielNg0729! Looking forward to trying it on our documents after the Summer! 😊 |
Add a usage guide and a runnable example for the section-header level inference shipped in docling-project#3633, docling-project#3688 and docling-project#3984, covering the bookmark, numbering and font-style signals, their precedence, and the generate_parsed_pages requirement of the style fallback. Signed-off-by: Daniel Nguyen <danielnguyenh07@gmail.com>
* docs: document PDF heading-level inference Add a usage guide and a runnable example for the section-header level inference shipped in #3633, #3688 and #3984, covering the bookmark, numbering and font-style signals, their precedence, and the generate_parsed_pages requirement of the style fallback. Signed-off-by: Daniel Nguyen <danielnguyenh07@gmail.com> * Remove PDF heading recovery from features list Removed PDF heading recovery feature from the list. Signed-off-by: Nguyen Hoang Duong <danielnguyenh07@gmail.com> --------- Signed-off-by: Daniel Nguyen <danielnguyenh07@gmail.com> Signed-off-by: Nguyen Hoang Duong <danielnguyenh07@gmail.com>
This PR is the follow-up to #3633 discussed in #3676.
Issue resolved by this Pull Request:
Resolves #3676
What this does
PR #3633 inferred PDF heading levels from numbering and font style, but deliberately left out the document's own bookmarks / table-of-contents. This is usually the most reliable hierarchy signal when a PDF has one. This PR wires bookmarks in as the authoritative source, so the precedence becomes bookmarks (when confidently matched) > numbering > style. Because the chunkers already read hierarchy from the document, the retrieval improvement flows through
automatically.
Approach
pypdfium2helper reads the outline viaget_toc(), exposed asPdfDocumentBackend.get_document_outline(). Since both the pypdfium2 and the default docling-parse backends hold a PDFium handle, it's implemented once on their shared base, so the signal is available under the default backend too. The outline is surfaced onConversionResult.pdf_outline(excluded from serialization, so the finalDoclingDocumentisn't bloated).difflib), comparing titles with and without their leading numbering marker, with a containment boost for truncated bookmarks, gated by a page constraint and a configurable threshold.SectionHeaderItemat the bookmark's level.Configuration
This PR adds two options to
HeadingHierarchyOptions; no existing options change behavior.use_bookmarksboolTrueFalseto keep the previous #3633 behavior (numbering + style only).bookmark_match_thresholdfloat(0–1)0.8These sit alongside the existing options (unchanged):
enabled(still off by default — must beTruefor any inference to run),use_numbering,use_style,numbering_schemes,max_level.Precedence when enabled:
bookmarks (confidently matched) > numbering > style.Usage
Notes / scope
heading_source,bookmark_path, match confidence) would need adocling-coreschema change so i would happy to open a separate issue for it so it doesn't block this pipeline change.Thanks for the guidance on this Dr @PeterStaar-IBM . I'm happy to adjust anything!
Checklist:
tests/test_heading_hierarchy_bookmarks.py)