Skip to content

feat(pdf): infer heading levels from PDF bookmarks/ToC - #3688

Merged
cau-git merged 9 commits into
docling-project:mainfrom
DanielNg0729:feat/pdf-bookmark-hierarchy
Jul 1, 2026
Merged

cau-git merged 9 commits into
docling-project:mainfrom
DanielNg0729:feat/pdf-bookmark-hierarchy

Conversation

@DanielNg0729

@DanielNg0729 DanielNg0729 commented Jun 24, 2026 •

Copy link
Copy Markdown
Member

This PR is the follow-up to #3633 discussed in #3676.

Issue resolved by this Pull Request:
Resolves #3676

What this does

PR #3633 inferred PDF heading levels from numbering and font style, but deliberately left out the document's own bookmarks / table-of-contents. This is usually the most reliable hierarchy signal when a PDF has one. This PR wires bookmarks in as the authoritative source, so the precedence becomes bookmarks (when confidently matched) > numbering > style. Because the chunkers already read hierarchy from the document, the retrieval improvement flows through
automatically.

Approach

  • Backend extraction. A shared pypdfium2 helper reads the outline via get_toc(), exposed as PdfDocumentBackend.get_document_outline(). Since both the pypdfium2 and the default docling-parse backends hold a PDFium handle, it's implemented once on their shared base, so the signal is available under the default backend too. The outline is surfaced on ConversionResult.pdf_outline (excluded from serialization, so the final DoclingDocument isn't bloated).
  • Matching
    • Bookmark titles often differ from the on-page text, so matching is fuzzy (stdlib difflib), comparing titles with and without their leading numbering marker, with a containment boost for truncated bookmarks, gated by a page constraint and a configurable threshold.
    • Headings the layout model gets wrong are often classified as list-items, so the matcher also considers list-items and promotes a confidently matched one to a SectionHeaderItem at the bookmark's level.
  • Safety. Only confidently matched entries act; partial/noisy outlines never degrade the feat(pdf): infer PDF heading levels so the hierarchy isn't flattened #3633 numbering result, unmatched headings fall back to numbering/style. PDF-only; OCR/VLM PDFs without an embedded outline simply lack the signal.

Configuration

This PR adds two options to HeadingHierarchyOptions; no existing options change behavior.

Option Type Default What it controls
use_bookmarks bool True New. Use PDF bookmarks/ToC as the authoritative heading signal when present. Set to False to keep the previous #3633 behavior (numbering + style only).
bookmark_match_threshold float (0–1) 0.8 New. Minimum fuzzy title similarity for a bookmark to match a detected heading/list-item. Higher = stricter; below it the bookmark is ignored and the heading falls back to numbering/style.

These sit alongside the existing options (unchanged): enabled (still off by default — must be True for any inference to run), use_numbering, use_style, numbering_schemes, max_level.

Precedence when enabled: bookmarks (confidently matched) > numbering > style.

Usage

from docling.datamodel.base_models import InputFormat
from docling.datamodel.pipeline_options import HeadingHierarchyOptions, PdfPipelineOptions
from docling.document_converter import DocumentConverter, PdfFormatOption

pipeline_options = PdfPipelineOptions()
pipeline_options.heading_hierarchy_options = HeadingHierarchyOptions(
    enabled=True,                  # turn heading-level inference on (off by default)
    use_bookmarks=True,            # new: PDF bookmarks/ToC as the authoritative signal
    bookmark_match_threshold=0.8,  # new: fuzzy title-match strictness (0..1)
)

converter = DocumentConverter(
    format_options={InputFormat.PDF: PdfFormatOption(pipeline_options=pipeline_options)}
)
result = converter.convert("contract.pdf")
print(result.document.export_to_markdown())

Notes / scope

  • Richer provenance metadata suggested in the discussion (heading_source, bookmark_path, match confidence) would need a docling-core schema change so i would happy to open a separate issue for it so it doesn't block this pipeline change.
  • Known limitation: a promoted heading keeps its original tree position, so it may sit inside its former list group. This is fine for hierarchy/chunking; can be refined in a follow-up.

Thanks for the guidance on this Dr @PeterStaar-IBM . I'm happy to adjust anything!

Checklist:

  • Documentation has been updated, if necessary. (option docstrings)
  • Examples have been added, if necessary.
  • Tests have been added, if necessary. (tests/test_heading_hierarchy_bookmarks.py)

Follow-up to docling-project#3633, which inferred PDF heading levels from numbering and font style but left out the document's own bookmarks/table-of-contents - usually the most reliable hierarchy signal. This wires bookmarks in as the authoritative source, with precedence bookmarks > numbering > style.

- Extract the outline via a shared pypdfium2 helper on PdfDocumentBackend (active under both the pypdfium2 and the default docling-parse backends); surfaced on ConversionResult.pdf_outline.

- Match bookmarks to detected headings by fuzzy title (difflib) + page, comparing with and without leading numbering markers; a confident match is authoritative, and a confidently matched list-item is promoted to a heading (layout models often mis-classify headings as list-items).

- Partial/noisy outlines never degrade the numbering result: unmatched entries fall back to numbering/style. New options use_bookmarks (default true) and bookmark_match_threshold; no chunker changes needed.

Signed-off-by: Daniel Nguyen <danielnguyenh07@gmail.com>
@github-actions

Copy link
Copy Markdown
Contributor

✅ DCO Check Passed

Thanks @DanielNg0729, all your commits are properly signed off. 🎉

@mergify

mergify Bot commented Jun 24, 2026 •

Copy link
Copy Markdown
Contributor

Merge Protections

🟢 All 2 merge protections satisfied — ready to merge.

Show 2 satisfied protections

🟢 Enforce conventional commit

Make sure that we follow https://www.conventionalcommits.org/en/v1.0.0/

  • title ~= ^(fix|feat|docs|style|refactor|perf|test|build|ci|chore|revert)(?:\(.+\))?(!)?:

🟢 Require two reviewer for test updates

When test data is updated, we require two reviewers

  • #approved-reviews-by >= 2

@codecov

codecov Bot commented Jun 24, 2026 •

Copy link
Copy Markdown

@PeterStaar-IBM

Copy link
Copy Markdown
Member

@DanielNg0729 Love the work: I would suggest that we add at least 1 example pdf with an existing outline in the tests.

Signed-off-by: Daniel Nguyen <danielnguyenh07@gmail.com>
@DanielNg0729

Copy link
Copy Markdown
Member Author

Thanks, Dr. @PeterStaar-IBM! I've added a committed sample PDF (tests/data/pdf/bookmark_sample.pdf) with a real, nested outline (PART → section → subsection across 3 pages), plus a test (test_backend_extracts_nested_outline_from_sample_pdf) that runs PdfDocumentBackend.get_document_outline() against it and asserts the full bookmark tree (title/level/page) and vertical positions.

@cau-git

cau-git commented Jun 25, 2026 •

Copy link
Copy Markdown
Member

@DanielNg0729 great to see this continuation PR. I have not reviewed deeply yet, however one feedback already below:

  • You are using pypdfium to retrieve the Table of Contents, for both the PyPdfiumDocumentBackend and the DoclingParseDocumentBackend (since internally it also owns a pypdfium handle). We would like the DoclingParseDocumentBackend to handle this with the native way to get the table-of-contents (no pypdfium dependency). You can find the hints here. Could you please check if that gives you the matching information needed? Thanks.
  • If the above is addressed, then we should be able to use it also with the (non-default) ThreadedDoclingParseDocumentBackend, which has no pypdfium handle.

…contents

Per review: DoclingParseDocumentBackend (+v2/v4) and ThreadedDoclingParseDocumentBackend now read the outline via docling-parse's native get_table_of_contents() instead of the internal pypdfium handle, so the signal works without a pypdfium dependency (incl. the threaded backend, which has none). The pypdfium2 backend keeps its richer extraction (title + page + position). The native ToC carries title + hierarchy only, so page_no/y_top are left None and the matcher falls back to title-only matching for those entries.

Signed-off-by: Daniel Nguyen <danielnguyenh07@gmail.com>
Signed-off-by: Nguyen Hoang Duong <danielnguyenh07@gmail.com>
@DanielNg0729

Copy link
Copy Markdown
Member Author

Thank you very much for the feedback Dr @cau-git

I've implemented it as requested: DoclingParseDocumentBackend (and the v2/v4 + threaded variants) now read the outline via docling-parse's native get_table_of_contents(), with no pypdfium2 dependency. I also added a sample PDF + tests covering both extractors. Thanks again!

get_table_of_contents() returns None for PDFs with no embedded outline, so outline_from_docling_parse crashed on None.children and failed heading-hierarchy conversion (use_bookmarks defaults on). Return [] in that case and fall back to numbering/style. Adds a regression test.

Signed-off-by: Daniel Nguyen <danielnguyenh07@gmail.com>
@DanielNg0729

Copy link
Copy Markdown
Member Author

Hi Dr. @cau-git, Dr. @PeterStaar-IBM

Quick update on the failing run-tests-ml (pdf-model) lane: the root cause was that docling-parse's get_table_of_contents() returns None for PDFs without an embedded outline, so the extractor crashed on None.children during heading-hierarchy conversion. I've fixed it to return an empty outline in that case (falling back to numbering/style as before) and added a regression test. Lint, type checks, and the core tests are green locally.

Could you kindly re-trigger the CI when you have a moment? Thank you very much!

@PeterStaar-IBM

Copy link
Copy Markdown
Member

Hi Dr. @cau-git, Dr. @PeterStaar-IBM

Quick update on the failing run-tests-ml (pdf-model) lane: the root cause was that docling-parse's get_table_of_contents() returns None for PDFs without an embedded outline, so the extractor crashed on None.children during heading-hierarchy conversion. I've fixed it to return an empty outline in that case (falling back to numbering/style as before) and added a regression test. Lint, type checks, and the core tests are green locally.

Could you kindly re-trigger the CI when you have a moment? Thank you very much!

of course!

Comment thread docling/datamodel/base_models.py Outdated
…dule

Per review: PdfOutlineItem is an internal backend->heading-stage data structure (already excluded from serialization), not public datamodel. Move it out of base_models into docling/utils/pdf_outline.py, next to the extractors that produce it. Kept as our own type rather than reusing docling-parse's PdfTocEntry, which has no vertical-position (y_top) field and would couple the pypdfium2 backend to docling-parse.

Signed-off-by: Daniel Nguyen <danielnguyenh07@gmail.com>
@PeterStaar-IBM

Copy link
Copy Markdown
Member

@cau-git Let's review this one internally!

@cau-git cau-git left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@DanielNg0729 Thanks for the updates! I went through the code and posted a few remarks and questions below.

Comment thread docling/utils/pdf_outline.py Outdated
Comment thread docling/models/stages/heading_hierarchy/heading_hierarchy_model.py
Comment thread docling/utils/pdf_outline.py Outdated
Comment thread docling/datamodel/document.py Outdated
Comment thread docling/utils/pdf_outline.py
…ternal

Signed-off-by: Daniel Nguyen <danielnguyenh07@gmail.com>
…rom_docling_parse for naming consistency

Signed-off-by: Daniel Nguyen <danielnguyenh07@gmail.com>
…line), reset after the heading stage consumes it

Signed-off-by: Daniel Nguyen <danielnguyenh07@gmail.com>
@DanielNg0729
DanielNg0729 requested a review from cau-git July 1, 2026 07:13

@PeterStaar-IBM PeterStaar-IBM left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

lgtm!

@cau-git
cau-git merged commit 9cbef42 into docling-project:main Jul 1, 2026
26 checks passed
@hadrien

hadrien commented Jul 3, 2026

Copy link
Copy Markdown

Thanks @DanielNg0729 and reviewing team.

@PierreMesure

Copy link
Copy Markdown

This is amazing work @DanielNg0729! Looking forward to trying it on our documents after the Summer! 😊

DanielNg0729 added a commit to DanielNg0729/docling that referenced this pull request Aug 21, 2026
Add a usage guide and a runnable example for the section-header level
inference shipped in docling-project#3633, docling-project#3688 and docling-project#3984, covering the bookmark, numbering
and font-style signals, their precedence, and the generate_parsed_pages
requirement of the style fallback.

Signed-off-by: Daniel Nguyen <danielnguyenh07@gmail.com>
@DanielNg0729 DanielNg0729 mentioned this pull request Aug 21, 2026
2 of 3 tasks
dolfim-ibm pushed a commit that referenced this pull request Aug 24, 2026
* docs: document PDF heading-level inference

Add a usage guide and a runnable example for the section-header level
inference shipped in #3633, #3688 and #3984, covering the bookmark, numbering
and font-style signals, their precedence, and the generate_parsed_pages
requirement of the style fallback.

Signed-off-by: Daniel Nguyen <danielnguyenh07@gmail.com>

* Remove PDF heading recovery from features list

Removed PDF heading recovery feature from the list.

Signed-off-by: Nguyen Hoang Duong <danielnguyenh07@gmail.com>

---------

Signed-off-by: Daniel Nguyen <danielnguyenh07@gmail.com>
Signed-off-by: Nguyen Hoang Duong <danielnguyenh07@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Discussion: PDF bookmarks/ToC as an authoritative heading-hierarchy signal for hierarchy-aware chunking (follow-up #3633)

5 participants