feat(pdf): infer PDF heading levels so the hierarchy isn't flattened - #3633
PeterStaar-IBM merged 8 commits into
Conversation
…nd style The PDF/image pipeline previously emitted every detected SECTION_HEADER at level 1, flattening document hierarchy: Roman-numeral parts and Arabic-numeral subsections collapsed to the same Markdown heading depth. Add an opt-in heading-level inference step to the shared reading-order stage that assigns SectionHeaderItem.level from: - numbering (primary): legal/outline schemes such as PART I -> 1. -> 1.1 -> (a) -> (i), with Roman/Arabic/alpha disambiguation and relative, compressed levels so a document that starts at "1." is not forced to start at depth 2; - style (fallback): font size approximated from parsed PDF cell heights, used only for headings without recognizable numbering (requires generate_parsed_pages=True). Gated behind PdfPipelineOptions.heading_hierarchy_options and disabled by default, so existing conversions and reference outputs are unchanged. PDF bookmark/outline inference is left as a documented follow-up extension point. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Signed-off-by: Daniel Nguyen <danielnguyenh07@gmail.com>
assign_heading_levels and the per-signal helpers accept conv_res=None so the numbering-only path (which needs no ConversionResult) is type-correct; style inference returns early when conv_res is absent. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Signed-off-by: Daniel Nguyen <danielnguyenh07@gmail.com>
|
✅ DCO Check Passed Thanks @DanielNg0729, all your commits are properly signed off. 🎉 |
Merge ProtectionsYour pull request matches the following merge protections and will not be merged until they are valid. 🟢 Enforce conventional commitWonderful, this rule succeeded.Make sure that we follow https://www.conventionalcommits.org/en/v1.0.0/
|
|
@DanielNg0729 thank you so much for this PR, really wonderful initiative! |
Codecov Report❌ Patch coverage is
📢 Thoughts on this report? Let us know! |
|
Hi Dr @PeterStaar-IBM. This is ready for review whenever you or the team have a moment. I left a few open questions at the bottom of the description (default-off vs. on, the post-assembly pass vs. setting the level at the add_heading call sites, and whether style should auto-keep parsed_page), and I'm happy to adjust any of it. |
|
@cau-git Let's discuss this later and give @DanielNg0729 some feedback. |
|
@DanielNg0729 Thanks for this proposal. We definitely appreciate seeing this topic addressed. Regarding your code and questions:
Proposal to move / refactor the code
|
Addresses review feedback on docling-project#3633: keep the heading-level inference out of the already-impure reading-order model. - Move the logic into its own stage model (docling/models/stages/heading_hierarchy/heading_hierarchy_model.py). HeadingHierarchyModel accepts a ConversionResult and returns the (in-place modified) DoclingDocument. - Invoke it right after the reading-order model in _assemble_document of both StandardPdfPipeline and LegacyStandardPdfPipeline. - Un-nest the options: ReadingOrderOptions no longer carries heading_hierarchy; HeadingHierarchyOptions stays standalone on PdfPipelineOptions. - Factor the core so it is reusable outside the pipeline: assign_heading_levels works on a bare DoclingDocument, with the font-based fallback taking the parsed pages explicitly (page height read from SegmentedPdfPage.dimension). - Add a style-fallback test (font size -> level) to cover that path. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Signed-off-by: Daniel Nguyen <danielnguyenh07@gmail.com>
… follow-up So nothing dead ships in this PR: remove the `use_bookmarks` option and the no-op `_infer_from_bookmarks` stub. Bookmark / PDF-outline inference (which needs new backend plumbing) will be introduced in a separate follow-up PR. Signed-off-by: Daniel Nguyen <danielnguyenh07@gmail.com>
|
Thanks for the detailed review, Dr @cau-git, Dr @PeterStaar-IBM. I've pushed a refactor addressing all the points:
For the bookmarks/ToC signal, before it's only a reserved no-op hook (
Happy to get started on it once this one lands. Thank you very much! |
The new `docling.models.stages.heading_hierarchy` package was not declared in tach.toml, so it inherited the generic `docling.models.stages` boundary, which forbids importing `docling.datamodel`. That failed `tach check` -- and the lint job too, since `prek run --all-files` runs the same tach hook (fail_fast). Register it as its own module (like `reading_order`) with a `docling.datamodel` dependency, and add it to the `depends_on` of the standard and legacy PDF pipelines that import it. Signed-off-by: Daniel Nguyen <danielnguyenh07@gmail.com>
|
Hi Dr @cau-git. I've pushed a fix for the failing checks (the new |
Signed-off-by: Christoph Auer <cau@zurich.ibm.com>
Signed-off-by: Christoph Auer <cau@zurich.ibm.com>
|
@DanielNg0729 Thanks for the updates. I added a small test method that uses an actual PDF from our test data and checks assertions about the levels. Seems to work fine! |
Signed-off-by: Christoph Auer <cau@zurich.ibm.com>
Follow-up to docling-project#3633, which inferred PDF heading levels from numbering and font style but left out the document's own bookmarks/table-of-contents - usually the most reliable hierarchy signal. This wires bookmarks in as the authoritative source, with precedence bookmarks > numbering > style. - Extract the outline via a shared pypdfium2 helper on PdfDocumentBackend (active under both the pypdfium2 and the default docling-parse backends); surfaced on ConversionResult.pdf_outline. - Match bookmarks to detected headings by fuzzy title (difflib) + page, comparing with and without leading numbering markers; a confident match is authoritative, and a confidently matched list-item is promoted to a heading (layout models often mis-classify headings as list-items). - Partial/noisy outlines never degrade the numbering result: unmatched entries fall back to numbering/style. New options use_bookmarks (default true) and bookmark_match_threshold; no chunker changes needed. Signed-off-by: Daniel Nguyen <danielnguyenh07@gmail.com>
* feat(pdf): infer heading levels from PDF bookmarks/ToC Follow-up to #3633, which inferred PDF heading levels from numbering and font style but left out the document's own bookmarks/table-of-contents - usually the most reliable hierarchy signal. This wires bookmarks in as the authoritative source, with precedence bookmarks > numbering > style. - Extract the outline via a shared pypdfium2 helper on PdfDocumentBackend (active under both the pypdfium2 and the default docling-parse backends); surfaced on ConversionResult.pdf_outline. - Match bookmarks to detected headings by fuzzy title (difflib) + page, comparing with and without leading numbering markers; a confident match is authoritative, and a confidently matched list-item is promoted to a heading (layout models often mis-classify headings as list-items). - Partial/noisy outlines never degrade the numbering result: unmatched entries fall back to numbering/style. New options use_bookmarks (default true) and bookmark_match_threshold; no chunker changes needed. Signed-off-by: Daniel Nguyen <danielnguyenh07@gmail.com> * test(pdf): add sample PDF with nested bookmarks for outline extraction Signed-off-by: Daniel Nguyen <danielnguyenh07@gmail.com> * refactor(pdf): extract docling-parse outline via native get_table_of_contents Per review: DoclingParseDocumentBackend (+v2/v4) and ThreadedDoclingParseDocumentBackend now read the outline via docling-parse's native get_table_of_contents() instead of the internal pypdfium handle, so the signal works without a pypdfium dependency (incl. the threaded backend, which has none). The pypdfium2 backend keeps its richer extraction (title + page + position). The native ToC carries title + hierarchy only, so page_no/y_top are left None and the matcher falls back to title-only matching for those entries. Signed-off-by: Daniel Nguyen <danielnguyenh07@gmail.com> * fix(pdf): handle PDFs without an outline in the docling-parse extractor get_table_of_contents() returns None for PDFs with no embedded outline, so outline_from_docling_parse crashed on None.children and failed heading-hierarchy conversion (use_bookmarks defaults on). Return [] in that case and fall back to numbering/style. Adds a regression test. Signed-off-by: Daniel Nguyen <danielnguyenh07@gmail.com> * refactor(pdf): make PdfOutlineItem an internal type in the outline module Per review: PdfOutlineItem is an internal backend->heading-stage data structure (already excluded from serialization), not public datamodel. Move it out of base_models into docling/utils/pdf_outline.py, next to the extractors that produce it. Kept as our own type rather than reusing docling-parse's PdfTocEntry, which has no vertical-position (y_top) field and would couple the pypdfium2 backend to docling-parse. Signed-off-by: Daniel Nguyen <danielnguyenh07@gmail.com> * refactor(pdf): rename PdfOutlineItem to _PdfOutlineItem to mark it internal Signed-off-by: Daniel Nguyen <danielnguyenh07@gmail.com> * refactor(pdf): rename outline_from_docling_parse to extract_outline_from_docling_parse for naming consistency Signed-off-by: Daniel Nguyen <danielnguyenh07@gmail.com> * refactor(pdf): make ConversionResult outline a private attr (_pdf_outline), reset after the heading stage consumes it Signed-off-by: Daniel Nguyen <danielnguyenh07@gmail.com> --------- Signed-off-by: Daniel Nguyen <danielnguyenh07@gmail.com> Signed-off-by: Nguyen Hoang Duong <danielnguyenh07@gmail.com>
…ocling-project#3633) * feat(reading_order): infer PDF section-header levels from numbering and style The PDF/image pipeline previously emitted every detected SECTION_HEADER at level 1, flattening document hierarchy: Roman-numeral parts and Arabic-numeral subsections collapsed to the same Markdown heading depth. Add an opt-in heading-level inference step to the shared reading-order stage that assigns SectionHeaderItem.level from: - numbering (primary): legal/outline schemes such as PART I -> 1. -> 1.1 -> (a) -> (i), with Roman/Arabic/alpha disambiguation and relative, compressed levels so a document that starts at "1." is not forced to start at depth 2; - style (fallback): font size approximated from parsed PDF cell heights, used only for headings without recognizable numbering (requires generate_parsed_pages=True). Gated behind PdfPipelineOptions.heading_hierarchy_options and disabled by default, so existing conversions and reference outputs are unchanged. PDF bookmark/outline inference is left as a documented follow-up extension point. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Signed-off-by: Daniel Nguyen <danielnguyenh07@gmail.com> * fix(reading_order): make conv_res optional in heading-level inference assign_heading_levels and the per-signal helpers accept conv_res=None so the numbering-only path (which needs no ConversionResult) is type-correct; style inference returns early when conv_res is absent. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Signed-off-by: Daniel Nguyen <danielnguyenh07@gmail.com> * refactor(heading-hierarchy): extract a standalone HeadingHierarchyModel Addresses review feedback on docling-project#3633: keep the heading-level inference out of the already-impure reading-order model. - Move the logic into its own stage model (docling/models/stages/heading_hierarchy/heading_hierarchy_model.py). HeadingHierarchyModel accepts a ConversionResult and returns the (in-place modified) DoclingDocument. - Invoke it right after the reading-order model in _assemble_document of both StandardPdfPipeline and LegacyStandardPdfPipeline. - Un-nest the options: ReadingOrderOptions no longer carries heading_hierarchy; HeadingHierarchyOptions stays standalone on PdfPipelineOptions. - Factor the core so it is reusable outside the pipeline: assign_heading_levels works on a bare DoclingDocument, with the font-based fallback taking the parsed pages explicitly (page height read from SegmentedPdfPage.dimension). - Add a style-fallback test (font size -> level) to cover that path. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Signed-off-by: Daniel Nguyen <danielnguyenh07@gmail.com> * refactor(heading-hierarchy): drop the bookmarks placeholder until the follow-up So nothing dead ships in this PR: remove the `use_bookmarks` option and the no-op `_infer_from_bookmarks` stub. Bookmark / PDF-outline inference (which needs new backend plumbing) will be introduced in a separate follow-up PR. Signed-off-by: Daniel Nguyen <danielnguyenh07@gmail.com> * fix(tach): register heading_hierarchy as a tach module The new `docling.models.stages.heading_hierarchy` package was not declared in tach.toml, so it inherited the generic `docling.models.stages` boundary, which forbids importing `docling.datamodel`. That failed `tach check` -- and the lint job too, since `prek run --all-files` runs the same tach hook (fail_fast). Register it as its own module (like `reading_order`) with a `docling.datamodel` dependency, and add it to the `depends_on` of the standard and legacy PDF pipelines that import it. Signed-off-by: Daniel Nguyen <danielnguyenh07@gmail.com> * Add test method using a real PDF Signed-off-by: Christoph Auer <cau@zurich.ibm.com> * Run pre-commit toolchain Signed-off-by: Christoph Auer <cau@zurich.ibm.com> * Separate tests to allow full-unit CI markers Signed-off-by: Christoph Auer <cau@zurich.ibm.com> --------- Signed-off-by: Daniel Nguyen <danielnguyenh07@gmail.com> Signed-off-by: Christoph Auer <cau@zurich.ibm.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> Co-authored-by: Christoph Auer <cau@zurich.ibm.com>
Add a usage guide and a runnable example for the section-header level inference shipped in docling-project#3633, docling-project#3688 and docling-project#3984, covering the bookmark, numbering and font-style signals, their precedence, and the generate_parsed_pages requirement of the style fallback. Signed-off-by: Daniel Nguyen <danielnguyenh07@gmail.com>
* docs: document PDF heading-level inference Add a usage guide and a runnable example for the section-header level inference shipped in #3633, #3688 and #3984, covering the bookmark, numbering and font-style signals, their precedence, and the generate_parsed_pages requirement of the style fallback. Signed-off-by: Daniel Nguyen <danielnguyenh07@gmail.com> * Remove PDF heading recovery from features list Removed PDF heading recovery feature from the list. Signed-off-by: Nguyen Hoang Duong <danielnguyenh07@gmail.com> --------- Signed-off-by: Daniel Nguyen <danielnguyenh07@gmail.com> Signed-off-by: Nguyen Hoang Duong <danielnguyenh07@gmail.com>
Infer PDF heading levels so the hierarchy isn't flattened
The problem
I hit this while converting legal contracts to Markdown for an LLM pipeline: every heading Docling pulls out of a PDF comes back as a level-1
#, so the whole document goes flat. Something structured likeends up as a flat run of
#headings, and the parent/child relationship between the parts, the numbered sections and the sub-clauses is just gone. Anything downstream that leans on heading depth to understand the document (RAG, summarisation, navigation) loses that structure.This is the issue reported in #3555, and the same thing has come up before in #386, #1170 and #2774. There were also two earlier attempts to infer levels from font information (#2676, #2421) that never landed. The standalone post-processor
krrome/docling-hierarchical-pdf(MIT) already does this after conversion — this PR brings the same idea into the pipeline itself.Why it happens
It isn't really a detection bug — Docling finds the headings fine, it just never gives them a level. In
readingorder_model.py, both places that create a heading calladd_heading()with nolevel=, so it falls back to1:DoclingDocument.add_heading()already accepts alevel(1–100) and the Markdown/HTML exporters already render it. The DOCX, LaTeX and HTML backends pass it because they have real heading markup to read; the PDF path has none, so nothing ever computes a level. So this is closer to a missing feature than a regression.What I changed
I added a small, self-contained step (
docling/models/stages/reading_order/heading_hierarchy.py) that runs fromReadingOrderModel.__call__once the document is assembled. It walks the headings in reading order and assigns each one a level. It only rewrites levels — it never adds, removes or reorders anything.Doing this inside the pipeline (instead of as a post-processor on the finished document) matters: the font data lives on
conv_res.pages[i].parsed_pageand never makes it into the finalDoclingDocument, so only an in-pipeline step can actually use it. It also fixes every PDF/image backend at once, since they all share this stage.There are two signals, with numbering as the primary one.
Numbering
A small grammar reads each heading's leading marker and maps it to a scheme:
PART I,TITLE IIArticle 1,Section 2,§ 1.2I.,II.,III.1.,2.1.1,1.1.1(a),A.(i),(ii)Levels are worked out relatively from the schemes that actually appear, so a document that starts at
1.isn't forced to start at depth 2, and1.1/1.1.1nest by how many segments they have. I went with numbering as the main signal on purpose — in legal and regulatory documents the numbering is far more reliable than the styling, which is often completely uniform.Single letters that are also valid Roman numerals (
I,V,C, …) are genuinely ambiguous, so they're resolved from the surrounding markers:I. / II. / III.reads as Roman, whileA. / B. / C.reads as alphabetic. There's also anumbering_schemesoption to override the ordering for house styles.Style (fallback)
For headings that have no recognisable number, it falls back to font size: it matches the heading's box against the parsed PDF cells and uses their median height as a size proxy — bigger text becomes a higher level. There's no explicit font-size field on the cells, so height is the best proxy available.
One caveat worth knowing: the parsed pages get dropped right before the reading-order stage unless you keep them, so the style fallback needs
generate_parsed_pages=True. Without it, style just no-ops and numbering still works.How to use it
It's opt-in and off by default, so nothing changes for anyone unless they turn it on:
The new options:
Tests
tests/test_heading_hierarchy.pycovers the headline case (Roman sections staying above Arabic subsections), the full legal stack (PART → 1. → 1.1 → (a) → (i)), relative/compressed levels, dotted-decimal depth, custom scheme order,max_levelclamping, the Roman-vs-alpha disambiguation, the negative cases (plain words and bare numbers shouldn't be treated as markers), and the end-to-end effect onexport_to_markdown.No reference data changes, since the feature is off by default and existing fixtures stay byte-identical.
What I deliberately left out
_infer_from_bookmarks()is already there as a no-op extension point, so it can be added later without disturbing this structure.Honest limitations
(i)with no sibling(ii)next to an(a)list can be misread — it needs a real sequence to anchor the Roman-vs-alpha call.A couple of things I'd like your take on
add_headingcall sites — happy to move it if you'd prefer.parsed_pagewhen it's enabled, instead of asking the user to also setgenerate_parsed_pages=True?krrome/docling-hierarchical-pdfapproach (with attribution) fine?Closes #3555.