Skip to content

feat(pdf): infer PDF heading levels so the hierarchy isn't flattened - #3633

Merged
PeterStaar-IBM merged 8 commits into
docling-project:mainfrom
DanielNg0729:feat/pdf-heading-level-inference
Jun 23, 2026
Merged

PeterStaar-IBM merged 8 commits into
docling-project:mainfrom
DanielNg0729:feat/pdf-heading-level-inference

Conversation

@DanielNg0729

Copy link
Copy Markdown
Member

Infer PDF heading levels so the hierarchy isn't flattened

The problem

I hit this while converting legal contracts to Markdown for an LLM pipeline: every heading Docling pulls out of a PDF comes back as a level-1 #, so the whole document goes flat. Something structured like

PART I
1. Definitions
1.1 Interpretation
(a) ...

ends up as a flat run of # headings, and the parent/child relationship between the parts, the numbered sections and the sub-clauses is just gone. Anything downstream that leans on heading depth to understand the document (RAG, summarisation, navigation) loses that structure.

This is the issue reported in #3555, and the same thing has come up before in #386, #1170 and #2774. There were also two earlier attempts to infer levels from font information (#2676, #2421) that never landed. The standalone post-processor krrome/docling-hierarchical-pdf (MIT) already does this after conversion — this PR brings the same idea into the pipeline itself.

Why it happens

It isn't really a detection bug — Docling finds the headings fine, it just never gives them a level. In readingorder_model.py, both places that create a heading call add_heading() with no level=, so it falls back to 1:

# _handle_text_element (the SECTION_HEADER branch)
new_item = out_doc.add_heading(
    text=cap_text, prov=prov, hyperlink=element.hyperlink
)  # no level= -> defaults to 1

# _add_child_elements
doc.add_heading(parent=doc_item, text=c_text, prov=c_prov)  # no level=

DoclingDocument.add_heading() already accepts a level (1–100) and the Markdown/HTML exporters already render it. The DOCX, LaTeX and HTML backends pass it because they have real heading markup to read; the PDF path has none, so nothing ever computes a level. So this is closer to a missing feature than a regression.

What I changed

I added a small, self-contained step (docling/models/stages/reading_order/heading_hierarchy.py) that runs from ReadingOrderModel.__call__ once the document is assembled. It walks the headings in reading order and assigns each one a level. It only rewrites levels — it never adds, removes or reorders anything.

Doing this inside the pipeline (instead of as a post-processor on the finished document) matters: the font data lives on conv_res.pages[i].parsed_page and never makes it into the final DoclingDocument, so only an in-pipeline step can actually use it. It also fixes every PDF/image backend at once, since they all share this stage.

There are two signals, with numbering as the primary one.

Numbering

A small grammar reads each heading's leading marker and maps it to a scheme:

Example marker Scheme
PART I, TITLE II part
Article 1, Section 2, § 1.2 article
I., II., III. upper roman
1., 2. arabic
1.1, 1.1.1 dotted decimal (depth from segment count)
(a), A. alpha
(i), (ii) lower roman

Levels are worked out relatively from the schemes that actually appear, so a document that starts at 1. isn't forced to start at depth 2, and 1.1 / 1.1.1 nest by how many segments they have. I went with numbering as the main signal on purpose — in legal and regulatory documents the numbering is far more reliable than the styling, which is often completely uniform.

Single letters that are also valid Roman numerals (I, V, C, …) are genuinely ambiguous, so they're resolved from the surrounding markers: I. / II. / III. reads as Roman, while A. / B. / C. reads as alphabetic. There's also a numbering_schemes option to override the ordering for house styles.

Style (fallback)

For headings that have no recognisable number, it falls back to font size: it matches the heading's box against the parsed PDF cells and uses their median height as a size proxy — bigger text becomes a higher level. There's no explicit font-size field on the cells, so height is the best proxy available.

One caveat worth knowing: the parsed pages get dropped right before the reading-order stage unless you keep them, so the style fallback needs generate_parsed_pages=True. Without it, style just no-ops and numbering still works.

How to use it

It's opt-in and off by default, so nothing changes for anyone unless they turn it on:

from docling.datamodel.base_models import InputFormat
from docling.datamodel.pipeline_options import (
    HeadingHierarchyOptions,
    PdfPipelineOptions,
)
from docling.document_converter import DocumentConverter, PdfFormatOption

opts = PdfPipelineOptions()
opts.heading_hierarchy_options = HeadingHierarchyOptions(enabled=True)
opts.generate_parsed_pages = True  # only needed for the style fallback

converter = DocumentConverter(
    format_options={InputFormat.PDF: PdfFormatOption(pipeline_options=opts)}
)
doc = converter.convert("contract.pdf").document
print(doc.export_to_markdown())

The new options:

enabled: bool = False          # master switch
use_numbering: bool = True     # primary signal
use_style: bool = True         # fallback; needs generate_parsed_pages=True
use_bookmarks: bool = False    # reserved for the follow-up below
numbering_schemes: list[str] | None = None
max_level: int = 6

Tests

tests/test_heading_hierarchy.py covers the headline case (Roman sections staying above Arabic subsections), the full legal stack (PART → 1. → 1.1 → (a) → (i)), relative/compressed levels, dotted-decimal depth, custom scheme order, max_level clamping, the Roman-vs-alpha disambiguation, the negative cases (plain words and bare numbers shouldn't be treated as markers), and the end-to-end effect on export_to_markdown.

No reference data changes, since the feature is off by default and existing fixtures stay byte-identical.

What I deliberately left out

  • Bookmarks / PDF outline. This is usually the most authoritative signal, but it needs new plumbing in the PDF backends (the outline isn't surfaced anywhere today), so I kept it out to keep this PR focused. _infer_from_bookmarks() is already there as a no-op extension point, so it can be added later without disturbing this structure.
  • Bold/italic from font names. A natural extension to the style signal, but not needed for the legal use case and easy to bolt on afterwards.

Honest limitations

  • The style fallback relies on geometric matching between the heading box and the parsed cells, which has rough edges (multi-line headings, rotated text, OCR cells with no font), so numbering carries the weight.
  • On scanned / VLM-parsed PDFs the font data is often missing, so it degrades to numbering-only.
  • A lone (i) with no sibling (ii) next to an (a) list can be misread — it needs a real sequence to anchor the Roman-vs-alpha call.

A couple of things I'd like your take on

  • Off by default (what I went with) vs. on with regenerated references?
  • I assign the level in a post-assembly pass rather than at the add_heading call sites — happy to move it if you'd prefer.
  • Should style auto-keep parsed_page when it's enabled, instead of asking the user to also set generate_parsed_pages=True?
  • Is building on the krrome/docling-hierarchical-pdf approach (with attribution) fine?

Closes #3555.

DanielNg0729 and others added 2 commits June 17, 2026 08:52
…nd style

The PDF/image pipeline previously emitted every detected SECTION_HEADER at level
1, flattening document hierarchy: Roman-numeral parts and Arabic-numeral
subsections collapsed to the same Markdown heading depth.

Add an opt-in heading-level inference step to the shared reading-order stage that
assigns SectionHeaderItem.level from:

- numbering (primary): legal/outline schemes such as PART I -> 1. -> 1.1 ->
  (a) -> (i), with Roman/Arabic/alpha disambiguation and relative, compressed
  levels so a document that starts at "1." is not forced to start at depth 2;
- style (fallback): font size approximated from parsed PDF cell heights, used
  only for headings without recognizable numbering (requires
  generate_parsed_pages=True).

Gated behind PdfPipelineOptions.heading_hierarchy_options and disabled by
default, so existing conversions and reference outputs are unchanged. PDF
bookmark/outline inference is left as a documented follow-up extension point.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Signed-off-by: Daniel Nguyen <danielnguyenh07@gmail.com>
assign_heading_levels and the per-signal helpers accept conv_res=None so the
numbering-only path (which needs no ConversionResult) is type-correct; style
inference returns early when conv_res is absent.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Signed-off-by: Daniel Nguyen <danielnguyenh07@gmail.com>
@github-actions

github-actions Bot commented Jun 17, 2026 •

Copy link
Copy Markdown
Contributor

✅ DCO Check Passed

Thanks @DanielNg0729, all your commits are properly signed off. 🎉

@mergify

mergify Bot commented Jun 17, 2026 •

Copy link
Copy Markdown
Contributor

Merge Protections

Your pull request matches the following merge protections and will not be merged until they are valid.

🟢 Enforce conventional commit

Wonderful, this rule succeeded.

Make sure that we follow https://www.conventionalcommits.org/en/v1.0.0/

  • title ~= ^(fix|feat|docs|style|refactor|perf|test|build|ci|chore|revert)(?:\(.+\))?(!)?:

@DanielNg0729 DanielNg0729 changed the title Infer PDF heading levels so the hierarchy isn't flattened feat(reading_order): infer PDF heading levels so the hierarchy isn't flattened Jun 17, 2026
@PeterStaar-IBM

Copy link
Copy Markdown
Member

@DanielNg0729 thank you so much for this PR, really wonderful initiative!

@codecov

codecov Bot commented Jun 17, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 92.07317% with 13 lines in your changes missing coverage. Please review.

Files with missing lines Patch % Lines
...tages/heading_hierarchy/heading_hierarchy_model.py 91.39% 13 Missing ⚠️

📢 Thoughts on this report? Let us know!

@DanielNg0729

Copy link
Copy Markdown
Member Author

Hi Dr @PeterStaar-IBM. This is ready for review whenever you or the team have a moment. I left a few open questions at the bottom of the description (default-off vs. on, the post-assembly pass vs. setting the level at the add_heading call sites, and whether style should auto-keep parsed_page), and I'm happy to adjust any of it.

@PeterStaar-IBM

Copy link
Copy Markdown
Member

@cau-git Let's discuss this later and give @DanielNg0729 some feedback.

@cau-git

cau-git commented Jun 22, 2026 •

Copy link
Copy Markdown
Member

@DanielNg0729 Thanks for this proposal. We definitely appreciate seeing this topic addressed.

Regarding your code and questions:

  • The provided functionality and logic looks good!
  • Post-modifying an already constructed DoclingDocument is fine in terms of approach
  • Off-by-default is a sensible choice
  • I would like to suggest moving it outside of the reading-order model, since that model is already impure in terms of separation of concerns and we don't want to add more debt there. Find some detailed comments below.
  • Building on krrome/docling-hierarchical-pdf should be fine as it is also MIT licensed.
  • We would be very interested if you take this further in a downstream iteration with PDF bookmarks / ToC as a signal

Proposal to move / refactor the code

  • Ideally, make it an own HeadingHierarchy model which accepts ConversionResult and emits (modified) DoclingDocument object.
  • Place the call to it right after the reading-order model in these places:
  • The options class can work as proposed, analog to the ReadingOrderOptions, not nested inside.
  • Ideally, the code would be internally factored such that it can be re-used also outside the pipeline code (e.g. providing a way to also call it with a bare DoclingDocument, or DoclingDocument + list of SegmentedPage (second iteration).

@cau-git cau-git changed the title feat(reading_order): infer PDF heading levels so the hierarchy isn't flattened feat(pdf): infer PDF heading levels so the hierarchy isn't flattened Jun 22, 2026
DanielNg0729 and others added 2 commits June 22, 2026 21:31
Addresses review feedback on docling-project#3633: keep the heading-level inference out of the
already-impure reading-order model.

- Move the logic into its own stage model
  (docling/models/stages/heading_hierarchy/heading_hierarchy_model.py).
  HeadingHierarchyModel accepts a ConversionResult and returns the (in-place
  modified) DoclingDocument.
- Invoke it right after the reading-order model in _assemble_document of both
  StandardPdfPipeline and LegacyStandardPdfPipeline.
- Un-nest the options: ReadingOrderOptions no longer carries heading_hierarchy;
  HeadingHierarchyOptions stays standalone on PdfPipelineOptions.
- Factor the core so it is reusable outside the pipeline: assign_heading_levels
  works on a bare DoclingDocument, with the font-based fallback taking the parsed
  pages explicitly (page height read from SegmentedPdfPage.dimension).
- Add a style-fallback test (font size -> level) to cover that path.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Signed-off-by: Daniel Nguyen <danielnguyenh07@gmail.com>
… follow-up

So nothing dead ships in this PR: remove the `use_bookmarks` option and the
no-op `_infer_from_bookmarks` stub. Bookmark / PDF-outline inference (which needs
new backend plumbing) will be introduced in a separate follow-up PR.

Signed-off-by: Daniel Nguyen <danielnguyenh07@gmail.com>
@DanielNg0729

Copy link
Copy Markdown
Member Author

Thanks for the detailed review, Dr @cau-git, Dr @PeterStaar-IBM. I've pushed a refactor addressing all the points:

  • Own model: extracted a standalone HeadingHierarchyModel (docling/models/stages/heading_hierarchy/heading_hierarchy_model.py) that takes a ConversionResult and returns the modified DoclingDocument.
  • Out of reading-order: the call now happens right after the reading-order model in _assemble_document of both StandardPdfPipeline and LegacyStandardPdfPipeline. The reading-order model is fully decoupled again.
  • Options not nested: ReadingOrderOptions no longer carries it; HeadingHierarchyOptions stays standalone on PdfPipelineOptions.
  • Reusable outside the pipeline: assign_heading_levels(document, parsed_pages=None) works on a bare DoclingDocument; the style fallback takes the parsed pages explicitly (page height from SegmentedPdfPage.dimension), so it no longer depends on ConversionResult. This sets up the DoclingDocument + list[SegmentedPage] path you mentioned for the next iteration.
  • Also added a style-fallback test to cover the font-size path. Let me know if the placement/naming looks right!

For the bookmarks/ToC signal, before it's only a reserved no-op hook (_infer_from_bookmarks returns {}, and use_bookmarks defaults to False and does nothing), since surfacing the PDF outline needs new plumbing in the backends. So i drop it for now and I'd be happy to open a separate PR for it as the downstream iteration that extracting the outline from the PDF backends and feeding bookmark depth + page/title matching into the same model. One quick question i have:

  • Is there any preference on where outline extraction should live: a method on PdfDocumentBackend (e.g. get_document_outline()), implemented per backend (pypdfium2 / docling-parse)?

Happy to get started on it once this one lands. Thank you very much!

The new `docling.models.stages.heading_hierarchy` package was not declared in
tach.toml, so it inherited the generic `docling.models.stages` boundary, which
forbids importing `docling.datamodel`. That failed `tach check` -- and the lint
job too, since `prek run --all-files` runs the same tach hook (fail_fast).

Register it as its own module (like `reading_order`) with a `docling.datamodel`
dependency, and add it to the `depends_on` of the standard and legacy PDF
pipelines that import it.

Signed-off-by: Daniel Nguyen <danielnguyenh07@gmail.com>
@DanielNg0729

Copy link
Copy Markdown
Member Author

Hi Dr @cau-git. I've pushed a fix for the failing checks (the new heading_hierarchy package wasn't registered in tach.toml, which tripped the tach boundary check and, through it, the lint job). Could you kindly re-trigger the CI run when you have a moment? Thanks alot!

cau-git added 2 commits June 23, 2026 11:34
Signed-off-by: Christoph Auer <cau@zurich.ibm.com>
Signed-off-by: Christoph Auer <cau@zurich.ibm.com>
@cau-git

cau-git commented Jun 23, 2026

Copy link
Copy Markdown
Member

@DanielNg0729 Thanks for the updates. I added a small test method that uses an actual PDF from our test data and checks assertions about the levels. Seems to work fine!

Signed-off-by: Christoph Auer <cau@zurich.ibm.com>

@PeterStaar-IBM PeterStaar-IBM left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

wonderful!

@PeterStaar-IBM
PeterStaar-IBM merged commit 5237d7f into docling-project:main Jun 23, 2026
46 checks passed
DanielNg0729 added a commit to DanielNg0729/docling that referenced this pull request Jun 24, 2026
Follow-up to docling-project#3633, which inferred PDF heading levels from numbering and font style but left out the document's own bookmarks/table-of-contents - usually the most reliable hierarchy signal. This wires bookmarks in as the authoritative source, with precedence bookmarks > numbering > style.

- Extract the outline via a shared pypdfium2 helper on PdfDocumentBackend (active under both the pypdfium2 and the default docling-parse backends); surfaced on ConversionResult.pdf_outline.

- Match bookmarks to detected headings by fuzzy title (difflib) + page, comparing with and without leading numbering markers; a confident match is authoritative, and a confidently matched list-item is promoted to a heading (layout models often mis-classify headings as list-items).

- Partial/noisy outlines never degrade the numbering result: unmatched entries fall back to numbering/style. New options use_bookmarks (default true) and bookmark_match_threshold; no chunker changes needed.

Signed-off-by: Daniel Nguyen <danielnguyenh07@gmail.com>
cau-git pushed a commit that referenced this pull request Jul 1, 2026
* feat(pdf): infer heading levels from PDF bookmarks/ToC

Follow-up to #3633, which inferred PDF heading levels from numbering and font style but left out the document's own bookmarks/table-of-contents - usually the most reliable hierarchy signal. This wires bookmarks in as the authoritative source, with precedence bookmarks > numbering > style.

- Extract the outline via a shared pypdfium2 helper on PdfDocumentBackend (active under both the pypdfium2 and the default docling-parse backends); surfaced on ConversionResult.pdf_outline.

- Match bookmarks to detected headings by fuzzy title (difflib) + page, comparing with and without leading numbering markers; a confident match is authoritative, and a confidently matched list-item is promoted to a heading (layout models often mis-classify headings as list-items).

- Partial/noisy outlines never degrade the numbering result: unmatched entries fall back to numbering/style. New options use_bookmarks (default true) and bookmark_match_threshold; no chunker changes needed.

Signed-off-by: Daniel Nguyen <danielnguyenh07@gmail.com>

* test(pdf): add sample PDF with nested bookmarks for outline extraction

Signed-off-by: Daniel Nguyen <danielnguyenh07@gmail.com>

* refactor(pdf): extract docling-parse outline via native get_table_of_contents

Per review: DoclingParseDocumentBackend (+v2/v4) and ThreadedDoclingParseDocumentBackend now read the outline via docling-parse's native get_table_of_contents() instead of the internal pypdfium handle, so the signal works without a pypdfium dependency (incl. the threaded backend, which has none). The pypdfium2 backend keeps its richer extraction (title + page + position). The native ToC carries title + hierarchy only, so page_no/y_top are left None and the matcher falls back to title-only matching for those entries.

Signed-off-by: Daniel Nguyen <danielnguyenh07@gmail.com>

* fix(pdf): handle PDFs without an outline in the docling-parse extractor

get_table_of_contents() returns None for PDFs with no embedded outline, so outline_from_docling_parse crashed on None.children and failed heading-hierarchy conversion (use_bookmarks defaults on). Return [] in that case and fall back to numbering/style. Adds a regression test.

Signed-off-by: Daniel Nguyen <danielnguyenh07@gmail.com>

* refactor(pdf): make PdfOutlineItem an internal type in the outline module

Per review: PdfOutlineItem is an internal backend->heading-stage data structure (already excluded from serialization), not public datamodel. Move it out of base_models into docling/utils/pdf_outline.py, next to the extractors that produce it. Kept as our own type rather than reusing docling-parse's PdfTocEntry, which has no vertical-position (y_top) field and would couple the pypdfium2 backend to docling-parse.

Signed-off-by: Daniel Nguyen <danielnguyenh07@gmail.com>

* refactor(pdf): rename PdfOutlineItem to _PdfOutlineItem to mark it internal

Signed-off-by: Daniel Nguyen <danielnguyenh07@gmail.com>

* refactor(pdf): rename outline_from_docling_parse to extract_outline_from_docling_parse for naming consistency

Signed-off-by: Daniel Nguyen <danielnguyenh07@gmail.com>

* refactor(pdf): make ConversionResult outline a private attr (_pdf_outline), reset after the heading stage consumes it

Signed-off-by: Daniel Nguyen <danielnguyenh07@gmail.com>

---------

Signed-off-by: Daniel Nguyen <danielnguyenh07@gmail.com>
Signed-off-by: Nguyen Hoang Duong <danielnguyenh07@gmail.com>
jumbomuffin101 pushed a commit to jumbomuffin101/docling that referenced this pull request Jul 5, 2026
…ocling-project#3633)

* feat(reading_order): infer PDF section-header levels from numbering and style

The PDF/image pipeline previously emitted every detected SECTION_HEADER at level
1, flattening document hierarchy: Roman-numeral parts and Arabic-numeral
subsections collapsed to the same Markdown heading depth.

Add an opt-in heading-level inference step to the shared reading-order stage that
assigns SectionHeaderItem.level from:

- numbering (primary): legal/outline schemes such as PART I -> 1. -> 1.1 ->
  (a) -> (i), with Roman/Arabic/alpha disambiguation and relative, compressed
  levels so a document that starts at "1." is not forced to start at depth 2;
- style (fallback): font size approximated from parsed PDF cell heights, used
  only for headings without recognizable numbering (requires
  generate_parsed_pages=True).

Gated behind PdfPipelineOptions.heading_hierarchy_options and disabled by
default, so existing conversions and reference outputs are unchanged. PDF
bookmark/outline inference is left as a documented follow-up extension point.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Signed-off-by: Daniel Nguyen <danielnguyenh07@gmail.com>

* fix(reading_order): make conv_res optional in heading-level inference

assign_heading_levels and the per-signal helpers accept conv_res=None so the
numbering-only path (which needs no ConversionResult) is type-correct; style
inference returns early when conv_res is absent.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Signed-off-by: Daniel Nguyen <danielnguyenh07@gmail.com>

* refactor(heading-hierarchy): extract a standalone HeadingHierarchyModel

Addresses review feedback on docling-project#3633: keep the heading-level inference out of the
already-impure reading-order model.

- Move the logic into its own stage model
  (docling/models/stages/heading_hierarchy/heading_hierarchy_model.py).
  HeadingHierarchyModel accepts a ConversionResult and returns the (in-place
  modified) DoclingDocument.
- Invoke it right after the reading-order model in _assemble_document of both
  StandardPdfPipeline and LegacyStandardPdfPipeline.
- Un-nest the options: ReadingOrderOptions no longer carries heading_hierarchy;
  HeadingHierarchyOptions stays standalone on PdfPipelineOptions.
- Factor the core so it is reusable outside the pipeline: assign_heading_levels
  works on a bare DoclingDocument, with the font-based fallback taking the parsed
  pages explicitly (page height read from SegmentedPdfPage.dimension).
- Add a style-fallback test (font size -> level) to cover that path.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Signed-off-by: Daniel Nguyen <danielnguyenh07@gmail.com>

* refactor(heading-hierarchy): drop the bookmarks placeholder until the follow-up

So nothing dead ships in this PR: remove the `use_bookmarks` option and the
no-op `_infer_from_bookmarks` stub. Bookmark / PDF-outline inference (which needs
new backend plumbing) will be introduced in a separate follow-up PR.

Signed-off-by: Daniel Nguyen <danielnguyenh07@gmail.com>

* fix(tach): register heading_hierarchy as a tach module

The new `docling.models.stages.heading_hierarchy` package was not declared in
tach.toml, so it inherited the generic `docling.models.stages` boundary, which
forbids importing `docling.datamodel`. That failed `tach check` -- and the lint
job too, since `prek run --all-files` runs the same tach hook (fail_fast).

Register it as its own module (like `reading_order`) with a `docling.datamodel`
dependency, and add it to the `depends_on` of the standard and legacy PDF
pipelines that import it.

Signed-off-by: Daniel Nguyen <danielnguyenh07@gmail.com>

* Add test method using a real PDF

Signed-off-by: Christoph Auer <cau@zurich.ibm.com>

* Run pre-commit toolchain

Signed-off-by: Christoph Auer <cau@zurich.ibm.com>

* Separate tests to allow full-unit CI markers

Signed-off-by: Christoph Auer <cau@zurich.ibm.com>

---------

Signed-off-by: Daniel Nguyen <danielnguyenh07@gmail.com>
Signed-off-by: Christoph Auer <cau@zurich.ibm.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: Christoph Auer <cau@zurich.ibm.com>
DanielNg0729 added a commit to DanielNg0729/docling that referenced this pull request Aug 21, 2026
Add a usage guide and a runnable example for the section-header level
inference shipped in docling-project#3633, docling-project#3688 and docling-project#3984, covering the bookmark, numbering
and font-style signals, their precedence, and the generate_parsed_pages
requirement of the style fallback.

Signed-off-by: Daniel Nguyen <danielnguyenh07@gmail.com>
@DanielNg0729 DanielNg0729 mentioned this pull request Aug 21, 2026
2 of 3 tasks
dolfim-ibm pushed a commit that referenced this pull request Aug 24, 2026
* docs: document PDF heading-level inference

Add a usage guide and a runnable example for the section-header level
inference shipped in #3633, #3688 and #3984, covering the bookmark, numbering
and font-style signals, their precedence, and the generate_parsed_pages
requirement of the style fallback.

Signed-off-by: Daniel Nguyen <danielnguyenh07@gmail.com>

* Remove PDF heading recovery from features list

Removed PDF heading recovery feature from the list.

Signed-off-by: Nguyen Hoang Duong <danielnguyenh07@gmail.com>

---------

Signed-off-by: Daniel Nguyen <danielnguyenh07@gmail.com>
Signed-off-by: Nguyen Hoang Duong <danielnguyenh07@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Heading level detection for Roman numeral headings (I, II) in PDF to Markdown conversion

3 participants