Skip to content

feat(pdf): infer heading levels from font weight, slant and case - #3984

Merged
PeterStaar-IBM merged 1 commit into
docling-project:mainfrom
DanielNg0729:feat/pdf-heading-font-style
Aug 14, 2026
Merged

PeterStaar-IBM merged 1 commit into
docling-project:mainfrom
DanielNg0729:feat/pdf-heading-font-style

Conversation

@DanielNg0729

Copy link
Copy Markdown
Member

Hi maintainers, this PR is to further follow-up to #3633 (numbering/style heading inference) and #3688 (bookmarks/ToC).

The style fallback of HeadingHierarchyModel, the only signal that fires for unnumbered headings in PDFs without an outline, ranked headings by exact rounded font size. Two problems, both measured on the PDFs in tests/data/pdf/sources/:

  1. The size proxy is noisy, and exact bucketing turns that noise into levels. A heading's size
    is the median height of the cells overlapping it, which measures the glyphs actually on the
    line rather than the font size. In redp5110_sampled.pdf, Securing and protecting IBM DB2 data measures 24 and Contents measures 22 in the same font — descenders. Every heading
    therefore landed in its own bucket: 10 distinct sizes over 20 headings, 12 of them clamped to
    the maximum level.
  2. Weight, slant and letter case were not read at all, although PdfTextCell.font_name
    carries them (/Helvetica-Bold, /NKDKGK+HelveticaNeueLTPro-Bd, /KIDKQO+Times-Italic).

Because of (1), fixing (2) alone is a no-op: with every heading alone in its size bucket, a tie-breaker never gets consulted. Both change together here.

Changes

  • docling/utils/font_style.py (new, pure): parse_font_style() reads weight and slant from a PDF font name; weight_class() buckets to light/regular, medium/semibold, bold+. Two rules keep it conservative, since a false "bold" silently rewrites a heading level while a miss only falls back to font size:

    • style words match as whole tokens after splitting on separators and camel case (Avenir-Book is a weight; the family Bookman is not; HelveticaNeue-BoldCond is bold);
    • abbreviations count only as a complete separator-delimited part (-Bd is bold; the TB in LinLibertineTB and the LT in HelveticaNeueLTPro are not, since foundry tags glued to a family name are indistinguishable from weight abbreviations).

    Names that carry no recognizable styling — /F1, "null", ArialMT, bare families, resolve to regular and upright, so those documents keep the previous font-size behavior.

  • heading_hierarchy_model.py: _heading_font_size becomes _heading_style, returning size, weight class, italic and all-caps. Weight and slant are a character-weighted vote across the cells of the heading, since a heading can mix a regular 1.1 marker with a bold title; cells with no font information (OCR, the pypdfium2 backend) contribute their size only.
    _cluster_sizes merges near-equal sizes, and headings are then ranked by (size cluster, weight, upright before italic, capitals before mixed case). A signal that does not vary across a document's headings adds no levels.

  • HeadingHierarchyOptions: use_font_style (default True, sub-switch of use_style) and style_size_tolerance (default 0.05). Both are exposed by the service API automatically, since docling/datamodel/service/options.py embeds the model directly.

No pipeline change: the model already receives the parsed pages, and the same cells carry the font names. No new dependency

Usage

The stage stays opt-in. generate_parsed_pages=True is what makes the parsed cells — and
therefore the fonts — available to it:

from docling.datamodel.base_models import InputFormat
from docling.datamodel.pipeline_options import (
    HeadingHierarchyOptions,
    PdfPipelineOptions,
)
from docling.document_converter import DocumentConverter, PdfFormatOption

pipeline_options = PdfPipelineOptions(generate_parsed_pages=True)
pipeline_options.heading_hierarchy_options = HeadingHierarchyOptions(
    enabled=True,
    # use_font_style=False,      # rank by font size alone, as before this PR
    # style_size_tolerance=0.0,  # never merge two measured sizes
)

converter = DocumentConverter(
    format_options={
        InputFormat.PDF: PdfFormatOption(pipeline_options=pipeline_options)
    }
)
print(converter.convert("report.pdf").document.export_to_markdown())

Through the service API, where the new fields need no separate wiring:

{
  "do_pdf_heading_hierarchy": true,
  "pdf_heading_hierarchy_options": {
    "use_font_style": true,
    "style_size_tolerance": 0.05
  }
}

Measured effect

Isolating the signal (use_bookmarks=False, use_numbering=False, use_style=True), against main:

2206.01062.pdf (DocLayNet, ACM template) — 7 of 18 headings change, all corrections:

heading main this PR
1 INTRODUCTION, 2 RELATED WORK, … REFERENCES 2 2
Baselines for Object Detection, Learning Curve, Impact of Class Labels, … 2 3
ACMReference Format: 3 4

The mixed-case subsection heads were being flattened into their all-caps parent sections. The discriminator here is letter case, not weight: this document is set in Linux Libertine, i.e. the LinLibertineTB convention the parser deliberately declines to read.

redp5110_sampled.pdf (IBM Redbook, Helvetica) — 11 of 20 change; headings pinned at the level cap drop from 12/20 to 8/20. 1.1 Security fundamentals now sits above 1.3.1 Existing row and column control; both were at the cap before. The italic
DB2 for i Center of Excellence now ranks below the same-size bold Contents/Preface.

amt_handbook_sample.pdf, 2305.03393v1-pg9.pdf — no changes: uniform heading styling, nothing to separate.

Compatibility

The stage is opt-in and off by default (heading_hierarchy_options.enabled=False), so no reference data changed. For users who had enabled it, use_font_style=False restores the previous ranking, and style_size_tolerance=0.0 restores exact size bucketing, worth noting that size clustering applies regardless of use_font_style, because it is a correctness fix for the size
proxy rather than a new signal.

Known gap: some PDFs report a font resource key (/F1) instead of a font name, so weight is unavailable there and the ranking falls back to size. The root cause is upstream in docling-parse (init_font_name() prefers the deprecated /Name key over /BaseFont) and is left for a separate fix.

Checklist:

  • Documentation has been updated, if necessary.
  • Examples have been added, if necessary.
  • Tests have been added, if necessary.

The style fallback ranked headings by exact rounded font size. A heading's size
is the median height of its cells, so the same font measures taller on a heading
with descenders, and every heading landed in its own bucket: on
redp5110_sampled.pdf, 10 buckets over 20 headings with 12 clamped to the maximum
level. Font weight and slant were never read at all, although PdfTextCell carries
the font name.

Merge near-equal sizes, then rank headings within a size by the weight and slant
read from the embedded font names and by letter case. Font names that carry no
recognizable styling fall back to font size, so the previous behavior stands
wherever the signal is unavailable.

Signed-off-by: Daniel Nguyen <danielnguyenh07@gmail.com>
@github-actions

Copy link
Copy Markdown
Contributor

✅ DCO Check Passed

Thanks @DanielNg0729, all your commits are properly signed off. 🎉

@mergify

mergify Bot commented Aug 13, 2026

Copy link
Copy Markdown
Contributor

Merge Protections

🟢 Merge protection satisfied — ready to merge.

Show 1 satisfied protection

🟢 Enforce conventional commit

Make sure that we follow https://www.conventionalcommits.org/en/v1.0.0/

  • title ~= ^(fix|feat|docs|style|refactor|perf|test|build|ci|chore|revert)(?:\(.+\))?(!)?:

@wittjeff

Copy link
Copy Markdown
Contributor

Very glad to see this land on top of #3633/#3688 — heading levels are the single highest-value structure signal for document accessibility (a flat H1 forest fails the intent of PDF/UA §7.4), and the font weight/slant/case signal here addresses exactly the noise that made size-only bucketing unreliable.

Related: we've just posted an RFC in Discussions — #3988 — proposing PDF/UA-motivated structure-type coverage, with heading levels as the top ask. It includes an offer of human-verified, multi-annotator-adjudicated heading-level ground truth captured from accessibility remediation sessions (openly licensed, DoclingDocument-native). That could give this stage's heuristics a measured accuracy target — and eventually let the trained models learn levels directly (DocTags already reserves section_header_level_0..5).

If it would help, we're happy to run this branch over our benchmark slices and report heading-level accuracy against adjudicated labels.

@dolfim-ibm dolfim-ibm left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

lgtm

@codecov

codecov Bot commented Aug 14, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 99.11504% with 1 line in your changes missing coverage. Please review.

Files with missing lines Patch % Lines
docling/utils/font_style.py 98.27% 1 Missing ⚠️

📢 Thoughts on this report? Let us know!

@PeterStaar-IBM PeterStaar-IBM left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

lgtm!

@PeterStaar-IBM
PeterStaar-IBM merged commit e74a905 into docling-project:main Aug 14, 2026
26 checks passed
@PierreMesure

Copy link
Copy Markdown

@DanielNg0729 ❤️

image

DanielNg0729 added a commit to DanielNg0729/docling that referenced this pull request Aug 21, 2026
Add a usage guide and a runnable example for the section-header level
inference shipped in docling-project#3633, docling-project#3688 and docling-project#3984, covering the bookmark, numbering
and font-style signals, their precedence, and the generate_parsed_pages
requirement of the style fallback.

Signed-off-by: Daniel Nguyen <danielnguyenh07@gmail.com>
@DanielNg0729 DanielNg0729 mentioned this pull request Aug 21, 2026
2 of 3 tasks
dolfim-ibm pushed a commit that referenced this pull request Aug 24, 2026
* docs: document PDF heading-level inference

Add a usage guide and a runnable example for the section-header level
inference shipped in #3633, #3688 and #3984, covering the bookmark, numbering
and font-style signals, their precedence, and the generate_parsed_pages
requirement of the style fallback.

Signed-off-by: Daniel Nguyen <danielnguyenh07@gmail.com>

* Remove PDF heading recovery from features list

Removed PDF heading recovery feature from the list.

Signed-off-by: Nguyen Hoang Duong <danielnguyenh07@gmail.com>

---------

Signed-off-by: Daniel Nguyen <danielnguyenh07@gmail.com>
Signed-off-by: Nguyen Hoang Duong <danielnguyenh07@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants