feat(pdf): infer heading levels from font weight, slant and case - #3984
PeterStaar-IBM merged 1 commit into
Conversation
The style fallback ranked headings by exact rounded font size. A heading's size is the median height of its cells, so the same font measures taller on a heading with descenders, and every heading landed in its own bucket: on redp5110_sampled.pdf, 10 buckets over 20 headings with 12 clamped to the maximum level. Font weight and slant were never read at all, although PdfTextCell carries the font name. Merge near-equal sizes, then rank headings within a size by the weight and slant read from the embedded font names and by letter case. Font names that carry no recognizable styling fall back to font size, so the previous behavior stands wherever the signal is unavailable. Signed-off-by: Daniel Nguyen <danielnguyenh07@gmail.com>
|
✅ DCO Check Passed Thanks @DanielNg0729, all your commits are properly signed off. 🎉 |
Merge Protections🟢 Merge protection satisfied — ready to merge. Show 1 satisfied protection🟢 Enforce conventional commitMake sure that we follow https://www.conventionalcommits.org/en/v1.0.0/
|
|
Very glad to see this land on top of #3633/#3688 — heading levels are the single highest-value structure signal for document accessibility (a flat H1 forest fails the intent of PDF/UA §7.4), and the font weight/slant/case signal here addresses exactly the noise that made size-only bucketing unreliable. Related: we've just posted an RFC in Discussions — #3988 — proposing PDF/UA-motivated structure-type coverage, with heading levels as the top ask. It includes an offer of human-verified, multi-annotator-adjudicated heading-level ground truth captured from accessibility remediation sessions (openly licensed, DoclingDocument-native). That could give this stage's heuristics a measured accuracy target — and eventually let the trained models learn levels directly (DocTags already reserves If it would help, we're happy to run this branch over our benchmark slices and report heading-level accuracy against adjudicated labels. |
Codecov Report❌ Patch coverage is
📢 Thoughts on this report? Let us know! |
Add a usage guide and a runnable example for the section-header level inference shipped in docling-project#3633, docling-project#3688 and docling-project#3984, covering the bookmark, numbering and font-style signals, their precedence, and the generate_parsed_pages requirement of the style fallback. Signed-off-by: Daniel Nguyen <danielnguyenh07@gmail.com>
* docs: document PDF heading-level inference Add a usage guide and a runnable example for the section-header level inference shipped in #3633, #3688 and #3984, covering the bookmark, numbering and font-style signals, their precedence, and the generate_parsed_pages requirement of the style fallback. Signed-off-by: Daniel Nguyen <danielnguyenh07@gmail.com> * Remove PDF heading recovery from features list Removed PDF heading recovery feature from the list. Signed-off-by: Nguyen Hoang Duong <danielnguyenh07@gmail.com> --------- Signed-off-by: Daniel Nguyen <danielnguyenh07@gmail.com> Signed-off-by: Nguyen Hoang Duong <danielnguyenh07@gmail.com>

Hi maintainers, this PR is to further follow-up to #3633 (numbering/style heading inference) and #3688 (bookmarks/ToC).
The style fallback of
HeadingHierarchyModel, the only signal that fires for unnumbered headings in PDFs without an outline, ranked headings by exact rounded font size. Two problems, both measured on the PDFs intests/data/pdf/sources/:is the median height of the cells overlapping it, which measures the glyphs actually on the
line rather than the font size. In
redp5110_sampled.pdf,Securing and protecting IBM DB2 datameasures 24 andContentsmeasures 22 in the same font — descenders. Every headingtherefore landed in its own bucket: 10 distinct sizes over 20 headings, 12 of them clamped to
the maximum level.
PdfTextCell.font_namecarries them (
/Helvetica-Bold,/NKDKGK+HelveticaNeueLTPro-Bd,/KIDKQO+Times-Italic).Because of (1), fixing (2) alone is a no-op: with every heading alone in its size bucket, a tie-breaker never gets consulted. Both change together here.
Changes
docling/utils/font_style.py(new, pure):parse_font_style()reads weight and slant from a PDF font name;weight_class()buckets to light/regular, medium/semibold, bold+. Two rules keep it conservative, since a false "bold" silently rewrites a heading level while a miss only falls back to font size:Avenir-Bookis a weight; the familyBookmanis not;HelveticaNeue-BoldCondis bold);-Bdis bold; theTBinLinLibertineTBand theLTinHelveticaNeueLTProare not, since foundry tags glued to a family name are indistinguishable from weight abbreviations).Names that carry no recognizable styling —
/F1,"null",ArialMT, bare families, resolve to regular and upright, so those documents keep the previous font-size behavior.heading_hierarchy_model.py:_heading_font_sizebecomes_heading_style, returning size, weight class, italic and all-caps. Weight and slant are a character-weighted vote across the cells of the heading, since a heading can mix a regular1.1marker with a bold title; cells with no font information (OCR, the pypdfium2 backend) contribute their size only._cluster_sizesmerges near-equal sizes, and headings are then ranked by(size cluster, weight, upright before italic, capitals before mixed case). A signal that does not vary across a document's headings adds no levels.HeadingHierarchyOptions:use_font_style(defaultTrue, sub-switch ofuse_style) andstyle_size_tolerance(default0.05). Both are exposed by the service API automatically, sincedocling/datamodel/service/options.pyembeds the model directly.No pipeline change: the model already receives the parsed pages, and the same cells carry the font names. No new dependency
Usage
The stage stays opt-in.
generate_parsed_pages=Trueis what makes the parsed cells — andtherefore the fonts — available to it:
Through the service API, where the new fields need no separate wiring:
{ "do_pdf_heading_hierarchy": true, "pdf_heading_hierarchy_options": { "use_font_style": true, "style_size_tolerance": 0.05 } }Measured effect
Isolating the signal (
use_bookmarks=False, use_numbering=False, use_style=True), againstmain:2206.01062.pdf(DocLayNet, ACM template) — 7 of 18 headings change, all corrections:1 INTRODUCTION,2 RELATED WORK, …REFERENCESBaselines for Object Detection,Learning Curve,Impact of Class Labels, …ACMReference Format:The mixed-case subsection heads were being flattened into their all-caps parent sections. The discriminator here is letter case, not weight: this document is set in Linux Libertine, i.e. the
LinLibertineTBconvention the parser deliberately declines to read.redp5110_sampled.pdf(IBM Redbook, Helvetica) — 11 of 20 change; headings pinned at the level cap drop from 12/20 to 8/20.1.1 Security fundamentalsnow sits above1.3.1 Existing row and column control; both were at the cap before. The italicDB2 for i Center of Excellencenow ranks below the same-size boldContents/Preface.amt_handbook_sample.pdf,2305.03393v1-pg9.pdf— no changes: uniform heading styling, nothing to separate.Compatibility
The stage is opt-in and off by default (
heading_hierarchy_options.enabled=False), so no reference data changed. For users who had enabled it,use_font_style=Falserestores the previous ranking, andstyle_size_tolerance=0.0restores exact size bucketing, worth noting that size clustering applies regardless ofuse_font_style, because it is a correctness fix for the sizeproxy rather than a new signal.
Known gap: some PDFs report a font resource key (
/F1) instead of a font name, so weight is unavailable there and the ranking falls back to size. The root cause is upstream in docling-parse (init_font_name()prefers the deprecated/Namekey over/BaseFont) and is left for a separate fix.Checklist: