Skip to content

feat(mhtml): add MHTML input backend - #4184

Merged
ceberam merged 7 commits into
docling-project:mainfrom
egorhowtocode:feat/mhtml-support
Sep 11, 2026
Merged

ceberam merged 7 commits into
docling-project:mainfrom
egorhowtocode:feat/mhtml-support

Conversation

@egorhowtocode

Copy link
Copy Markdown
Contributor

Resolves #659.

Adds a thin MHTML adapter based on Python's standard-library MIME parser.
The backend selects the root HTML part, resolves embedded raster images from
Content-Location and cid: MIME parts without network/filesystem access, then
delegates document semantics to HTMLDocumentBackend.

The design follows docling.rs's established MHTML architecture while adapting
to Python Docling's backend and safety conventions.

Tests cover .mhtml/.mht detection, root/start selection, nested MIME,
quoted-printable and base64 HTML, Unicode, HTML structure, Content-Location
and cid images, missing/unsupported images, text-only/unusable input,
path/stream input, origin preservation, image-size limits, and no outbound
network access.

@github-actions

github-actions Bot commented Sep 6, 2026

Copy link
Copy Markdown
Contributor

✅ DCO Check Passed

Thanks @egorhowtocode, all your commits are properly signed off. 🎉

@mergify

mergify Bot commented Sep 6, 2026 •

Copy link
Copy Markdown
Contributor

Merge Protections

🟢 All 2 merge protections satisfied — ready to merge.

Show 2 satisfied protections

🟢 Enforce conventional commit

Make sure that we follow https://www.conventionalcommits.org/en/v1.0.0/

  • title ~= ^(fix|feat|docs|style|refactor|perf|test|build|ci|chore|revert)(?:\(.+\))?(!)?:

🟢 Require two reviewer for test updates

When test data is updated, we require two reviewers

  • #approved-reviews-by >= 2

@codecov

codecov Bot commented Sep 7, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 86.34361% with 31 lines in your changes missing coverage. Please review.

Files with missing lines Patch % Lines
docling/backend/html_backend.py 86.16% 17 Missing and 14 partials ⚠️

📢 Thoughts on this report? Let us know!

@ceberam ceberam added enhancement New feature or request html issue related to html backend labels Sep 7, 2026

@ceberam ceberam left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks @egorhowtocode for taking the initiative and suggesting a PR to address a quite old feature request.

You suggest adding a separate declarative backend that internally creates a second one, a subclass of HTMLDocumentBackend, and delegates convert() to it. This is a two-layer design.

Looking at how we have created declarative backends in Docling, the right model here would be to extend HTMLDocumentBackend directly and override supported_formats() to return {InputFormat.HTML, InputFormat.MHTML}, exactly how other backends accept multiple extensions. After all, MHTML is HTML: it is just HTML with embedded resources packed into a MIME envelope.

The most natural structure would be a single HTMLDocumentBackend that, when the format is MHTML, pre-processes the MIME envelope in __init__ (extracting the HTML bytes and building the resources dict), then proceeds with the existing convert() path but with its image resolver redirected. There would be no need for a second internal InputDocument, no second backend instantiation, and no private _MhtmlHTMLDocumentBackend class at all.

The current design introduces unnecessary indirection and complexity:

  • Two backend classes where one would do.
  • An _MhtmlHTMLDocumentBackend with set_archive_resources() that must be called immediately after construction (a fragile two-phase init pattern).
  • A new InputDocument and in_doc._backend access (inspecting a private attribute) inside convert(), which is brittle.
  • A new MHTMLFormatOption in document_converter.py that is structurally identical to HTMLFormatOption minus its backend_options_for_input override.

There are other specific issues that I have seen in the PR (_load_image_data dead code, cid lookup redundancy, ruff formatting failure,...), but I would prefer that we first agree on the design (a single backend vs 2 backends) before entering into implementation details. What do you think?

@egorhowtocode

Copy link
Copy Markdown
Contributor Author

Thanks @ceberam for the detailed review. I agree that the previous two-layer design was unnecessarily complex, and I have reworked the PR around Docling’s existing declarative-backend model.

The updated change is rebased onto the current main and now:

  • extends HTMLDocumentBackend to support both HTML and MHTML through the same initialization and conversion path;
  • reuses HTMLFormatOption;
  • parses the MIME envelope and initializes the selected multipart/related scope and its resources during normal backend initialization;
  • follows multipart/related root semantics, including start, the default first entity, and multipart/alternative;
  • applies the root part’s declared MIME charset before HTML parsing;
  • limits embedded-resource lookup to the selected related scope and uses deterministic first-part-wins handling for duplicate identifiers;
  • supports both Content-Location and normalized cid: image references;
  • gives archive resources precedence over explicitly permitted external loading while preserving the existing image-fetch controls;
  • safely resolves relative, POSIX, Windows, and file: locations without allowing archive-controlled paths to escape the source directory;
  • preserves the original MHTML filename, MIME type, and binary hash without mutating caller-supplied options;
  • explicitly rejects unsupported browser rendering and malformed or rootless archives;
  • removes the separate MHTML backend and format option, synthetic InputDocument, private _backend access, two-phase resource setter, dead helper, and redundant CID handling.

The tests cover .mhtml and .mht detection, path and stream input, a real Blink-style quoted-printable fixture, non-UTF-8 MIME charsets, standard HTML structures, related-root selection, multipart/alternative, CID and Content-Location images, duplicate and cross-scope resources, image-fetch permissions and precedence, relative and Windows/file-URI path confinement, invalid input, origin metadata, and rendering rejection.

Validation completed:

  • 39 dedicated MHTML tests passed;
  • 72 additional HTML, EPUB, email, and input-detection regression tests passed, for 111 tests in total;
  • repository-wide Ruff formatting and lint checks passed;
  • Tach dependency checks passed;
  • the new MHTML test module passed type checking;
  • changed-line coverage for the HTML-backend additions is 95.09%;
  • a separate read-only LLM review of the final diff found no remaining actionable architectural issues.

For context, I am an active Docling user and have encountered several unsupported formats relevant to my work, including MHTML. I opened this PR to help bring attention and a possible implementation to this feature request.

I also want to be transparent that the implementation was produced through an LLM-driven workflow based on my requirements, examples, and validation criteria. I can test the resulting behavior, incorporate concrete review feedback, and address localized implementation problems. However, I do not have sufficiently deep familiarity with Docling’s internal architecture to lead a broader architectural redesign without maintainer guidance. For this revision, I iterated specifically against your architectural guidance, expanded tests, and a separate review of the resulting diff.

If the revised implementation is generally sound but needs minor corrections or additional tests, I will be happy to address them. If it still has fundamental architectural or logical problems, I would also be completely comfortable with this PR being superseded by an implementation from a maintainer or another contributor. My main goal is to help get the missing feature resolved, rather than to retain ownership of the implementation.

@ceberam

ceberam commented Sep 10, 2026

Copy link
Copy Markdown
Member

Thanks @egorhowtocode for the refactoring! The design is now much more robust.
There are a couple of things that still need to be addressed. One of them is the MIME types for this new input format. Instead of registering the new types into Python's process-wide mimetypes, we should extend the list of extra MIME types in docling-core. I have opened a PR for that (docling-project/docling-core#763 ) and until it gets approved and a new version of docling-core is released, we need to put this PR on hold.

egorhowtocode and others added 7 commits September 10, 2026 17:12
Signed-off-by: Egor Ivanov <ivanovegor1648@gmail.com>
…ration

Signed-off-by: Cesar Berrospi Ramis <ceb@zurich.ibm.com>
Signed-off-by: Cesar Berrospi Ramis <ceb@zurich.ibm.com>
…limitation

Signed-off-by: Cesar Berrospi Ramis <ceb@zurich.ibm.com>
Signed-off-by: Cesar Berrospi Ramis <ceb@zurich.ibm.com>
Signed-off-by: Cesar Berrospi Ramis <ceb@zurich.ibm.com>
Signed-off-by: Cesar Berrospi Ramis <ceb@zurich.ibm.com>

@ceberam ceberam left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@egorhowtocode we finally released docling-core and we can now properly handle the MIME types for MHTML.
I have added a commit to address this point and others minor wants that I had identified to strengthen this PR.
I think we are ready to merge this new feature 🚀

@PeterStaar-IBM PeterStaar-IBM left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

lgtm!

@ceberam
ceberam merged commit 7e05ed6 into docling-project:main Sep 11, 2026
26 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

enhancement New feature or request html issue related to html backend

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[FEAT] Support MHTML Conversion

3 participants