Conversation
… charset When the HTML part of an archive declares a charset, it was decoded and re-encoded as UTF-8 before parsing. BeautifulSoup then followed the page's own <meta charset> and decoded those UTF-8 bytes a second time, so text in windows-1252, iso-8859-1 and other single-byte charsets came out as mojibake. Hand the decoded text to the parser instead, so the charset declared on the part is the only one applied. Signed-off-by: Rohit Behera <126186063+r0h1tb@users.noreply.github.com>
|
✅ DCO Check Passed Thanks @r0h1tb, all your commits are properly signed off. 🎉 |
Merge Protections🟢 Merge protection satisfied — ready to merge. Show 1 satisfied protection🟢 Enforce conventional commitMake sure that we follow https://www.conventionalcommits.org/en/v1.0.0/
|
Codecov Report✅ All modified and coverable lines are covered by tests. 📢 Thoughts on this report? Let us know! |
chrikrah
left a comment
There was a problem hiding this comment.
@r0h1tb I would merge this. At 81ea2e8 the html and mhtml suites pass, and reverting html_backend.py to the merge base 2d5c590 with your tests kept fails the two meta cases:
$ python -m pytest tests/test_backend_mhtml.py tests/test_backend_html.py -q
84 passed
# html_backend.py at 2d5c590, tests kept
2 failed, 82 passed # test_declared_non_utf8_charset_is_decoded[meta-charset], [meta-http-equiv]
$ python -m pytest tests/test_backend_email.py tests/test_backend_epub.py -q
34 passed
The fix also covers a case the parametrization does not:
$ python docling-4388_mhtml_probe.py # at 2d5c590, then at 81ea2e8
# each archive: multipart/related, one base64 text/html part holding <p>Привет</p>,
# converted from a DocumentStream by DocumentConverter(allowed_formats=[InputFormat.MHTML])
MIME part charset / <meta> / body bytes 2d5c590 81ea2e8
utf-8 / windows-1251 / utf-8 'Привет' 'Привет'
windows-1251 / utf-8 / windows-1251 'Привет' 'Привет'
x-bogus / windows-1251 / windows-1251 'Привет' 'Привет'
none / windows-1251 / windows-1251 'Привет' 'Привет'
non-blocking: the first row, a UTF-8 part whose <meta> still says windows-1251, is not covered. A sibling test with a utf-8 part and that meta would pin it. The branch merges cleanly onto d6f0307.
@ceberam, you reviewed and merged the MHTML backend in #4184; could you take a look at this one?
Problem
If an archive's HTML part declares a charset in its MIME header and the page also has its own
<meta charset>, non-ASCII text in a single-byte charset comes out garbled:_decode_mhtml_htmldecoded the part with its declared charset and re-encoded it as UTF-8. BeautifulSoup then picked up the<meta>declaration and decoded those UTF-8 bytes again as windows-1251. The docstring already noted this as a known limitation.Fix
Return the decoded text and give it to the parser as
str, so there's no second decode. The part's charset wins over the page's own declaration, the same precedence HTML gives aContent-Typecharset over<meta>. Parts without a charset still reach the parser as bytes and get detected from the<meta>tag as before (the Blink fixture is that case).Tests
Parametrized
test_declared_non_utf8_charset_is_decodedover no meta tag,<meta charset>and<meta http-equiv="Content-Type">. Without the fix:html/mhtml/epub/email/jats backend tests: 173 passed, 7 failed before; 175 passed, 5 failed after. The 5 are remote-image tests that fail on both sides locally because example.com doesn't resolve here. ruff, ty and tach are clean, with no new ty diagnostics.
Checklist: