Repository navigation
fix(i18n): clear 1,188 fuzzy translations with broken placeholders - #44611
Conversation
|
Bito Automatic Review Skipped - Large PR |
These translations carry placeholders that do not match their English source, so none of them can format. At runtime they raise an exception, show a literal placeholder such as "Error: %s", drop a value, show a blank one, or fail silently because the frontend Translator catches the error and returns the English key. Superset serves fuzzy translations on purpose, so a broken fuzzy entry is as live as a confirmed one: `pybabel compile --use-fuzzy` reports 1,090 errors and still writes every .mo file, and po2json.sh passes --fuzzy. Clearing the msgstr makes each entry untranslated, so the UI falls back to the English source with its values. Confirmed (non-fuzzy) translations are untouched; the 8 broken confirmed entries are fixed by hand separately. The Slovak SQL Lab progress message is deliberately left alone: apache#44587 rewrites that entry properly and promotes it out of fuzzy, so this change stays clear of it. The 97 entries that carried a `Machine-translated via backfill_po.py` attribution lose it along with the translation it described, so they read as plainly untranslated and backfill_po.py sees clean candidates. Measured at master 2ac6e24, across 18 catalogs: - compile errors under --use-fuzzy: 1,090 -> 155 - 36 of the 37 entries that raised now format cleanly (the 37th is apache#44587's) - placeholder checker findings: 1,227 -> 38 - check_translation_regression.py: no regressions; confirmed-translation counts unchanged in all 29 catalogs, fuzzy down exactly 1,189 Refs: apache#44551
✅ Deploy Preview for superset-docs-preview ready!
To edit notification comments on pull requests, go to your Netlify project configuration. |
a47bfcb to
4569826
Compare
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## master #44611 +/- ##
==========================================
- Coverage 81.51% 81.50% -0.01%
==========================================
Files 2973 2973
Lines 180202 180280 +78
Branches 41720 41737 +17
==========================================
+ Hits 146883 146945 +62
- Misses 30602 30612 +10
- Partials 2717 2723 +6
Flags with carried forward coverage won't be shown. Click here to find out more. ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
rusackas
left a comment
There was a problem hiding this comment.
This is exactly the plan from #44551, clearing the ones that can't format cleanly instead of letting them crash SQL Lab. Spot-checked a few catalogs and the diff only touches the fuzzy entries called out, confirmed translations are untouched. LGTM, approving!
| msgstr[0] "۱ ستون در مجموعه داده مجازی تغییر یافت" | ||
| msgstr[1] "%s ستون در مجموعه داده مجازی تغییر یافت" | ||
| msgstr[0] "" | ||
| msgstr[1] "" |
There was a problem hiding this comment.
Clearing this drops a translation whose placeholders actually match: msgid "Modified 1 column in the virtual dataset" carries no placeholder, and the removed msgstr[0] ("۱ ستون در مجموعه داده مجازی تغییر یافت") had none either; msgid_plural "Modified %s columns in the virtual dataset" carries %s, and the removed msgstr[1] ("%s ستون در مجموعه داده مجازی تغییر یافت") had it too. Persian users modifying dataset columns now see the English fallback instead of a previously-correct translation.
The same pattern shows up at superset/translations/zh_TW/LC_MESSAGES/messages.po:1654 ("Added to 1 dashboard", where the removed msgstr[0] "增加到看板" also had no placeholder mismatch). Since the same detection script produced both misses, could you restore these two entries and recheck the rest of the 52 cleared plural entries for the same false positive before merging?
There was a problem hiding this comment.
Good catch on the Persian entry. It is restored in ff91653. Persian's plural rule is n > 1, so msgstr[0] serves both 0 and 1, and it writes the count as the Persian digit ۱. The check looked for the value at n=1 and didn't recognise ۱ as it, so it counted the entry as broken at 0 and 1. That kept it out of the "fails only at 0" exclusion. The call site only toasts when columnChanges.modified.length is non-zero, so the translation is correct in practice.
zh_TW is different, and I'd keep it cleared. Its header is nplurals=1; plural=0;, so msgstr[0] "增加到看板" is used for every count. "Added to 5 dashboards" renders as "Added to dashboard", with the number gone. That's the "drops a value" case from #44551.
I rechecked all 52 cleared plural entries by each language's plural rule (n = 0, 1, 2, 3, 5, 11, 21, 22, 101, 111). Persian was the only one that fails just at n=0. The other 51 each drop the count for some n ≥ 2. The PR is 1,188 entries now; the title and description are updated.
| # Machine-translated via backfill_po.py (claude-sonnet-4-6) [refs: ca, cs, de, | ||
| # es, fr, ja, lv, mi, ro, ru, sk, sr, sr_Latn, tr, uk] | ||
| #, fuzzy, python-format | ||
| #, python-format |
There was a problem hiding this comment.
Removing the Machine-translated via backfill_po.py attribution above only deleted the first line of the wrapped two-line comment, leaving this now-meaningless continuation attached to the entry. The same one-line-only removal repeats for all 97 comment removals in this PR across the touched catalogs, leaving 97 orphaned reference-list fragments behind. Could this line (and the equivalent stray line in the other 96 spots) also be deleted?
| #, python-format |
There was a problem hiding this comment.
You're right, and thanks for spotting it. My comparison checked msgids, headers, locations and extracted comments, but not translator comments, so the stray lines slipped through. Fixed in ff91653: 113 continuation lines on 96 entries (some attributions wrapped to three lines). I matched every entry against master to remove exactly the lines that continued each attribution there. The comparison now also covers translator comments. It flags all 97 leftovers on the previous head and none on this one. The Persian entry that's restored keeps its attribution, since its translation stays.
Two corrections from review. The Persian "Modified 1 column in the virtual dataset" entry was cleared by mistake. Persian's plural rule (n > 1) sends both 0 and 1 to msgstr[0], which writes the count as the Persian digit; the check read that as a missing value at n=1 and so did not treat it as a zero-only case. The call site only shows the message when at least one column changed, so n is never 0 and the translation is correct. It is restored as it was. The other 51 cleared plural entries were rechecked by plural rule: each drops the count for n >= 2. Removing the `# Machine-translated via backfill_po.py` attributions deleted only the first line of each wrapped comment and left 96 continuation lines (113 lines, some comments wrapped to three lines) attached to the cleared entries. They are removed. 1,188 entries are cleared. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
SUMMARY
This is PR 2 of the plan in #44551. @rusackas approved the approach there:
It clears 1,188 fuzzy translations in 18 catalogs. Their placeholders do not match the English source, so none of them can format correctly. Clearing
msgstrmakes each entry untranslated, so the UI shows the English text with its values instead of a crash or a literal%s.This PR does not touch the Slovak SQL Lab entry. #44587 rewrites that one properly and promotes it out of fuzzy, which is a better outcome than falling back to English. I left it out so the two changes stay independent and neither can undo the other.
Confirmed (non-fuzzy) translations are not touched. The 8 broken confirmed entries (4
mi, 2es, 1fr, 1nl) are PR 3. They are fixed by hand so native speakers can review the wording.What gets cleared
KeyError/TypeError)Error: %sTranslatorcatches the error and returns the English keyThe last row needs a word of caution, because it is easy to overstate this bug.
Translator.translatewrapsfetchin atry/catchand returns the input when formatting fails. So a broken frontend entry only shows a%son screen if the English source itself has a placeholder. 26 of them do, and the screenshots below are one of those 26. The other 220 already fall back to English today.They are still worth clearing. They are translations that can never render. They survive only because of a
catch, and they count as "translated" in every coverage number.Why these entries are live
Superset serves fuzzy translations on purpose, so a broken fuzzy entry ships like any other:
pybabel compile --use-fuzzyreports 1,090 errors, exits 1, and still writes all 30.mofiles. TheDockerfileends that command with|| true, andflask fab babel-compiledoes not check the exit status.superset-frontend/scripts/po2json.shpasses--fuzzy, so the frontend packs carry them too.Three things to look at
# Machine-translated via backfill_po.py (...)comment. The comment described the translation being removed, so it goes too.backfill_po.pypicks entries by emptymsgstr, so these are now clean candidates for a future backfill..potext directly instead of rewriting the files with Babel or polib, because both reformat unrelated entries. Babel adds apython-formatflag next tono-python-format; polib rewraps thousands of lines. I compared every entry in all 29 catalogs againstmaster: msgid sets, headers, locations, extracted comments and translator comments are identical everywhere else. Each changed entry is one of the 1,188, and its only changes are the clearedmsgstr, the droppedfuzzyflag and, where present, the whole attribution comment.BEFORE/AFTER SCREENSHOTS OR ANIMATED GIF
Duplicating a role in Dutch. Same build, same locale. Only the catalogs differ.
Before. The dialog title reads
Duplicate role %(name)s. Thenltranslation isDubbele kolomnaam (of -namen): %(columns)s, which asks for acolumnsvalue the call site does not pass, so formatting throws,Translatorfalls back to the English key, and the key's own placeholder is left on screen.After. The title reads Duplicate role Admin. The entry is untranslated, so the English source is used and the role name lands where it belongs.
The same thing on a server-rendered page, in
pt_BR.Before. The Reset my password button reads
{} SENHA. The catalog hasmsgstr "%s SENHA"for a source string with no placeholder, and Flask-AppBuilder's label handling turns the%sinto{}.After. The button reads Redefinir minha senha. Flask-AppBuilder has a correct translation for this string in its own catalog, and Superset's broken entry was overriding it. Clearing ours brings back the right Portuguese.
A second case, through the API. Creating a dataset that already exists, in Dutch:
masternlHTTP 500 {"message": "Fatal error"}HTTP 422 {"table": ["Dataset main.fruit already exists"]}en(control)HTTP 422 {"table": ["Dataset main.fruit already exists"]}Dataset %(table)s already existsis translated with%(name)sinnl. Building the validation message raisesKeyError, so a clean 422 becomes a 500.TESTING INSTRUCTIONS
Measured at master
2ac6e2468d. The flagged set is the same one reported in #44551 at269b9f99fe: 1,227 entries, no drift.1. Compile errors drop
The 155 left are a different problem:
pybabel's compile check is stricter than the runtime. 98 are strings with a prose percent sign and no real format code (% of total,% calculation; most already carryno-python-format), 44 are plural forms where one form spells its single number out, and 13 are others.Only 8 of the 155 match an entry the runtime check still flags, and I can name them. One is the Slovak SQL Lab message #44587 fixes. Six are the "fails only when the count is 0" plural forms left alone on purpose (
%s day agoand%s min agoinplandpt_BR,%s out of %s column/metricinar). One is a confirmedfrtranslation that PR 3 fixes (There was an issue deleting the selected %s).2. The backend exceptions are gone
I re-checked this way every backend entry that raises. On
masterall 37 raise. Here 36 of them format cleanly, and the 37th is the Slovak SQL Lab message that #44587 fixes.3. No translation regression
Confirmed-translation counts do not move in any of the 29 catalogs (
+0each). Only the fuzzy count drops, by exactly 1,189 across the 18 touched catalogs.4. In a running app
Enable the affected languages and compile the catalogs. You need
pybabel compile -d superset/translationsand a Jed 1.xmessages.jsonper locale, the filepo2json.shproduces. Released images ship neither, so the UI stays English until you build them.masterreturns 500Fatal error, English returns 422Dataset main.fruit already exists. With this PR, Dutch returns the same 422./users/userinfo/inpt_BR: the button reads{} SENHAonmasterand Redefinir minha senha with this PR.Three things that cost me time:
?_l_=<lang>or the session, never fromAccept-Language. An API probe that only setsAccept-Languageruns in English and everything looks fine./superset/language_pack/<lang>/and the browser caches the response, so swapping catalogs under a running app needs a cache-bypassing reload. A normal reload keeps showing the old strings.flask_babel.gettextraisesKeyError: 'statement_num'in a Slovak request context). But I could not trigger it by running a query through a released 6.1.0 image: that path logged the message in English while the same request returned Slovak error text elsewhere. I have not worked out why, so check that one at the catalog level rather than through the UI.ADDITIONAL INFORMATION
Order. #44551 also plans a CI check (PR 1) that fails when a translation's placeholders cannot format. It should land after this PR and PR 3 so it starts green. This PR takes the flagged set from 1,227 to 38. #44587 takes one of those, and PR 3 takes 8 more.
Conflicts. #43022 also edits
ptandpt_BR, and already conflicts with master. Whichever of the two lands second needs a rebase. I am happy to do that in either order.