Skip to content

fix(i18n): clear 1,188 fuzzy translations with broken placeholders - #44611

Merged
rusackas merged 3 commits into
apache:masterfrom
glaterza:fix/i18n-clear-broken-placeholder-fuzzies
Sep 29, 2026
Merged

rusackas merged 3 commits into
apache:masterfrom
glaterza:fix/i18n-clear-broken-placeholder-fuzzies

Conversation

@glaterza

@glaterza glaterza commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor

SUMMARY

Updated after review (ff916530bc). @sadpandajoe caught two problems, both fixed:

  • The Persian Modified 1 column in the virtual dataset entry was cleared by mistake, and is restored. Persian sends both 0 and 1 to the first plural form, which writes the count as the Persian digit ۱. The check read that as a missing value at n=1. The call site never passes 0, so the translation is correct. I rechecked the other 51 cleared plural entries by each language's plural rule: every one drops the count for n ≥ 2.
  • Removing the attribution comments left the second (and sometimes third) line of each wrapped comment behind: 113 stray lines on 96 entries. They are gone.

The count is 1,188. The numbers under TESTING INSTRUCTIONS were measured at 2ac6e2468d, before #44587 merged. On current master, pybabel compile --use-fuzzy reports 1,089 errors and this branch 154. The one-error drop on each side is #44587's Slovak fix.

This is PR 2 of the plan in #44551. @rusackas approved the approach there:

On PR 2: clearing the entries that error out or drop values seems right to me, a translation that crashes SQL Lab is worse than falling back to English. Go ahead and open it.

It clears 1,188 fuzzy translations in 18 catalogs. Their placeholders do not match the English source, so none of them can format correctly. Clearing msgstr makes each entry untranslated, so the UI shows the English text with its values instead of a crash or a literal %s.

This PR does not touch the Slovak SQL Lab entry. #44587 rewrites that one properly and promotes it out of fuzzy, which is a better outcome than falling back to English. I left it out so the two changes stay independent and neither can undo the other.

Confirmed (non-fuzzy) translations are not touched. The 8 broken confirmed entries (4 mi, 2 es, 1 fr, 1 nl) are PR 3. They are fixed by hand so native speakers can review the wording.

What gets cleared

Effect at runtime Entries
Raises an exception (backend KeyError / TypeError) 36
Drops a value: a name, a count, an error detail 838
Shows a blank or wrong value 68
Shows a literal placeholder, such as Error: %s 26
Shows nothing wrong: Translator catches the error and returns the English key 220
Total 1,188

The last row needs a word of caution, because it is easy to overstate this bug. Translator.translate wraps fetch in a try/catch and returns the input when formatting fails. So a broken frontend entry only shows a %s on screen if the English source itself has a placeholder. 26 of them do, and the screenshots below are one of those 26. The other 220 already fall back to English today.

They are still worth clearing. They are translations that can never render. They survive only because of a catch, and they count as "translated" in every coverage number.

Why these entries are live

Superset serves fuzzy translations on purpose, so a broken fuzzy entry ships like any other:

  • pybabel compile --use-fuzzy reports 1,090 errors, exits 1, and still writes all 30 .mo files. The Dockerfile ends that command with || true, and flask fab babel-compile does not check the exit status.
  • superset-frontend/scripts/po2json.sh passes --fuzzy, so the frontend packs carry them too.

Three things to look at

  1. 30 flagged entries are left alone: 18 fail only when a count is 0, which no call site passes, and 12 render correctly in practice. 27 of those 30 are fuzzy, so a simple "fuzzy and flagged" filter would have cleared entries that work. The selection uses the runtime outcome, not the static check.
  2. 96 entries had a # Machine-translated via backfill_po.py (...) comment. The comment described the translation being removed, so it goes too. backfill_po.py picks entries by empty msgstr, so these are now clean candidates for a future backfill.
  3. The diff only touches those 1,188 entries. I edited the .po text directly instead of rewriting the files with Babel or polib, because both reformat unrelated entries. Babel adds a python-format flag next to no-python-format; polib rewraps thousands of lines. I compared every entry in all 29 catalogs against master: msgid sets, headers, locations, extracted comments and translator comments are identical everywhere else. Each changed entry is one of the 1,188, and its only changes are the cleared msgstr, the dropped fuzzy flag and, where present, the whole attribution comment.

BEFORE/AFTER SCREENSHOTS OR ANIMATED GIF

Duplicating a role in Dutch. Same build, same locale. Only the catalogs differ.

Before. The dialog title reads Duplicate role %(name)s. The nl translation is Dubbele kolomnaam (of -namen): %(columns)s, which asks for a columns value the call site does not pass, so formatting throws, Translator falls back to the English key, and the key's own placeholder is left on screen.

before

After. The title reads Duplicate role Admin. The entry is untranslated, so the English source is used and the role name lands where it belongs.

after

The same thing on a server-rendered page, in pt_BR.

Before. The Reset my password button reads {} SENHA. The catalog has msgstr "%s SENHA" for a source string with no placeholder, and Flask-AppBuilder's label handling turns the %s into {}.

before

After. The button reads Redefinir minha senha. Flask-AppBuilder has a correct translation for this string in its own catalog, and Superset's broken entry was overriding it. Clearing ours brings back the right Portuguese.

after

A second case, through the API. Creating a dataset that already exists, in Dutch:

master this PR
nl HTTP 500 {"message": "Fatal error"} HTTP 422 {"table": ["Dataset main.fruit already exists"]}
en (control) HTTP 422 {"table": ["Dataset main.fruit already exists"]} same

Dataset %(table)s already exists is translated with %(name)s in nl. Building the validation message raises KeyError, so a clean 422 becomes a 500.

TESTING INSTRUCTIONS

Measured at master 2ac6e2468d. The flagged set is the same one reported in #44551 at 269b9f99fe: 1,227 entries, no drift.

1. Compile errors drop

cp -r superset/translations /tmp/after
pybabel compile --use-fuzzy -d /tmp/after > /tmp/after.log 2>&1; grep -c '^error:' /tmp/after.log
# master: 1090      this PR: 155

The 155 left are a different problem: pybabel's compile check is stricter than the runtime. 98 are strings with a prose percent sign and no real format code (% of total, % calculation; most already carry no-python-format), 44 are plural forms where one form spells its single number out, and 13 are others.

Only 8 of the 155 match an entry the runtime check still flags, and I can name them. One is the Slovak SQL Lab message #44587 fixes. Six are the "fails only when the count is 0" plural forms left alone on purpose (%s day ago and %s min ago in pl and pt_BR, %s out of %s column/metric in ar). One is a confirmed fr translation that PR 3 fixes (There was an issue deleting the selected %s).

2. The backend exceptions are gone

python -c '
from babel.support import Translations
t = Translations.load("/tmp/after", ["nl"], "messages")
print(t.ugettext("Dataset %(table)s already exists") % {"table": "fruit"})
'
# master: KeyError: 'name'      this PR: Dataset fruit already exists

I re-checked this way every backend entry that raises. On master all 37 raise. Here 36 of them format cleanly, and the 37th is the Slovak SQL Lab message that #44587 fixes.

3. No translation regression

python scripts/translations/check_translation_regression.py --count \
  --translations-dir /path/to/master/superset/translations > /tmp/before.json
python scripts/translations/check_translation_regression.py --compare /tmp/before.json
# "No translation regressions."

Confirmed-translation counts do not move in any of the 29 catalogs (+0 each). Only the fuzzy count drops, by exactly 1,189 across the 18 touched catalogs.

4. In a running app

Enable the affected languages and compile the catalogs. You need pybabel compile -d superset/translations and a Jed 1.x messages.json per locale, the file po2json.sh produces. Released images ship neither, so the UI stays English until you build them.

  • Create a dataset that already exists, in Dutch: master returns 500 Fatal error, English returns 422 Dataset main.fruit already exists. With this PR, Dutch returns the same 422.
  • Open /users/userinfo/ in pt_BR: the button reads {} SENHA on master and Redefinir minha senha with this PR.

Three things that cost me time:

  • Flask-AppBuilder picks the locale from ?_l_=<lang> or the session, never from Accept-Language. An API probe that only sets Accept-Language runs in English and everything looks fine.
  • The frontend fetches /superset/language_pack/<lang>/ and the browser caches the response, so swapping catalogs under a running app needs a cache-bypassing reload. A normal reload keeps showing the old strings.
  • The Slovak SQL Lab failure reproduces at the catalog level every time (flask_babel.gettext raises KeyError: 'statement_num' in a Slovak request context). But I could not trigger it by running a query through a released 6.1.0 image: that path logged the message in English while the same request returned Slovak error text elsewhere. I have not worked out why, so check that one at the catalog level rather than through the UI.

ADDITIONAL INFORMATION

Order. #44551 also plans a CI check (PR 1) that fails when a translation's placeholders cannot format. It should land after this PR and PR 3 so it starts green. This PR takes the flagged set from 1,227 to 38. #44587 takes one of those, and PR 3 takes 8 more.

Conflicts. #43022 also edits pt and pt_BR, and already conflicts with master. Whichever of the two lands second needs a rebase. I am happy to do that in either order.

@bito-code-review

bito-code-review Bot commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor

Bito Automatic Review Skipped - Large PR

Bito didn't auto-review this change because the pull request exceeded the line limit. No action is needed if you didn't intend for the agent to review it. Otherwise, to manually trigger a review, type /review in a comment and save.

@github-actions github-actions Bot added i18n Namespace | Anything related to localization i18n:chinese Translation related to Chinese language i18n:korean Translation related to Korean language i18n:dutch i18n:slovak i18n:ukrainian i18n:brazilian i18n:traditional-chinese i18n:persian i18n:czech i18n:latvian labels Sep 24, 2026
These translations carry placeholders that do not match their English
source, so none of them can format. At runtime they raise an exception,
show a literal placeholder such as "Error: %s", drop a value, show a blank
one, or fail silently because the frontend Translator catches the error and
returns the English key.

Superset serves fuzzy translations on purpose, so a broken fuzzy entry is
as live as a confirmed one: `pybabel compile --use-fuzzy` reports 1,090
errors and still writes every .mo file, and po2json.sh passes --fuzzy.

Clearing the msgstr makes each entry untranslated, so the UI falls back to
the English source with its values. Confirmed (non-fuzzy) translations are
untouched; the 8 broken confirmed entries are fixed by hand separately.

The Slovak SQL Lab progress message is deliberately left alone: apache#44587
rewrites that entry properly and promotes it out of fuzzy, so this change
stays clear of it.

The 97 entries that carried a `Machine-translated via backfill_po.py`
attribution lose it along with the translation it described, so they read
as plainly untranslated and backfill_po.py sees clean candidates.

Measured at master 2ac6e24, across 18 catalogs:

- compile errors under --use-fuzzy: 1,090 -> 155
- 36 of the 37 entries that raised now format cleanly (the 37th is apache#44587's)
- placeholder checker findings: 1,227 -> 38
- check_translation_regression.py: no regressions; confirmed-translation
  counts unchanged in all 29 catalogs, fuzzy down exactly 1,189

Refs: apache#44551
@netlify

netlify Bot commented Sep 24, 2026 •

Copy link
Copy Markdown

✅ Deploy Preview for superset-docs-preview ready!

Name Link
🔨 Latest commit 47ee7ab
🔍 Latest deploy log https://app.netlify.com/projects/superset-docs-preview/deploys/6abb40cab59a3f0008a97640
😎 Deploy Preview https://deploy-preview-44611--superset-docs-preview.netlify.app
📱 Preview on mobile
Toggle QR Code...

QR Code

Use your smartphone camera to open QR code link.
🤖 Make changes Run an agent on this branch

To edit notification comments on pull requests, go to your Netlify project configuration.

@glaterza
glaterza force-pushed the fix/i18n-clear-broken-placeholder-fuzzies branch from a47bfcb to 4569826 Compare September 24, 2026 15:12
@glaterza glaterza changed the title fix(i18n): clear 1,190 fuzzy translations with broken placeholders fix(i18n): clear 1,189 fuzzy translations with broken placeholders Sep 24, 2026
@codecov

codecov Bot commented Sep 24, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 81.50%. Comparing base (b4f79f9) to head (ff91653).
⚠️ Report is 16 commits behind head on master.

Additional details and impacted files
@@            Coverage Diff             @@
##           master   #44611      +/-   ##
==========================================
- Coverage   81.51%   81.50%   -0.01%     
==========================================
  Files        2973     2973              
  Lines      180202   180280      +78     
  Branches    41720    41737      +17     
==========================================
+ Hits       146883   146945      +62     
- Misses      30602    30612      +10     
- Partials     2717     2723       +6     
Flag Coverage Δ
hive 36.70% <ø> (-0.01%) ⬇️
mysql 55.99% <ø> (-0.03%) ⬇️
postgres 55.99% <ø> (-0.03%) ⬇️
presto 38.62% <ø> (-0.01%) ⬇️
python 85.65% <ø> (-0.01%) ⬇️
sqlite 55.72% <ø> (-0.03%) ⬇️
unit 78.38% <ø> (+<0.01%) ⬆️

Flags with carried forward coverage won't be shown. Click here to find out more.

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

@rusackas rusackas left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is exactly the plan from #44551, clearing the ones that can't format cleanly instead of letting them crash SQL Lab. Spot-checked a few catalogs and the diff only touches the fuzzy entries called out, confirmed translations are untouched. LGTM, approving!

@rusackas rusackas added the merge-if-green If approved and tests are green, please go ahead and merge it for me label Sep 29, 2026
msgstr[0] "۱ ستون در مجموعه داده مجازی تغییر یافت"
msgstr[1] "%s ستون در مجموعه داده مجازی تغییر یافت"
msgstr[0] ""
msgstr[1] ""

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Clearing this drops a translation whose placeholders actually match: msgid "Modified 1 column in the virtual dataset" carries no placeholder, and the removed msgstr[0] ("۱ ستون در مجموعه داده مجازی تغییر یافت") had none either; msgid_plural "Modified %s columns in the virtual dataset" carries %s, and the removed msgstr[1] ("%s ستون در مجموعه داده مجازی تغییر یافت") had it too. Persian users modifying dataset columns now see the English fallback instead of a previously-correct translation.

The same pattern shows up at superset/translations/zh_TW/LC_MESSAGES/messages.po:1654 ("Added to 1 dashboard", where the removed msgstr[0] "增加到看板" also had no placeholder mismatch). Since the same detection script produced both misses, could you restore these two entries and recheck the rest of the 52 cleared plural entries for the same false positive before merging?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Good catch on the Persian entry. It is restored in ff91653. Persian's plural rule is n > 1, so msgstr[0] serves both 0 and 1, and it writes the count as the Persian digit ۱. The check looked for the value at n=1 and didn't recognise ۱ as it, so it counted the entry as broken at 0 and 1. That kept it out of the "fails only at 0" exclusion. The call site only toasts when columnChanges.modified.length is non-zero, so the translation is correct in practice.

zh_TW is different, and I'd keep it cleared. Its header is nplurals=1; plural=0;, so msgstr[0] "增加到看板" is used for every count. "Added to 5 dashboards" renders as "Added to dashboard", with the number gone. That's the "drops a value" case from #44551.

I rechecked all 52 cleared plural entries by each language's plural rule (n = 0, 1, 2, 3, 5, 11, 21, 22, 101, 111). Persian was the only one that fails just at n=0. The other 51 each drop the count for some n ≥ 2. The PR is 1,188 entries now; the title and description are updated.

# Machine-translated via backfill_po.py (claude-sonnet-4-6) [refs: ca, cs, de,
# es, fr, ja, lv, mi, ro, ru, sk, sr, sr_Latn, tr, uk]
#, fuzzy, python-format
#, python-format

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Removing the Machine-translated via backfill_po.py attribution above only deleted the first line of the wrapped two-line comment, leaving this now-meaningless continuation attached to the entry. The same one-line-only removal repeats for all 97 comment removals in this PR across the touched catalogs, leaving 97 orphaned reference-list fragments behind. Could this line (and the equivalent stray line in the other 96 spots) also be deleted?

Suggested change
#, python-format

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

You're right, and thanks for spotting it. My comparison checked msgids, headers, locations and extracted comments, but not translator comments, so the stray lines slipped through. Fixed in ff91653: 113 continuation lines on 96 entries (some attributions wrapped to three lines). I matched every entry against master to remove exactly the lines that continued each attribution there. The comparison now also covers translator comments. It flags all 97 leftovers on the previous head and none on this one. The Persian entry that's restored keeps its attribution, since its translation stays.

Two corrections from review.

The Persian "Modified 1 column in the virtual dataset" entry was cleared by
mistake. Persian's plural rule (n > 1) sends both 0 and 1 to msgstr[0], which
writes the count as the Persian digit; the check read that as a missing value
at n=1 and so did not treat it as a zero-only case. The call site only shows
the message when at least one column changed, so n is never 0 and the
translation is correct. It is restored as it was. The other 51 cleared plural
entries were rechecked by plural rule: each drops the count for n >= 2.

Removing the `# Machine-translated via backfill_po.py` attributions deleted
only the first line of each wrapped comment and left 96 continuation lines
(113 lines, some comments wrapped to three lines) attached to the cleared
entries. They are removed.

1,188 entries are cleared.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@glaterza glaterza changed the title fix(i18n): clear 1,189 fuzzy translations with broken placeholders fix(i18n): clear 1,188 fuzzy translations with broken placeholders Sep 29, 2026
@rusackas
rusackas merged commit cba2574 into apache:master Sep 29, 2026
74 checks passed
villebro pushed a commit that referenced this pull request Sep 30, 2026
…44611)

Co-authored-by: Evan Rusackas <evan@preset.io>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
(cherry picked from commit cba2574)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

i18n:brazilian i18n:chinese Translation related to Chinese language i18n:czech i18n:dutch i18n:korean Translation related to Korean language i18n:latvian i18n:persian i18n:slovak i18n:traditional-chinese i18n:ukrainian i18n Namespace | Anything related to localization merge-if-green If approved and tests are green, please go ahead and merge it for me size/XXL

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants