Skip to content

Count the prefix penalty in codepoints, not UTF-16 code units - #78

Merged
matt-edmondson merged 1 commit into
mainfrom
claude/prefix-penalty-codepoints
Sep 17, 2026
Merged

matt-edmondson merged 1 commit into
mainfrom
claude/prefix-penalty-codepoints

Conversation

@matt-edmondson

Copy link
Copy Markdown
Contributor

Fixes #77

The bug

CalculateScoreCore passed the raw UTF-16 offset strIdx to PenalizeNonPatternCharacters, which computes Math.Max(strIdx * unmatchedPrefixLetterPenalty, maxPrefixPenalty) — a formula that means its argument as a count of skipped characters. Since the codepoint-aware refactor in #71, the scoring loop advances strIdx by 1 or 2 code units per iteration depending on codepoint width, so a supplementary-plane prefix was charged twice per character.

Every other penalty in the function already runs once per codepoint; only this call site kept the code-unit offset.

Concretely, before this change:

subject pattern unmatched prefix chars score
😁😁😁y y 3 -8
abcy y 3 -6

Identical character counts, different scores — purely because the emoji prefix occupies 6 UTF-16 units rather than 3. That silently skews ranking order for any list containing emoji or CJK Extension B+ text ahead of the match.

The fix

Track a strCodepointIdx counter alongside strIdx in the scoring loop and pass that to PenalizeNonPatternCharacters. The parameter is renamed precedingCodepointCount so the contract states which unit it expects.

Behaviour for basic-plane text is unchanged — for BMP-only subjects the two counters are identical.

Tests

Two tests added to the surrogate-pair region of FuzzyTests.cs:

  • Contains_WithScore_SupplementaryPlanePrefix_PenalizedPerCharacterNotPerCodeUnit — the case from the issue: 😁😁😁y and abcy must score equally against y.
  • Contains_WithScore_SupplementaryPlanePrefix_DoesNotExceedTheUncappedPenalty — a two-character prefix, whose 4 code units stay under the maxPrefixPenalty cap, so the defect shows through rather than being masked by clamping.

Verified by reverting the one-line call-site change and re-running: both new tests fail (-8 vs -6, -4 vs -2) and the other 42 pass. With the fix restored, all 44 pass.

total: 44   failed: 0   succeeded: 44

An unrelated .gitignore update that the SDK wrote during the build was deliberately left out of this branch.

🤖 Generated with Claude Code

https://claude.ai/code/session_01QGbFXFrsMi1r2B64mxE2mp


Generated by Claude Code

…atch]

CalculateScoreCore passed the raw UTF-16 offset strIdx to
PenalizeNonPatternCharacters, but that formula multiplies its argument by
unmatchedPrefixLetterPenalty and so means it as a count of skipped
characters. Since the codepoint-aware refactor the loop advances strIdx by
1 or 2 code units per iteration, so a supplementary-plane prefix was
penalized twice over: "😁😁😁y" matching "y" scored -8 where the
character-count-equivalent "abcy" scored -6.

Track a codepoint counter alongside strIdx and pass that instead, and
rename the parameter to precedingCodepointCount so the contract states
which unit it wants.

Fixes #77

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QGbFXFrsMi1r2B64mxE2mp
@sonarqubecloud

Copy link
Copy Markdown

@matt-edmondson
matt-edmondson merged commit ddcc79f into main Sep 17, 2026
12 checks passed
@matt-edmondson
matt-edmondson deleted the claude/prefix-penalty-codepoints branch September 17, 2026 00:08
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Prefix penalty is computed from UTF-16 code-unit offset, not codepoint count, over-penalizing supplementary-plane prefixes

2 participants