Count the prefix penalty in codepoints, not UTF-16 code units - #78
Merged
Merged
Conversation
…atch] CalculateScoreCore passed the raw UTF-16 offset strIdx to PenalizeNonPatternCharacters, but that formula multiplies its argument by unmatchedPrefixLetterPenalty and so means it as a count of skipped characters. Since the codepoint-aware refactor the loop advances strIdx by 1 or 2 code units per iteration, so a supplementary-plane prefix was penalized twice over: "😁😁😁y" matching "y" scored -8 where the character-count-equivalent "abcy" scored -6. Track a codepoint counter alongside strIdx and pass that instead, and rename the parameter to precedingCodepointCount so the contract states which unit it wants. Fixes #77 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QGbFXFrsMi1r2B64mxE2mp
|
This was referenced Sep 17, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.



Fixes #77
The bug
CalculateScoreCorepassed the raw UTF-16 offsetstrIdxtoPenalizeNonPatternCharacters, which computesMath.Max(strIdx * unmatchedPrefixLetterPenalty, maxPrefixPenalty)— a formula that means its argument as a count of skipped characters. Since the codepoint-aware refactor in #71, the scoring loop advancesstrIdxby 1 or 2 code units per iteration depending on codepoint width, so a supplementary-plane prefix was charged twice per character.Every other penalty in the function already runs once per codepoint; only this call site kept the code-unit offset.
Concretely, before this change:
😁😁😁yyabcyyIdentical character counts, different scores — purely because the emoji prefix occupies 6 UTF-16 units rather than 3. That silently skews ranking order for any list containing emoji or CJK Extension B+ text ahead of the match.
The fix
Track a
strCodepointIdxcounter alongsidestrIdxin the scoring loop and pass that toPenalizeNonPatternCharacters. The parameter is renamedprecedingCodepointCountso the contract states which unit it expects.Behaviour for basic-plane text is unchanged — for BMP-only subjects the two counters are identical.
Tests
Two tests added to the surrogate-pair region of
FuzzyTests.cs:Contains_WithScore_SupplementaryPlanePrefix_PenalizedPerCharacterNotPerCodeUnit— the case from the issue:😁😁😁yandabcymust score equally againsty.Contains_WithScore_SupplementaryPlanePrefix_DoesNotExceedTheUncappedPenalty— a two-character prefix, whose 4 code units stay under themaxPrefixPenaltycap, so the defect shows through rather than being masked by clamping.Verified by reverting the one-line call-site change and re-running: both new tests fail (
-8vs-6,-4vs-2) and the other 42 pass. With the fix restored, all 44 pass.An unrelated
.gitignoreupdate that the SDK wrote during the build was deliberately left out of this branch.🤖 Generated with Claude Code
https://claude.ai/code/session_01QGbFXFrsMi1r2B64mxE2mp
Generated by Claude Code