Skip to content

Prefix penalty is computed from UTF-16 code-unit offset, not codepoint count, over-penalizing supplementary-plane prefixes #77

Description

@matt-edmondson

What's wrong

Fuzzy.CalculateScoreCore (Fuzzy.cs:213) calls PenalizeNonPatternCharacters(score, patternIdx, strIdx) when the first pattern character matches, and PenalizeNonPatternCharacters (Fuzzy.cs:395) computes Math.Max(strIdx * unmatchedPrefixLetterPenalty, maxPrefixPenalty). Since the surrogate-pair fix in #71, the surrounding loop advances strIdx by 1 or 2 UTF-16 code units per iteration depending on codepoint width (CodepointLengthAt), but this penalty formula still treats strIdx as a count of skipped characters. Every other unmatched-character penalty in the function (the unmatchedLetterPenalty branch) correctly runs once per codepoint regardless of width — only this prefix-penalty call site was missed when the codepoint-aware refactor went in.

Why it matters (concrete failure scenario)

Fuzzy.Contains("😁😁😁y", "y", out int emojiScore) yields emojiScore == -8, while the character-count-equivalent Fuzzy.Contains("abcy", "y", out int asciiScore) yields asciiScore == -6 (hand-traced through the current algorithm). Both subjects have exactly 3 unmatched characters before the match, but the emoji case is penalized 2 points more purely because its 3 codepoints occupy 6 UTF-16 units, not 3. This is a pure implementation-encoding leak into scores meant to reflect character counts, and it will silently skew ranking/sort order for any UI list containing supplementary-plane text (emoji, some CJK Extension B+ characters, etc.) with unmatched lead-in text.

Suggested fix / acceptance criteria

Track a codepoint count (or reuse patternIdx's counterpart — a running codepoint index) alongside strIdx and pass that to PenalizeNonPatternCharacters instead of the raw UTF-16 offset. Add a test asserting Fuzzy.Contains("😁😁😁y", "y", out int score) and Fuzzy.Contains("abcy", "y", out int score2) produce equal scores.

Files: FuzzySearch/Fuzzy.cs:213 (call site), FuzzySearch/Fuzzy.cs:395 (PenalizeNonPatternCharacters)

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Labels

No labels
No labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions