What's wrong
Fuzzy.CalculateScoreCore (Fuzzy.cs:213) calls PenalizeNonPatternCharacters(score, patternIdx, strIdx) when the first pattern character matches, and PenalizeNonPatternCharacters (Fuzzy.cs:395) computes Math.Max(strIdx * unmatchedPrefixLetterPenalty, maxPrefixPenalty). Since the surrogate-pair fix in #71, the surrounding loop advances strIdx by 1 or 2 UTF-16 code units per iteration depending on codepoint width (CodepointLengthAt), but this penalty formula still treats strIdx as a count of skipped characters. Every other unmatched-character penalty in the function (the unmatchedLetterPenalty branch) correctly runs once per codepoint regardless of width — only this prefix-penalty call site was missed when the codepoint-aware refactor went in.
Why it matters (concrete failure scenario)
Fuzzy.Contains("😁😁😁y", "y", out int emojiScore) yields emojiScore == -8, while the character-count-equivalent Fuzzy.Contains("abcy", "y", out int asciiScore) yields asciiScore == -6 (hand-traced through the current algorithm). Both subjects have exactly 3 unmatched characters before the match, but the emoji case is penalized 2 points more purely because its 3 codepoints occupy 6 UTF-16 units, not 3. This is a pure implementation-encoding leak into scores meant to reflect character counts, and it will silently skew ranking/sort order for any UI list containing supplementary-plane text (emoji, some CJK Extension B+ characters, etc.) with unmatched lead-in text.
Suggested fix / acceptance criteria
Track a codepoint count (or reuse patternIdx's counterpart — a running codepoint index) alongside strIdx and pass that to PenalizeNonPatternCharacters instead of the raw UTF-16 offset. Add a test asserting Fuzzy.Contains("😁😁😁y", "y", out int score) and Fuzzy.Contains("abcy", "y", out int score2) produce equal scores.
Files: FuzzySearch/Fuzzy.cs:213 (call site), FuzzySearch/Fuzzy.cs:395 (PenalizeNonPatternCharacters)
What's wrong
Fuzzy.CalculateScoreCore(Fuzzy.cs:213) callsPenalizeNonPatternCharacters(score, patternIdx, strIdx)when the first pattern character matches, andPenalizeNonPatternCharacters(Fuzzy.cs:395) computesMath.Max(strIdx * unmatchedPrefixLetterPenalty, maxPrefixPenalty). Since the surrogate-pair fix in #71, the surrounding loop advancesstrIdxby 1 or 2 UTF-16 code units per iteration depending on codepoint width (CodepointLengthAt), but this penalty formula still treatsstrIdxas a count of skipped characters. Every other unmatched-character penalty in the function (theunmatchedLetterPenaltybranch) correctly runs once per codepoint regardless of width — only this prefix-penalty call site was missed when the codepoint-aware refactor went in.Why it matters (concrete failure scenario)
Fuzzy.Contains("😁😁😁y", "y", out int emojiScore)yieldsemojiScore == -8, while the character-count-equivalentFuzzy.Contains("abcy", "y", out int asciiScore)yieldsasciiScore == -6(hand-traced through the current algorithm). Both subjects have exactly 3 unmatched characters before the match, but the emoji case is penalized 2 points more purely because its 3 codepoints occupy 6 UTF-16 units, not 3. This is a pure implementation-encoding leak into scores meant to reflect character counts, and it will silently skew ranking/sort order for any UI list containing supplementary-plane text (emoji, some CJK Extension B+ characters, etc.) with unmatched lead-in text.Suggested fix / acceptance criteria
Track a codepoint count (or reuse
patternIdx's counterpart — a running codepoint index) alongsidestrIdxand pass that toPenalizeNonPatternCharactersinstead of the raw UTF-16 offset. Add a test assertingFuzzy.Contains("😁😁😁y", "y", out int score)andFuzzy.Contains("abcy", "y", out int score2)produce equal scores.Files:
FuzzySearch/Fuzzy.cs:213(call site),FuzzySearch/Fuzzy.cs:395(PenalizeNonPatternCharacters)