What happens
Fuzzy.Contains/CalculateScore (FuzzySearch/Fuzzy.cs, lines 46-69 and 99-200) compare subject and pattern per UTF-16 char with no awareness of surrogate pairs. A lone/unpaired high surrogate in the pattern is treated as an ordinary matchable character, so it can match the high-surrogate half of an unrelated supplementary-plane character (e.g. an emoji) in the subject.
Reproduction
// subject contains 😁 (U+1F601, surrogate pair 😁)
Fuzzy.Contains("x😁y", "\uD83D"); // → true
The lone high surrogate \uD83D (not a valid standalone character) is reported as found inside the subject purely because it matches half of the emoji's surrogate pair.
Why it matters
This is a narrow edge case — it requires a malformed/truncated pattern or subject, which can happen when supplementary-plane text elsewhere is naively Substring'd or truncated and split a surrogate pair. It's a genuine correctness gap shared by many naive char-by-char string matchers, and worth tracking as a known limitation even though it's lower priority than whole-codepoint normalization issues.
Suggested fix
When enumerating characters for matching, iterate by Unicode codepoint (e.g. via System.Globalization.StringInfo / rune enumeration, or by detecting and matching surrogate pairs as a unit) rather than by raw UTF-16 char, so a lone surrogate can't spuriously match half of an unrelated supplementary-plane character.
Acceptance criteria
- A lone/unpaired surrogate in the pattern does not match against half of a valid surrogate pair in the subject.
- Add a test covering a pattern/subject pair with a supplementary-plane character (e.g. an emoji) adjacent to other matchable characters.
What happens
Fuzzy.Contains/CalculateScore(FuzzySearch/Fuzzy.cs, lines 46-69 and 99-200) compare subject and pattern per UTF-16charwith no awareness of surrogate pairs. A lone/unpaired high surrogate in the pattern is treated as an ordinary matchable character, so it can match the high-surrogate half of an unrelated supplementary-plane character (e.g. an emoji) in the subject.Reproduction
The lone high surrogate
\uD83D(not a valid standalone character) is reported as found inside the subject purely because it matches half of the emoji's surrogate pair.Why it matters
This is a narrow edge case — it requires a malformed/truncated pattern or subject, which can happen when supplementary-plane text elsewhere is naively
Substring'd or truncated and split a surrogate pair. It's a genuine correctness gap shared by many naive char-by-char string matchers, and worth tracking as a known limitation even though it's lower priority than whole-codepoint normalization issues.Suggested fix
When enumerating characters for matching, iterate by Unicode codepoint (e.g. via
System.Globalization.StringInfo/ rune enumeration, or by detecting and matching surrogate pairs as a unit) rather than by raw UTF-16char, so a lone surrogate can't spuriously match half of an unrelated supplementary-plane character.Acceptance criteria