What's wrong
The README promises "Case is always ignored", but letters outside the Basic Multilingual Plane (encoded as surrogate pairs) are compared case-sensitively.
Fuzzy.CodepointsEqual (FuzzySearch/Fuzzy.cs, ~line 333) handles 2-code-unit codepoints with a raw == on both halves:
return leftLength == 2
? left[leftIndex] == right[rightIndex] && left[leftIndex + 1] == right[rightIndex + 1]
: char.ToLowerInvariant(left[leftIndex]) == char.ToLowerInvariant(right[rightIndex]);
The remark explains this by saying char.ToLowerInvariant has no mapping for a surrogate. That is true for one char, but the codepoint itself does have a case mapping. For example, Rune.ToLowerInvariant(U+10400) returns U+10428.
Repro
In each case below, the subject is uppercase and the pattern is the same letters in lowercase. All of them should be True:
[π] ~ [π¨] => False -1 (Deseret U+10400 vs U+10428)
[π€] ~ [π€’] => False -1 (Adlam U+1E900 vs U+1E922, used to write Fulani)
[π°π±] ~ [ππ] => False -2 (Osage)
In practice, a user searching a list of Adlam or Osage names finds nothing unless they type the exact case. The same query in Latin script works.
This is separate from #88, which covers BMP letters with more than one lowercase form, such as sigma and the micro sign. A randomized check over a BMP alphabet (with the #88 letters excluded) found no case mismatches, so this gap is confined to the surrogate-pair branch.
Suggested fix
CodepointsEqual, netstandard2.1 / net5+: for leftLength == 2, compare Rune.ToLowerInvariant(Rune.GetRuneAt(...)) (or Rune.DecodeFromUtf16) on both sides.
CodepointsEqual, netstandard2.0: fall back to comparing char.ConvertFromUtf32(char.ConvertToUtf32(hi, lo)).ToLowerInvariant(), or compare the invariant-lowercased two-char strings.
- camelCase bonus: it reads case from the lead surrogate only (
strLower/strUpper around line 206), so supplementary-plane uppercase letters never earn the bonus. Consider applying the same Rune-based case test there.
- Remark: correct the XML remark on
CodepointsEqual.
Acceptance:
Fuzzy.Contains("π", "π¨") and Fuzzy.Contains("π¨", "π") both return true on every target framework.
- Their scores equal the scores for the exact-case match.
- Regression tests cover Deseret and Adlam.
What's wrong
The README promises "Case is always ignored", but letters outside the Basic Multilingual Plane (encoded as surrogate pairs) are compared case-sensitively.
Fuzzy.CodepointsEqual(FuzzySearch/Fuzzy.cs, ~line 333) handles 2-code-unit codepoints with a raw==on both halves:The remark explains this by saying
char.ToLowerInvarianthas no mapping for a surrogate. That is true for onechar, but the codepoint itself does have a case mapping. For example,Rune.ToLowerInvariant(U+10400)returnsU+10428.Repro
In each case below, the subject is uppercase and the pattern is the same letters in lowercase. All of them should be
True:In practice, a user searching a list of Adlam or Osage names finds nothing unless they type the exact case. The same query in Latin script works.
This is separate from #88, which covers BMP letters with more than one lowercase form, such as sigma and the micro sign. A randomized check over a BMP alphabet (with the #88 letters excluded) found no case mismatches, so this gap is confined to the surrogate-pair branch.
Suggested fix
CodepointsEqual, netstandard2.1 / net5+: forleftLength == 2, compareRune.ToLowerInvariant(Rune.GetRuneAt(...))(orRune.DecodeFromUtf16) on both sides.CodepointsEqual, netstandard2.0: fall back to comparingchar.ConvertFromUtf32(char.ConvertToUtf32(hi, lo)).ToLowerInvariant(), or compare the invariant-lowercased two-char strings.strLower/strUpperaround line 206), so supplementary-plane uppercase letters never earn the bonus. Consider applying the same Rune-based case test there.CodepointsEqual.Acceptance:
Fuzzy.Contains("π", "π¨")andFuzzy.Contains("π¨", "π")both returntrueon every target framework.