Skip to content

Case-insensitive matching fails for supplementary-plane letters: Contains("𐐀", "𐐨") is false (Deseret, Osage, Adlam, …)Β #90

Description

@matt-edmondson

What's wrong

The README promises "Case is always ignored", but letters outside the Basic Multilingual Plane (encoded as surrogate pairs) are compared case-sensitively.

Fuzzy.CodepointsEqual (FuzzySearch/Fuzzy.cs, ~line 333) handles 2-code-unit codepoints with a raw == on both halves:

return leftLength == 2
    ? left[leftIndex] == right[rightIndex] && left[leftIndex + 1] == right[rightIndex + 1]
    : char.ToLowerInvariant(left[leftIndex]) == char.ToLowerInvariant(right[rightIndex]);

The remark explains this by saying char.ToLowerInvariant has no mapping for a surrogate. That is true for one char, but the codepoint itself does have a case mapping. For example, Rune.ToLowerInvariant(U+10400) returns U+10428.

Repro

In each case below, the subject is uppercase and the pattern is the same letters in lowercase. All of them should be True:

[𐐀]   ~ [𐐨]   => False -1   (Deseret U+10400 vs U+10428)
[πž€€]   ~ [𞀒]   => False -1   (Adlam U+1E900 vs U+1E922, used to write Fulani)
[𐒰𐒱] ~ [π“˜π“™] => False -2   (Osage)

In practice, a user searching a list of Adlam or Osage names finds nothing unless they type the exact case. The same query in Latin script works.

This is separate from #88, which covers BMP letters with more than one lowercase form, such as sigma and the micro sign. A randomized check over a BMP alphabet (with the #88 letters excluded) found no case mismatches, so this gap is confined to the surrogate-pair branch.

Suggested fix

  • CodepointsEqual, netstandard2.1 / net5+: for leftLength == 2, compare Rune.ToLowerInvariant(Rune.GetRuneAt(...)) (or Rune.DecodeFromUtf16) on both sides.
  • CodepointsEqual, netstandard2.0: fall back to comparing char.ConvertFromUtf32(char.ConvertToUtf32(hi, lo)).ToLowerInvariant(), or compare the invariant-lowercased two-char strings.
  • camelCase bonus: it reads case from the lead surrogate only (strLower/strUpper around line 206), so supplementary-plane uppercase letters never earn the bonus. Consider applying the same Rune-based case test there.
  • Remark: correct the XML remark on CodepointsEqual.

Acceptance:

  • Fuzzy.Contains("𐐀", "𐐨") and Fuzzy.Contains("𐐨", "𐐀") both return true on every target framework.
  • Their scores equal the scores for the exact-case match.
  • Regression tests cover Deseret and Adlam.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't workingreadyFully specified; implement as written

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions