What's wrong
NaturalStringComparer.cs detects a "numeric" chunk via char.IsDigit(...) and a regex \d+ (line 16), both of which match any Unicode Nd-category digit, not just ASCII 0-9. But CompareNumericChunks (lines 83-104) normalizes leading zeros with TrimStart('0') (ASCII zero only) and falls back to string.Compare(..., StringComparison.Ordinal) — both of which assume the digit characters are ASCII code points 0x30-0x39. When a "numeric" chunk contains a non-ASCII digit, the comparison silently degrades to a raw UTF-16 code-point comparison, which does not reflect numeric magnitude.
Why it matters (concrete failure scenario)
Verified directly against the compiled class:
Compare("٥", "9") (Arabic-Indic digit for 5 vs. ASCII 9) returns positive — the comparer ranks the digit "5" as greater than "9", the opposite of the numeric order the class's own XML doc promises ("embedded numbers are compared as numeric values").
Compare("٠0", "0") (Arabic-Indic zero followed by ASCII zero, vs. plain ASCII "0") also returns positive, because TrimStart('0') doesn't strip the non-ASCII zero, so the leading-zero normalization the numeric-length comparison relies on doesn't apply.
Existing tests (Compare_StringsWithLeadingZeros_HandledCorrectly, Compare_LargeNumbers_HandledCorrectly) only exercise ASCII digits, so this is currently invisible to CI.
Suggested fix / acceptance criteria
Either restrict "numeric chunk" detection to ASCII digits only (e.g. check value[0] is >= '0' and <= '9' instead of char.IsDigit, or constrain the regex to [0-9]+), or, if Unicode digit scripts should genuinely be supported, normalize each digit character to its numeric value (CharUnicodeInfo.GetDecimalDigitValue) before comparing instead of doing raw ordinal/length comparison. Acceptance: a test with at least one non-ASCII digit input sorts consistently with its actual numeric value.
File: Sorting/NaturalStringComparer.cs (lines 61, 83-104)
What's wrong
NaturalStringComparer.csdetects a "numeric" chunk viachar.IsDigit(...)and a regex\d+(line 16), both of which match any UnicodeNd-category digit, not just ASCII0-9. ButCompareNumericChunks(lines 83-104) normalizes leading zeros withTrimStart('0')(ASCII zero only) and falls back tostring.Compare(..., StringComparison.Ordinal)— both of which assume the digit characters are ASCII code points 0x30-0x39. When a "numeric" chunk contains a non-ASCII digit, the comparison silently degrades to a raw UTF-16 code-point comparison, which does not reflect numeric magnitude.Why it matters (concrete failure scenario)
Verified directly against the compiled class:
Compare("٥", "9")(Arabic-Indic digit for 5 vs. ASCII 9) returns positive — the comparer ranks the digit "5" as greater than "9", the opposite of the numeric order the class's own XML doc promises ("embedded numbers are compared as numeric values").Compare("٠0", "0")(Arabic-Indic zero followed by ASCII zero, vs. plain ASCII "0") also returns positive, becauseTrimStart('0')doesn't strip the non-ASCII zero, so the leading-zero normalization the numeric-length comparison relies on doesn't apply.Existing tests (
Compare_StringsWithLeadingZeros_HandledCorrectly,Compare_LargeNumbers_HandledCorrectly) only exercise ASCII digits, so this is currently invisible to CI.Suggested fix / acceptance criteria
Either restrict "numeric chunk" detection to ASCII digits only (e.g. check
value[0] is >= '0' and <= '9'instead ofchar.IsDigit, or constrain the regex to[0-9]+), or, if Unicode digit scripts should genuinely be supported, normalize each digit character to its numeric value (CharUnicodeInfo.GetDecimalDigitValue) before comparing instead of doing raw ordinal/length comparison. Acceptance: a test with at least one non-ASCII digit input sorts consistently with its actual numeric value.File:
Sorting/NaturalStringComparer.cs(lines 61, 83-104)