Skip to content

fix: compare non-ASCII digits by numeric value in NaturalStringComparer [patch] - #51

Merged
matt-edmondson merged 1 commit into
mainfrom
fix/natural-comparer-unicode-digits
Sep 22, 2026
Merged

matt-edmondson merged 1 commit into
mainfrom
fix/natural-comparer-unicode-digits

Conversation

@matt-edmondson

Copy link
Copy Markdown
Contributor

Fixes #50

The bug

NaturalStringComparer detects a "numeric" chunk with char.IsDigit(...) and the regex \d+ — both of which match any Unicode Nd digit, not just ASCII 0-9. But CompareNumericChunks then normalized leading zeros with TrimStart('0') (ASCII zero only) and fell back to string.Compare(..., StringComparison.Ordinal), both of which assume the digits are ASCII code points 0x30-0x39.

So a chunk containing a non-ASCII digit silently degraded to a raw UTF-16 code-point comparison, which has nothing to do with numeric magnitude:

Comparison Before Expected
Compare("٥", "9") — Arabic-Indic 5 vs. ASCII 9 positive (5 ranked above 9) negative
Compare("٠0", "0") — leading Arabic-Indic zero positive zero

This contradicts the class's own XML doc, which promises "embedded numbers are compared as numeric values".

The fix

Of the two options the issue offers, this takes the one that honours the documented contract rather than narrowing it: the chunking already treats a run of Nd digits as one number, so compare it by the number it spells.

CompareNumericChunks now:

  1. Skips leading zeros by digit value rather than by the ASCII '0' character (SkipLeadingZeros), so "٠٠" reduces to a single zero just as "00" does.
  2. Compares significant digit counts — more digits means a larger number, as before.
  3. On a tie, walks the chunks digit by digit comparing CharUnicodeInfo.GetDecimalDigitValue.

ASCII behaviour is unchanged. For equal-length ASCII chunks, comparing digit values gives exactly the same result as the ordinal compare it replaces, since ASCII digit code points are already in numeric order. All 11 pre-existing tests pass untouched.

Both chunks reaching this method come from the \d+ alternative of the chunk regex, so every character is category Nd and GetDecimalDigitValue is guaranteed to return 0-9 — noted in a remark on the method.

The type's XML doc now states that a run of Unicode decimal digits is a number whatever script it is written in.

Tests

Four new tests in NaturalStringComparerTests.cs:

  • Compare_NonAsciiDigits_ComparedByNumericValue — the issue's "٥" vs "9" case, plus within-script and Devanagari 5 < 30.
  • Compare_NonAsciiDigits_EqualValuesAreEqual — "٥" equals "5", bare and embedded.
  • Compare_NonAsciiLeadingZeros_NormalizedLikeAsciiZeros — the issue's "٠0" vs "0" case, plus all-zero chunks.
  • Compare_MixedScriptDigits_ComparedByNumericValue — a single chunk mixing scripts ("1٥" = 15) still spells one number.

Verified by reverting only the NaturalStringComparer.cs change and re-running: 4 failed / 11 passed. With the fix: 15 passed, 0 failed.

🤖 Generated with Claude Code

https://claude.ai/code/session_01FyZyutu7Xna2FUQK6o8KAC


Generated by Claude Code

…er [patch]

Chunk detection used char.IsDigit and \d+, which match any Unicode Nd
digit, but CompareNumericChunks then normalized with TrimStart('0') and
fell back to an ordinal compare — both of which assume ASCII 0x30-0x39.
A chunk containing a non-ASCII digit silently degraded to a raw UTF-16
code-point comparison, so Compare("٥", "9") ranked five above nine,
contradicting the type's own documented numeric ordering.

Compare digit chunks by the value each digit spells, via
CharUnicodeInfo.GetDecimalDigitValue, and skip leading zeros by value
rather than by the ASCII '0' character. ASCII ordering is unchanged:
for equal-length ASCII chunks, comparing digit values gives the same
result as the ordinal compare it replaces.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FyZyutu7Xna2FUQK6o8KAC
@sonarqubecloud

Copy link
Copy Markdown

@matt-edmondson
matt-edmondson merged commit 7f9e0c6 into main Sep 22, 2026
14 checks passed
@matt-edmondson
matt-edmondson deleted the fix/natural-comparer-unicode-digits branch September 22, 2026 00:23
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

NaturalStringComparer sorts non-ASCII decimal digits incorrectly, contradicting its own documented numeric-ordering guarantee

2 participants