Skip to content

Keep NaturalStringComparer transitive across digit scripts - #57

Merged
matt-edmondson merged 2 commits into
mainfrom
claude/sorting-55-transitive-unicode-digits
Sep 27, 2026
Merged

matt-edmondson merged 2 commits into
mainfrom
claude/sorting-55-transitive-unicode-digits

Conversation

@matt-edmondson

Copy link
Copy Markdown
Contributor

Fixes #55

What changed

Mixed digit/text chunks. When a digit chunk met a text chunk, the comparer compared raw code points. That let ٥ (U+0665) sort after letters while 5, which it equals, sorts before them. The result was a cycle (٥ < 10 < z < ٥), and sort output depended on input order. Digit chunks are now stored as ASCII digits of the same value. A text chunk never starts with a digit, so a number compared with text is decided by its first character, the same way for every script.

I chose this over the issue's "numeric always before text" suggestion because it leaves ASCII ordering exactly as it is today. For example, "-5"-style punctuation still sorts before numbers and letters still sort after them. The only change is where non-ASCII digits land.

Chunking by code point. The \d+|\D+ regex is gone. It worked on UTF-16 code units, and those treat surrogate halves as non-digits. In its place is a small splitter that walks code points using CharUnicodeInfo.GetUnicodeCategory(string, int) and GetDecimalDigitValue(string, int), which works back to netstandard2.0. Digits outside the BMP, such as the mathematical digits, now form numbers: "file𝟗" < "file𝟏𝟎".

Tests

New tests:

  • Compare_NonAsciiDigitAgainstText_OrdersLikeTheAsciiDigit
  • Compare_SortOfMixedScriptDigitsAndText_DoesNotDependOnInputOrder
  • Compare_IsTransitiveOverMixedScriptDigitsAndText, a check of antisymmetry and transitivity over every triple from a mixed set of ASCII, Arabic-Indic, Devanagari and mathematical digits, letters, punctuation and emoji
  • Compare_DigitsOutsideTheBasicMultilingualPlane_ComparedByNumericValue

All four fail with the fix stashed and pass with it. Full suite: 19/19 locally. The library builds clean for net10.0, net9.0, net8.0, netstandard2.1 and netstandard2.0.

🤖 Generated with Claude Code

https://claude.ai/code/session_01SZUfnvhMFzj96bfFNzbfpn


Generated by Claude Code

A digit chunk compared with a text chunk used its raw code points, so a
non-ASCII digit such as the Arabic-Indic five sorted after letters while
the ASCII five it equals sorted before them. Sort output then depended on
input order. Digit chunks are now held as ASCII digits, so a number
against text orders the same way whatever script spells it, and ASCII
input orders exactly as before.

Chunking also walks code points instead of UTF-16 units, so digits outside
the BMP, such as the mathematical digits, form numbers as the remarks say.

Fixes #55

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SZUfnvhMFzj96bfFNzbfpn
Brings its cognitive complexity under Sonar's S3776 limit without changing
the chunks it produces.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SZUfnvhMFzj96bfFNzbfpn
@sonarqubecloud

Copy link
Copy Markdown

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

NaturalStringComparer is intransitive for non-ASCII digits (٥ < 10 < z < ٥), so sort output depends on input order

2 participants