Description
string.ToUpperOrdinal() documents a canonicalization relationship with
StringComparison.OrdinalIgnoreCase: two strings are ordinal-ignore-case equal
if and only if their ordinal-uppercase forms are ordinally equal.
That contract does not hold for a Deseret supplementary-character case when
.NET uses Windows NLS:
U+10428 DESERET SMALL LETTER LONG I
U+10400 DESERET CAPITAL LETTER LONG I
Under NLS, StringComparison.OrdinalIgnoreCase reports these strings as
different, while ToUpperOrdinal() maps both strings to U+10400.
This blocks consumers such as ASP.NET Core Output Caching from using
ToUpperOrdinal() as a collision-free canonical representation for keys whose
equivalence must match routing's ordinal-ignore-case semantics.
Reproduction
using System.Runtime.InteropServices;
using System.Text;
var cases = new (string Name, string Left, string Right)[]
{
("ASCII case control", "s", "S"),
("long-s BMP control", "s", "\u017F"),
("Kelvin BMP control", "k", "\u212A"),
("sigma BMP control", "\u03C2", "\u03C3"),
("Deseret supplementary regression", "\U00010428", "\U00010400"),
};
Console.WriteLine($"Framework: {RuntimeInformation.FrameworkDescription}");
Console.WriteLine($"OS: {RuntimeInformation.OSDescription}");
Console.WriteLine($"DOTNET_SYSTEM_GLOBALIZATION_USENLS: {Environment.GetEnvironmentVariable("DOTNET_SYSTEM_GLOBALIZATION_USENLS") ?? "<unset>"}");
Console.WriteLine($"DOTNET_SYSTEM_GLOBALIZATION_INVARIANT: {Environment.GetEnvironmentVariable("DOTNET_SYSTEM_GLOBALIZATION_INVARIANT") ?? "<unset>"}");
var failures = 0;
foreach (var (name, left, right) in cases)
{
var leftCanonical = left.ToUpperOrdinal();
var rightCanonical = right.ToUpperOrdinal();
var originalsEqual = left.Equals(right, StringComparison.OrdinalIgnoreCase);
var canonicalFormsEqual = leftCanonical.Equals(rightCanonical, StringComparison.Ordinal);
var passed = originalsEqual == canonicalFormsEqual;
Console.WriteLine($"{(passed ? "PASS" : "FAIL")}: {name}");
Console.WriteLine($" originals OIC-equal: {originalsEqual}");
Console.WriteLine($" canonical ordinal-equal: {canonicalFormsEqual}");
Console.WriteLine($" canonical values: {Format(leftCanonical)} / {Format(rightCanonical)}");
failures += passed ? 0 : 1;
}
return failures == 0 ? 0 : 1;
static string Format(string value) =>
string.Join(" ", value.EnumerateRunes().Select(static rune => $"U+{rune.Value:X4}"));
Run on Windows with NLS explicitly enabled:
$env:DOTNET_SYSTEM_GLOBALIZATION_USENLS = "1"
dotnet run
Actual result
On .NET 11.0.0-rc.1.26420.103, Windows 10.0.26200 x64:
PASS: ASCII case control
PASS: long-s BMP control
PASS: Kelvin BMP control
PASS: sigma BMP control
FAIL: Deseret supplementary regression
originals OIC-equal: False
canonical ordinal-equal: True
canonical values: U+10400 / U+10400
The process exits with code 1. The same complete test passes under default ICU
and invariant globalization, where the Deseret originals are
ordinal-ignore-case equal and both canonicalize to U+10400.
Expected result
For every pair of strings and every supported globalization mode:
left.Equals(right, StringComparison.OrdinalIgnoreCase)
==
left.ToUpperOrdinal().Equals(right.ToUpperOrdinal(), StringComparison.Ordinal)
Runtime regression coverage should include supplementary Unicode scalars in
addition to the existing exhaustive BMP checks, and should execute under
Windows NLS.
Additional context
The public API was approved in #90999 and implemented in #130140 specifically
to provide this iff relationship. Current tests exhaust the BMP and separately
verify Rune casing for Deseret, but do not test the canonicalization iff
contract for supplementary pairs under NLS.
Under NLS, ordinal-ignore-case comparison delegates to Windows
CompareStringOrdinal, while ordinal casing maps the supplementary Deseret
Rune to its capital form. These surfaces need to agree before
ToUpperOrdinal() can safely serve as an ordinal-ignore-case key
canonicalizer.
Description
string.ToUpperOrdinal()documents a canonicalization relationship withStringComparison.OrdinalIgnoreCase: two strings are ordinal-ignore-case equalif and only if their ordinal-uppercase forms are ordinally equal.
That contract does not hold for a Deseret supplementary-character case when
.NET uses Windows NLS:
U+10428 DESERET SMALL LETTER LONG IU+10400 DESERET CAPITAL LETTER LONG IUnder NLS,
StringComparison.OrdinalIgnoreCasereports these strings asdifferent, while
ToUpperOrdinal()maps both strings to U+10400.This blocks consumers such as ASP.NET Core Output Caching from using
ToUpperOrdinal()as a collision-free canonical representation for keys whoseequivalence must match routing's ordinal-ignore-case semantics.
Reproduction
Run on Windows with NLS explicitly enabled:
Actual result
On .NET 11.0.0-rc.1.26420.103, Windows 10.0.26200 x64:
The process exits with code 1. The same complete test passes under default ICU
and invariant globalization, where the Deseret originals are
ordinal-ignore-case equal and both canonicalize to U+10400.
Expected result
For every pair of strings and every supported globalization mode:
Runtime regression coverage should include supplementary Unicode scalars in
addition to the existing exhaustive BMP checks, and should execute under
Windows NLS.
Additional context
The public API was approved in #90999 and implemented in #130140 specifically
to provide this iff relationship. Current tests exhaust the BMP and separately
verify Rune casing for Deseret, but do not test the canonicalization iff
contract for supplementary pairs under NLS.
Under NLS, ordinal-ignore-case comparison delegates to Windows
CompareStringOrdinal, while ordinal casing maps the supplementary DeseretRune to its capital form. These surfaces need to agree before
ToUpperOrdinal()can safely serve as an ordinal-ignore-case keycanonicalizer.