Description
TextInfo.ToTitleCase does not handle the ’ aka U+2019 RIGHT SINGLE QUOTATION MARK correctly as a non-word separator. It is classified in Unicode's word-break data as MidNumLet which forbids being treated as a word break when flanked by letters or numbers on both sides.
U+2019 is often used as a typographic curly apostrophe (with a long history of fonts and character sets that mapped them to the same code point and generally made this into a mess).
Reproduction Steps
CultureInfo.GetCultureInfo("en").TextInfo.ToTitleCase("Grandma’s pictures")
// ⛔ "Grandma’S Pictures"
Expected behavior
U+2019 should not be treated as a word separator and should not result in the next character being treated as the beginning of a new word.
Actual behavior
U+2019 is being treated as a word separator and results in the next character being treated as the beginning of a new word.
Regression?
No response
Known Workarounds
None.
Configuration
.NET version: v10.0.11
Other information
Note: when flanked by - is strictly speaking load-bearing and non-optional. Unicode doesn't define word break simply via separator characters. It may require such characters to have a particular ante/post context. U+2019 is non-breaking when it is flanked on both sides, ante and post, by a letter or number. And it is breaking in other cases. That's iirc how the MidNumLet-type word-break is classified. TextInfo is using an oversimplification and is wrong. And not just on U+2019 - it's just one of the more directly overt cases.
Description
TextInfo.ToTitleCasedoes not handle the’akaU+2019 RIGHT SINGLE QUOTATION MARKcorrectly as a non-word separator. It is classified in Unicode's word-break data asMidNumLetwhich forbids being treated as a word break when flanked by letters or numbers on both sides.U+2019is often used as a typographic curly apostrophe (with a long history of fonts and character sets that mapped them to the same code point and generally made this into a mess).Reproduction Steps
Expected behavior
U+2019 should not be treated as a word separator and should not result in the next character being treated as the beginning of a new word.
Actual behavior
U+2019 is being treated as a word separator and results in the next character being treated as the beginning of a new word.
Regression?
No response
Known Workarounds
None.
Configuration
.NET version: v10.0.11
Other information
Note: when flanked by - is strictly speaking load-bearing and non-optional. Unicode doesn't define word break simply via separator characters. It may require such characters to have a particular ante/post context. U+2019 is non-breaking when it is flanked on both sides, ante and post, by a letter or number. And it is breaking in other cases. That's iirc how the
MidNumLet-type word-break is classified.TextInfois using an oversimplification and is wrong. And not just on U+2019 - it's just one of the more directly overt cases.