Skip to content

TextInfo.ToTitleCase does not handle U+2019 correctly #133034

Description

@rjgotten

Description

TextInfo.ToTitleCase does not handle the ’ aka U+2019 RIGHT SINGLE QUOTATION MARK correctly as a non-word separator. It is classified in Unicode's word-break data as MidNumLet which forbids being treated as a word break when flanked by letters or numbers on both sides.

U+2019 is often used as a typographic curly apostrophe (with a long history of fonts and character sets that mapped them to the same code point and generally made this into a mess).

Reproduction Steps

CultureInfo.GetCultureInfo("en").TextInfo.ToTitleCase("Grandma’s pictures")
// ⛔ "Grandma’S Pictures"

Expected behavior

U+2019 should not be treated as a word separator and should not result in the next character being treated as the beginning of a new word.

Actual behavior

U+2019 is being treated as a word separator and results in the next character being treated as the beginning of a new word.

Regression?

No response

Known Workarounds

None.

Configuration

.NET version: v10.0.11

Other information

Note: when flanked by - is strictly speaking load-bearing and non-optional. Unicode doesn't define word break simply via separator characters. It may require such characters to have a particular ante/post context. U+2019 is non-breaking when it is flanked on both sides, ante and post, by a letter or number. And it is breaking in other cases. That's iirc how the MidNumLet-type word-break is classified. TextInfo is using an oversimplification and is wrong. And not just on U+2019 - it's just one of the more directly overt cases.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions