Skip to content

Proposal: emit UserWarning from Split.pre_tokenize_str() when it produces zero matches (catches silent str vs Regex confusion) #2109

Description

@vamsin07

Context. The Split pre-tokenizer accepts both str (literal substring) and Regex (regex pattern) for pattern. This is correctly documented since #1264 and #1565, and a behavior change has been considered and declined (#1369, "Not planned"). I'm not proposing any change to encode() behavior, I'd like to add a small debug-time surface that catches the most common silent-failure mode that motivates these recurring reports.

The silent failure. If a user passes a str containing regex metacharacters ([, \, |, +, etc.), Split searches for that literal substring, which usually doesn't appear in any natural text => 0 matches => the entire input is returned as one pre-token. encode() runs without errors, the tokenizer trains, and downstream BPE silently collapses to byte-level BPE on raw input. The user has no signal anything is wrong until they inspect token outputs by hand.

Real-world impact (recent example): a multi-language ASR fine-tuning project ran on byte-level BPE for months under the belief that script-aware pre-tokenization was active. Diagnosed via the chain in #1369 (linked by my comment there).

Proposal => Emit a UserWarning only when pre_tokenize_str(text) is called on an input longer than 1 character AND the result is exactly one piece spanning the whole input:

def pre_tokenize_str(self, text):
    out = self._inner.pre_tokenize_str(text)
    if (
        isinstance(self._pattern, str)  # only warn for str-pattern Split
        and len(text) > 1
        and len(out) == 1
        and out[0][1] == (0, len(text))
    ):
        import warnings
        warnings.warn(
            f"Split(pattern={self._pattern!r}, ...).pre_tokenize_str() "
            f"returned the entire input as a single piece (no matches). "
            f"If you intended regex semantics, wrap the pattern in "
            f"tokenizers.Regex(); a plain string is matched literally.",
            UserWarning, stacklevel=2,
        )
    return out

Why this is safe to merge?

Touches only pre_tokenize_str() encode() and training are unchanged. Only fires for str-pattern Split, not Regex(...), not other pre-tokenizers.
Suppressible the standard way (warnings.filterwarnings("ignore")). Catches the failure exactly when a user is most likely to be debugging tokenization.
Repro that would trigger the warning:

from tokenizers.pre_tokenizers import Split
Split(pattern="[a-z]+", behavior="isolated").pre_tokenize_str("hello world")
# current: returns [("hello world", (0, 11))], silent
# proposed: same return + UserWarning about Regex() wrapping

Happy to send a PR if there's directional interest. References: #1369 (behavior change rejected => agreed, no proposal to revisit that), #1264, #1565 (docs improvements already merged => this is the runtime-debug complement).

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions