Context. The Split pre-tokenizer accepts both str (literal substring) and Regex (regex pattern) for pattern. This is correctly documented since #1264 and #1565, and a behavior change has been considered and declined (#1369, "Not planned"). I'm not proposing any change to encode() behavior, I'd like to add a small debug-time surface that catches the most common silent-failure mode that motivates these recurring reports.
The silent failure. If a user passes a str containing regex metacharacters ([, \, |, +, etc.), Split searches for that literal substring, which usually doesn't appear in any natural text => 0 matches => the entire input is returned as one pre-token. encode() runs without errors, the tokenizer trains, and downstream BPE silently collapses to byte-level BPE on raw input. The user has no signal anything is wrong until they inspect token outputs by hand.
Real-world impact (recent example): a multi-language ASR fine-tuning project ran on byte-level BPE for months under the belief that script-aware pre-tokenization was active. Diagnosed via the chain in #1369 (linked by my comment there).
Proposal => Emit a UserWarning only when pre_tokenize_str(text) is called on an input longer than 1 character AND the result is exactly one piece spanning the whole input:
def pre_tokenize_str(self, text):
out = self._inner.pre_tokenize_str(text)
if (
isinstance(self._pattern, str) # only warn for str-pattern Split
and len(text) > 1
and len(out) == 1
and out[0][1] == (0, len(text))
):
import warnings
warnings.warn(
f"Split(pattern={self._pattern!r}, ...).pre_tokenize_str() "
f"returned the entire input as a single piece (no matches). "
f"If you intended regex semantics, wrap the pattern in "
f"tokenizers.Regex(); a plain string is matched literally.",
UserWarning, stacklevel=2,
)
return out
Why this is safe to merge?
Touches only pre_tokenize_str() encode() and training are unchanged. Only fires for str-pattern Split, not Regex(...), not other pre-tokenizers.
Suppressible the standard way (warnings.filterwarnings("ignore")). Catches the failure exactly when a user is most likely to be debugging tokenization.
Repro that would trigger the warning:
from tokenizers.pre_tokenizers import Split
Split(pattern="[a-z]+", behavior="isolated").pre_tokenize_str("hello world")
# current: returns [("hello world", (0, 11))], silent
# proposed: same return + UserWarning about Regex() wrapping
Happy to send a PR if there's directional interest. References: #1369 (behavior change rejected => agreed, no proposal to revisit that), #1264, #1565 (docs improvements already merged => this is the runtime-debug complement).
Context. The Split pre-tokenizer accepts both str (literal substring) and Regex (regex pattern) for pattern. This is correctly documented since #1264 and #1565, and a behavior change has been considered and declined (#1369, "Not planned"). I'm not proposing any change to encode() behavior, I'd like to add a small debug-time surface that catches the most common silent-failure mode that motivates these recurring reports.
The silent failure. If a user passes a str containing regex metacharacters ([, \, |, +, etc.), Split searches for that literal substring, which usually doesn't appear in any natural text => 0 matches => the entire input is returned as one pre-token. encode() runs without errors, the tokenizer trains, and downstream BPE silently collapses to byte-level BPE on raw input. The user has no signal anything is wrong until they inspect token outputs by hand.
Real-world impact (recent example): a multi-language ASR fine-tuning project ran on byte-level BPE for months under the belief that script-aware pre-tokenization was active. Diagnosed via the chain in #1369 (linked by my comment there).
Proposal => Emit a UserWarning only when pre_tokenize_str(text) is called on an input longer than 1 character AND the result is exactly one piece spanning the whole input:
Why this is safe to merge?
Touches only pre_tokenize_str() encode() and training are unchanged. Only fires for str-pattern Split, not Regex(...), not other pre-tokenizers.
Suppressible the standard way (warnings.filterwarnings("ignore")). Catches the failure exactly when a user is most likely to be debugging tokenization.
Repro that would trigger the warning:
Happy to send a PR if there's directional interest. References: #1369 (behavior change rejected => agreed, no proposal to revisit that), #1264, #1565 (docs improvements already merged => this is the runtime-debug complement).