What's wrong
ExtractGlobFilterTokens and ExtractTextTokens (TextFilter/TextFilter.cs ~261-276) both split only on the ' ' character and then call Trim(). Trim() removes all Unicode whitespace. A space-separated chunk that contains only other whitespace ("\t", "\r\n", " ") therefore survives RemoveEmptyEntries and then trims down to "". This causes two separate failures.
1. Glob filters throw
ExtractGlobFilterTokens then calls t.First() on the empty token inside GroupBy, which throws InvalidOperationException: Sequence contains no elements. The IsNullOrWhiteSpace(filter) guard does not catch this, because the filter as a whole is not blank. It is the same kind of keystroke/paste crash as #106: text pasted with a trailing " \r\n" is enough.
TextFilter.IsMatch("foo", "foo \t"); // throws InvalidOperationException
TextFilter.IsMatch("foo", "foo \r\n"); // throws
TextFilter.Filter(new[]{"foo"}, "foo \t").ToList(); // throws
// Expected: true / ["foo"]
2. Regex ByWordAll sees an empty word, which undoes the #107 fix
ExtractTextTokens produces a "" word for the same kind of chunks. Under ByWordAll, every word must match, so ordinary text fails when it contains a tab. Blank text with a tab also produces {""} instead of no tokens, so the #107 textTokens.Count == 0 guard never runs.
TextFilter.IsMatch("hello \t world", "o", TextFilterType.Regex, TextFilterMatchOptions.ByWordAll); // False, expected True
TextFilter.IsMatch(" \t ", "a*", TextFilterType.Regex, TextFilterMatchOptions.ByWordAll); // True, expected False (#107)
All of the above were reproduced with temporary MSTest cases.
Suggested fix
In both tokenizers, split on all whitespace and drop empty entries:
text.Split((char[]?)null, StringSplitOptions.RemoveEmptyEntries)
Keeping the current split and adding .Where(s => s.Length != 0) after Trim() also works.
Acceptance: the cases above return the expected values without throwing, and they are added as regression tests.
What's wrong
ExtractGlobFilterTokensandExtractTextTokens(TextFilter/TextFilter.cs~261-276) both split only on the' 'character and then callTrim().Trim()removes all Unicode whitespace. A space-separated chunk that contains only other whitespace ("\t","\r\n"," ") therefore survivesRemoveEmptyEntriesand then trims down to"". This causes two separate failures.1. Glob filters throw
ExtractGlobFilterTokensthen callst.First()on the empty token insideGroupBy, which throwsInvalidOperationException: Sequence contains no elements. TheIsNullOrWhiteSpace(filter)guard does not catch this, because the filter as a whole is not blank. It is the same kind of keystroke/paste crash as #106: text pasted with a trailing" \r\n"is enough.2. Regex ByWordAll sees an empty word, which undoes the #107 fix
ExtractTextTokensproduces a""word for the same kind of chunks. Under ByWordAll, every word must match, so ordinary text fails when it contains a tab. Blank text with a tab also produces{""}instead of no tokens, so the #107textTokens.Count == 0guard never runs.All of the above were reproduced with temporary MSTest cases.
Suggested fix
In both tokenizers, split on all whitespace and drop empty entries:
Keeping the current split and adding
.Where(s => s.Length != 0)afterTrim()also works.Acceptance: the cases above return the expected values without throwing, and they are added as regression tests.