Repository navigation
Detection unit: byte-period matching or real tokenizer? #54
Description
Activity
- added a parent issue
on Sep 1, 2026 Phone-friendly options (reply with a letter):
- A (recommended): ship byte-period detection. Prototype catches single-token stutter and multi-token chains today; zero new deps, zero tokenizer latency. False-positive risk is negligible at realistic limits (N≥20).
- B: token-based detection. Exact N-per-token semantics, but needs a tokenizer per provider family and adds per-chunk cost.
- C: hybrid. Byte-period now, tokenizer behind a flag later.
Grill note: byte-period at p≤256B covers every loop we've actually observed from real models (they repeat words/phrases, never random byte noise).
As we have tokennizer already in the gomodel dependencies for the audits this wouldnt matter but add real value!
B
Resolution in progress — one blocker. You picked B (token-based), reasoning a tokenizer is already in deps. I checked
go.modandinternal/: no tokenizer library is present (only a string mention inanthropicapi/request.go). Usage/audit token counts come from provider-reported usage, not local tokenization.So B means adding a tokenizer dependency. I've opened Tokenizer library selection for the repetition guard as the follow-up ticket; #54 stays open until that lands. Code stays byte-period in the meantime (still correct, just not token-aligned).
Resolved. Token-based detection (your B) via tiktoken per tokenizer selection: latest/fastest pure-Go port, own file, lazy-loaded only when the guard is active. Byte-period detection remains as the fallback path when a model vocabulary is missing. Detection operates on the token stream per chunk (per #55/#57: eager, no back-truncation).
Status
Options A/B/C posted in a follow-up comment. Waiting on user reply.
Question
Question
The prototype detects repetition on concatenated delta text bytes: the smallest period p in 1..256 bytes whose last p*N bytes are N identical copies. Catches single-token stutter and multi-token chains without a tokenizer. Trade-offs: a real tokenizer would align units to token boundaries (fewer false positives on byte-level coincidence, exact N semantics per token) at the cost of a tokenizer dependency per provider and latency. Is byte-period detection acceptable to ship, or do we need token-based detection?