Skip to content

Detection unit: byte-period matching or real tokenizer? #54

Description

@weselben

Status

Options A/B/C posted in a follow-up comment. Waiting on user reply.

Question

Question

The prototype detects repetition on concatenated delta text bytes: the smallest period p in 1..256 bytes whose last p*N bytes are N identical copies. Catches single-token stutter and multi-token chains without a tokenizer. Trade-offs: a real tokenizer would align units to token boundaries (fewer false positives on byte-level coincidence, exact N semantics per token) at the cost of a tokenizer dependency per provider and latency. Is byte-period detection acceptable to ship, or do we need token-based detection?

Activity

  1. weselben commented on Sep 1, 2026

    @weselben
    OwnerAuthor

    Phone-friendly options (reply with a letter):

    • A (recommended): ship byte-period detection. Prototype catches single-token stutter and multi-token chains today; zero new deps, zero tokenizer latency. False-positive risk is negligible at realistic limits (N≥20).
    • B: token-based detection. Exact N-per-token semantics, but needs a tokenizer per provider family and adds per-chunk cost.
    • C: hybrid. Byte-period now, tokenizer behind a flag later.

    Grill note: byte-period at p≤256B covers every loop we've actually observed from real models (they repeat words/phrases, never random byte noise).

  2. weselben commented on Sep 1, 2026

    @weselben
    OwnerAuthor

    As we have tokennizer already in the gomodel dependencies for the audits this wouldnt matter but add real value!

    B

  3. weselben commented on Sep 1, 2026

    @weselben
    OwnerAuthor

    Resolution in progress — one blocker. You picked B (token-based), reasoning a tokenizer is already in deps. I checked go.mod and internal/: no tokenizer library is present (only a string mention in anthropicapi/request.go). Usage/audit token counts come from provider-reported usage, not local tokenization.

    So B means adding a tokenizer dependency. I've opened Tokenizer library selection for the repetition guard as the follow-up ticket; #54 stays open until that lands. Code stays byte-period in the meantime (still correct, just not token-aligned).

  4. weselben commented on Sep 1, 2026

    @weselben
    OwnerAuthor

    Resolved. Token-based detection (your B) via tiktoken per tokenizer selection: latest/fastest pure-Go port, own file, lazy-loaded only when the guard is active. Byte-period detection remains as the fallback path when a model vocabulary is missing. Detection operates on the token stream per chunk (per #55/#57: eager, no back-truncation).

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions