Skip to content

Add obvious-code and tool-call skip heuristic #71

Description

@weselben

Derived from: #57, #65

Work:

Done when: a chat that legitimately repeats code blocks / tool-call deltas does not trigger; a real text loop does.

Activity

  1. weselben commented on Sep 1, 2026

    @weselben
    OwnerAuthor

    I have all the research I need. Let me synthesize into the phone-readable markdown deliverable. The user explicitly requested:

    1. Short answer — recommended v1 heuristic set for a Go streaming guard
    2. Comparison table of heuristics with source links
    3. Explicit recommendation whether to guard at the tool-call layer vs text layer
    4. What to skip/do-not-skip, with the Ollama base64 bug as the cautionary case
    5. Phone-readable markdown
    6. "Do not modify code or issues; just return the markdown"

    Let me write the report. I should keep it tight, scannable on a phone, and use the research findings.

    Key insights to convey:

    Short answer (v1 heuristic set):

    • Post-hoc SSE filter on choices[i].delta.content only (skip tool_calls channel at v1)
    • Pre-filter exemptions: skip when inside code fence (track fence parity), when symbol density high + entropy low (likely base64/hex), when last N bytes parse as valid JSON/YAML
    • Main detector: byte-level repeating-unit detection (current detectRun with maxUnit=256, limit≥3) — already in the codebase
    • Add a "varied-token" semantic detector for phrase loops not caught by exact-byte match
    • Set limit ≥ 3 (not 2) to avoid killing legitimate repetition like """"""

    Tool-call layer: do NOT guard at the tool-call layer in v1. The tool-call JSON is structured and the loop is detected at the text layer anyway. Cross-turn tool-loop detection is a separate, future concern.

    Skip list: code fences (track parity), base64/hex blobs (symbol density > 80% alphanumeric/punct + entropy < threshold), valid JSON/YAML blocks (parse attempt), tool_calls deltas (different channel), table rows (lines starting with |), indented code (leading whitespace ≥ 4).

    Do-not-skip: pure prose text, single-token stutter, multi-token chain, thinking-loop signatures.

    Let me write the report. I need to be careful about:

    • Not creating files (user instruction)
    • Returning it inline
    • Keeping phone-readable (short lines, bullets, tables)
    • Including comparison table
    • Including the Ollama#17562 cautionary case

    I'll structure it with:

    • Title

    • TL;DR / Short answer

    • Comparison table of heuristics

    • Tool-call vs text layer

    • Skip / do-not-skip

    • Ollama#17562 cautionary case

    • Sources

    Let me draft it.

    The user's primary language seems to be English (the ticket is in English). The user explicitly asked for "phone-readable markdown" — so I'll prioritize plain text readability with bullets and short paragraphs.

    Also, this is a "phone-readable" markdown — meaning readable on a phone. So:

    • Short lines
    • Avoid wide tables (or make them narrow)
    • Use bullet lists
    • No deep nesting
    • Important content at top

    Let me write it now.

    Note: I should NOT modify code or issues, per user instruction. Just return the markdown in my response.

    Wayfinder Ticket #65 — Legitimate-repetition heuristics

    Status: research complete. No files modified. No issues opened.

    TL;DR — recommended v1 heuristic set

    Guard only choices[].delta.content (skip delta.tool_calls). Three-stage pipeline, each O(1)-ish per chunk:

    1. Channel split. Tool-call deltas pass through opaque. Text deltas feed the guard.
    2. Skip pre-filter. Skip detection when ANY of:
      • Inside a fenced code block (track ``` parity).
      • Inside a Markdown table row (line starts with |).
      • Symbol-density ≥ 0.85 AND Shannon entropy ≤ 4.5 bits/char → looks like base64/hex.
      • Last holdback bytes parse as balanced JSON {}/[].
      • Pure whitespace run ≥ 8 chars.
    3. Repetition detector. Keep the current byte-suffix scan, but:
      • Raise min limit from 2 → 3 (Ollama #17562 lesson: 2 fires on " " and "").
      • Cap unit at 32 chars, not 256 (maxUnitBytes=32).
      • Require total run length ≥ 96 chars (minRunBytes = limit × maxUnit).
      • Run Z-array exact-cycle scan on the rolling 4096-char tail every 128 chars (oh-my-pi trick — catches "the the the the" within ~600 chars).

    This is post-hoc, allocation-free, regex-free, no per-token work. vLLM PR #40099 covers grammar-constrained gen separately (upstream concern, not us).

    Comparison table

    System Layer Detector Defaults Code/table/base64 skip Tool-call handling Source
    GoModel (current) Post-hoc SSE filter detectRun byte-suffix, periods 1–256 B limit=2, maxUnit=256 None — FP on base64/hex/tables Skipped (opaque) internal/streaming/repetition_guard_stream.go:55–487
    vLLM Sampling + post-hoc check_sequence_repetition N-gram + apply_penalties max=20, min=3, count=4 (auto on grammar) None — token-level only None — runs on full output vllm#40099, vllm/v1/core/sched/utils.py
    oh-my-pi Post-hoc stream wrapper Z-array suffix + trigram Jaccard + novelty-stall + Gemini header 4× ≤60 / 3× ≥1024; Jaccard ≥0.8; novelty ≤0.2 Title-strip + anchor reset only — no b64/hex skip Cross-turn sha256(name+JSON(args)), threshold 5 thinking-loop.ts
    gemini-cli Post-hoc turn-level 50-char sha256 chunks; tool-call hash; LLM double-check threshold 10; cycles 1–5×5; LLM @30 turns Skips inside code blocks (fence parity toggle); resets on tables/lists/headings Skipped — separate hash on (name, args) loopDetectionService.ts
    OpenCrabs Post-hoc stream 2048-byte rolling window, n-gram per-docs unspecified Excludes fenced code + table rows (v0.3.77) Separate tool-loop guard OpenCrabs docs
    OpenClaw Post-hoc turn-level sha256(tool, args, result) triple + post-compaction warn-then-block Strips volatile runtime fields Yes — that's its job OpenClaw tool-loop detection
    BeeLlama Sampling-time repeat_line segment sampler window=1024, max-period=128, min-coverage=256, min-tokens=512 Delimiter-based (\n.!?:) n/a beellama features
    llama.cpp DRY Sampling-time Token-N-gram penalty dry_multiplier=0.8, dry_base=1.75, dry_allowed_length=2 Sequence-breakers \n, :, ", * n/a llama-dry-docs
    Ollama#15212 Sampling-time repeat_line segment sampler min_length=20, delimiters=\n.!?, temp_boost=0.5 Delimiter-based n/a ollama#15212
    Ollama #17562 (anti-pattern) Post-hoc guard 31 identical tokens → abort repeat_last_n=64 None — kills base64/hex/indented tables n/a (silent abort, no done) ollama#17562
    OpenFang / hermes-agent Post-hoc sha256(tool_name + sorted_args) + ping-pong threshold 3, sliding window n/a Yes — primary purpose hermes-agent#481
    LangChain Sampling-time NoRepeatNGramLogitsProcessor, repeat_penalty n-gram size passed in n/a n/a transformers logits_process.py

    Tool-call layer vs text layer — recommendation

    Guard the text layer only in v1. Three reasons:

    1. Tool-call deltas are already structurally repetitive. OpenAI tool-call streaming chunks look like {"index":0,"function":{"arguments":"{\\"pa…\\"}}} — high bracket density, low entropy, lots of literal characters. Running a suffix-detect on those would FP constantly.
    2. Cross-turn tool-call loops need turn-level context, not streaming. oh-my-pi, OpenFang, OpenClaw all do this at the tool-call-loop-guard boundary with sha256(name + canonicalize(args)) — it needs a tool-call finished boundary, not a delta.
    3. The text-layer guard already catches "model emits fake tool-call deltas in a loop" because the content delta still flows through the guard. Belt-and-braces.

    Skip the tool-call channel at the SSE layer. Defer cross-turn tool-loop detection to v2 (out of scope here — that's an agent-runtime concern, not a streaming-gateway concern).

    Skip vs do-not-skip

    Skip (legitimate repetition — exempt the guard)

    • Inside fenced code block — track ``` parity per choice (gemini-cli pattern, OpenCrabs v0.3.77).
    • Markdown table row — line starts with | after optional whitespace.
    • Base64/hex blob — symbol density (alnum + /+= or 0-9a-f) ≥ 0.85 over a 64-char window AND Shannon entropy < 4.5 bits/char.
    • Valid JSON/YAML block — quick parse attempt on the last holdback bytes (encoding/json Valid() is O(n) and allocation-free).
    • Pure whitespace run — 8+ consecutive spaces/tabs/newlines.
    • Tool-call deltas — never inspect delta.tool_calls, delta.function_call.arguments.
    • Thinking-channel blocks — when SSE channel = thought (Anthropic), suppress entirely.

    Do NOT skip (real loops)

    • Plain prose repetition ("the the the the").
    • Chain repetition with varied tokens ("Let me check. Let me check. Let me check.") — caught by Z-array exact-cycle, not byte-equal.
    • Gemini-style header runaway (consecutive ## or **bold title** lines).
    • Whitespace mix with non-blank tokens ("x x x x x x x").
    • Repetition inside a content delta that accompanies a tool_calls delta — only the content bytes are inspected.

    Ollama#17562 — the cautionary case

    ollama#17562 is the canonical anti-pattern:

    Generation is aborted after 31 identical tokens. Base64 of a binary file, a hex dump, a long run of indentation or a wide table of repeated values all reach that in normal, correct output.

    Three specific bugs in one issue:

    1. Silent abort. Abort reported as success (HTTP 200, no done message). Caller sees Did not receive done or success response in stream. Fix: emit done_reason:"repeat" (OpenAI streaming spec already supports this).
    2. Token-count gate instead of character budget. 31 identical tokens is a hard, low ceiling — base64 chunks reach it trivially. Fix: detect a repeating unit of up to 32 tokens and measure against a character budget no real payload reaches.
    3. No skip list. No code-fence, base64, hex, table, or JSON exemptions.

    GoModel's current impl has bug #2 (limit=2 is even more aggressive) and bug #3 (no skip list). Bug #1 is partially mitigated — we append data: [DONE]\n\n, but we don't emit done_reason:"repeat". v1 should add done_reason to the final event so clients can distinguish a guard-triggered abort from a real finish.

    v1 implementation outline (Go)

    // Per-choice state
    type choiceGuard struct {
        text        []byte    // ring buffer of last 4096 bytes
        fenceParity int       // ``` count mod 2 (toggle)
        inTable     bool      // last line started with |
        skipUntil   int       // bytes to skip (in skip mode)
    }
    
    // Hot path — O(1) per delta
    func (g *choiceGuard) observe(delta []byte) (action action) {
        g.text = append(g.text[:min(len(g.text), 4096-len(delta))], delta...)
        if g.skipUntil > 0 { g.skipUntil -= len(delta); return actionPass }
    
        // Fence parity (cheap, byte scan)
        for _, b := range delta {
            // detect ``` on its own line — cheap 3-byte match
        }
        if g.fenceParity%2 == 1 { return actionPass }
    
        // Base64/hex quick screen on the tail
        if isLikelyBase64OrHex(g.tail(64)) { return actionSkip(64) }
    
        // Pure whitespace
        if isWhitespaceRun(g.tail(8)) { return actionPass }
    
        // Main detector — current detectRun with raised thresholds
        if detectRun(g.text, limit=3, maxUnit=32, minRunBytes=96) {
            return actionCut
        }
        return actionPass
    }

    Defaults for v1: STREAM_REPETITION_LIMIT=3, STREAM_REPETITION_MAX_UNIT_BYTES=32, STREAM_REPETITION_MIN_RUN_BYTES=96. Disable with 0 (current behaviour).

    Sources

    Open questions for ticket owner

    1. Tool-call cross-turn guard — v1 or v2? Recommend v2 — not a streaming-gateway concern; needs the agent runtime.
    2. Emit done_reason:"repeat" in final SSE event? Yes, per Ollama#17562 lesson. OAI spec already has the field.
    3. Sampling-time penalty (DRY / repeat_line) — out of scope? Yes — that's an upstream provider concern. We only see post-hoc.
    4. LLM semantic double-check (gemini-cli pattern)? No for v1 — costs extra model call per stream, not deterministic.
  2. weselben commented on Sep 1, 2026

    @weselben
    OwnerAuthor

    Pending user decision (one letter):

    The research output suggests emitting done_reason:"repeat" in the final SSE event so clients can distinguish a guard cut from a real finish (Ollama#17562 lesson). Your #57 answer said "client stays intact, no extra chunks" — done_reason is an additional field on the final chunk, not an extra chunk, so it is borderline.

    Reply:

    • A: keep the pure [DONE] (no done_reason field).
    • B: add finish_reason:"stop" on the final content event when the cut hits a content delta, then [DONE].
    • C: add done_reason:"repeat" only — minimal client-visible signal.

    Until you reply, A ships by default.

    Repo-conformant notes (independent of A/B/C):

    • Skip heuristics live in internal/streaming/repetition_guard_stream.go next to the detector (single file per Add lazy tiktoken loader with byte-period fallback #67 rationale).
    • Three pure functions, each O(1) per delta:
      • looksLikeBase64OrHex([]byte) bool — symbol-density ≥ 0.85 AND Shannon entropy < 4.5 bits/char over a 64-byte window.
      • isFencedCode([]byte) (parity int, inside bool) — ``` parity tracker per choice; resets on newline.
      • looksLikeMarkdownTableRow([]byte) bool — first non-whitespace byte is |.
    • Pure whitespace run ≥ 8 chars → skip.
    • choices[].delta.tool_calls deltas never enter the detector (channel split at the parser, mirror the existing contentDeltas extraction).
    • All skip decisions logged at slog.Debug (not Warn) — operators see counters, not every skip.

    Tests: table-driven cases for each skip rule (base64 hex, fenced code in/out, table rows, whitespace run, tool-call delta). Repo norm: testing only, t.Run subtests, no testify. Brown-test style reserved for #72.

    This runs on 400+ machines — skip heuristics must be deterministic and allocation-free; the three functions are pure byte scans. No regex, no strings.Contains on hot path.

  3. weselben commented on Sep 2, 2026

    @weselben
    OwnerAuthor

    Resolved on feat/stream-repetition-canceller (commit 9707690).

    What changed in internal/streaming/repetition_guard_stream.go inspectDelta:

    • Fenced code: ``` toggles parity per choice; deltas inside a fenced block are never inspected.
    • Markdown table rows: lines starting with | are skipped.
    • Long whitespace runs: runs of >= 8 consecutive whitespace bytes are skipped.
    • Encoded blobs: trailing 64-byte window with symbol density >= 0.85 and Shannon entropy < 4.5 bits/char is skipped (catches base64/hex).
    • tool_calls / function_call deltas: contentDeltas ignores choices whose delta carries tool_calls or function_call, so those never enter the detector.

    Evidence: tests TestRepetitionGuardStream_FencedCodeNoTrigger, _MarkdownTableNoTrigger, _WhitespaceRunNoTrigger, _Base64BlobNoTrigger, _ToolCallsNeverTrigger, _CleanStreamIdentical all green.

    🤖 Written by Kimi Code (AI agent) from the stream-repetition-canceller worktree.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions