Skip to content

Documentation, YAML example, and end-to-end smoke test #72

Description

@weselben

Approach (from user comment 2026-09-01)

Build a fake HTTP SSE server harness in /tmp/ (or a sibling temp dir) that produces random repetition streams from JSON fixtures, drives the guard end-to-end, and fails loud on any regression. Then port the harness into the project as the real test suite (99% coverage goal).

What the harness produces

  • Multiple stream shapes: single-token loops, multi-token chain loops, no-loop control streams, mixed loops, edge cases (exactly N-1 then N repeats, varying pattern length, varying unit boundaries).
  • Random parameters within sane ranges: token count, pattern length, repeat count, stream length.
  • Real text/event-stream over HTTP via httptest (or net/http on a free port).

What the harness asserts

  • For every non-loop input: output byte-identical to source.
  • For every loop input with guard active: client receives [DONE], the upstream server sees a client cancel, and the Prometheus counter increments by one.
  • For every loop input with guard off (limit=0): nothing happens — bytes pass through, no counter change.

Files (in /tmp until ported)

  • fake_server.go — configurable loop producer over HTTP.
  • fixtures/*.json — input scenarios.
  • driver.go — runs each fixture, asserts behavior, fails loud.
  • Makefile (or shell) target: make brown / ./brown.sh.

Port to project

  • Move fake-server + driver into internal/streaming/repetition_guard_test.go (table-driven).
  • Use httptest so it runs in CI.
  • Aim for 99% coverage on the new code (tokenization path, skip heuristics, trigger, counter increment, done_reason branch).

Status: awaiting "implement" signal. Will land as the last build step.

Activity

  1. weselben commented on Sep 1, 2026

    @weselben
    OwnerAuthor

    Use Fake Server instead and produce random repetition based on fake inference using json data all in tmp file to fail Proof everything propperly from this Brown Test WE then can build propper Test Files in Project with 99% coverage

  2. weselben commented on Sep 1, 2026

    @weselben
    OwnerAuthor

    Repo-conformant notes (already updated above):

    • Test files inspired by the brown test (/tmp/stream-repetition-brown/zz_brown_manual_test.go) — fake SSE server, table-driven assertions — but not a copy. Match the in-project style of internal/streaming/observed_sse_stream_test.go and internal/streaming/slowdown_stream_test.go:
      • testing only, no testify.
      • recordingReadCloser helper (already exists in repetition_guard_stream_test.go — extend, do not duplicate).
      • TestRepetitionGuardStream_* table-driven, one t.Run per case.
      • No top-level httptest.NewServer per case; build the producer as a small helper that returns an io.ReadCloser (mirrors the existing pattern).
    • Harness in /tmp/ lives only during development; it is not committed.
    • Real test file lands as internal/streaming/repetition_guard_stream_test.go (extend the existing file; the brown-test scaffolding stays in the prototype as commented-out references at most — preferred: just delete the brown-test file once the in-project tests exist).
    • Coverage: go test -coverprofile=cover.out ./internal/streaming/... then go tool cover -func=cover.out — aim for 99% on repetition_guard_stream.go (excluding unreachable error branches documented as such).

    Cases the in-project tests must cover (minimum):

    • disabled passthrough (limit=0, identity preserved, bytes identical)
    • single-token stutter at limit=N
    • multi-token chain loop at limit=2
    • non-repeating stream untouched
    • base64 blob skipped (no trigger)
    • fenced code block skipped (no trigger)
    • markdown table row skipped (no trigger)
    • tool-call delta channel never inspected
    • per-request effective limit resolution: per-model 0 → off, per-model unset + global N → N, per-model M + global N → M
    • counter increments once per trigger
    • Close() is idempotent and terminates the stream on EOF

    This runs on 400+ machines in production — every public code path tested, no flaky timing (no real time.Sleep; use httptest + injected clocks if needed, mirroring slowdown_stream_test.go).

  3. weselben commented on Sep 2, 2026

    @weselben
    OwnerAuthor

    Status update on feat/stream-repetition-canceller (commits pushed to fork).

    Done:

    • In-project table-driven tests: internal/streaming/repetition_guard_stream_test.go, internal/streaming/tokenizer_test.go, internal/gateway/inference_execute_repetition_test.go, internal/server/passthrough_repetition_guard_test.go, internal/virtualmodels/repetition_test.go.
    • End-to-end brown harness moved from /tmp to worktree .tmp/stream-repetition-brown/ (gitignored nested module). Runs against real HTTP SSE upstreams with gpt-4o tokenizer; latest run: guard triggered, upstream cut after 17 events, client saw 4 loop phrases, output ends with data: [DONE].

    Still open / why #72 is not closed yet:

    • Coverage: go test -coverprofile ./internal/streaming/... reports 85.2% statements, below the 99% target. Main gaps (functions <100%):
      • repetition_guard_stream.go Read 70.8%, observe 72.7%, inspectEvent 85.0%, eventPayload 30.8%, contentDeltas 81.8%, looksLikeEncodedBlob 83.3%, isEncodedSymbol 50.0%, WithTriggerCallback 0.0%
      • stream_buffer.go AppendBytes/AppendString/Read/Consume/Release 75-86%
      • slowdown_stream.go / observed_sse_stream.go legacy lines pulled in by guard tests
    • Docs: user-facing doc page / README note about STREAM_REPETITION_LIMIT and per-model fields still needed.

    Next step to close #72: add tests to hit the remaining branches and write the docs/YAML example section.

    🤖 Written by Kimi Code (AI agent) from the stream-repetition-canceller worktree.

  4. added a commit that references this issue on Sep 3, 2026
    d6d49ab
  5. weselben commented on Sep 3, 2026

    @weselben
    OwnerAuthor

    Resolved on feat/stream-repetition-canceller (commit d6d49ab).

    Coverage goal met:

    • repetition_guard_stream.go: 197/197 = 100.0%
    • tokenizer.go: 8/8 = 100.0%
    • Combined new code: 205/205 = 100.0% (was 85.2%)

    End-to-end smoke tests ported in-repo (CI-safe, httptest):

    • internal/streaming/repetition_guard_e2e_test.go — real HTTP SSE upstreams driving the production guard: token stutter, token chain loop, clean answer byte-identical, upstream cancel propagation, empty buffer read, partial event boundaries.
    • Brown scratch harness kept at .tmp/stream-repetition-brown/ (gitignored).

    Docs:

    • docs/features/stream-repetition-guard.mdx — detection, skip heuristics, STREAM_REPETITION_LIMIT / STREAM_REPETITION_MAX_PATTERN, per-model override semantics, Prometheus counter. Registered in docs/docs.json.

    Also fixed on the way: two test lint issues (errcheck, ineffassign) and one dead continue in RepetitionGuardStream.Read.

    Verification: go build ./..., targeted tests across streaming/gateway/server/virtualmodels/observability/config all green; golangci-lint 0 issues on the streaming package.


    This comment was generated with AI assistance.

  6. added a commit that references this issue on Sep 3, 2026
    df58536
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions