Repository navigation
Documentation, YAML example, and end-to-end smoke test #72
Description
Activity
- added a parent issue
on Sep 1, 2026 Use Fake Server instead and produce random repetition based on fake inference using json data all in tmp file to fail Proof everything propperly from this Brown Test WE then can build propper Test Files in Project with 99% coverage
Repo-conformant notes (already updated above):
- Test files inspired by the brown test (
/tmp/stream-repetition-brown/zz_brown_manual_test.go) — fake SSE server, table-driven assertions — but not a copy. Match the in-project style ofinternal/streaming/observed_sse_stream_test.goandinternal/streaming/slowdown_stream_test.go:testingonly, notestify.recordingReadCloserhelper (already exists inrepetition_guard_stream_test.go— extend, do not duplicate).TestRepetitionGuardStream_*table-driven, onet.Runper case.- No top-level
httptest.NewServerper case; build the producer as a small helper that returns anio.ReadCloser(mirrors the existing pattern).
- Harness in
/tmp/lives only during development; it is not committed. - Real test file lands as
internal/streaming/repetition_guard_stream_test.go(extend the existing file; the brown-test scaffolding stays in the prototype as commented-out references at most — preferred: just delete the brown-test file once the in-project tests exist). - Coverage:
go test -coverprofile=cover.out ./internal/streaming/...thengo tool cover -func=cover.out— aim for 99% onrepetition_guard_stream.go(excluding unreachable error branches documented as such).
Cases the in-project tests must cover (minimum):
- disabled passthrough (limit=0, identity preserved, bytes identical)
- single-token stutter at limit=N
- multi-token chain loop at limit=2
- non-repeating stream untouched
- base64 blob skipped (no trigger)
- fenced code block skipped (no trigger)
- markdown table row skipped (no trigger)
- tool-call delta channel never inspected
- per-request effective limit resolution: per-model 0 → off, per-model unset + global N → N, per-model M + global N → M
- counter increments once per trigger
Close()is idempotent and terminates the stream on EOF
This runs on 400+ machines in production — every public code path tested, no flaky timing (no real
time.Sleep; usehttptest+ injected clocks if needed, mirroringslowdown_stream_test.go).- Test files inspired by the brown test (
Status update on feat/stream-repetition-canceller (commits pushed to fork).
Done:
- In-project table-driven tests: internal/streaming/repetition_guard_stream_test.go, internal/streaming/tokenizer_test.go, internal/gateway/inference_execute_repetition_test.go, internal/server/passthrough_repetition_guard_test.go, internal/virtualmodels/repetition_test.go.
- End-to-end brown harness moved from /tmp to worktree .tmp/stream-repetition-brown/ (gitignored nested module). Runs against real HTTP SSE upstreams with gpt-4o tokenizer; latest run: guard triggered, upstream cut after 17 events, client saw 4 loop phrases, output ends with data: [DONE].
Still open / why #72 is not closed yet:
- Coverage: go test -coverprofile ./internal/streaming/... reports 85.2% statements, below the 99% target. Main gaps (functions <100%):
- repetition_guard_stream.go Read 70.8%, observe 72.7%, inspectEvent 85.0%, eventPayload 30.8%, contentDeltas 81.8%, looksLikeEncodedBlob 83.3%, isEncodedSymbol 50.0%, WithTriggerCallback 0.0%
- stream_buffer.go AppendBytes/AppendString/Read/Consume/Release 75-86%
- slowdown_stream.go / observed_sse_stream.go legacy lines pulled in by guard tests
- Docs: user-facing doc page / README note about STREAM_REPETITION_LIMIT and per-model fields still needed.
Next step to close #72: add tests to hit the remaining branches and write the docs/YAML example section.
🤖 Written by Kimi Code (AI agent) from the stream-repetition-canceller worktree.
- added a commit that references this issue
on Sep 3, 2026 Resolved on feat/stream-repetition-canceller (commit d6d49ab).
Coverage goal met:
- repetition_guard_stream.go: 197/197 = 100.0%
- tokenizer.go: 8/8 = 100.0%
- Combined new code: 205/205 = 100.0% (was 85.2%)
End-to-end smoke tests ported in-repo (CI-safe, httptest):
- internal/streaming/repetition_guard_e2e_test.go — real HTTP SSE upstreams driving the production guard: token stutter, token chain loop, clean answer byte-identical, upstream cancel propagation, empty buffer read, partial event boundaries.
- Brown scratch harness kept at .tmp/stream-repetition-brown/ (gitignored).
Docs:
- docs/features/stream-repetition-guard.mdx — detection, skip heuristics, STREAM_REPETITION_LIMIT / STREAM_REPETITION_MAX_PATTERN, per-model override semantics, Prometheus counter. Registered in docs/docs.json.
Also fixed on the way: two test lint issues (errcheck, ineffassign) and one dead continue in RepetitionGuardStream.Read.
Verification: go build ./..., targeted tests across streaming/gateway/server/virtualmodels/observability/config all green; golangci-lint 0 issues on the streaming package.
This comment was generated with AI assistance.
- added a commit that references this issue
on Sep 3, 2026
Approach (from user comment 2026-09-01)
Build a fake HTTP SSE server harness in
/tmp/(or a sibling temp dir) that produces random repetition streams from JSON fixtures, drives the guard end-to-end, and fails loud on any regression. Then port the harness into the project as the real test suite (99% coverage goal).What the harness produces
text/event-streamover HTTP viahttptest(ornet/httpon a free port).What the harness asserts
[DONE], the upstream server sees a client cancel, and the Prometheus counter increments by one.Files (in /tmp until ported)
fake_server.go— configurable loop producer over HTTP.fixtures/*.json— input scenarios.driver.go— runs each fixture, asserts behavior, fails loud.Makefile(or shell) target:make brown/./brown.sh.Port to project
internal/streaming/repetition_guard_test.go(table-driven).httptestso it runs in CI.Status: awaiting "implement" signal. Will land as the last build step.