ops(gate): the chaos legs — api-down and engine-reconnect (STORY-447, gh-#777) - #792
Merged
Merged
Conversation
…#777) With capture running (F178.8a): stop/start api around GATE_OUTAGE_SECS, recovery polled through the engine metadata against GATE_RECOVERY_SECS, silence during the outage attributed to the chaos leg (not capture) under --chaos, chaos report block in md+json, GATE_OUTAGE_SECS/GATE_RECOVERY_SECS knob validation. Story447 api-down facts green.
After the api-down scenario passes, --chaos restarts the engine container, waits for api health then first on-air within GATE_RECONNECT_SECS (60), records a fixed 60 s capture-reconnect.wav and measures it: a timeout fails the leg "engine-reconnect on-air", a recording failure "engine-reconnect capture", a measure error "engine-reconnect measure", any silence event "engine-reconnect silence" (loudness is not judged). The report gains engine-reconnect on-air seconds and silence events in the md and json; the new knob is validated like GATE_POLL_SECS. Story447 facts for the engine scenario are un-skipped; the harness stub passes GATE_RECONNECT_SECS=4 and answers the restart marker.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to subscribe to this conversation on GitHub.
Already have an account?
Sign in.
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
PR-7 of the gh-#777 epic (STORY-447, SPEC F178.8). The gate gains
--chaos: two scenarios that break the running stack on purpose and measure how the stream and the stack come back.💥 What
docker compose stop api, sleepGATE_OUTAGE_SECS(90),start api, then poll for a non-safetrack_idwithinGATE_RECOVERY_SECS(120). The capture is measured once and attributed: silence during a chaos run fails the chaos legapi-down silence(the loudness check is skipped, a silent stretch masks it), silence without chaos fails the capture leg as before. A recovery timeout failsapi-down recovery.docker compose restart engine, wait for api health then first on-air withinGATE_RECONNECT_SECS(60, gap counted from the restart command), then a fixed 60 scapture-reconnect.wavmeasured for silence only. Failures:engine-reconnect on-air / capture / measure / silence.--chaosrequires--capture(exit 2, like--capturerequires--fresh). The three knobs validate likeGATE_POLL_SECS.api-down outage / recovery,engine-reconnect on-air / silence eventsin the md; json twins underlegs.chaos.measurements(nulls when a scenario never ran).Story447_ChaosLegs.cs: 17 facts green; the harness stub answersstop api/start api/restart enginewith markers and can be told to stay dark after a stop or a restart.🔌 Real runs against v5.8.3 (dev box, dev station stopped)
30 s outage (
GATE_OUTAGE_SECS=30 tools/gate/stack_gate.sh --tag v5.8.3 --fresh --capture --chaos): exit 0, wall 497 s.Gate defaults, 90 s outage (T498): exit 1,
chaos | ran/failed | api-down silence, silence events 2, recovery 0 s. Twice, reproducibly. The engine log puts the silence on the timeline: the main queue drains 63 s into the outage, the safe branch plays its single prefetched track, then the 7 s gap, thenmksafeblank until the api answers/internal/safe-trackagain (about 23 s of silence). Filed as #791. So the gate did its job: F178.8(a) "mksafe/safe loop holds" is not true of v5.8.3 once the queue drains, and T498 stays open until #791 is decided (fix the product, tolerate the designed gap in the gate overlay, or amend the spec).🔍 Review rounds worth knowing
-tand thedown -v).track_idover a fresh TCP client. The header comment now says so.📝 For later tasks
docker logs -f engine.track_id; a track already queued before the outage gives ~0 s. A stricter probe would wait for atrack_idpushed afterstart api.actions/checkoutneedsfetch-depth: 0orfetch-tags: truefor the upgrade leg.✅ Gate
Full solution,
Category!=Integration: 9 projects, 0 failed (Host 3008 passed / 110 skipped, Architecture 150 / 10).bash -nand shellcheck clean.