Skip to content

test(device): a raw-capture test was failing on CI for a reason that had nothing to do with the PR under review - #517

Merged
tylerkron merged 1 commit into
mainfrom
fix/rawcapture-hammer-flake-516
Aug 13, 2026
Merged

tylerkron merged 1 commit into
mainfrom
fix/rawcapture-hammer-flake-516

Conversation

@tylerkron

Copy link
Copy Markdown
Contributor

What was wrong. One of the raw-capture tests (RawCapture_UnderASendHammer_...) would go red on GitHub Actions on pull requests that had not touched any of the code it covers — it failed twice in a row on #515, a branch that only changes Windows discovery. The test fires a stream of Send() calls at a device while a raw capture holds it, then checks that the very first of those messages (HAMMER-0) was parked and replayed rather than thrown away. But Send()'s hold-back backlog is capped, and when it overflows it discards its oldest entry — so HAMMER-0 is by definition the first casualty. The hammer ran for as long as the capture window stayed open, so on a fast machine it parked a couple of hundred messages and passed, and on a loaded CI runner it parked thousands, evicted the message the test was waiting for, and failed. Nothing was wrong with the library; the test was asserting something the design does not promise.

How it was fixed. The hammer now sends a fixed number of messages instead of running until told to stop, and that number is derived from the backlog cap (DefaultMaxDeferredSends / 8) rather than picked by hand, so the two cannot drift apart silently. To keep the test as strong as it was, the capture now holds the stream until the hammer has finished, which makes "every one of those sends happened while the capture owned the device" true by construction rather than by hoping the window outlasts the sender. A DroppedDeferredSendCount == 0 assertion sits next to the original wait, so if anyone does raise the budget past the cap the failure says "112 messages were dropped" instead of the misleading "HAMMER-0 never reached the wire". No production code changed.

The one thing worth pushing back on: the capture task now waits on the hammer, which is a dependency the old version did not have. It is signalled from a finally, so a hammer that throws surfaces as its own failure rather than hanging the capture.

Verified. Simulating a loaded runner by serving the capture payload one byte at a time (a ~10 s window) reproduces the CI failure on the old test with the exact same message, and the new test passes 3/3 under those same conditions. Two more mutations confirm the assertions are real catchers: shrinking the backlog cap to 16 trips the new dropped-count assertion, and disabling send deferral entirely still fails the test — now catching all 128 sends instead of a timing-dependent subset. Full suite green on net9.0 (3045 Core + 86 Mcp) and net10.0 (3045), 0 warnings. Bench health check on the Nq1 (fw 3.7.2, /dev/cu.usbmodem1101), non-destructive: discovery, 3 s @ 500 Hz on channels 0-2 → 1186 samples (this unit's usual ratio), SD storage 7.80 GB, clean disconnect — unchanged, as expected for a test-only change.

closes #516

Not merging — opened for your review.

The hammer sent for as long as the raw capture held the device, then the test
waited for HAMMER-0 to reach the wire. But the deferred-send backlog is capped
at 1024 and overflows drop-OLDEST, so once the hammer parked more than the cap
HAMMER-0 was the first thing evicted — and the test failed with "HAMMER-0 never
reached the wire". That only happened when the capture window ran long, which is
why it passed locally (~130 ms) and reddened unrelated PRs on loaded runners
(6-12 s), twice in a row on #515.

The hammer now fires a fixed budget derived from the cap itself
(DefaultMaxDeferredSends / 8), so nothing can be evicted however long the window
stays open, and the capture holds the stream until the hammer is done, so every
one of those sends is issued inside the window by construction rather than by
timing. DroppedDeferredSendCount == 0 is asserted alongside, so if the budget and
the cap ever drift apart the test says so instead of failing on HAMMER-0.

Verified by mutation: serving the payload one byte at a time (a ~10 s window,
the loaded-runner shape) fails the old test with the exact CI message and passes
the new one 3/3; capping the backlog at 16 trips the new dropped-count
assertion; disabling TryDeferSend still fails the test, now catching all 128
sends rather than a timing-dependent subset.

closes #516

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@tylerkron
tylerkron requested a review from a team as a code owner August 13, 2026 16:17
@tylerkron

Copy link
Copy Markdown
Contributor Author

/agentic_review

@qodo-code-review

Copy link
Copy Markdown

PR Summary by Qodo

Fix RawCapture send-hammer CI flake by budgeting deferred sends

🧪 Tests 🐞 Bug fix 🕐 20-40 Minutes

Grey Divider

AI Description

• Bound the raw-capture send hammer to a fixed budget derived from the deferral cap.
• Hold the capture window open until the hammer completes to remove timing dependence.
• Add explicit assertions for dropped deferred sends and in-window hammer coverage.
Diagram

sequenceDiagram
participant T as "RawCapture lock test"
participant D as "DaqifiDevice"
participant S as "Capture stream"
participant Q as "Deferred backlog"
participant W as "Transport writes"
T->>D: RunRawCaptureAsync(...)
D->>S: Open capture window
par Send hammer (budgeted)
T->>D: Send HAMMER-0..HAMMER-(N-1)
D->>Q: Defer sends while capture owns device
end
T->>S: Read payload until hammerFinished
T->>D: Assert DroppedDeferredSendCount == 0
D->>Q: Replay deferred sends after capture
Q->>W: Write HAMMER-0...
T->>W: WaitForWrite("HAMMER-0")
Loading
High-Level Assessment

The following are alternative approaches to this PR:

1. Slow down / throttle the hammer until it stays under the cap
  • ➕ Minimal structural change to the test logic
  • ➖ Still timing-dependent and machine-load sensitive; can re-flake as CI characteristics change
  • ➖ Doesn’t guard against cap/budget drift explicitly
2. Assert on the newest surviving hammer message instead of HAMMER-0
  • ➕ Avoids dependence on drop-oldest semantics
  • ➖ Weakens the original intent (proves “some deferred messages replayed”, not that the earliest message survived)
  • ➖ Can mask regressions where early commands are preferentially lost
3. Increase DefaultMaxDeferredSends to avoid overflow in tests
  • ➕ Would reduce likelihood of overflow across the suite
  • ➖ Production behavior change (memory + replay burst) for a test problem
  • ➖ Doesn’t prevent future flakes if the test keeps sending unboundedly

Recommendation: Keep the PR’s approach: derive a fixed HammerSendBudget from DefaultMaxDeferredSends and hold the capture open until the hammer completes. This removes timing dependence, preserves the strength of the original assertion (HAMMER-0 must replay), and adds an explicit DroppedDeferredSendCount==0 signal that explains failures when cap/budget assumptions drift.

Files changed (1) +42 / -11

Tests (1) +42 / -11
DaqifiDeviceRawCaptureLockTests.csDe-flake RawCapture send-hammer by fixed budget + dropped-count assertion +42/-11

De-flake RawCapture send-hammer by fixed budget + dropped-count assertion

• Introduces a HammerSendBudget derived from DaqifiDevice.DefaultMaxDeferredSends to keep the hammer safely under the deferred-send cap. Updates the raw-capture/hammer test to keep the capture window open until the hammer finishes, asserts all hammer sends occurred during the capture window, and adds DroppedDeferredSendCount==0 alongside the existing HAMMER-0 replay check.

src/Daqifi.Core.Tests/Device/DaqifiDeviceRawCaptureLockTests.cs

@qodo-code-review

Copy link
Copy Markdown

Code Review by Qodo

🐞 Bugs (0) 📘 Rule violations (0) 📎 Requirement gaps (0)

Grey Divider

Great, no issues found!

Qodo reviewed your code and found no material issues that require review

Grey Divider

Tip of the day
💡 Did you know, you can type 'qodo, fix this' on a finding and the fix lands right on your PR

More tips ↗ | Customize Qodo ↗ | Qodo docs ↗

Grey Divider

Qodo Logo

@tylerkron

Copy link
Copy Markdown
Contributor Author

Qodo-clean, CI green — ready for review. (1 round on head fb3c8ec: Bugs (0) / Rule violations (0) / Requirement gaps (0), zero review threads; settle re-check ~5.5 min later returned a byte-identical summary, still 0 threads, head unchanged. No fixes were needed, so no production code changed after the bench run.)

@tylerkron
tylerkron added this pull request to the merge queue Aug 13, 2026
Merged via the queue into main with commit d49fc51 Aug 13, 2026
1 check passed
@tylerkron
tylerkron deleted the fix/rawcapture-hammer-flake-516 branch August 13, 2026 17:16
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

test(device): RawCapture_UnderASendHammer flakes on slow runners — it waits for the one message the backlog cap drops first

1 participant