fix(tests): stop the serial-probe teardown tests racing a 1s response window - #609
Merged
Merged
Conversation
… window SerialProbeTeardownTests answered its scripted probe stream only after `await stream.WaitForFirstReadAsync()`, which dispatched a `Task.Run` and then resumed the test method — two thread-pool queue hops inside the probe's fixed 1000ms `ResponseTimeoutMs`. On a loaded CI runner those hops can outlast the window, so `RequestDeviceStatusAsync` gave up and returned null before the test ever got to answer. The wait is now synchronous on the calling thread. The consumer's reader is a dedicated Thread, not a pool work item, so it reaches its first read regardless of pool pressure and the wait returns promptly — the device still answers while the reader is parked in Read(), and no assertion changed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Contributor
Author
|
/agentic_review |
PR Summary by QodoDeflake SerialProbeTeardownTests by making first-read wait synchronous
AI Description
Diagram
High-Level Assessment
Files changed (1)
|
Code Review by Qodo🐞 Bugs (0) 📘 Rule violations (0) 📎 Requirement gaps (0)
Great, no issues found!Qodo reviewed your code and found no material issues that require reviewTip of the day💡 Did you know, you can turn these tips off under Display preferences |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Produced by the FLAKY-TEST-FIXER maintenance routine. Test-only change: no production code is touched and no assertion was weakened, so the only thing it can affect is how reliably the existing assertions get a chance to run.
What was wrong.
SerialProbeTeardownTests.RequestDeviceStatusAsync_WhenItReturns_ReaderThreadHasExitedfailed on CI on net10.0 while the same commit passed on net9.0 (run 31724429734, merge-queue run for #514) withAssert.NotNull() Failure: Value is null— the probe returned "no device" instead of the status it was scripted to receive. The cause is a wall-clock race, not a bug in the probe.SerialDeviceFinder.RequestDeviceStatusAsyncgives a device a fixed 1000 ms (ResponseTimeoutMs) to answer, and that clock starts the moment the test calls it. The test then let the device answer only afterawait stream.WaitForFirstReadAsync(), whose body wasawait Task.Run(() => _firstRead.Wait(token))— two thread-pool queue hops (one to dispatch theTask.Runbody, one to resume the awaiting test method) spent inside that 1000 ms budget. CI runs both target frameworks' test processes at once, each running xunit collections in parallel; when the pool is saturated it injects new workers only about twice a second, so those two hops can easily outlast the window. The probe gave up and returned null before the test ever got to callRespondWithStatus(). Two sibling tests in the same file used the same helper and carried the same latent race.How it was fixed.
WaitForFirstReadAsyncbecomes a plain synchronousWaitForFirstRead()that waits theManualResetEventSlimdirectly on the calling thread. That is what removes the race rather than hiding it: the thing being waited for is the consumer's reader, whichStreamMessageConsumerruns on a dedicatedThreadrather than a pool work item, so it reaches its first read regardless of pool pressure — while the old helper needed the pool to schedule two continuations before the test was allowed to answer. Nothing about the scenario changes: the device still answers while the reader is genuinely parked inRead(), theContinueWithis still attached before any answer can arrive (so it still samples reader-liveness at completion rather than running vacuously inline), and every assertion is byte-for-byte the same. Widening a timeout or relaxingAssert.NotNullwould have been the hiding fix; this deletes the pool dependency instead. TheWaitis still bounded at 5 s and now asserts on expiry, so a genuine regression that stops the reader from reading fails loudly instead of hanging.Verified: full suite green on both target frameworks locally (3730 passed / 0 failed on net9.0 and net10.0, plus Daqifi.Mcp.Tests 217/0), build clean with 0 warnings, and the touched class passes 11/11 on both.
Not merging — opened for review.