perf(device): wake the SCPI text exchange on the line, not on a 50ms tick - #666
Conversation
…lability Both properties an event-driven rewrite of the loop could silently break, asserted against the unchanged polling implementation first. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The reply wait loop polled the collected-line count every 50ms, so an exchange finished up to a tick after its inactivity window actually expired and noticed the first reply up to a tick late. It now waits on a signal raised by the parse handler, registered under the same gate the handler appends under so no arrival can be missed. Both phases, both timeouts and the overall ceiling are unchanged. Deadlines move from DateTime.UtcNow to the exchange's Stopwatch: they are elapsed-time budgets, and a stepping wall clock would move them mid-request. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
/agentic_review |
|
Follow-up for the deferred items opened as #667. |
PR Summary by QodoPerf: make SCPI text exchanges wake on line arrival (remove 50ms polling)
AI Description
Diagram
High-Level Assessment
Files changed (2)
|
Code Review by Qodo
1.
|
|
/agentic_review |
|
Code review by qodo was updated up to the latest commit f70738d |
Ready for review — Qodo loop settled clean on
|
The new test in this branch was sized so tightly that the merge queue could not get it through: a 100 ms fragment gap inside a 300 ms completion window is barely 3x, and on the macOS and Windows runners the exchange came back with "one" alone -- the first gap by itself had outrun the window. It is the gap that also pays for JIT-ing the read and parse path and for the freshly started consumer thread's first scheduling, on a machine already running the rest of the suite in parallel. The gap is now 5 idle reads (50 ms nominal, 55 ms measured) inside a 500 ms window, so ~450 ms of cold-start delay has to land in one gap before the test lies about the exchange. Twelve fragments instead of five keep the property the test exists for: 11 gaps still outlast one completion window in total, so a wait loop that never restarts its window still fails. Both sides hold in the direction a slow machine pushes them. The elapsed time is asserted now rather than left to that arithmetic. If the pacing ever collapses, the reply arrives inside a single window and a loop that never restarted anything would look correct -- the test would pass while testing nothing. It fails instead, and says so. Sizing measured rather than assumed: on this 12-core M-series Mac the pacing is accurate to ~10% and does not degrade at 4x CPU oversubscription, so the flake does not reproduce locally at all. The margin is what carries it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…ot on sleep counts The staged reply's gap between fragments was counted in the reader thread's own idle reads -- five Thread.Sleep(10)s -- which makes the gap the SUM of five scheduling delays. On the macOS runner each of those sleeps came back at roughly 100 ms under the load of the rest of the suite, so the nominal 50 ms gap arrived at ~500 ms, filled the whole completion window, and the exchange returned six of the twelve lines. Widening the window 300 ms -> 500 ms did not help, because the stretch scales with it. The stream now hands the next fragment over on the first read that finds StageGapMs elapsed on its own Stopwatch since the previous one drained, so a stretched sleep no longer accumulates: a gap is StageGapMs plus at most ONE late read. A single 450 ms stall now has to land inside one gap before the test lies about the exchange, instead of five 90 ms ones. Fragment count goes 12 -> 16 so the nominal span (15 x 50 ms = 750 ms) still outlasts the window. Reproduced the old failure locally 2/6 with the test process at nice 20 against 48 spinners; the new pacing is 6/6 on the same harness, and 8/8 under plain CPU oversubscription. The vacuous elapsed-time guard goes with it. It asserted that the whole call took longer than one completion window, but the wait loop cannot return until a completion window has passed with no new line, so that was true however the pacing behaved. The stream now records the span from its first fragment draining to its last, and the test asserts on that instead -- the quantity that actually has to outlast the window. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Every SCPI text exchange — every SD listing, every diagnostics query, every LAN-info call, and the three exchanges a connect makes — decided that the device had finished answering by waking up every 50 ms and comparing line counts. That polling is the smaller half of the ~700 ms of fixed overhead issue #485 describes, but it is the part that is pure waste: an exchange whose inactivity window expired 1 ms after a tick still sat there until the next tick, and a reply that landed 1 ms after a tick was not noticed until the next one. It also woke a thread-pool thread twenty times a second for the whole life of every exchange, on an otherwise idle device.
The wait loop now waits on a signal instead. The parse handler that already appends each line to the collected list raises a
TaskCompletionSourceunder the same lock it appends under, and the loop registers for that signal under the same lock — so a line that arrives between the loop's last look and its next wait either is already visible to it or wakes it, never neither. Both phases, both timeouts and the overallresponseTimeoutMs * 5ceiling are exactly what they were; only the 50 ms quantisation is gone. The loop's deadlines also move fromDateTime.UtcNowto the exchange's existingStopwatch, because they are elapsed-time budgets and a stepping wall clock would otherwise move them under a request already in flight.What a reviewer should push back on. Two threading assumptions. First, the
TaskCompletionSourceis deliberately not aSemaphoreSlim: the signal is raised on the text consumer's reader thread, which can outlive the exchange (its stop and its dispose are both time-bounded), so anything needing disposal could be signalled after disposal — a TCS never needs disposing andTrySetResulton a stale one is a no-op. Second, it is constructed withRunContinuationsAsynchronously; without that the reader thread would run the exchange's continuation itself, i.e. resume the exchange — including the consumer stop that joins that very thread — on the thread being joined. Neither the signal nor its absence can hang the exchange: every wait is bounded by the same inactivity window the poll loop used, so a signal that never fires degrades to exactly today's timeout rather than to a deadlock.Measurements. Bench Nq1 on
/dev/cu.usbmodem1101, USB serial,Daqifi.Core.Cli --serial /dev/cu.usbmodem1101 --lan-chip-info(connect +InitializeAsync's three exchanges + one LAN-info text exchange + disconnect), Release build against this branch vs.origin/main, timed withtimeon the whole process:origin/mainThat is honestly a wash, and it is worth saying so plainly: the saving here is roughly 25–75 ms per exchange (up to one tick of detection latency plus up to one tick of tail), and four exchanges of that is well inside the run-to-run noise of a connect. The costs that would actually move this number are the ones this PR deliberately does not touch — see below. What this change does deliver measurably is the loop's shape: the exchange now ends when its inactivity window ends rather than at the next tick, and an idle device stops paying 20 wakeups/second per in-flight exchange (relevant to #491 item 1, which is not mine and is not otherwise addressed here).
Verification. Full suite green on net9.0 and net10.0 under
-warnaserror(3858 + 217 passing). Two guarding tests were added toTextExchangeLineFramingTestsbefore the refactor and confirmed passing against the unchanged polling code: one pins that every arriving line restarts the inactivity window (a device that dribbles five fragments out over more than one completion timeout must still return all five — a loop that measured its deadline from when it started waiting would drop the tail), and one pins that the wait stays cancellable. The fragments are paced by the reader thread's own idle reads rather than by a delay on the test thread, so a starved thread pool cannot stretch a gap into a false completion. The existingOnReplyWaitCompletedseam from #650 is what proves the branch taken is unchanged in every framing case.Deliberately deferred, and none of it is a claim that it does not matter:
Thread.Sleep(100) × 3on the init path (success criterion 3). Switching that call site to the async setup overload is two lines, but it moves the init exchange off theActionoverload that 27 test doubles across 10 files intercept — a mechanical but large diff that does not belong in the same PR as a wait-loop rewrite.I will open a follow-up issue covering the deferred items.
closes #485
Not merging — for review.