Skip to content

fix(device): make background failures visible and stop the last silent read-loop spin (closes #377, #394, #378) - #415

Merged
tylerkron merged 5 commits into
mainfrom
fix/connection-loss-detection-and-error-surface
Jul 31, 2026
Merged

tylerkron merged 5 commits into
mainfrom
fix/connection-loss-detection-and-error-surface

Conversation

@tylerkron

Copy link
Copy Markdown
Contributor

Why

When a DAQiFi device stops sending data, the library used to give you nothing to go on. A cable
pull, a dead WiFi link, or a decoder that fails on every frame all looked identical from the
outside: no samples, no error, no status change. Apps had no way to tell "the connection died"
from "the device is quiet", and no way to answer "why am I getting no samples".

Most of the detection landed in #403. This finishes the job by adding the missing piece — a place
for those failures to actually show up — and by closing the one path where the reader could still
spin forever in silence.

What

  • A new ErrorOccurred event on every device. Failures on the background threads — a failed
    read, an unparseable frame, a subscriber that threw, a frame that wouldn't decode — now reach
    your code, tagged with which stage failed. They also go to the logger, so they're visible even
    with nobody subscribed.
  • A decode-failure counter. DaqifiStreamingDevice.DecodeFailureCount tells you how many
    frames were dropped in the current streaming session. Zero on a healthy stream.
  • No more silent spinning. A stream that has gone permanently unreadable is now reported and
    escalated instead of being retried forever with nothing logged.
  • Docs: a new "Error Surface" section in docs/DEVICE_INTERFACES.md.

The event is diagnostics only. It never tears anything down, never retries, and never changes
connection status. A single bad frame is still dropped on its own without disturbing the stream,
exactly as before. Declaring a connection actually dead stays the transports' job and still arrives
as ConnectionStatus.Lost.

How

A broken thing usually breaks over and over — thousands of times a second at high sample rates — so
events are collapsed rather than fired for every occurrence. The first failure of a given kind is
reported immediately; after that, the same kind is reported at most once every five seconds, and
each report says how many were folded into it. A different kind of failure never waits behind an
ongoing storm. Reconnecting resets this, so a fresh session always reports its first problem right
away.

Handlers run on a background thread, and one that throws is caught and ignored — it can't disturb
reading or streaming.

Note for implementers

IDevice gained an event, so anything implementing that interface directly (test doubles, mocks)
needs to add it. DaqifiDevice and its subclasses are unaffected.

Testing

  • Full suite green on net9.0 and net10.0 (2213 passing each), plus the MCP tests; Release build
    with zero warnings.
  • New coverage: read failure reaching a subscriber, a decode that fails on every frame staying
    observable while the stream survives, the throttle under a 5000-failure storm, an unreadable
    stream being escalated, and an intentional disconnect still never reporting Lost.
  • Bench (Nyquist 1, FW 3.7.2): 60 s USB stream and 45 s WiFi stream, both clean — zero error
    events, zero decode failures, no spurious Lost, and Disconnect() reporting Disconnected.
    TCP keep-alive did not disturb WiFi connect or streaming.

Still to verify by hand: a physical mid-stream cable pull (someone has to actually unplug it).

closes #377
closes #394
closes #378

Not merging — for review.

🤖 Generated with Claude Code

… the last silent-spin path

Adds a device-level error event so read, parse, dispatch and per-frame decode
failures are observable instead of silent, and closes the one remaining place
the reader loop could spin forever with no data, no error and no status change.

- IDevice.ErrorOccurred + DeviceErrorEventArgs / DeviceErrorSource
- DaqifiDevice subscribes to the message consumer's ErrorOccurred (protobuf and
  text consumers), logs every failure, and raises it under a documented throttle
  (first occurrence immediate, then at most one per 5s per source+exception type,
  with the collapsed count reported)
- DaqifiStreamingDevice keeps per-frame decode isolation but now counts failures
  (DecodeFailureCount, reset per streaming session) and raises the event
- StreamMessageConsumer escalates a permanently unreadable stream instead of
  backing off silently forever

Purely observational: nothing here changes stream behaviour, retry policy or
ConnectionStatus. The transports keep sole ownership of declaring a link lost.

closes #377
closes #394
closes #378

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@tylerkron
tylerkron requested a review from a team as a code owner July 31, 2026 19:36
@tylerkron

Copy link
Copy Markdown
Contributor Author

/agentic_review

@qodo-code-review

Copy link
Copy Markdown

PR Summary by Qodo

Surface device background failures via ErrorOccurred + decode counter; stop unreadable spin

🐞 Bug fix ✨ Enhancement 🧪 Tests 📝 Documentation 🕐 40+ Minutes

Grey Divider

AI Description

• Add IDevice.ErrorOccurred to surface background read/parse/decode failures to apps.
• Throttle repeated failures and always log them to avoid event storms.
• Track per-session streaming decode drops via DecodeFailureCount.
• Escalate permanently unreadable streams instead of silently spinning forever.
• Document the new error surface and add end-to-end and throttle policy tests.
Diagram

graph TD
  iface["IDevice"] --> device["DaqifiDevice"] --> event["ErrorOccurred"] --> app["App diagnostics"]
  consumer["StreamMessageConsumer"] --> device --> throttle["DeviceErrorThrottle"]
  streaming["DaqifiStreamingDevice"] --> device --> event
  device --> logger["ILogger warnings"] --> app
Loading
High-Level Assessment

The following are alternative approaches to this PR:

1. Expose IMessageConsumer.ErrorOccurred directly to callers
  • ➕ Avoids expanding IDevice surface area
  • ➕ Keeps responsibility local to the failing component
  • ➖ Leaks internal implementation details (consumer choice/swap) into the public API
  • ➖ Doesn't naturally cover streaming decode failures (separate pipeline)
  • ➖ Harder to ensure consistent logging/throttling semantics across failure sources
2. Model failures as status transitions (e.g., Error/Degraded) instead of an event
  • ➕ Single place to watch for 'something is wrong'
  • ➕ Can be polled and stored as state
  • ➖ Conflates diagnostics with lifecycle/connection semantics (explicitly avoided here)
  • ➖ Risky behavior change for existing apps relying on current ConnectionStatus meaning
  • ➖ Hard to represent multiple simultaneous failure modes without additional structure
3. Only log failures (no public event) and rely on ILogger sinks
  • ➕ No API change to IDevice / implementers
  • ➕ Naturally throttled by logging config/filters in some environments
  • ➖ Apps cannot react programmatically (telemetry, UI, metrics, alerts)
  • ➖ Logging may be disabled or redirected; failures can remain effectively invisible
  • ➖ Harder to attach raw context (e.g., bytes) in a structured way

Recommendation: Keep the PR’s approach: a device-level diagnostic event with explicit source classification plus throttling. It provides a stable API that covers both consumer and decode pipelines, preserves existing connection semantics (Lost remains transport-owned), and remains safe under high-rate systematic faults by collapsing repeats and isolating subscriber exceptions.

Files changed (12) +1526 / -4

Enhancement (6) +461 / -2
DaqifiDevice.csAdd device-level ErrorOccurred event with logging and throttled raising +125/-0

Add device-level ErrorOccurred event with logging and throttled raising

• Adds 'ErrorOccurred' plus a 'DeviceErrorThrottle', resets throttling on connect, and wires message consumer errors (including temporary text consumer usage) into a unified device error surface. Introduces 'RaiseDeviceError' to log warnings and safely invoke subscribers without allowing handler failures to disrupt background loops.

src/Daqifi.Core/Device/DaqifiDevice.cs

DaqifiStreamingDevice.csTrack per-session decode failures and surface decode errors via device error event +34/-2

Track per-session decode failures and surface decode errors via device error event

• Adds 'DecodeFailureCount', resets it on 'StartStreaming()', and updates the per-frame decode catch to increment the counter and forward the exception through 'RaiseDeviceError(DeviceErrorSource.StreamDecode)' while preserving best-effort frame isolation.

src/Daqifi.Core/Device/DaqifiStreamingDevice.cs

DeviceErrorEventArgs.csIntroduce DeviceErrorEventArgs for background failure reporting +84/-0

Introduce DeviceErrorEventArgs for background failure reporting

• Defines the event payload including source classification, exception, suppressed occurrence count, optional raw bytes, and timestamp; enforces argument validation for error and suppressed-count.

src/Daqifi.Core/Device/DeviceErrorEventArgs.cs

DeviceErrorSource.csIntroduce DeviceErrorSource enum for pipeline-stage classification +35/-0

Introduce DeviceErrorSource enum for pipeline-stage classification

• Adds a small enum to distinguish failures originating from the message consumer vs stream decode (plus Unknown), enabling meaningful diagnostics without implying connection-loss semantics.

src/Daqifi.Core/Device/DeviceErrorSource.cs

DeviceErrorThrottle.csAdd per-bucket throttling to collapse high-rate repeated background failures +172/-0

Add per-bucket throttling to collapse high-rate repeated background failures

• Implements a throttler keyed by (source, exception runtime type) with a default 5-second interval, suppressed-count accumulation, reset-on-connect support, and a bounded bucket table with an overflow bucket for safety.

src/Daqifi.Core/Device/DeviceErrorThrottle.cs

IDevice.csExtend IDevice contract with ErrorOccurred event +11/-0

Extend IDevice contract with ErrorOccurred event

• Adds 'event EventHandler<DeviceErrorEventArgs> ErrorOccurred' to the public interface, documenting its diagnostic-only semantics and throttling guarantees.

src/Daqifi.Core/Device/IDevice.cs

Bug fix (1) +10 / -2
StreamMessageConsumer.csEscalate permanently unreadable streams instead of sleeping and looping +10/-2

Escalate permanently unreadable streams instead of sleeping and looping

• Treats '_stream.CanRead == false' as a persistent I/O fault by reporting it to the transport health sink and raising the consumer error event, preventing the reader loop from silently spinning forever on an unreadable/disposed stream.

src/Daqifi.Core/Communication/Consumers/StreamMessageConsumer.cs

Tests (4) +977 / -0
ConnectionLossEscalationTests.csAdd end-to-end tests for loss escalation and unreadable-stream reporting +355/-0

Add end-to-end tests for loss escalation and unreadable-stream reporting

• Introduces tests that drive a device through mid-stream read failures to assert 'ConnectionStatus.Lost' and the new 'ErrorOccurred' surface. Also adds coverage for the previously-silent unreadable stream loop in 'StreamMessageConsumer'.

src/Daqifi.Core.Tests/Communication/Transport/ConnectionLossEscalationTests.cs

DeviceErrorSurfaceTests.csAdd tests for device-level ErrorOccurred behavior and decode observability +429/-0

Add tests for device-level ErrorOccurred behavior and decode observability

• Verifies message-consumer failures reach subscribers, idle timeouts remain quiet, throwing handlers do not break the reader loop, and disconnect detaches error forwarding. Adds streaming tests asserting systematic decode failures keep the stream running while incrementing 'DecodeFailureCount' and emitting throttled errors.

src/Daqifi.Core.Tests/Device/DeviceErrorSurfaceTests.cs

DeviceErrorThrottleTests.csAdd unit tests pinning the error-throttle policy and bucket cap +179/-0

Add unit tests pinning the error-throttle policy and bucket cap

• Covers first-occurrence passthrough, collapsing within interval, suppressed-count reporting, independence across sources/types, zero-interval behavior, reset semantics, and bounded bucket growth via overflow bucketing.

src/Daqifi.Core.Tests/Device/DeviceErrorThrottleTests.cs

FirmwareUpdateServiceTests.csUpdate IDevice test doubles to implement ErrorOccurred +14/-0

Update IDevice test doubles to implement ErrorOccurred

• Extends several fake device implementations used by firmware update tests to include the new 'IDevice.ErrorOccurred' event, keeping compilation and behavior consistent with the updated interface.

src/Daqifi.Core.Tests/Firmware/FirmwareUpdateServiceTests.cs

Documentation (1) +78 / -0
DEVICE_INTERFACES.mdDocument the new device error surface and decode failure counter +78/-0

Document the new device error surface and decode failure counter

• Adds 'ErrorOccurred' to the IDevice interface documentation and introduces a new "Error Surface" section describing sources, non-escalating semantics, throttle policy, and 'DecodeFailureCount' usage.

docs/DEVICE_INTERFACES.md

@qodo-code-review

qodo-code-review Bot commented Jul 31, 2026 •

Copy link
Copy Markdown

Code Review by Qodo

🐞 Bugs (0) 📘 Rule violations (0) 📎 Requirement gaps (0) 🎨 UX issues (0) 🔗 Cross-repo conflicts (0) 📜 Skill insights (0)

Grey Divider


Action required

1. Fault callbacks not isolated ✓ Resolved 🐞 Bug ☼ Reliability
Description
StreamMessageConsumer.ReportStreamFault calls ITransportHealthSink.ReportIoFault and invokes
ErrorOccurred without try/catch; if either callback throws, the exception can escape the reader
thread and stop message consumption (and may be process-fatal depending on runtime settings). This
undermines the PR’s goal of preventing background-loop exceptions from escaping while handling fault
paths.
Code

src/Daqifi.Core/Communication/Consumers/StreamMessageConsumer.cs[R533-536]

+        _healthSink?.ReportIoFault(error);
+        OnErrorOccurred(error);
+        Thread.Sleep(backoffMs);
+        return true;
Relevance

●●● Strong

Team repeatedly isolates subscriber/callback exceptions on background paths (accepted in PRs #260,
#323, #354).

PR-#260
PR-#323
PR-#354

ⓘ Recommendations generated based on similar findings in past PRs

Evidence
The new helper unconditionally calls into user/transport code without a protective catch, and
OnErrorOccurred simply invokes the event; therefore a thrown exception can propagate out of
ReportStreamFault on the consumer background thread.

src/Daqifi.Core/Communication/Consumers/StreamMessageConsumer.cs[526-537]
src/Daqifi.Core/Communication/Consumers/StreamMessageConsumer.cs[605-608]
src/Daqifi.Core/Communication/Transport/ITransportHealthSink.cs[33-47]
PR-#260

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

### Issue description
`ReportStreamFault` invokes two extension points (`_healthSink.ReportIoFault` and `ErrorOccurred` via `OnErrorOccurred`) from the background reader thread without isolation. If either throws, the exception can escape the reader thread, terminating the loop (and potentially the host).

### Issue Context
This PR adds/centralizes additional fault paths through `ReportStreamFault`, and also adds commentary in `ProcessMessages` about preventing escaping exceptions. To fully deliver that guarantee, the callbacks invoked by the fault helper must be best-effort.

### Fix Focus Areas
- src/Daqifi.Core/Communication/Consumers/StreamMessageConsumer.cs[526-537]

### Suggested fix
- Wrap `_healthSink?.ReportIoFault(error)` in its own `try/catch` and swallow (optionally: log via a best-effort logger if available, but do not throw).
- Wrap `OnErrorOccurred(error)` (or the `ErrorOccurred?.Invoke(...)` call) in its own `try/catch` and swallow.
- Consider re-checking `_isRunning` before sleeping to reduce stop latency if stop is requested immediately after reporting.
- Add a regression test (new or extending existing backoff tests) where a throwing `ITransportHealthSink` and/or a throwing `StreamMessageConsumer.ErrorOccurred` subscriber does not kill the reader loop.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools



Remediation recommended

2. Test stream drops bytes ✓ Resolved 🐞 Bug ☼ Reliability ⭐ New
Description
RecoveringStream.Read clears _pending after copying only Math.Min(_pending.Length, count) bytes,
so any unread suffix is silently discarded. This can make the new backoff/continuation tests brittle
or misleading if the consumer’s bufferSize is ever configured smaller than the test payload
(partial-read scenario).
Code

src/Daqifi.Core.Tests/Communication/Consumers/StreamMessageConsumerBackoffTests.cs[R364-367]

+                var length = Math.Min(_pending.Length, count);
+                Array.Copy(_pending, 0, buffer, offset, length);
+                _pending = null;
+                return length;
Relevance

●●● Strong

Similar fix accepted: test stream helper updated to honor Stream.Read count/partial reads (PR #384).

PR-#384

ⓘ Recommendations generated based on similar findings in past PRs

Evidence
The test stream currently drops data on partial reads, which violates normal Stream semantics and
can break these tests if the consumer’s read count is smaller than the recovered payload. The
consumer’s read size is configurable, so the test double should correctly handle partial reads.

src/Daqifi.Core.Tests/Communication/Consumers/StreamMessageConsumerBackoffTests.cs[341-369]
src/Daqifi.Core/Communication/Consumers/StreamMessageConsumer.cs[75-83]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

### Issue description
`RecoveringStream.Read(...)` is not a correct `Stream` test double: it discards unread bytes when the caller-provided buffer is smaller than the pending payload. This can cause the tests to pass/fail for the wrong reason and reduces coverage for split-read framing.

### Issue Context
`StreamMessageConsumer<T>` supports configurable `bufferSize` (default 4096). A well-behaved stream may return fewer bytes than requested; a test stream that returns partial chunks must retain the remainder for subsequent reads.

### Fix Focus Areas
- src/Daqifi.Core.Tests/Communication/Consumers/StreamMessageConsumerBackoffTests.cs[354-369]

### Suggested fix
Update `RecoveringStream.Read` to only remove the bytes actually returned (e.g., keep an offset into `_pending`, or replace `_pending` with the remaining slice when `length < _pending.Length`, and only set `_pending = null` once fully drained).

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


3. Stop-time faults still surface ✓ Resolved 🐞 Bug ☼ Reliability
Description
StreamMessageConsumer.ProcessMessages reports CanRead exceptions and CanRead==false as I/O faults
even after StopSafely() has cleared _isRunning, so intentional teardown can still emit
ErrorOccurred/ReportIoFault (and add a 100ms delay) as if the connection died. This contradicts the
new outer catch’s explicit behavior of treating stop-time exceptions as teardown noise and can
pollute custom transport health sinks or consumer-level diagnostics.
Code

src/Daqifi.Core/Communication/Consumers/StreamMessageConsumer.cs[R334-363]

+                // Probing readability is itself I/O against a handle that may be coming apart, and
+                // a CanRead getter is under no obligation not to throw on one. A throwing probe is
+                // a stream fault and is handled exactly like a failing read; letting it reach the
+                // outer catch instead would neither tell the transport nor back off.
+                bool canRead;
+                try
+                {
+                    canRead = _stream.CanRead;
+                }
+                catch (Exception ex)
                {
-                    Thread.Sleep(10);
+                    _healthSink?.ReportIoFault(ex);
+                    OnErrorOccurred(ex);
+                    Thread.Sleep(100);
+                    continue;
+                }
+
+                // A stream that reports itself unreadable never becomes readable again — that is a
+                // closed or disposed stream, not a momentary lull. Report it instead of spinning
+                // here forever producing no data, no error and no status change, which is the
+                // failure mode issue #377 was filed for. Backed off at the same cadence as a
+                // failing read so the escalation timing matches.
+                if (!canRead)
+                {
+                    var unreadable = new IOException(
+                        "The stream is no longer readable; the underlying connection has been closed.");
+                    _healthSink?.ReportIoFault(unreadable);
+                    OnErrorOccurred(unreadable);
+                    Thread.Sleep(100);
                    continue;
Relevance

●●● Strong

Team treats stop-time as teardown noise; accepted multiple shutdown-window race fixes in
StreamMessageConsumer (PRs #350,#384).

PR-#350
PR-#384
PR-#403

ⓘ Recommendations generated based on similar findings in past PRs

Evidence
The new CanRead probe and unreadable-stream branches unconditionally report faults and sleep, with
no _isRunning guard; but the outer catch explicitly breaks without reporting once _isRunning is
false. Since StopSafely clears _isRunning without aborting the current loop iteration, these new
branches can still run after stop is requested, generating teardown-time fault signals.

src/Daqifi.Core/Communication/Consumers/StreamMessageConsumer.cs[334-363]
src/Daqifi.Core/Communication/Consumers/StreamMessageConsumer.cs[430-458]
src/Daqifi.Core/Communication/Consumers/StreamMessageConsumer.cs[212-219]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
`StreamMessageConsumer.ProcessMessages()` now treats `Stream.CanRead` probe failures and `CanRead == false` as I/O faults, reporting them via `_healthSink?.ReportIoFault(...)` and `OnErrorOccurred(...)`. If `StopSafely()` has already cleared `_isRunning` while the consumer thread is mid-iteration (common during teardown), these new paths can still fire and sleep 100ms before exit, producing spurious “device fault” diagnostics during intentional shutdown.

## Issue Context
The outer `catch (Exception ex)` was updated to explicitly suppress reporting when `_isRunning` is false (teardown noise), but the new CanRead/unreadable branches do not apply the same rule.

## Fix Focus Areas
- src/Daqifi.Core/Communication/Consumers/StreamMessageConsumer.cs[334-363]
- src/Daqifi.Core/Communication/Consumers/StreamMessageConsumer.cs[430-458]

## Suggested fix
- In the `catch (Exception ex)` around `_stream.CanRead`, and in the `if (!canRead)` branch:
 - If `_isRunning` is already `false`, `break` (or `return`) immediately without reporting to `_healthSink` or raising `ErrorOccurred`, and without the 100ms backoff sleep.
 - Otherwise, keep the current reporting + backoff behavior.
- Optionally, factor the stop-aware reporting into a small helper (e.g., `ReportFaultUnlessStopping(Exception ex)`), so future fault paths don’t accidentally bypass the teardown suppression policy.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


4. Text consumer retains device ✓ Resolved 🐞 Bug ☼ Reliability
Description
ExecuteTextCommandCoreAsync subscribes the temporary textConsumer.ErrorOccurred to
DaqifiDevice.OnConsumerErrorOccurred but never unsubscribes it, so if the consumer’s reader thread
outlives the using-scope (e.g., StopSafely times out on a stuck read) the consumer can keep the
DaqifiDevice object graph alive. This retention chain is new with the added error-surface wiring and
can turn a stuck-reader scenario into a long-lived memory leak.
Code

src/Daqifi.Core/Device/DaqifiDevice.cs[R1090-1094]

+                    // The protobuf consumer is stopped for the duration of this exchange, so
+                    // without this a read failure during a text command (an unplug mid-SD-listing,
+                    // say) would be the one background failure with nowhere to go (issue #378).
+                    textConsumer.ErrorOccurred += OnConsumerErrorOccurred;
+
Relevance

●●● Strong

Similar “unsubscribe to avoid device retention via event delegates” was raised and partially
accepted around producer events in #413.

PR-#413
PR-#384

ⓘ Recommendations generated based on similar findings in past PRs

Evidence
The text exchange creates a temporary StreamMessageConsumer and subscribes its ErrorOccurred event
to a device instance method; StreamMessageConsumer.Dispose explicitly may return while the reader
thread is still alive after a timed-out stop, which keeps the consumer object (and its subscriber
delegate list) reachable and therefore can retain the device.

src/Daqifi.Core/Device/DaqifiDevice.cs[1077-1095]
src/Daqifi.Core/Communication/Consumers/StreamMessageConsumer.cs[511-533]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

### Issue description
`ExecuteTextCommandCoreAsync` attaches `textConsumer.ErrorOccurred += OnConsumerErrorOccurred` but never detaches it. If `StreamMessageConsumer.Dispose()` returns while the reader thread is still alive (StopSafely join timeout), that temporary consumer instance can remain reachable and its event delegate list retains the `DaqifiDevice` instance.

### Issue Context
This is most visible in the already-supported “stuck read” scenario: stop/dispose is intentionally time-bounded and may return while a reader thread is still alive in an in-flight `Stream.Read`.

### Fix
Detach the device handler deterministically (in a `finally`) before the `using` scope ends, so even if the consumer object outlives the method, it does not retain the device.

### Fix Focus Areas
- src/Daqifi.Core/Device/DaqifiDevice.cs[1079-1095]
- src/Daqifi.Core/Communication/Consumers/StreamMessageConsumer.cs[511-533]

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


View more (1)
5. CanRead exception spins loop ✓ Resolved 🐞 Bug ☼ Reliability
Description
StreamMessageConsumer.ProcessMessages now reads _stream.CanRead outside the guarded read-failure
path, but exceptions from CanRead fall into the outer catch which has no backoff and does not call
_healthSink.ReportIoFault. If a stream’s CanRead getter throws, the loop can hot-spin issuing
repeated errors and also bypass connection-loss escalation for that failure mode.
Code

src/Daqifi.Core/Communication/Consumers/StreamMessageConsumer.cs[R334-346]

+                // A stream that reports itself unreadable never becomes readable again — that is a
+                // closed or disposed stream, not a momentary lull. Report it instead of spinning
+                // here forever producing no data, no error and no status change, which is the
+                // failure mode issue #377 was filed for. Backed off at the same cadence as a
+                // failing read so the escalation timing matches.
                if (!_stream.CanRead)
                {
-                    Thread.Sleep(10);
+                    var unreadable = new IOException(
+                        "The stream is no longer readable; the underlying connection has been closed.");
+                    _healthSink?.ReportIoFault(unreadable);
+                    OnErrorOccurred(unreadable);
+                    Thread.Sleep(100);
                    continue;
Relevance

●● Moderate

Team fixes spin/backoff paths often (e.g., watchdog exception handling in #403), but no prior
CanRead-throws precedent.

PR-#403
PR-#33

ⓘ Recommendations generated based on similar findings in past PRs

Evidence
ProcessMessages introduces an early _stream.CanRead probe; the only backoff/health reporting is in
the explicit CanRead == false branch or the read() exception branch. Any exception thrown
before/during the CanRead probe is handled by the outer catch which neither sleeps nor reports to
the health sink.

src/Daqifi.Core/Communication/Consumers/StreamMessageConsumer.cs[334-417]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

### Issue description
The new `if (!_stream.CanRead)` branch correctly reports a fault and sleeps when `CanRead` returns `false`, but if `CanRead` throws (e.g., due to a disposed/invalid stream implementation), control goes to the outer `catch (Exception ex) when (_isRunning)` which only calls `OnErrorOccurred(ex)` and immediately retries.

This can create a tight loop (CPU burn / repeated error raising) and it also skips `_healthSink?.ReportIoFault(ex)`, preventing the transport watchdog from escalating the drop.

### Issue Context
Read failures already follow the intended pattern: report to `_healthSink`, raise `ErrorOccurred`, and `Thread.Sleep(100)` backoff. The `CanRead` probe should behave the same on exceptions.

### Fix
Wrap the `CanRead` probe in its own try/catch (or expand the outer catch) so a thrown `CanRead` is treated like a read fault: report to `_healthSink`, raise `ErrorOccurred`, and apply the same backoff.

### Fix Focus Areas
- src/Daqifi.Core/Communication/Consumers/StreamMessageConsumer.cs[334-417]

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


Grey Divider

To customize comments, go to the Qodo configuration screen, or learn more in the docs.

Previous review results

Review updated until commit 048d29e

Results up to commit 0dc2327 ⚖️ Balanced


🐞 Bugs (0) 📘 Rule violations (0) 📎 Requirement gaps (0) 🎨 UX issues (0) 🔗 Cross-repo conflicts (0) 📜 Skill insights (0)


Remediation recommended
1. Text consumer retains device ✓ Resolved 🐞 Bug ☼ Reliability
Description
ExecuteTextCommandCoreAsync subscribes the temporary textConsumer.ErrorOccurred to
DaqifiDevice.OnConsumerErrorOccurred but never unsubscribes it, so if the consumer’s reader thread
outlives the using-scope (e.g., StopSafely times out on a stuck read) the consumer can keep the
DaqifiDevice object graph alive. This retention chain is new with the added error-surface wiring and
can turn a stuck-reader scenario into a long-lived memory leak.
Code

src/Daqifi.Core/Device/DaqifiDevice.cs[R1090-1094]

+                    // The protobuf consumer is stopped for the duration of this exchange, so
+                    // without this a read failure during a text command (an unplug mid-SD-listing,
+                    // say) would be the one background failure with nowhere to go (issue #378).
+                    textConsumer.ErrorOccurred += OnConsumerErrorOccurred;
+
Relevance

●●● Strong

Similar “unsubscribe to avoid device retention via event delegates” was raised and partially
accepted around producer events in #413.

PR-#413
PR-#384

ⓘ Recommendations generated based on similar findings in past PRs

Evidence
The text exchange creates a temporary StreamMessageConsumer and subscribes its ErrorOccurred event
to a device instance method; StreamMessageConsumer.Dispose explicitly may return while the reader
thread is still alive after a timed-out stop, which keeps the consumer object (and its subscriber
delegate list) reachable and therefore can retain the device.

src/Daqifi.Core/Device/DaqifiDevice.cs[1077-1095]
src/Daqifi.Core/Communication/Consumers/StreamMessageConsumer.cs[511-533]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

### Issue description
`ExecuteTextCommandCoreAsync` attaches `textConsumer.ErrorOccurred += OnConsumerErrorOccurred` but never detaches it. If `StreamMessageConsumer.Dispose()` returns while the reader thread is still alive (StopSafely join timeout), that temporary consumer instance can remain reachable and its event delegate list retains the `DaqifiDevice` instance.

### Issue Context
This is most visible in the already-supported “stuck read” scenario: stop/dispose is intentionally time-bounded and may return while a reader thread is still alive in an in-flight `Stream.Read`.

### Fix
Detach the device handler deterministically (in a `finally`) before the `using` scope ends, so even if the consumer object outlives the method, it does not retain the device.

### Fix Focus Areas
- src/Daqifi.Core/Device/DaqifiDevice.cs[1079-1095]
- src/Daqifi.Core/Communication/Consumers/StreamMessageConsumer.cs[511-533]

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


2. CanRead exception spins loop ✓ Resolved 🐞 Bug ☼ Reliability
Description
StreamMessageConsumer.ProcessMessages now reads _stream.CanRead outside the guarded read-failure
path, but exceptions from CanRead fall into the outer catch which has no backoff and does not call
_healthSink.ReportIoFault. If a stream’s CanRead getter throws, the loop can hot-spin issuing
repeated errors and also bypass connection-loss escalation for that failure mode.
Code

src/Daqifi.Core/Communication/Consumers/StreamMessageConsumer.cs[R334-346]

+                // A stream that reports itself unreadable never becomes readable again — that is a
+                // closed or disposed stream, not a momentary lull. Report it instead of spinning
+                // here forever producing no data, no error and no status change, which is the
+                // failure mode issue #377 was filed for. Backed off at the same cadence as a
+                // failing read so the escalation timing matches.
                if (!_stream.CanRead)
                {
-                    Thread.Sleep(10);
+                    var unreadable = new IOException(
+                        "The stream is no longer readable; the underlying connection has been closed.");
+                    _healthSink?.ReportIoFault(unreadable);
+                    OnErrorOccurred(unreadable);
+                    Thread.Sleep(100);
                    continue;
Relevance

●● Moderate

Team fixes spin/backoff paths often (e.g., watchdog exception handling in #403), but no prior
CanRead-throws precedent.

PR-#403
PR-#33

ⓘ Recommendations generated based on similar findings in past PRs

Evidence
ProcessMessages introduces an early _stream.CanRead probe; the only backoff/health reporting is in
the explicit CanRead == false branch or the read() exception branch. Any exception thrown
before/during the CanRead probe is handled by the outer catch which neither sleeps nor reports to
the health sink.

src/Daqifi.Core/Communication/Consumers/StreamMessageConsumer.cs[334-417]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

### Issue description
The new `if (!_stream.CanRead)` branch correctly reports a fault and sleeps when `CanRead` returns `false`, but if `CanRead` throws (e.g., due to a disposed/invalid stream implementation), control goes to the outer `catch (Exception ex) when (_isRunning)` which only calls `OnErrorOccurred(ex)` and immediately retries.

This can create a tight loop (CPU burn / repeated error raising) and it also skips `_healthSink?.ReportIoFault(ex)`, preventing the transport watchdog from escalating the drop.

### Issue Context
Read failures already follow the intended pattern: report to `_healthSink`, raise `ErrorOccurred`, and `Thread.Sleep(100)` backoff. The `CanRead` probe should behave the same on exceptions.

### Fix
Wrap the `CanRead` probe in its own try/catch (or expand the outer catch) so a thrown `CanRead` is treated like a read fault: report to `_healthSink`, raise `ErrorOccurred`, and apply the same backoff.

### Fix Focus Areas
- src/Daqifi.Core/Communication/Consumers/StreamMessageConsumer.cs[334-417]

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


Results up to commit 3908260 ⚖️ Balanced


🐞 Bugs (0) 📘 Rule violations (0) 📎 Requirement gaps (0) 🎨 UX issues (0) 🔗 Cross-repo conflicts (0) 📜 Skill insights (0)


Remediation recommended
1. Stop-time faults still surface ✓ Resolved 🐞 Bug ☼ Reliability
Description
StreamMessageConsumer.ProcessMessages reports CanRead exceptions and CanRead==false as I/O faults
even after StopSafely() has cleared _isRunning, so intentional teardown can still emit
ErrorOccurred/ReportIoFault (and add a 100ms delay) as if the connection died. This contradicts the
new outer catch’s explicit behavior of treating stop-time exceptions as teardown noise and can
pollute custom transport health sinks or consumer-level diagnostics.
Code

src/Daqifi.Core/Communication/Consumers/StreamMessageConsumer.cs[R334-363]

+                // Probing readability is itself I/O against a handle that may be coming apart, and
+                // a CanRead getter is under no obligation not to throw on one. A throwing probe is
+                // a stream fault and is handled exactly like a failing read; letting it reach the
+                // outer catch instead would neither tell the transport nor back off.
+                bool canRead;
+                try
+                {
+                    canRead = _stream.CanRead;
+                }
+                catch (Exception ex)
                {
-                    Thread.Sleep(10);
+                    _healthSink?.ReportIoFault(ex);
+                    OnErrorOccurred(ex);
+                    Thread.Sleep(100);
+                    continue;
+                }
+
+                // A stream that reports itself unreadable never becomes readable again — that is a
+                // closed or disposed stream, not a momentary lull. Report it instead of spinning
+                // here forever producing no data, no error and no status change, which is the
+                // failure mode issue #377 was filed for. Backed off at the same cadence as a
+                // failing read so the escalation timing matches.
+                if (!canRead)
+                {
+                    var unreadable = new IOException(
+                        "The stream is no longer readable; the underlying connection has been closed.");
+                    _healthSink?.ReportIoFault(unreadable);
+                    OnErrorOccurred(unreadable);
+                    Thread.Sleep(100);
                    continue;
Relevance

●●● Strong

Team treats stop-time as teardown noise; accepted multiple shutdown-window race fixes in
StreamMessageConsumer (PRs #350,#384).

PR-#350
PR-#384
PR-#403

ⓘ Recommendations generated based on similar findings in past PRs

Evidence
The new CanRead probe and unreadable-stream branches unconditionally report faults and sleep, with
no _isRunning guard; but the outer catch explicitly breaks without reporting once _isRunning is
false. Since StopSafely clears _isRunning without aborting the current loop iteration, these new
branches can still run after stop is requested, generating teardown-time fault signals.

src/Daqifi.Core/Communication/Consumers/StreamMessageConsumer.cs[334-363]
src/Daqifi.Core/Communication/Consumers/StreamMessageConsumer.cs[430-458]
src/Daqifi.Core/Communication/Consumers/StreamMessageConsumer.cs[212-219]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
`StreamMessageConsumer.ProcessMessages()` now treats `Stream.CanRead` probe failures and `CanRead == false` as I/O faults, reporting them via `_healthSink?.ReportIoFault(...)` and `OnErrorOccurred(...)`. If `StopSafely()` has already cleared `_isRunning` while the consumer thread is mid-iteration (common during teardown), these new paths can still fire and sleep 100ms before exit, producing spurious “device fault” diagnostics during intentional shutdown.

## Issue Context
The outer `catch (Exception ex)` was updated to explicitly suppress reporting when `_isRunning` is false (teardown noise), but the new CanRead/unreadable branches do not apply the same rule.

## Fix Focus Areas
- src/Daqifi.Core/Communication/Consumers/StreamMessageConsumer.cs[334-363]
- src/Daqifi.Core/Communication/Consumers/StreamMessageConsumer.cs[430-458]

## Suggested fix
- In the `catch (Exception ex)` around `_stream.CanRead`, and in the `if (!canRead)` branch:
 - If `_isRunning` is already `false`, `break` (or `return`) immediately without reporting to `_healthSink` or raising `ErrorOccurred`, and without the 100ms backoff sleep.
 - Otherwise, keep the current reporting + backoff behavior.
- Optionally, factor the stop-aware reporting into a small helper (e.g., `ReportFaultUnlessStopping(Exception ex)`), so future fault paths don’t accidentally bypass the teardown suppression policy.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


Results up to commit 4b19d59 ⚖️ Balanced


🐞 Bugs (0) 📘 Rule violations (0) 📎 Requirement gaps (0) 🎨 UX issues (0) 🔗 Cross-repo conflicts (0) 📜 Skill insights (0)


Action required
1. Fault callbacks not isolated ✓ Resolved 🐞 Bug ☼ Reliability
Description
StreamMessageConsumer.ReportStreamFault calls ITransportHealthSink.ReportIoFault and invokes
ErrorOccurred without try/catch; if either callback throws, the exception can escape the reader
thread and stop message consumption (and may be process-fatal depending on runtime settings). This
undermines the PR’s goal of preventing background-loop exceptions from escaping while handling fault
paths.
Code

src/Daqifi.Core/Communication/Consumers/StreamMessageConsumer.cs[R533-536]

+        _healthSink?.ReportIoFault(error);
+        OnErrorOccurred(error);
+        Thread.Sleep(backoffMs);
+        return true;
Relevance

●●● Strong

Team repeatedly isolates subscriber/callback exceptions on background paths (accepted in PRs #260,
#323, #354).

PR-#260
PR-#323
PR-#354

ⓘ Recommendations generated based on similar findings in past PRs

Evidence
The new helper unconditionally calls into user/transport code without a protective catch, and
OnErrorOccurred simply invokes the event; therefore a thrown exception can propagate out of
ReportStreamFault on the consumer background thread.

src/Daqifi.Core/Communication/Consumers/StreamMessageConsumer.cs[526-537]
src/Daqifi.Core/Communication/Consumers/StreamMessageConsumer.cs[605-608]
src/Daqifi.Core/Communication/Transport/ITransportHealthSink.cs[33-47]
PR-#260

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

### Issue description
`ReportStreamFault` invokes two extension points (`_healthSink.ReportIoFault` and `ErrorOccurred` via `OnErrorOccurred`) from the background reader thread without isolation. If either throws, the exception can escape the reader thread, terminating the loop (and potentially the host).

### Issue Context
This PR adds/centralizes additional fault paths through `ReportStreamFault`, and also adds commentary in `ProcessMessages` about preventing escaping exceptions. To fully deliver that guarantee, the callbacks invoked by the fault helper must be best-effort.

### Fix Focus Areas
- src/Daqifi.Core/Communication/Consumers/StreamMessageConsumer.cs[526-537]

### Suggested fix
- Wrap `_healthSink?.ReportIoFault(error)` in its own `try/catch` and swallow (optionally: log via a best-effort logger if available, but do not throw).
- Wrap `OnErrorOccurred(error)` (or the `ErrorOccurred?.Invoke(...)` call) in its own `try/catch` and swallow.
- Consider re-checking `_isRunning` before sleeping to reduce stop latency if stop is requested immediately after reporting.
- Add a regression test (new or extending existing backoff tests) where a throwing `ITransportHealthSink` and/or a throwing `StreamMessageConsumer.ErrorOccurred` subscriber does not kill the reader loop.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


Qodo Logo

Comment thread src/Daqifi.Core/Device/DaqifiDevice.cs
Comment thread src/Daqifi.Core/Communication/Consumers/StreamMessageConsumer.cs
…ts scope

Addresses Qodo review on #415, plus a process-crash hazard the regression test
exposed while verifying the fix.

- Text exchange: the temporary consumer's error forwarding is now scope-bound
  (ConsumerErrorSubscription), so a reader that outlives the exchange's bounded
  stop/dispose can neither retain the device nor keep raising errors on it.
- CanRead probe: a throwing readability getter is a stream fault and is handled
  like a failing read (health sink + error + backoff) instead of falling to the
  outer catch, which neither reported nor backed off.
- Outer catch: added a backoff. A parser that throws on the bytes it holds
  throws on the same bytes next iteration, so retrying at full speed was a hot
  spin — measured at 59k error raises in 700ms. Deliberately not reported to the
  health sink: a parse failure is not evidence the link is gone.
- Outer catch is now unconditional. `when (_isRunning)` left a hole where a stop
  landing mid-try made the exception escape a background thread and terminate
  the host process (it crashed the test host). Only reporting was ever meant to
  be conditional.

Six regression tests added, each verified to fail on the pre-fix code.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@tylerkron

Copy link
Copy Markdown
Contributor Author

/agentic_review

Comment thread src/Daqifi.Core/Communication/Consumers/StreamMessageConsumer.cs
@qodo-code-review

Copy link
Copy Markdown

Code review by qodo was updated up to the latest commit 3908260

Addresses Qodo round 2 on #415. The CanRead paths reported I/O faults even
after StopSafely() cleared the running flag, contradicting the outer catch's
teardown-noise rule added in the previous commit — the same event got two
different answers depending on which path caught it.

All stream-fault sites (CanRead throw, CanRead false, read exception, socket
EOF) now route through one ReportStreamFault helper that states the rule once:
nothing is reported once a stop has been requested, and the loop exits at once
instead of sleeping out a backoff it no longer needs. Parse/dispatch failures
deliberately stay outside it — they are not evidence the link is gone.

Checked whether this could breach #377's "intentional Disconnect() never
reports Lost": it could not. A continue re-tests the loop condition, so at most
ONE fault could ever be reported after a stop, against an escalation threshold
of five consecutive — measured at exactly 1 with the guard removed. The
transports also disarm their watchdog before touching the handle, and
DaqifiDevice._isDisconnecting independently suppresses Lost. Diagnostic noise,
not a hole in the guarantee — but noise the new device-level error event would
have made user-visible on every disconnect.

Two regression tests, both verified to fail on the pre-fix code, plus a
teardown-silence assertion on the existing device-level disconnect test.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@tylerkron

Copy link
Copy Markdown
Contributor Author

/agentic_review

Comment thread src/Daqifi.Core/Communication/Consumers/StreamMessageConsumer.cs Outdated
@qodo-code-review

Copy link
Copy Markdown

Code review by qodo was updated up to the latest commit 4b19d59

…the reader loop

Addresses Qodo round 3 on #415. ReportStreamFault invoked ITransportHealthSink
and ErrorOccurred without isolation, so a throwing callback escaped the reader
thread — process-fatal, the same shape as the escaping-catch defect fixed last
round, one layer up. Concentrating the four fault sites into one helper made a
single unguarded callback affect all of them at once.

The route is two hops, and the crash trace confirms it exactly: the handler
throws on the fault path, the loop's outer catch reports that failure by
calling the same handler, and the second throw is inside a catch block with
nothing above it. Verified against pre-fix code — the test host dies with the
unhandled exception surfacing from ProcessMessages' outer catch.

Every callback out of the loop is now isolated via SafeReportIoFault /
SafeReportIoSuccess / SafeRaiseError: the two ReportStreamFault callbacks, the
per-read success report, the outer catch's error raise, and the error raise in
ProcessMessageBuffer's dispatch handler. Plain methods rather than a lambda
helper so the once-per-read success path allocates no closure. Swallowed rather
than logged, matching the convention already used for a throwing
MessageReceived subscriber (#180) and mirrored across RaiseClassifiedEvent
(#323), AllTransportsDeviceFinder (#354), DeviceFinderBase and
RaiseGapDetected.

Two regression tests assert the reader keeps consuming — a real message
delivered after the throwing phase — not merely that nothing surfaced. Both
verified failing pre-fix: the subscriber case crashes the host, the health-sink
case silently stops delivering messages.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@tylerkron

Copy link
Copy Markdown
Contributor Author

/agentic_review

@qodo-code-review

Copy link
Copy Markdown

Code review by qodo was updated up to the latest commit aebb59d

Addresses Qodo round 4 on #415. RecoveringStream.Read copied
min(payload, count) bytes and then discarded the unread suffix, so the tests
that depend on it were silently coupled to the consumer's read buffer being
larger than the payload — they would have kept passing for a reason unrelated
to the code under test, and would have started failing on an unrelated
bufferSize change. A test double that violates the contract is a latent
false-negative generator, and these tests are the evidence for this PR's
claims.

RecoveringStream now retains the remainder across calls and clears the payload
only once fully drained. DeviceErrorSurfaceTests.ScriptedStream had the same
defect in a queue nothing enqueues to any more, so the queue is deleted rather
than fixed.

Added TheRecoveringStreamHelper_DeliversAWholePayloadAcrossPartialReads, which
drives the helper with a one-byte read buffer so the partial-read path is
actually exercised; verified it fails against the old helper ("the payload
never arrived in full").

Re-verified with the corrected helper that the round-3 tests still fail against
pre-fix production code: the throwing-subscriber case still crashes the host,
the throwing-health-sink case still stops consuming.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@tylerkron

Copy link
Copy Markdown
Contributor Author

/agentic_review

@qodo-code-review

Copy link
Copy Markdown

Code review by qodo was updated up to the latest commit 048d29e

@tylerkron
tylerkron merged commit fcdbb7e into main Jul 31, 2026
1 check passed
@tylerkron
tylerkron deleted the fix/connection-loss-detection-and-error-surface branch July 31, 2026 22:27
@tylerkron
tylerkron restored the fix/connection-loss-detection-and-error-surface branch July 31, 2026 22:28
@tylerkron
tylerkron deleted the fix/connection-loss-detection-and-error-surface branch July 31, 2026 22:28
tylerkron added a commit that referenced this pull request Jul 31, 2026
Resolves against #415 (connection-loss detection and the device ErrorOccurred
surface) and #417 (SD->LAN restore inside the exchange lock, which added a
finalizeAsync phase to ExecuteTextCommandAsync).

One real conflict, in DaqifiDeviceInitializeTests: #417 wrapped the testable
device's ExecuteTextCommandAsync body in try/finally to honor the new finalize
phase and re-indented it, while this branch had inserted a
MutateDuringInitialization hook between the prepare phase and setupAction.
Kept both — the hook now sits inside the new try block, still after prepare and
before setupAction, so the two overlapping-initialization tests still mutate
state at the intended point.

The two test doubles this branch added (OverlappingInitDevice and
CancelDuringCapabilityReadDevice) also had to widen their
ExecuteTextCommandAsync overrides for the finalizeAsync parameter, and now honor
the finalize phase the way the other doubles do. That compile break is the seam
from #406 working as designed.

Verified nothing from main was dropped: the only deletions relative to
origin/main are this branch's three intended OnDeviceInitializingAsync signature
changes. #415's ErrorOccurred wiring, OnConsumerErrorOccurred subscription and
the Connected->Lost transition, and #417's finalizeAsync phase are all intact,
as are this branch's PreserveActiveStream command skipping and the pre-Ready
cancellation guard.

Full suite green on net9.0 and net10.0 (2246 Core + 23 MCP).
tylerkron added a commit that referenced this pull request Jul 31, 2026
Resolves against #415 (background-error surface, connection-loss escalation)
and #417 (SD->LAN restore inside the exchange lock).

Three conflicts, all where #415 edited the same connect/disconnect bodies this
branch factored into shared sync/async step sets:

- IDevice.cs: #415's ErrorOccurred event landed immediately before Connect(),
  whose doc comment this branch rewrote. Kept both.
- DaqifiDevice.Connect(): #415's consumer ErrorOccurred subscription moved into
  the shared CompleteConnect(), so the async path wires it too.
- DaqifiDevice.Disconnect(): #415's ErrorOccurred unsubscribe moved into the
  shared StopMessagePumps(), reached by both Disconnect() and DisconnectAsync().

_errorThrottle.Reset() moved from Connect() into the shared BeginConnect().
Leaving it on the sync path alone would have quietly dropped #415's per-session
reset from the primary connect path, since the factory now connects through
ConnectAsync. Nothing in #415's suite covers that reset, so it would have
survived a fully green build.

Added two regression tests for that seam — every #415 test drives Connect(),
because ConnectAsync() did not exist when they were written. Both verified to
fail when the connect-side wiring is dropped.

OnTransportStatusChanged, the _isDisconnecting guard and #417's finalizeAsync
plumbing are byte-identical to main.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment