Skip to content

Raise CAN device and bus health alerts, with a chain break hint - #57

Open
nlaverdure wants to merge 8 commits into
main-2027-alpha7from
can-health
Open

nlaverdure wants to merge 8 commits into
main-2027-alpha7from
can-health

Conversation

@nlaverdure

@nlaverdure nlaverdure commented Sep 27, 2026 •

Copy link
Copy Markdown
Member

Summary

This PR merges PR 54 (Phoenix device and CAN bus alerts) and PR 55 (CAN chain break hint and PD connection) into one change. It follows section 4 of the 2026-09-27 simplification review. The two PRs built overlapping per-device and per-bus logic, and they conflicted in ModuleIOTalonFXBase, LoggedCANBus and Robot. Here that logic is built once. When this merges, PR 54 and PR 55 close in its favor.

Every alert is computed from logged inputs, so a replayed log shows the same alerts.

Basis marks: ✔ checked in code, sim or tests · R read, not rechecked · D derived · 🤖 needs a robot.

What it does

Per-device health (per module)

  • turnEncoderDisconnected joins the drive and turn disconnect alerts.
  • setupFailed: one alert that lists every failed setup step: drive config and position reset, turn config, and CANcoder config read and write.
  • firmwareBlocked: raised when setControl returns FirmwareTooOld or ApiTooOld, which is how Phoenix reports that it is blocking output (R: ParentDevice.setControlPrivate).
  • A failed CANcoder config read protects the calibration. After a failed read, cancoderConfig holds default values. Three guards follow from that:
    • The constructor skips the config write.
    • Module never seeds or overwrites the turn-zero Preference: not on the first loop, not from the zero-encoders action, and not before the first loop has read the inputs.
    • ModuleIOTalonFXBase.setTurnZero also refuses to apply, as a backstop.
  • Firmware versions are read without blocking and without error reports, until each first arrives, and then logged as inputs. Phoenix sends the version signal at 4 Hz (✔ Phoenix 26.70.0-alpha-2 Javadoc), so it usually isn't available yet at construction.
  • Phoenix .hoot logging is off unless FeatureFlags.HOOT_LOGGING_ENABLED is set.

CAN topology and PD

  • Chain declaration. Each bus in Constants declares a CANChain, then one CANChain.Device per line in daisy-chain order. Java runs static initializers in the order written (JLS §12.4.2), so line order is chain order. Each bus class holds the single CANBus instance for its port.
  • Break detection. CANChainMonitor raises a HIGH alert naming the broken link once a clean split holds for 2.5 s. A clean split means devices 0..k-1 respond and at least two devices from k on don't.
    • It reads one map from device to connection state, which Robot builds from Drive and LoggedPowerDistribution.
    • Both buses start untraced, so this half is off and silent until the wiring is traced.
  • PD connection. LoggedPowerDistribution logs a debounced voltage > 0 as a replayable connected input and raises PD/disconnected. While the PDH is missing, it skips the other reads and reads the voltage only once a second, timed with Timer.advanceIfElapsed as RobotStats does. Each failed read sends a Driver Station error (✔ PowerDistributionJNI.cpp:153), so that limits the errors to about one per second instead of one per loop.

One bus-health step per loop (LoggedCANBus.log())

  1. Log the status. Copy the latest status sample into logged inputs, including the new Log Phoenix CAN bus health fields that AdvantageKit SystemStats doesn't cover #50 fields.
    • A WPILib Notifier reads CANBus.getStatus() every 400 ms in REAL mode, like the other background readers. CTRE documents that call as blocking for up to 1 ms (R: CTRE Javadoc). This is a documented worst case, not a measurement.
    • The robot checklist below times the call. SampleCount lands because the stale-reader check needs it. Revisit both after that timing (Refs Log Phoenix CAN bus health fields that AdvantageKit SystemStats doesn't cover #50).
  2. Raise bus alerts. CANBusHealth raises HIGH for ErrorPassive, BusOff, Stopped, a failed status read, a rising bus-off or restart count, or a stalled reader. It raises MEDIUM for ErrorWarning. Each alert is held for 0.5 s.
  3. Update the chain hint, gated by bus health. While the bus alert is HIGH, a break further along the chain is not shown, because a bus fault can drop devices anywhere. A break at the SystemCore end still shows: "check the SystemCore port and plug" is also the right advice for a bus fault.
    • D, 🤖: a SystemCore-end unplug should leave the controller in ErrorPassive, because missing ACKs raise the transmit error count to that level and no further. This comes from the CAN standard and still needs a robot check.

Simplifications from the review

  • CANChain has no freeze and no late-add throw. Class initialization already runs every add before any other class can read the chain.
  • Connection sources: there is no per-port filtering and no duplicate-source check. Nothing can report one device twice today (✔).
  • Alerts: six setup and firmware alerts per module become two, and the "set text if changed" logic lives in one helper, Util.setAlert.
  • Removed: firstError, Drive/ConstructMs (apply it as a temporary patch for the robot timing), and the per-loop firmware refreshes.

Verified

  • Unit tests. ./gradlew build runs 254 tests (✔). The only failure is VisionFilterTest > yawConsistency, which fails the same way on main-2027-alpha7 (review bug 4). The new logic is tested without the HAL:
    • ModuleHealthTest: turn-zero guard, setup text, blocked text.
    • CANBusHealthTest: a parameterized severity table plus the time-based cases.
    • CANChainTest: a table of findBreak cases, the hint text, and validate.
    • CANChainMonitorTest: hold time, the skew case, and two bus-fault cases (k=3 hidden, k=0 shown).
    • DriveConstantsTest: every module ID equals main's.
  • Sim, committed code (✔, about 4,450 loops):
    • No alerts. /PD/Connected stays true. No .hoot session directory. No chain entries.
  • Sim with temporary injected failures on FrontLeft (not committed; ModuleIOSimTalonFX): drive ConfigFailed and CANcoder refresh TxFailed.
    • ✔ One alert: "Setup failed on module FrontLeft: drive motor setup (ConfigFailed); turn encoder config read (TxFailed); turn zero not saved."
    • ✔ With FrontLeft's Preference deleted first, it stays unset. A zero request before the first loop, and another 200 loops later, both skip FrontLeft and zero the other three modules.
    • The first run found that a zero request before the first loop still wrote the Preference from default inputs. The last commit fixes that.
    • ✔ Firmware reads 26.70.0.0 on the Phoenix sim devices.
  • Sim with the CAN reader started and the PD forced to 0 V (✔, temporary patches):
    • SampleCount rose 2.508 per second on both buses, with no alerts.
    • The PD was read every loop through the 0.5 s debounce, then every 1.000 s (± 0.02 s).
  • Sim with both chains temporarily traced (✔): there is no validation warning, and ChainBreakIndex is -1 on both buses. On SC0 only the gyro is down, and a single device at the end is not a break.

Robot session

Termination and tracing (disabled, on blocks)

  • With power off, measure CAN-H to CAN-L at each bus's SystemCore end. Expect about 60 Ω, which confirms a terminator at both ends.
  • Trace SC0 and SC1, reorder each bus's device lines to match the wiring, and set CHAIN_ORDER_TRACED.

Chain and bus behavior

  • Unplug one mid-chain SC1 connector. The devices before it stay connected, and the hint names the right link. This tests the premise that the SystemCore side of a break keeps working.
    • Also record whether the bus alert goes HIGH. A HIGH alert would hide the hint.
  • Unplug SC1 at the SystemCore.
    • Record the bus State. Expect ErrorPassive (D, from the CAN standard).
    • The HIGH CAN alert appears within about 1 s.
    • The "no device responds" hint appears alongside it once the chain is traced.
  • SC0, which has no Phoenix devices, reports OK / ErrorActive. Record State, TEC and REC with the PD and the gyro unplugged.
  • Reseat everything. The alerts clear. Boot twice from power-off: no chain alert.
  • Count CAN alert activations over a full healthy session, including boot.
  • Check whether restartRose just duplicates busOffRose.
  • Note /SystemStats/Network/CAN<n>/FD for both buses. CAN FD is less tolerant of a missing terminator.

Devices

  • Unplug one CANcoder at boot. The setupFailed and turnEncoderDisconnected alerts appear, and no zero is saved.
  • Unplug the PDH's CAN. /PD/Connected goes false within about 0.6 s, PD/disconnected appears, and the DS console shows about one PD error per second rather than one per loop.
  • The firmware strings match Tuner X, and there is no firmwareBlocked alert.

Timing and logging

  • Time CANBus.getStatus() on each bus. Also compare /RealOutputs/LoggedRobot/UserCodeMS p50 and p99 with a log from before this change.
  • CANBus/SC*/SampleCount rises about 2.5 per second.
  • For Log Phoenix CAN bus health fields that AdvantageKit SystemStats doesn't cover #50:
    • Compare CANBus/SC<n>/BusUtilization (0–1) with SystemStats/Network/CAN<n>/Utilization (percent).
    • Compare BusErrorCount with RX/Errors and TX/Errors.
    • This checks the bus numbering and where the fields overlap.
  • Time the loop in which each module applies its turn zero: the first periodic(), and a zero-encoders press.
    • Both apply the CANcoder config on the main loop, and Phoenix's default config timeout is 0.100 s (✔ ParentConfigurator.java:31, used by CANcoderConfigurator.apply(configs)). A slow device could hold the first loop for up to about 0.4 s across four modules (D).
    • The apply only runs after a successful config read, so the full timeout needs a CANcoder that drops out after boot (D).
  • Apply the temporary Drive/ConstructMs timing patch. Record it with a healthy bus and with SC1 unplugged at boot.
  • A new /U/logs/session_N contains only .wpilog files.

Closes #52
Refs #50, #47

🤖 Generated with Claude Code

nlaverdure and others added 4 commits September 27, 2026 18:17
Each bus in Constants declares a CANChain and one CANChain.Device per
line, in daisy-chain order from the SystemCore. Java runs static
initializers in the order written (JLS 12.4.2), so line order is chain
order. Each bus class also holds the one Phoenix CANBus for its port.

CANChainMonitor raises a HIGH alert naming the broken link once a clean
split (at least two devices down) holds for 2.5 s. It reads a single
map from device to connection state, which Robot builds from
Drive.canConnections() and LoggedPowerDistribution.canConnections().
Both buses start untraced (CHAIN_ORDER_TRACED = null), so the monitor
stays off and silent until the wiring is traced.

LoggedPowerDistribution logs a debounced "voltage > 0" as a replayable
connected input, raises PD/disconnected, and skips the other reads while
the PDH is missing.

This carries PR 55 forward, simplified per the 2026-09-27 review: no
chain freezing, no per-port filtering or duplicate-source check, and
module hardware keeps using SC1.BUS with the IDs from its constants.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Phoenix 6 sends its automatic alerts straight to MrcLib, so AdvantageKit
never logs them (#47). Each module now raises WPILib alerts from logged
inputs, so replay shows the same alerts:

- turnEncoderDisconnected, next to the motor disconnect alerts.
- setupFailed, one per module, listing every failed setup step (drive
  config and position reset, turn config, CANcoder read and write).
- firmwareBlocked, when setControl returns FirmwareTooOld or ApiTooOld,
  which is how Phoenix reports that it is blocking output.

A failed CANcoder config read leaves cancoderConfig holding defaults.
The constructor then skips the write, and Module neither seeds nor
overwrites the turn-zero Preference, on the first loop or from the
zero-encoders button. ModuleIOTalonFXBase.setTurnZero also refuses to
apply, as a backstop.

Firmware versions are read without blocking or error reports until each
first arrives (Phoenix sends them at 4 Hz), then logged as inputs.
Phoenix .hoot logging is off unless FeatureFlags.HOOT_LOGGING_ENABLED.

The setup and firmware text builders are pure static functions, tested
in ModuleHealthTest without the HAL.

Refs #47, #52

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
LoggedCANBus now runs one bus-health step per loop:

1. Copy the latest status sample into logged inputs. A background thread
   reads CANBus.getStatus() every 400 ms in REAL mode, because CTRE
   documents that call as blocking for up to 1 ms. The new #50 fields
   (BusErrorCount, ArbitrationLostCount, RestartCount, State, Status)
   are logged with it, and SampleCount lets CANBusHealth detect a
   stalled reader.
2. CANBusHealth raises HIGH for ErrorPassive, BusOff, Stopped, a failed
   status read, a rising bus-off or restart count, or a stalled reader,
   and MEDIUM for ErrorWarning, each held 0.5 s.
3. The chain monitor runs with the bus fault flag. While the bus alert
   is HIGH, a break further along the chain is not shown, since a bus
   fault can drop devices anywhere. A break at the SystemCore end still
   shows: "check the SystemCore port and plug" fits a bus fault too.

Everything comes from logged inputs, so replay reproduces the alerts.

Refs #50

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Before the first Module.periodic(), the inputs hold their defaults, so
turnEncoderRefreshStatus reads "OK" even for an encoder whose config
read failed. A zero request in that window got past the refresh check
and wrote the turn-zero Preference from default values. The device
write was still refused by ModuleIOTalonFXBase.setTurnZero.

Found in sim with an injected CANcoder read failure on FrontLeft and a
zero call before the first loop. With this change, FrontLeft's
Preference stays unset and its setTurnZero is never called, both before
and after the first loop.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
nlaverdure and others added 4 commits September 27, 2026 18:48
A missing PDH still sent one Driver Station error per loop:
PowerDistributionJNI.getVoltage reports its own failed read
(PowerDistributionJNI.cpp:153, allwpilib v2027.0.0-alpha-7), and the
quiet getVoltageNoError needs the handle that PowerDistribution keeps
private. While the module is missing, LoggedPowerDistribution now reads
it once per second, so it sends about one error per second, and a
returning module is noticed at the next read. shouldRead is a pure
static function, tested without the HAL.

DriveConstants.MODULE_DEVICES is now an array in Drive's module order
(FL, FR, BL, BR). It replaces the IdentityHashMap lookup from each
module's SwerveModuleConstants and keeps the per-device buses.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Both follow the patterns already in the codebase:

- LoggedCANBus reads CANBus.getStatus() from a WPILib Notifier every
  0.4 s, like CanandgyroThread, SparkOdometryThread and VisionThread,
  instead of a hand-rolled Thread and sleep loop.
- LoggedPowerDistribution times its retry of a missing module with
  Timer.advanceIfElapsed, like RobotStats, instead of comparing
  timestamps by hand. The retry logic now lives in WPILib's Timer, so
  its unit test is gone rather than bootstrapping the HAL clock.

Verified in sim: with the reader started, SampleCount rose 2.508 per
second on both buses with no alerts. With the PD forced to 0 V, it was
read every loop through the 0.5 s debounce, then every 1.000 s
(+/- 0.02 s).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A public static final array fixes only the reference; any class could
still overwrite its entries. List.of makes the list itself immutable,
and a test checks that set() throws.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
CANChainMonitor.update fills one boolean[] field each loop instead of
allocating a new array. CANChain.hint uses name directly instead of a
local alias.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant