Skip to content

Hint where a CAN daisy chain is broken, and report PD connection - #55

Closed
nlaverdure wants to merge 9 commits into
main-2027-alpha7from
can-chain-hint
Closed

nlaverdure wants to merge 9 commits into
main-2027-alpha7from
can-chain-hint

Conversation

@nlaverdure

@nlaverdure nlaverdure commented Sep 27, 2026 •

Copy link
Copy Markdown
Member

Summary

When a CAN cable or connector fails, every device past the break drops at once, and the Driver Station fills with disconnect alerts that don't say where to look. This PR gives each device a CAN index, its position along the daisy chain from the SystemCore. When exactly one break explains the pattern, a single HIGH alert names the spot, for example:

CAN chain break on SC1 (order traced 2026-10-03 A.Student): #0–#4 respond, #5–#11 don't. Check #4 BackLeft turn (ID 11)'s outgoing connector, the cable, and #5 BackRight turn (ID 19)'s incoming connector.

It also gives the REV PDH a connection signal and a PD/disconnected alert.

How it works

  • Topology in Constants. Each bus declares a CANChain, which names the bus and its port, then one line per device:
    public static final CANChain CHAIN = new CANChain("SC1", CANPort.CAN_S1);
    public static final CANChain.Device BACK_LEFT_DRIVE = CHAIN.add(10, "BackLeft drive");
    ...
    • Java runs static initializers in the order written (JLS §12.4.2), so line order = wiring order, and a device's position is its CAN index. Each device is declared once, so none can be left out of the chain.
    • The chain freezes when the monitor first reads it, and an add after that throws.
    • Each constant is a CANChain.Device carrying its bus, ID and label, so the bus is stated once: by which bus class the device is declared in. Phoenix's SwerveModuleConstants hold only IDs, so DriveConstants builds each module from a ModuleDevices record and returns it from DriveConstants.devices(module). ModuleIOTalonFXBase, GyroIOBoron and LoggedPowerDistribution take their bus from their devices.
    • CHAIN_ORDER_TRACED (date and name) turns on a bus's hint. Both buses start untraced, because the current order is the old constant order, not the wiring. The CANBusPorts Javadoc explains how to trace a bus.
  • Connection sources. Each subsystem reports every CAN device it owns, on any bus, in one canConnections() map keyed by CANChain.Device (which carries the bus, since a CAN ID is only unique within a bus): Drive for the module devices and the gyro, LoggedPowerDistribution for the PD. Each bus's LoggedCANBus (built from its chain, new LoggedCANBus(SC1.CHAIN, SC1.CHAIN_ORDER_TRACED)) owns its monitor. Its monitorChain(...) takes every subsystem's map and keeps the entries on its own port. A device reported twice turns that bus's check off with one DS warning.
  • Break detection (CANChain, static and pure): a clean split with devices 0..k-1 up and k..n-1 down, with at least 2 down. With one device down, a failed device and its last cable look the same, and that device's own alert already names it.
  • Hold time (in CANChainMonitor): the break must hold for 2.5 s. Redux's isConnected() waits 2 s while the Phoenix flags drop after 0.5 s, so a shorter hold could briefly name a link one step too far along the chain.
  • Monitor (CANChainMonitor, owned by each bus's LoggedCANBus, which updates it at the end of log()):
    • Logs the raw index every cycle (CANBus/<bus>/ChainBreakIndex), so flickers show up in post-match review.
    • Turns off with one DS warning if the chain and its connection sources don't match one to one.
    • Reads only logged inputs, so replay reproduces the alert.
  • PD connection: WPILib has no connected check. On SystemCore, though, a PDH read with no status frame in the last 40 ms times out, and the HAL then returns exactly 0.0 V (REVPDH.cpp:465-467, allwpilib v2027.0.0-alpha-7). A responding PDH can't read 0 V, because it powers the SystemCore running the code. So voltage > 0 (debounced 0.5 s) is logged as the PD's connected input, and it:
    • raises PD/disconnected;
    • skips the other PD reads while the PDH is missing, since each one sends a DS error.

Verified

  • Unit tests: 30 new tests (CANChainTest, CANChainMonitorTest, DriveConstantsTest) cover:

    • every split case;
    • the hint text;
    • validation (missing, unknown and duplicate IDs);
    • chain order and freezing;
    • sorting connection sources by bus, and rejecting a device reported twice;
    • each module's IDs coming from its devices, equal to main's values, all on CAN_S1;
    • the hold time;
    • the slow-device skew scenario.

    ./gradlew build runs 218 tests. The only failure is VisionFilterTest > yawConsistency, which fails the same way on main-2027-alpha7.

  • ID refactor: a script compared every device ID with the int constants on main-2027-alpha7 (all 14 match, none missing), and DriveConstantsTest pins each module's drive, turn and encoder IDs. A sim run with Phoenix sim devices built all 12 module devices from the device lookup.

  • Sim, with an injected break at SC1 index 5 (temporary, not committed):

    • The raw index was 5 from the first loop.
    • The alert named BackLeft turn (ID 11) at index 4 and BackRight turn (ID 19) at index 5 2.501 s later, listed above the per-module disconnect alerts.
    • SC0 (only the sim gyro down) stayed silent.
  • Sim, committed code: no chain entries and no alerts. /PD/Connected is true, and the PD outputs are unchanged.

Robot session (disabled, on blocks)

  • Power off. Measure CAN-H to CAN-L at each bus's SystemCore end and expect about 60 Ω, which confirms termination at both ends.
  • Trace SC0 and SC1, reorder each bus's device lines to match, and set CHAIN_ORDER_TRACED.
  • Unplug one mid-chain SC1 connector: the devices before it stay connected, and the hint names the right link. This tests the premise that the SystemCore side of a break keeps working.
  • Unplug at the SystemCore: the "no device responds" hint appears.
  • Unplug the PDH's CAN: /PD/Connected goes false within about 0.6 s, PD/disconnected appears, and the DS console stops flooding.
  • Reseat everything: the alerts clear. Boot twice from power-off: no chain alert.
  • Note /SystemStats/Network/CAN<n>/FD for both buses. CAN FD is less tolerant of a missing terminator.

Relation to #54

This PR is independent of #54. Both touch Robot.java and LoggedCANBus.java, so whichever merges second gets rebased onto main-2027-alpha7.

🤖 Generated with Claude Code

nlaverdure and others added 3 commits September 27, 2026 10:10
When a CAN cable or connector fails, every device past the break
drops at once, and the Driver Station shows a list of disconnect
alerts with nothing pointing at the fault. Each bus now lists its
devices in daisy-chain order, and when exactly one break explains the
pattern, a single HIGH alert names the two devices on either side.

Constants: each bus's device IDs move into a Chain enum. Declaration
order is wiring order from the SystemCore, and a device's position is
its CAN index. Device code reads IDs from the enum, and alerts name
devices by their constant names. CHAIN_ORDER_TRACED records when and
by whom the order was traced from the wiring; while it is null, the
bus has no hint. Both buses start untraced.

CANChain finds a clean split: devices 0..k-1 connected, k..n-1 not,
with at least two down, since one device down can't be told apart
from its last cable. CANChainTracker reports a break only after it
has held 2.5 s, longer than the gap between the 2 s Redux timeout and
the 0.5 s Phoenix debounce, so a break before a slow device isn't
briefly reported one link too far. CANChainMonitor logs the raw index
every cycle, turns off with one DS warning if the chain and its
connection sources don't match one to one, and reads logged inputs so
replay reproduces the alert.

Power distribution: WPILib has no connected check, but on SystemCore
a PDH read with no status frame in 40 ms times out and the HAL returns
exactly 0.0 V (REVPDH.cpp, allwpilib v2027.0.0-alpha-7). A debounced
voltage > 0 is logged as the PD's connected input, raises a HIGH alert
when false, and skips the other PD reads, each of which would send a
DS error.

Verified: 19 new pure unit tests; every enum ID equals the old int
constant; in sim with an injected break at SC1 index 5 (temporary), the
raw index was 5 from the first loop and, 2.501 s later, the alert
named BACK_LEFT_TURN (ID 11) at index 4 and BACK_RIGHT_TURN (ID 19) at
index 5, while SC0 (only the gyro down) stayed silent. The committed
code shows no chain entries or alerts, and PD/Connected is true in sim.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The Chain enums needed an id field, constructor and accessor in each
bus class, and every call site read IDs through .id(). Each bus now
declares its devices as plain int constants through CANChainBuilder:

  public static final int FRONT_LEFT_DRIVE =
      CHAIN_BUILDER.add(28, "FrontLeft drive");
  ...
  public static final List<CANChainDevice> CHAIN = CHAIN_BUILDER.build();

Java runs static initializers in the order written (JLS 12.4.2), so the
line order is still the chain order, and each device is still declared
once. build() freezes the builder, so a device placed after CHAIN
throws while the class loads rather than leaving the chain silently.
CANChainDevice becomes a record with a friendly label, and call sites
return to the plain constants, so DriveConstants and GyroIOBoron are
unchanged from main.

Verified: all 14 IDs equal main's constants; 4 new builder tests; the
injected sim break at SC1 index 5 gives the same alert 2.501 s later,
now naming "BackLeft turn (ID 11)" and "BackRight turn (ID 19)".

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Five classes did the work of two. CANChain now holds a bus's devices
(a nested Device record, added in declaration order) along with the
pure findBreak, hint and validate methods. CANChainMonitor keeps the
hold-time state itself. CANChainDevice, CANChainBuilder and
CANChainTracker are gone.

Constants declares CHAIN first and then one CHAIN.add(...) line per
device, so there is no longer a rule that the chain comes after the
last device. The chain freezes when the monitor first reads it, and an
add after that throws.

Verified: 212 tests pass apart from the known VisionFilterTest failure,
including 6 monitor tests run directly (Alert and Logger work in JUnit
without HAL setup); the injected sim break at SC1 index 5 shows the
same alert 2.501 s later, listed above the per-module disconnect
alerts; the committed code shows no chain alerts.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
nlaverdure and others added 6 commits September 27, 2026 10:13
AdvantageKit records each bus's FD flag every loop as
/SystemStats/Network/CAN<n>/FD, from the interface info
(LoggedSystemStats.java:135 in 27.0.0-alpha-5), so the one-off
metadata from Phoenix's CANBus.isNetworkFD() added nothing.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Robot built both CANChainMonitors itself and updated them after the bus
logs. Each monitor watches exactly one bus, so LoggedCANBus now takes
the bus's chain and trace record in its constructor, builds the monitor
in monitorChain() once the devices that report connection states
exist, and updates it at the end of log(). Robot keeps only the two
monitorChain() calls.

powerDistribution.log() now runs before the bus logs, so SC0's chain
check reads this loop's PD connection state rather than last loop's.

Verified: build and tests unchanged (212, the known VisionFilterTest
failure only); in sim, the injected break at SC1 index 5 shows the
alert 2.500 s after the raw index appears, SC0 stays silent, and the
committed code shows no chain alerts.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Robot assembled SC0's connection map from Drive.isGyroConnected() and
the PD, while Drive handed over a separate SC1 map. Now each subsystem
reports every CAN device it owns, on any bus, through one
canConnections() map keyed by CANChain.Address (CAN port plus ID, since
a CAN ID is only unique within a bus). Each CANChain records its port,
and LoggedCANBus.monitorChain() takes every subsystem's map and keeps
the entries on its own port. CANChain.connectionsOn() does that sorting
as pure logic; the same address reported twice turns that bus's check
off with one DS warning.

Drive.canConnections() covers the module devices on SC1 and the gyro on
SC0, replacing sc1Connections() and isGyroConnected().
LoggedPowerDistribution.canConnections() reports the PD. Robot passes
both maps to each bus.

Verified: 215 tests (3 new for connectionsOn) with only the known
VisionFilterTest failure; in sim, SC0 validated with entries from both
Drive and the PD, the injected SC1 break showed the same alert 2.501 s
later, and the committed code shows no chain alerts.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A CAN ID is only unique within a bus, so the connection maps keyed
devices by CANChain.Address(port, id), and every call site restated the
bus: Address(SC1.BUS_ID, ...) three times per module in Drive, plus
SC1.BUS in each Phoenix constructor.

CHAIN.add() now returns a CANChain.Device (port, id, label) on the
chain's port, and the constants hold those devices, so each constant is
its own address and the bus is stated once, by which bus class the
device is declared in. CANChain.Address is gone.

Phoenix's SwerveModuleConstants hold only IDs, so DriveConstants builds
each module from a ModuleDevices record and remembers it;
DriveConstants.devices(module) returns them. ModuleIOTalonFXBase builds
its TalonFXs and CANcoder from those devices, and Drive's
canConnections() keys by them, so neither names a bus. The gyro and PD
constructors take their bus from their devices too.

Verified: 218 tests with only the known VisionFilterTest failure,
including new DriveConstantsTest checks that each module's IDs come
from its devices, equal main's values, and sit on CAN_S1; sim with
Phoenix sim devices builds all 12 from the device lookup; the injected
SC1 break shows the same alert 2.501 s later; the committed code shows
no chain alerts.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The key field was only read in one place, so log() now passes
"CANBus/" + name to Logger.processInputs directly. The log paths are
unchanged (/CANBus/SC0/..., /CANBus/SC1/..., checked in sim).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Each bus class held NAME, BUS_ID (CANPort) and BUS (Phoenix CANBus)
next to its CHAIN, which already stored the port. The chain now takes
the bus name and port, new CANChain("SC1", CANPort.CAN_S1), and the
three constants are gone.

LoggedCANBus takes (CHAIN, CHAIN_ORDER_TRACED) and builds its Phoenix
CANBus from chain.port(); a Phoenix CANBus is only a name derived from
the port. CANChainMonitor takes the name from the chain, and hint() is
now a chain method. DriveConstants sets the drivetrain network from
SC1.CHAIN.port().

Verified: 218 tests with only the known VisionFilterTest failure; in
sim the log paths are unchanged (/CANBus/SC0, /CANBus/SC1), the
injected SC1 break shows the same alert 2.501 s later, and the
committed code shows no chain alerts.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@nlaverdure

Copy link
Copy Markdown
Member Author

Superseded by PR 57 (#57), which merges this PR's CAN chain hint and PD connection signal with the device and CAN bus alerts from PR 54, simplified per the 2026-09-27 review.

@nlaverdure nlaverdure closed this Sep 27, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant