Repository navigation
[Bug]: Container subnet routing breaks when a competing default route appears (e.g. VPN/exit-node) and doesn't self-heal #1881
Description
Activity
Adding a second, related variant with a deterministic repro: outbound NAT from vmnet guests breaks under the same trigger (Tailscale exit node), but with different healing behavior than the host→container routing failure described above — which may help separate the two mechanisms.
Environment
- OS: macOS 27.0 beta (26A5368g) - container: 1.0.0 (release, commit: ee848e3) - Trigger: Tailscale 1.98.8 exit node (another Mac on the same tailnet) - Guests affected: a `container machine` VM (Ubuntu 24.04) and the buildkit builder VM, both on the default 192.168.64.0/24 network - Architecture: arm64Failure signature (guest → internet, i.e. the NAT/InternetSharing path):
- DNS still resolves inside the guest (the 192.168.64.1 gateway answers it locally), but every forwarded TCP/ICMP connection black-holes (
SYN-SENTforever inss -tnp). - The host itself stays fully online through the exit-node tunnel the entire time.
Behavior, reproduced in 2 full cycles (stable wired-equivalent WiFi; host connectivity to the same endpoint verified HTTP 200 immediately before and after every guest probe, to rule out host-side blips):
tailscale set --exit-node=<node>→ guest outbound black-holes immediately, even against freshly rebuilt NAT state.tailscale set --exit-node=→ guest does not recover. Once an exit node has been active, the NAT state stays broken.container machine stop <vm> && container system stop && container system startwith the exit node off → guest fully recovers (verifiedcurl https://pypi.org/→ 200 from inside the guest).
Step 3 is the notable difference from the OP's variant: the outbound-NAT breakage does self-heal on a
container systemrestart, provided the exit node is off when the network stack is rebuilt. (Rebuilding while the exit node is active leaves it broken from the start.)Ruled out as triggers (none of these broke the NAT, with the exit node never enabled): WiFi interface flaps of 20s and 45s; migrating the host between networks (phone hotspot → home WiFi) with sleep/wake in between; Tailscale merely being connected (
RouteAll=true, MagicDNS, no exit node) coexists fine.So from the guest-outbound side, the interference appears to be with the InternetSharing/NAT state rather than (or in addition to) the missing explicit subnet route — deactivating the exit node restores the host's default route to the physical interface, yet guest traffic stays black-holed until the NAT is rebuilt.
Happy to run further diagnostics on this setup if useful.
Testing and write-up assisted by Claude Code.
- DNS still resolves inside the guest (the 192.168.64.1 gateway answers it locally), but every forwarded TCP/ICMP connection black-holes (
Hit a variant of this while migrating a workload from OrbStack to
container. Sharing since the trigger and the recovery behavior differ a bit from what's documented above, and it looks like the corruption goes deeper than the host routing table.Environment
- OS: macOS 27.0 (26A428) — "Golden Gate", a few days post-release - container: 1.4.1 (release) - Architecture: arm64 (M4)Trigger — a stale/wedged interface, not an active exit-node
Tailscale had not been reconnected since upgrading to macOS 27.
tailscale statusreported "Tailscale is stopped", bututun4was stillUP/RUNNING, still bound to a Tailscale CGNAT address (100.x.x.x), and was still winning the system default route:$ route -n get 192.168.64.4 interface: utun4 # should be bridge100So this bug's trigger isn't limited to someone actively using an exit-node — a Tailscale daemon that thinks it's stopped but never tore down its network extension state (e.g. after an OS upgrade) reproduces the identical symptom, silently, with no VPN "on" from the user's perspective.
Only a full reboot fixed the routing — matching the pattern in the original report (step 7), disconnecting Tailscale in the app and a full
brew uninstall/reinstall of the Tailscale cask both leftutun4and the bad default route completely unchanged.container system stop && container system startwas not tried in isolation (went straight to reboot), so I can't confirm whether that alone would've been insufficient here too, but neither app-level disconnect nor package uninstall touched it. Only a full OS reboot clearedutun4and restored the correct default route viabridge100/en7.New finding: container-level state stays broken even after host routing is confirmably fixed
After the reboot,
route -n get <container-ip>resolved correctly viabridge100— by the reasoning in this issue, connectivity should have been restored. It wasn't:container start <existing-container>(created before the reboot): host→container published port still timed out.- Traced further —
/proc/net/tcpinside the container confirmed the app was correctlyLISTENing on0.0.0.0:<port>. - But a request to
127.0.0.1:<port>from a process inside that same container also hung indefinitely (container exec <name> node -e "fetch('http://127.0.0.1:<port>/', {signal: AbortSignal.timeout(5000)})..."never returned, even with an abort signal — had topkillthe exec process host-side). - Deleted the container entirely (
container rm) and recreated it fresh from the same image (container run -d ...) — identical symptom on the brand-new instance/IP. Ruled out a stale per-container VM resuming badly across the reboot.
So whatever gets corrupted isn't only the host's routing table entry for the container subnet (which by this point was verified correct) — something in the network stack
container system startrebuilds also comes up broken, to the point that a container can't even reach its own loopback, and recreating the container doesn't reset it. Only restarting the whole VM stack (not tried in isolation) or another full reboot might clear it — didn't get to testcontainer system stop/startalone as the last step before restoring service on OrbStack instead, since this was blocking a real migration.Happy to gather more diagnostics (vmnet logs,
container system logsoutput, etc.) if useful — I still have the affected container/image/volume around for a bit.Investigation and write-up assisted by Claude Code.
A cleanly-isolated third variant, distinct from both my earlier comment (stale Tailscale interface) and the "outbound NAT" comment above (Tailscale exit-node) — this one's triggered by NordVPN, and specifically does not show the routing-table symptom described in the original report.
Environment
- OS: macOS 27.0 (26A428) - container: 1.4.1 (release) - Trigger: NordVPN (NordLynx/WireGuard), no exit-node/subnet-routing config involved — just a normal connect - Architecture: arm64 (M4)Symptom — both directions break, but the routing table looks correct:
With NordVPN connected,
route -n get <container-ip>resolves viabridge100exactly as expected (not hijacked to the VPN'sutuninterface, unlike the original report's symptom). Loopback within a container works fine (fetch('http://127.0.0.1:<port>/')succeeds). But:- Host → published port:
curltimes out completely. - Container → real internet (egress through the vmnet NAT): also times out (
fetch('https://api.github.com')aborts on timeout).
So this isn't the "competing default route shadows the container subnet" mechanism from the original report — the subnet-level routing is fine. It looks like the same class of issue as the "outbound NAT from vmnet guests breaks" variant described above, just with a different, more common trigger (a plain consumer VPN connect, not an exit-node config) and additionally affecting the inbound port-forward path too, not just outbound.
Isolation testing (disconnect/reconnect one VPN client at a time, retest both directions each time):
State Host→container Container→internet NordVPN + Azure VPN both connected ❌ timeout ❌ timeout NordVPN alone ❌ timeout ❌ timeout Azure VPN alone (standard corporate VPN client) ✅ works ✅ works Neither connected ✅ works ✅ works Azure VPN's own
utuninterface doesn't even take over the system default route when it connects; NordVPN's does. Disconnecting NordVPN alone (Azure VPN still connected) restores full connectivity in both directions immediately, with nocontainer systemrestart or reboot needed — recovery here is much lighter-weight than the routing-table variant in the original report, which needed a full reboot to clear.Given how common NordVPN is, this seems likely to affect a fair number of
containerusers who also use a VPN day-to-day. Happy to gather more diagnostics (vmnet logs, pfctl state before/after NordVPN connects, etc.) if useful — this Mac still has NordVPN installed and reproducible on demand.Investigation and write-up assisted by Claude Code.
- Host → published port:
Summary
container's vmnet-backed network only ever installs an interface-scoped default route (RTF_IFSCOPE) for the container subnet (e.g.bridge100), never an explicit non-scoped route for the subnet itself (e.g.192.168.64.0/24). Because of that, routing to containers is entirely dependent on which global default route currently wins on the host. Any other tool that installs or changes the system's default route (a VPN client, a Tailscale exit node, etc.) can silently break all container connectivity — and once broken, it does not self-heal when the competing route is removed, or even whencontainer systemis fully restarted. Only manually adding an explicit subnet route restores it.This is a different bug from #856 — that one is about
container-apiserverfailing to bind a listener to the gateway IP (EADDRNOTAVAIL). This is about outbound routing to already-running containers breaking due to route table interference, discovered while investigating #856/#1813 (it isn't the same root cause — I confirmed binding to the gateway IP for listening still succeeds even while this routing bug is active).Update: I initially assumed this was a routing priority problem (a competing default route "winning" over the more-specific
192.168.64.0/24route). Testing further, that's not quite right — I manually installed an explicit192.168.64.0/24 -interface bridge100route and confirmed it correctly restores connectivity on its own, but when a Tailscale exit node is subsequently enabled/disabled, that route is removed from the table outright, not merely outranked. This matches tailscale/tailscale#18653 ("Exit node overrides existing connected RFC1918 routes"), which explicitly calls outapple/container'sbridge100interface and is already being investigated on their side — it looks like a Tailscale-side behavior at the System Extension layer, not somethingcontainercan prevent outright.Given that, this issue is really two things:
containernever installs an explicit subnet route in the first place, only relying on the interface-scoped default (fixable here, see below).containercan't unilaterally prevent — that's tracked upstream at Exit node overrides existing connected RFC1918 routes tailscale/tailscale#18653 for the Tailscale-specific case.Environment
Steps to reproduce
container systemand run a container on the default network:bridge100. After, it resolves via the other default route (in my caseen10via the LAN gateway), because there's never been an explicit192.168.64.0/24route — onlybridge100's own interface-scoped default, which isn't consulted unless something explicitly scopes to that interface.route -n get 192.168.64.5still resolves via the wrong interface.-interface 192.168.64.1for the default network'sbridge100gateway)Current behavior
Container subnet connectivity can be permanently broken by any third-party tool that manipulates the default route, and neither removing that tool's route nor restarting
container systemrepairs it.containeralso never installs an explicit subnet route on its own, so it has no baseline resilience against this class of interference at all.Expected behavior
containershould install an explicit, non-scoped subnet route (e.g.192.168.64.0/24 -interface bridge100) when the network is created, rather than relying solely on the interface-scoped default. This won't fully eliminate exposure to tools like Tailscale that actively remove routes they don't own (that's tracked at tailscale/tailscale#18653), but it does close the gap for the simpler and more common case — anything that merely adds a competing default route without deleting others — and givescontainera correct baseline instead of none.Workaround in the meantime
Tools built on top of
container/container-composethat need to survive this (e.g. avoiding an exit node's effects on local dev) currently work around it by explicitly binding sockets to the bridge gateway IP before connecting out, forcing the kernel to use the interface-scoped route directly rather than relying on the general routing table.