Conversation
mn_fini stops the sysmon watchdog by setting a flag and joining it, but the watchdog waited between scans in a bare nanosleep, so the join waited for the rest of the current tick. By the end of a run the hubs are idle and the adaptive idle backoff has stretched the tick to its cap (wedge_ns/2, 25 ms by default), so every runloom.run() paid ~20 ms of teardown latency: an empty run() took ~21.7 ms with the watchdog on and ~2 ms without it. Wait out the tick on a condvar instead and have stop_join signal it. The stop flag is re-checked and set under the same lock, so a stop that lands between the loop test and the wait cannot lose its wakeup. The lock and cond are initialised per spawn (no watchdog is alive then, and a fork child may have inherited the lock held) and destroyed after the join. Tick length, idle backoff, wedge detection and preemption are unchanged; the crash-handler freeze still just sets the flag (async-signal-safe) and the watchdog exits at its next tick as before. The ~20 ms showed up as throughput in short benchmark samples: the workflow benchmark times the whole run(), so a ~70 ms mixed sample lost 10-15% to it. Test: test_mn_teardown.py::test_fini_does_not_wait_out_the_sysmon_tick measures mn_fini with a long backed-off tick, sweeping the idle time across one tick (the ticks are phase-locked to mn_init). Unfixed: 43.9 ms, fails. Fixed: ~0.7 ms.
johng
marked this pull request as ready for review
September 23, 2026 20:51
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
mn_finistops the sysmon watchdog by settingrunloom_sysmon_stopand joining the thread. Between scans the watchdog sleeps in a barenanosleep, so the join waits out whatever is left of the current tick.By the end of a run the hubs are idle, and the adaptive idle backoff (9ad5f98) has stretched the tick to its cap (
wedge_ns/2, 25 ms by default). So everyrunloom.run()pays ~20 ms of teardown latency:runloom.run()(18 hubs, median of 15)The watchdog is on by default on free-threaded 3.13+ (preemption forces it on), and in migration mode since the deadlock-census change. The cost lands on any code that calls
run()repeatedly: the test suite, sync wrappers aroundrun(), and short benchmark samples.Fix
The watchdog now waits out its tick on a condvar (
runloom_cond_timedwait_ns) instead of sleeping, andstop_joinsignals it. This is the same patternmn_finialready uses for the hubs' idle waits (BUG #10).stop_joinunder the same lock. A stop that lands between the loop test and the wait can't lose its wakeup.runloom_sched_freeze_for_crash) still only sets the flag, so it stays async-signal-safe. The watchdog exits at its next tick as before.Test
tests/test_mn_teardown.py::test_fini_does_not_wait_out_the_sysmon_ticktimesmn_finiwithRUNLOOM_SYSMON_MS=2000, which gives an 80 ms backed-off tick.The watchdog's ticks are phase-locked to
mn_init, so a fixed idle time always lands at the same point in the tick. The first version of the test used a fixed 0.3 s idle and passed on unfixed code, measuring ~7 ms. The test therefore sweeps the idle time across one full tick and averages.Also green:
test_cov100_sysmon,test_sysmon_oracle,test_preempt_timeslicer,test_cov100_resume_preempt,test_cov100_init_fini,test_mn_teardown,test_teardown_park_matrix,test_mn_deadlock_detect,test_cov_go_deadlock_differential,test_fork_safety,test_fork_balance_lock,test_monkey_fork(viatests/run_isolated.py, 3.14.4t, macOS arm64).Benchmark impact
benchmark/workflows(from the workflow-compare branch), 18-core M-series, 3.14.4t PGO+LTO, 18 hubs. Mean of 2 interleaved runs × 5 samples, before → after on a tree with this change applied:This isn't a throughput change in the scheduler: the harness times the whole
run(), and the gain tracks how short each sample is. Mixed samples are ~70 ms, so a fixed 20 ms is ~13% of each one. With the fix, what those rows measure is the workload rather than teardown.It also explains an apparent migration-mode regression. Enabling sysmon under migration made that column start paying the same ~20 ms that stock already paid.