Conversation
johng
force-pushed
the
followup/migration-mode-failures
branch
from
September 16, 2026 07:56
b30a83b to
7c2b8c6
Compare
johng
marked this pull request as draft
September 16, 2026 08:00
CI builds the patched interpreter and asserts migration_available(), then runs the suite with migration OFF, so the migrating scheduler was never under a gate. Add a `migtests` phase to scripts/check_all.sh (default local phase list; skips on an unpatched interpreter) and a `migration-tests` phase in tools/ci/test_patched_cpython.sh, wired as its own step in the runloom-tests action. Hang ceiling 120s per file, load-flake retry off.
johng
force-pushed
the
followup/migration-mode-failures
branch
5 times, most recently
from
September 16, 2026 20:22
65a54d3 to
855841c
Compare
…posed All gated on per-g-tstate mode unless noted; the default scheduler's suite stays green on macOS and Linux. 1. runloom_g_decref: never touch a per-g tstate once Py_IsFinalizing(). Py_FinalizeEx removes and Clears (on 3.13, frees) every non-main tstate via _PyThreadState_DeleteList; a g deallocated from a hub tstate's biased-refcount merge inside that walk Cleared its per-g tstate a second time: SIGSEGV in Py_FinalizeEx (10/10 on the heavy-offload workload, now 0/100). 2. runloom_sysmon_config: keep the sysmon instrumentation on in per-g mode. Without its running-fiber signal the deadlock census reports "wakeable" unconditionally, so RUNLOOM_DEADLOCK could never fire under migration. 3. runloom_mn_fiber_n: take the loop path in per-g mode; the bulk-arena builder allocates no per-g tstates and hub_main skips such a g as dead. 4. hub_main: the anti-starvation turn alternates between the global run-queue and a fresh deque g. Woken migratable fibers live on that queue, which was only pulled once ready ring AND deque were empty, so a sched_yield loop starved them as soon as spinners >= hubs; alternating keeps a never-run g from starving behind a ping-pong pair in turn. 5. hub_main (all backends, not just kqueue): the periodic non-blocking self-pump now runs on every backend while anything is parked or in flight, and reaps this hub's io_uring ring. A sched_yield loop keeps a hub from ever idling, and the blocking pump only ran at idle, so a parker on that hub hung -- at H=1 on any backend, and at H=2 under migration once the global run-queue placed the loop on the parker's hub (test_iouring_cancel_close). 6. io_uring hub-ring park: after a signal wake the fiber submitted its ASYNC_CANCEL inline, assuming it resumed on the ring's owning hub. Under migration it may resume anywhere, which breaks SINGLE_ISSUER and the cancel never lands. Route through the owner's mailbox when not at home, retry if the single-slot mailbox is busy, and retract the deposit if the op completes first (op lives on the fiber's stack). The mailbox request now kicks the owner out of any wait mode (test_signal_recipient). 7. io_uring loop backend (RUNLOOM_IOURING_LOOP=1): off under per-g mode, with a one-time note. Its design pins fibers to their hub's ring (multishot buffers are returned by "its pinned echo fibers"); a migrated fiber lands on the wrong ring and exact-once echo failed (test_cov95_iouring_loopbuf). The CI migration-tests phase now refuses to run on an interpreter where migration_available() is False instead of skipping green. migtests lane after: macOS 229 passed / 6 not-passed, Linux 228 / 8 of which 3 are root-only resource-limit tests on the dev box. What is left needs preemption, which per-g mode stands down, plus the hub-tstate datastack sweep accounting. 8. tests/conftest.py: the ten tests that fail under migration by construction -- they encode the default mode's preemption, hub-tstate introspection and per-hub memory-footprint semantics, which per-g mode changes -- are skipped from one list, only under RUNLOOM_MIGRATION=1, each with its reason, so the lane is green and any new failure in it is a regression. Remove an entry when its gap closes; the set is the "preemption in per-g mode" design item. One OPEN intermittent is listed apart from them: test_signal_recipient's selector-outranks-sleeper case loses the signal ~1 run in 15 under migration (0/20 in default mode); not root-caused.
johng
force-pushed
the
followup/migration-mode-failures
branch
from
September 16, 2026 20:45
855841c to
a0ac4da
Compare
test_mn_fiber_core_coro_alloc_failure_releases_admission caps RLIMIT_AS at VmSize + 8 MiB and expects a 32 MiB stack reservation to fail. That is a sizing race, not a scheduler oracle: it failed 4/20 in DEFAULT mode on an 8-core Linux box and twice on the shared GitHub runners (once per lane). Skip it under GITHUB_ACTIONS only; locally it still runs, where a genuine regression in the coro==NULL cleanup path shows up deterministically.
johng
marked this pull request as ready for review
September 17, 2026 20:23
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The patches in
src/patchesexist so fibers can migrate between hubs, but CI only ever ran the suite with migration off. Two commits:ci: run the suite with RUNLOOM_MIGRATION=1 on the patched interpreteradds amigtestsphase toscripts/check_all.shand amigration-testsstep to CI. With only this commit the new lane is red on every job (15–18 files).sched: fix the migration-mode defects the RUNLOOM_MIGRATION=1 lane exposedfixes what it found. Nothing here is the patches or the PGO/LTO build; every failure reproduced on a non-optimized patched build.Py_FinalizeEx: a per-g tstate cleared a second time from a hub tstate's biased-refcount mergePy_IsFinalizing()fiber_nbulk-arena batch never runs (no per-g tstates)sched_yieldloopsched_yieldloop starves the netpoll pump on its hub (H=1 on any backend; H=2 under migration)Ten tests encode default-mode preemption, hub-tstate introspection and per-hub memory-footprint semantics, which per-g mode changes; they are skipped only under migration from one list in
tests/conftest.py, each with its reason, so the lane gates regressions. Closing that set is the "preemption in per-g mode" design item. One open intermittent is listed apart from them:test_signal_recipient's selector-outranks-sleeper case loses the signal about 1 run in 15 under migration (0 of 20 in default mode), not root-caused.Default-mode suite: green on all four CI jobs. Migration lane: green apart from the pre-existing
test_monkey_leakfailure on macOS.