Skip to content

Add emscripten_epoll_listener_add/remove for epoll readiness callbacks - #27547

Merged
guybedford merged 29 commits into
emscripten-core:mainfrom
guybedford:epoll-callback
Oct 3, 2026
Merged

guybedford merged 29 commits into
emscripten-core:mainfrom
guybedford:epoll-callback

Conversation

@guybedford

@guybedford guybedford commented Aug 15, 2026 •

Copy link
Copy Markdown
Collaborator

Follow-on to #27207.

Adds experimental emscripten_epoll_listener_add and emscripten_epoll_listener_remove under a new experimental <emscripten/epoll.h>, to allow registering for readiness on an epoll fd. The runtime then invokes the callback from the event loop whenever the set has uncollected ready events. The callback itself does not collect ready events, rather a zero-timeout epoll_wait within the callback itself is able to do this. Allows supporting epoll on a single thread without blocking the main thread or requiring JSPI.

  • Any number of listeners, identified by the (callback, userdata) pair as for the HTML5 event handlers; a duplicate pair is EEXIST, and emscripten_epoll_listener_remove(epfd, callback, userdata) removes by pair.
  • Every listener is signalled while uncollected ready events remain, and collectors race over the single shared ready list. EPOLLET and EPOLLONESHOT items are therefore collected by exactly one listener, matching distribution between multiple blocking epoll_wait callers.
  • Firing is gated on the same readiness derivation used by the epoll fd's own poll handler, so stale ready-list entries never spuriously fire. Delivery occurs on a later event-loop task, never re-entrantly under the frames that made the set ready.
  • A listener is an unref'd handle: it does not keep the runtime alive. A program whose only reason to stay alive is a listener holds the runtime itself with emscripten_runtime_keepalive_push()/pop(). The runtime only takes a temporary hold for scheduled delivery.
  • Listeners are main-thread-only. With pthreads, registration from another thread returns ENOTSUP; pthreads can use blocking epoll_wait() directly.
  • If a callback leaves the set ready, another delivery task is scheduled, yielding to I/O between level-triggered callbacks. A fatal error escaping the callback remains an uncaught task exception rather than an unhandled rejection.
  • Listeners are epoll-instance state: they are shared across dup'd fds and removed on the last close.

Tests cover the syscall surface and errors, edge and level triggering, MOD re-arm, multiple listeners, nesting, dup sharing, event-loop task ordering, callback errors, lifecycle and runtime keepalive behavior, real Node sockets, pthread rejection, and a blocking epoll_wait sharing one ready list with a listener under ASYNCIFY and JSPI.

Made with AI assistance under my review

@sbc100

sbc100 commented Aug 17, 2026

Copy link
Copy Markdown
Collaborator

Can you rebase/merge?

@guybedford

Copy link
Copy Markdown
Collaborator Author

This is now rebased, with the local PR change isolated to commit 9a7b5d3.

@guybedford
guybedford force-pushed the epoll-callback branch 4 times, most recently from d07dbaf to ee21de3 Compare August 18, 2026 22:21
guybedford added a commit to guybedford/tokio that referenced this pull request Sep 14, 2026
Follow-on to tokio-rs#8285 and tokio-rs#8438. A `LocalRuntime` driven by the host
JavaScript event loop instead of by a park, so it needs neither JSPI
nor pthreads and never blocks or suspends the host.

The scheduler drives to a fixed point, then waits. On this target a wait
is either a JSPI suspension (`block_on`) or a return to the host to be
called back; the event-loop runtime is the latter. The same
`current_thread` scheduler and drivers run one `event_interval` batch
per drive followed by the zero-duration driver turn a park would do
(due timers, I/O readiness, deferred wakers), then the wait is lowered
to arming host callbacks: an immediate drive if runnable work remains,
so the host loop gets a turn between batches as under a JSPI `block_on`,
and one `emscripten_set_timeout` for the soonest timer deadline. Timers
are fire-only with an epoch (`emscripten_clear_timeout` leaks its
keepalive) and kept across drives while the deadline is unchanged.
Socket readiness re-enters the drive through a persistent
`emscripten_epoll_add_listener` on the reactor's epoll fd, which also
carries the mio waker, so an external unpark of an I/O-enabled runtime
needs no extra arming; the I/O-less parker arms an immediate drive. A
wake during a drive is absorbed by that drive's own turn and re-arm.

Host callbacks run on an empty stack, so a drive re-enters inline; armed
callbacks hold `Weak` refs, so dropping the `EventLoopRuntime` has native
`Runtime::drop` semantics and a late callback upgrades to nothing.
`block_on` on an event-loop runtime is rejected eagerly like a nested
runtime: its wait is the host loop, so no stack can hold the result.
`EventLoopRuntime::schedule` takes a root and a completion callback
(`Err(JoinError)` on a root panic). The API is `tokio_unstable` and
gated to non-pthread builds; a drive from a host callback while another
runtime's `block_on` is suspended under JSPI still panics as a nested
runtime.

Tests: `rt_emscripten_event_loop` pins the on-stack contracts (schedule
queues only, one batch per drive, host-context wakes never drive inline,
cross-runtime wakes, `block_on` and nested drive rejected, drop cancels).
`rt_emscripten_event_loop_main` (`harness = false`) schedules roots on
two runtimes and returns into the host loop, completing them through
timer re-arms, a greedy yielding sibling, a cross-runtime oneshot and a
TCP round trip between the runtimes, with and without `-sJSPI`; the
pre-js fails the run if the loop drains before the runtime exits.

CI adds an unstable lane. TEMPORARY: it overlays the emscripten
`epoll-callback` branch (emscripten-core/emscripten#27547) for
`emscripten_epoll_add_listener` until released.
guybedford added a commit to guybedford/tokio that referenced this pull request Sep 14, 2026
Follow-on to tokio-rs#8285 and tokio-rs#8438. A `LocalRuntime` driven by the host
JavaScript event loop instead of by a park, so it needs neither JSPI
nor pthreads and never blocks or suspends the host.

The scheduler drives to a fixed point, then waits. On this target a wait
is either a JSPI suspension (`block_on`) or a return to the host to be
called back; the event-loop runtime is the latter. The same
`current_thread` scheduler and drivers run one `event_interval` batch
per drive followed by the zero-duration driver turn a park would do
(due timers, I/O readiness, deferred wakers), then the wait is lowered
to arming host callbacks: an immediate drive if runnable work remains,
so the host loop gets a turn between batches as under a JSPI `block_on`,
and one `emscripten_set_timeout` for the soonest timer deadline. Timers
are fire-only with an epoch (`emscripten_clear_timeout` leaks its
keepalive) and kept across drives while the deadline is unchanged.
Socket readiness re-enters the drive through a persistent
`emscripten_epoll_add_listener` on the reactor's epoll fd, which also
carries the mio waker, so an external unpark of an I/O-enabled runtime
needs no extra arming; the I/O-less parker arms an immediate drive. A
wake during a drive is absorbed by that drive's own turn and re-arm.

Host callbacks run on an empty stack, so a drive re-enters inline; armed
callbacks hold `Weak` refs, so dropping the `EventLoopRuntime` has native
`Runtime::drop` semantics and a late callback upgrades to nothing.
`block_on` on an event-loop runtime is rejected eagerly like a nested
runtime: its wait is the host loop, so no stack can hold the result.
`EventLoopRuntime::schedule` takes a root and a completion callback
(`Err(JoinError)` on a root panic). The API is `tokio_unstable` and
gated to non-pthread builds; a drive from a host callback while another
runtime's `block_on` is suspended under JSPI still panics as a nested
runtime.

Tests: `rt_emscripten_event_loop` pins the on-stack contracts (schedule
queues only, one batch per drive, host-context wakes never drive inline,
cross-runtime wakes, `block_on` and nested drive rejected, drop cancels).
`rt_emscripten_event_loop_main` (`harness = false`) schedules roots on
two runtimes and returns into the host loop, completing them through
timer re-arms, a greedy yielding sibling, a cross-runtime oneshot and a
TCP round trip between the runtimes, with and without `-sJSPI`; the
pre-js fails the run if the loop drains before the runtime exits.

CI adds an unstable lane. TEMPORARY: it overlays the emscripten
`epoll-callback` branch (emscripten-core/emscripten#27547) for
`emscripten_epoll_add_listener` until released.
guybedford added a commit to guybedford/tokio that referenced this pull request Sep 14, 2026
Follow-on to tokio-rs#8285 and tokio-rs#8438. A `LocalRuntime` driven by the host
JavaScript event loop instead of by a park, so it needs neither JSPI
nor pthreads and never blocks or suspends the host.

The scheduler drives to a fixed point, then waits. On this target a wait
is either a JSPI suspension (`block_on`) or a return to the host to be
called back; the event-loop runtime is the latter. The same
`current_thread` scheduler and drivers run one `event_interval` batch
per drive followed by the zero-duration driver turn a park would do
(due timers, I/O readiness, deferred wakers), then the wait is lowered
to arming host callbacks: an immediate drive if runnable work remains,
so the host loop gets a turn between batches as under a JSPI `block_on`,
and one `emscripten_set_timeout` for the soonest timer deadline. Timers
are fire-only with an epoch (`emscripten_clear_timeout` leaks its
keepalive) and kept across drives while the deadline is unchanged.
Socket readiness re-enters the drive through a persistent
`emscripten_epoll_add_listener` on the reactor's epoll fd, which also
carries the mio waker, so an external unpark of an I/O-enabled runtime
needs no extra arming; the I/O-less parker arms an immediate drive. A
wake during a drive is absorbed by that drive's own turn and re-arm.

Host callbacks run on an empty stack, so a drive re-enters inline; armed
callbacks hold `Weak` refs, so dropping the `EventLoopRuntime` has native
`Runtime::drop` semantics and a late callback upgrades to nothing.
`block_on` on an event-loop runtime is rejected eagerly like a nested
runtime: its wait is the host loop, so no stack can hold the result.
`EventLoopRuntime::schedule` takes a root and a completion callback
(`Err(JoinError)` on a root panic). The API is `tokio_unstable` and
gated to non-pthread builds; a drive from a host callback while another
runtime's `block_on` is suspended under JSPI still panics as a nested
runtime.

Tests: `rt_emscripten_event_loop` pins the on-stack contracts (schedule
queues only, one batch per drive, host-context wakes never drive inline,
cross-runtime wakes, `block_on` and nested drive rejected, drop cancels).
`rt_emscripten_event_loop_main` (`harness = false`) schedules roots on
two runtimes and returns into the host loop, completing them through
timer re-arms, a greedy yielding sibling, a cross-runtime oneshot and a
TCP round trip between the runtimes, with and without `-sJSPI`; the
pre-js fails the run if the loop drains before the runtime exits.

CI adds an unstable lane. TEMPORARY: it overlays the emscripten
`epoll-callback` branch (emscripten-core/emscripten#27547) for
`emscripten_epoll_add_listener` until released.
guybedford added a commit to guybedford/tokio that referenced this pull request Sep 14, 2026
Follow-on to tokio-rs#8285 and tokio-rs#8438. A `LocalRuntime` driven by the host
JavaScript event loop instead of by a park, so it needs neither JSPI
nor pthreads and never blocks or suspends the host.

The scheduler drives to a fixed point, then waits. On this target a wait
is either a JSPI suspension (`block_on`) or a return to the host to be
called back; the event-loop runtime is the latter. The same
`current_thread` scheduler and drivers run one `event_interval` batch
per drive followed by the zero-duration driver turn a park would do
(due timers, I/O readiness, deferred wakers), then the wait is lowered
to arming host callbacks: an immediate drive if runnable work remains,
so the host loop gets a turn between batches as under a JSPI `block_on`,
and one `emscripten_set_timeout` for the soonest timer deadline. Timers
are fire-only with an epoch (`emscripten_clear_timeout` leaks its
keepalive) and kept across drives while the deadline is unchanged.
Socket readiness re-enters the drive through a persistent
`emscripten_epoll_add_listener` on the reactor's epoll fd, which also
carries the mio waker, so an external unpark of an I/O-enabled runtime
needs no extra arming; the I/O-less parker arms an immediate drive. A
wake during a drive is absorbed by that drive's own turn and re-arm.

Host callbacks run on an empty stack, so a drive re-enters inline; armed
callbacks hold `Weak` refs, so dropping the `EventLoopRuntime` has native
`Runtime::drop` semantics and a late callback upgrades to nothing.
`block_on` on an event-loop runtime is rejected eagerly like a nested
runtime: its wait is the host loop, so no stack can hold the result.
`EventLoopRuntime::schedule` takes a root and a completion callback
(`Err(JoinError)` on a root panic). The API is `tokio_unstable` and
gated to non-pthread builds; a drive from a host callback while another
runtime's `block_on` is suspended under JSPI still panics as a nested
runtime.

Tests: `rt_emscripten_event_loop` pins the on-stack contracts (schedule
queues only, one batch per drive, host-context wakes never drive inline,
cross-runtime wakes, `block_on` and nested drive rejected, drop cancels).
`rt_emscripten_event_loop_main` (`harness = false`) schedules roots on
two runtimes and returns into the host loop, completing them through
timer re-arms, a greedy yielding sibling, a cross-runtime oneshot and a
TCP round trip between the runtimes, with and without `-sJSPI`; the
pre-js fails the run if the loop drains before the runtime exits.

CI adds an unstable lane. TEMPORARY: it overlays the emscripten
`epoll-callback` branch (emscripten-core/emscripten#27547) for
`emscripten_epoll_add_listener` until released.
guybedford added a commit to guybedford/tokio that referenced this pull request Sep 14, 2026
Follow-on to tokio-rs#8285 and tokio-rs#8438. A `LocalRuntime` driven by the host
JavaScript event loop instead of by a park, so it needs neither JSPI
nor pthreads and never blocks or suspends the host.

The scheduler drives to a fixed point, then waits. On this target a wait
is either a JSPI suspension (`block_on`) or a return to the host to be
called back; the event-loop runtime is the latter. The same
`current_thread` scheduler and drivers run one `event_interval` batch
per drive followed by the zero-duration driver turn a park would do
(due timers, I/O readiness, deferred wakers), then the wait is lowered
to arming host callbacks: an immediate drive if runnable work remains,
so the host loop gets a turn between batches as under a JSPI `block_on`,
and one `emscripten_set_timeout` for the soonest timer deadline. Timers
are fire-only with an epoch (`emscripten_clear_timeout` leaks its
keepalive) and kept across drives while the deadline is unchanged.
Socket readiness re-enters the drive through a persistent
`emscripten_epoll_add_listener` on the reactor's epoll fd, which also
carries the mio waker, so an external unpark of an I/O-enabled runtime
needs no extra arming; the I/O-less parker arms an immediate drive. A
wake during a drive is absorbed by that drive's own turn and re-arm.

Host callbacks run on an empty stack, so a drive re-enters inline; armed
callbacks hold `Weak` refs, so dropping the `EventLoopRuntime` has native
`Runtime::drop` semantics and a late callback upgrades to nothing.
`block_on` on an event-loop runtime is rejected eagerly like a nested
runtime: its wait is the host loop, so no stack can hold the result.
`EventLoopRuntime::schedule` takes a root and a completion callback
(`Err(JoinError)` on a root panic). The API is `tokio_unstable` and
gated to non-pthread builds; a drive from a host callback while another
runtime's `block_on` is suspended under JSPI still panics as a nested
runtime.

Tests: `rt_emscripten_event_loop` pins the on-stack contracts (schedule
queues only, one batch per drive, host-context wakes never drive inline,
cross-runtime wakes, `block_on` and nested drive rejected, drop cancels).
`rt_emscripten_event_loop_main` (`harness = false`) schedules roots on
two runtimes and returns into the host loop, completing them through
timer re-arms, a greedy yielding sibling, a cross-runtime oneshot and a
TCP round trip between the runtimes, with and without `-sJSPI`; the
pre-js fails the run if the loop drains before the runtime exits.

CI adds an unstable lane. TEMPORARY: it overlays the emscripten
`epoll-callback` branch (emscripten-core/emscripten#27547) for
`emscripten_epoll_add_listener` until released.
guybedford added a commit to guybedford/tokio that referenced this pull request Sep 14, 2026
Follow-on to tokio-rs#8285 and tokio-rs#8438. A `LocalRuntime` driven by the host
JavaScript event loop instead of by a park, so it needs neither JSPI
nor pthreads and never blocks or suspends the host.

The scheduler drives to a fixed point, then waits. On this target a wait
is either a JSPI suspension (`block_on`) or a return to the host to be
called back; the event-loop runtime is the latter. The same
`current_thread` scheduler and drivers run one `event_interval` batch
per drive followed by the zero-duration driver turn a park would do
(due timers, I/O readiness, deferred wakers), then the wait is lowered
to arming host callbacks: an immediate drive if runnable work remains,
so the host loop gets a turn between batches as under a JSPI `block_on`,
and one `emscripten_set_timeout` for the soonest timer deadline. Timers
are fire-only with an epoch (`emscripten_clear_timeout` leaks its
keepalive) and kept across drives while the deadline is unchanged.
Socket readiness re-enters the drive through a persistent
`emscripten_epoll_add_listener` on the reactor's epoll fd, which also
carries the mio waker, so an external unpark of an I/O-enabled runtime
needs no extra arming; the I/O-less parker arms an immediate drive. A
wake during a drive is absorbed by that drive's own turn and re-arm.

Host callbacks run on an empty stack, so a drive re-enters inline; armed
callbacks hold `Weak` refs, so dropping the `EventLoopRuntime` has native
`Runtime::drop` semantics and a late callback upgrades to nothing.
`block_on` on an event-loop runtime is rejected eagerly like a nested
runtime: its wait is the host loop, so no stack can hold the result.
`EventLoopRuntime::schedule` takes a root and a completion callback
(`Err(JoinError)` on a root panic). The API is `tokio_unstable` and
gated to non-pthread builds; a drive from a host callback while another
runtime's `block_on` is suspended under JSPI still panics as a nested
runtime.

Tests: `rt_emscripten_event_loop` pins the on-stack contracts (schedule
queues only, one batch per drive, host-context wakes never drive inline,
cross-runtime wakes, `block_on` and nested drive rejected, drop cancels).
`rt_emscripten_event_loop_main` (`harness = false`) schedules roots on
two runtimes and returns into the host loop, completing them through
timer re-arms, a greedy yielding sibling, a cross-runtime oneshot and a
TCP round trip between the runtimes, with and without `-sJSPI`; the
pre-js fails the run if the loop drains before the runtime exits.

CI adds an unstable lane. TEMPORARY: it overlays the emscripten
`epoll-callback` branch (emscripten-core/emscripten#27547) for
`emscripten_epoll_add_listener` until released.
@guybedford
guybedford force-pushed the epoll-callback branch 5 times, most recently from 41885a0 to 9d68e70 Compare September 23, 2026 06:19
@kripken

kripken commented Sep 23, 2026

Copy link
Copy Markdown
Member

I honestly do not know enough about epoll and related systems to review this...

@kripken
kripken removed their request for review September 23, 2026 18:08
@sbc100

sbc100 commented Sep 23, 2026

Copy link
Copy Markdown
Collaborator

No problem, I just added you in case you were interested. I should have just CC'ed you probably

This is an automatic change generated by tools/maint/rebaseline_tests.py.

The following (12) test expectation files were updated by
running the tests with `--rebaseline`:

```
codesize/test_codesize_cxx_ctors2.json: 152604 => 152604 [+0 bytes / +0.00%]
codesize/test_codesize_cxx_except_wasm.json: 168663 => 168663 [+0 bytes / +0.00%]
codesize/test_codesize_cxx_except_wasm_legacy.json: 166523 => 166523 [+0 bytes / +0.00%]
codesize/test_codesize_cxx_lto.json: 119632 => 119632 [+0 bytes / +0.00%]
codesize/test_codesize_hello_dylink.json: 43250 => 43250 [+0 bytes / +0.00%]
codesize/test_codesize_hello_dylink_all.json: 859704 => 859814 [+110 bytes / +0.01%]
codesize/test_codesize_mem_O3_grow_standalone.json: 9306 => 9306 [+0 bytes / +0.00%]
codesize/test_codesize_mem_O3_standalone.json: 9140 => 9140 [+0 bytes / +0.00%]
codesize/test_codesize_mem_O3_standalone_narg.json: 8455 => 8455 [+0 bytes / +0.00%]
codesize/test_codesize_mem_O3_standalone_narg_flto.json: 7386 => 7386 [+0 bytes / +0.00%]
codesize/test_codesize_minimal_pthreads.json: 26030 => 26109 [+79 bytes / +0.30%]
codesize/test_codesize_minimal_pthreads_memgrowth.json: 26489 => 26573 [+84 bytes / +0.32%]

Average change: +0.05% (+0.00% - +0.32%)
```
A listener never keeps the runtime or its registering thread alive; a
program holds the runtime itself with emscripten_runtime_keepalive_push/pop.
Drops the host-armed keepalive model, the cross-thread keepalive_holds in
struct pthread and the keepRuntimeAlive() change. The pending-delivery hold
stays. Tests rewritten to the new contract.
Per review: the cross-thread delivery moves to a follow-up so the initial
version is main-thread only (ENOTSUP from other threads), and listeners
are identified by the (callback, userdata) pair like the html5 event
handlers, so remove_listener takes userdata too and a duplicate pair is
EEXIST.
@guybedford

Copy link
Copy Markdown
Collaborator Author

Would it make sense to pass the epoll fd to the callback. e.g. typedef void (*em_epoll_callback)(int epfd, void *userdata); ? Without this caller would have to either store the fd globally or embed it in userdata somehow.

I've added this - in the process I had to refactor the listeners since we were attaching them to shared item and not per-fd (so fd wasn't unique under dup). This is now fixed by having the listeners on the stream type directly.

Comment thread ChangeLog.md Outdated
@sbc100

sbc100 commented Oct 2, 2026

Copy link
Copy Markdown
Collaborator

This change LGTM now. Feel for to land once all the tests are passing.

@guybedford
guybedford enabled auto-merge (squash) October 2, 2026 23:46
@guybedford
guybedford merged commit 30f57a0 into emscripten-core:main Oct 3, 2026
42 checks passed
@guybedford
guybedford deleted the epoll-callback branch October 3, 2026 00:06
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants