Repository navigation
io_uring threads - #20
Open
adamsitnik wants to merge 46 commits into
Open
adamsitnik wants to merge 46 commits into
adamsitnik wants to merge 46 commits into
Conversation
Adds a minimal, generic io_uring native API to System.Native as the shared foundation for future ThreadPool integration architectures: - SystemNative_IoRingIsAvailable / Create / Submit / WaitForCompletions / Close - Talks to the kernel via raw io_uring_setup/enter syscalls (no liburing dependency), mmap-based SQ/CQ/SQE ring setup, with acquire/release atomics on the shared head/tail indices. - Supports single-buffer and vectored read/write, both positional (offset) and non-positional (current file position / non-seekable fds). - Gated by a new HAVE_LINUX_IO_URING_H CMake detection (falls back to ENOTSUP on older kernels / non-Linux). - Adds matching managed Interop.Sys bindings (Interop.IoRing.cs). Out of scope for this iteration: RandomAccess.Unix.cs / SafeFileHandle wiring, ThreadPool integration, and cancellation - tracked as follow-up work. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Implements Option 3 from the io_uring design doc's Optimal Thread Count section: a single shared, plain io_uring ring guarded by a lock for submission, with a CAS-elected rotating driver thread that reaps completions and queues continuations as ordinary ThreadPool work items (never run inline on the driver thread). - PortableThreadPool.IoUring.Unix.cs: new IoUringThreadPool static class (ring creation/enablement, TrySubmit, TryBecomeDriverAndDrive, Dispatch) and IIoUringOperation interface. Enabled by default on Linux when the kernel supports io_uring; opt out via DOTNET_USE_IO_URING=0. Queue depth 1024. - PortableThreadPool.WorkerThread.cs: worker threads about to park now first try to become the io_uring driver. - SafeFileHandle.ThreadPoolValueTaskSource.cs: Read/Write/ReadScatter/ WriteGather now attempt submission via io_uring first (pinning buffers/vectors, ref-counting the SafeHandle for the in-flight duration), handling partial writes by resubmitting the remainder, and falling back transparently to the existing blocking-work-item implementation when io_uring is disabled/unavailable or a submission fails. - RandomAccess.Unix.cs required no changes; it already delegates to ThreadPoolValueTaskSource. Verified: clr.corelib+clr.nativecorelib+libs.pretest -rc checked builds clean. System.IO.FileSystem.Tests RandomAccess async test classes (ReadAsync, WriteAsync, ReadScatterAsync, WriteGatherAsync, NonSeekable_AsyncHandles) pass identically with io_uring on (default) and with DOTNET_USE_IO_URING=0. Out of scope for this iteration: cancellation of in-flight io_uring operations, SINGLE_ISSUER/ DEFER_TASKRUN, ring sharding. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
s_lock only ever needs to guard the SQ ring (TrySubmit's SQE pushes); it has no bearing on the CQ ring, which is a physically separate mmap'd region read exclusively by the elected driver thread (guarded by the s_isDriving CAS, not by s_lock). Wrapping the drain loop's subsequent IoRingWaitForCompletions calls in s_lock therefore added no correctness benefit and needlessly contended with concurrent TrySubmit callers while draining. Updated comments to make clear what s_lock actually protects. Verified: clr.corelib+clr.nativecorelib+libs.pretest -rc checked builds clean; RandomAccess async tests (87 tests) still pass identically with io_uring on and with DOTNET_USE_IO_URING=0. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
1. SystemNative_IoRingSubmit could report an entry as "not submitted" (and the managed caller would free its correlation GCHandle) even after the SQE was already durably published to the kernel via the SQ tail, if io_uring_enter's return value was negative/short. This could cause a later completion to reference freed/reused handle memory. Fixed by always reporting *submittedCount = queued once entries are tail-published, treating io_uring_enter's result as an advisory "kick" only. 2. IORING_SETUP_SINGLE_ISSUER | IORING_SETUP_DEFER_TASKRUN (added in the PAL-only iteration for their perf benefit) are incompatible with the ThreadPool integration's shared-ring design, where any worker thread may submit to or drain the ring: per io_uring_setup_flags(7), SINGLE_ISSUER pins the "issuer" role to whichever thread first calls io_uring_enter, and every other thread's call fails with -EEXIST. This caused a hang under real multi-threaded load. Removed both flags. 3. The driver-rotation hook only runs opportunistically, from a worker thread that is already about to park. Submitting an operation via io_uring does not go through the normal work-queue signaling path, so if every worker was already parked when an operation was submitted, nobody would ever wake up to reap its completion - a permanent hang. Fixed by calling WorkerThread.MaybeAddWorkingWorker after a successful submit (mirroring what enqueuing an ordinary work item already does) so a worker is guaranteed to loop back and attempt to become the driver. Also clamp IORING_OP_READV/WRITEV's vector count to IOV_MAX (matching the existing plain readv/writev behavior), since io_uring rejects requests with more vectors than that; the managed caller already handles the resulting short read/write by resubmitting the remainder. Verified: System.IO.FileSystem.Tests RandomAccess async test classes (ReadAsync/WriteAsync/ReadScatterAsync/WriteGatherAsync/ NonSeekable_AsyncHandles - 87 tests) pass repeatedly and reliably with io_uring both enabled and disabled (DOTNET_USE_IO_URING=0). Known remaining limitation (not fixed here, out of scope for this experimental iteration): concurrent, unlinked io_uring reads/writes on the same non-seekable (pipe/socket) fd are inherently racy per io_uring's documented semantics and can complete with -ECANCELED; this surfaced as failures in the broader FileStream-over-pipe conformance test suite (not in RandomAccess's own async test classes). Fixing this would require per-fd operation serialization or a different dispatch strategy for pollable fds, tracked as future work. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Two concrete scalability bottlenecks were found via benchmarking: 1. TrySubmit performed the entire IoRingSubmit P/Invoke - including the internal blocking io_uring_enter syscall - inside the shared lock, serializing every concurrent submitting thread behind one lock AND one syscall each. 2. TryBecomeDriverAndDrive drained completions one at a time (maxCompletions=1), requiring one syscall per completion even though the native function already supports batch-draining multiple completions in a single call. Fixes (single-ring, single-driver architecture preserved): - Split submission into SystemNative_IoRingSubmit (SQE fill + tail publish only, no syscall) and a new SystemNative_IoRingKick (calls io_uring_enter). The managed lock now only guards the cheap enqueue step; the kick happens after releasing the lock, and is safe to call concurrently since IORING_SETUP_SINGLE_ISSUER is not used. - TryBecomeDriverAndDrive now requests up to 64 completions per IoRingWaitForCompletions call instead of 1, draining many completions per syscall. Benchmark (concurrency 1/4/16/32/64, corerun, tmpfs-backed regular files): - Before: 16,190 / 22,750 / 24,084 / 29,363 (n/a) ops/sec - After: 15,790 / 45,977 / 81,028 / 88,305 / 87,962 ops/sec Still slower than the blocking-ThreadPool fallback for this workload, but scales meaningfully with concurrency now instead of flatlining. All 187 RandomAccess async tests pass with io_uring enabled and disabled. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Replaces the hand-rolled io_uring_setup(2)/io_uring_enter(2) syscall wrappers and manual mmap'd submission/completion ring-buffer plumbing (head/tail indices, acquire/release barriers, SQE array indexing) with liburing (https://github.com/axboe/liburing): - SystemNative_IoRingIsAvailable / IoRingCreate now use io_uring_queue_init(_params) / io_uring_queue_exit. - SystemNative_IoRingSubmit now uses io_uring_get_sqe + io_uring_prep_read/write/readv/writev + io_uring_sqe_set_data64 + io_uring_submit. - SystemNative_IoRingWaitForCompletions now uses io_uring_wait_cqe (blocking wait for the first completion) + io_uring_peek_batch_cqe (batched drain) + io_uring_cq_advance. - SystemNative_IoRingClose now uses io_uring_queue_exit. This removes ~150 lines of manual ring-buffer/syscall code, replacing it with liburing's own reviewed/tested implementation, at the cost of removing the previous SystemNative_IoRingKick split: liburing's io_uring_submit is documented as not thread-safe with itself, combining SQE flush and the io_uring_enter syscall into one non-splittable call, so the managed lock in PortableThreadPool.IoUring.Unix.cs now once again wraps the whole IoRingSubmit call (as it did before the io_uring_enter/lock-split optimization). Build changes: - configure.cmake / pal_config.h.in: HAVE_LINUX_IO_URING_H -> HAVE_LIBURING_H, probing for liburing.h and the liburing library via find_library instead of the raw uapi header and syscall numbers. - extra_libs.cmake: link liburing into System.Native's shared library when found. This is a no-op (does not fail the build) on any machine/image without liburing-dev installed, matching the existing graceful-degradation behavior (io_uring support is simply unavailable at runtime). Verified: all 187 RandomAccess async tests pass with io_uring enabled and with DOTNET_USE_IO_URING=0. Benchmark (concurrency 1/4/16/32/64, corerun, tmpfs-backed regular files, 4KB buffers): io_uring ON regresses from the syscall-split version (15,790/45,977/81,028/88,305/87,962) back to roughly the original single-lock numbers (9,372/17,080/20,276/25,714/28,053), since liburing's io_uring_submit cannot be split the way the raw syscall could. io_uring OFF (unaffected fallback path) remains at 19,485/87,281/317,757/482,335/660,303. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…ssion The initial liburing rewrite (c5c0425) combined SQE publish and the io_uring_enter(2) syscall into a single io_uring_submit() call under the shared lock, since liburing documents io_uring_get_sqe/io_uring_submit as not thread-safe with themselves. This silently reintroduced the concurrency-scaling bottleneck that an earlier session had specifically fixed for the raw-syscall implementation (a ~3x regression at concurrency 64: 28,053 vs 87,962 ops/sec), and the previous commit was pushed without catching or flagging this - it should not have been. The actual non-thread-safe part of liburing's submission API is only the *local* (non-atomic) sq.sqe_head/sq.sqe_tail bookkeeping touched by io_uring_get_sqe and io_uring_submit's internal flush step - the io_uring_enter(2) syscall itself is safe to call concurrently from multiple threads for a ring without IORING_SETUP_SINGLE_ISSUER (the kernel serializes it internally). This restores the same split as before, but on top of liburing: - SystemNative_IoRingSubmit now only fills SQEs via io_uring_get_sqe/io_uring_prep_* and publishes the SQ tail via a new IoRingFlushSq helper - a direct, minimal reimplementation of liburing's own internal __io_uring_flush_sq (5 lines, using only liburing's public struct fields), instead of calling io_uring_submit(). No syscall happens under the lock. - SystemNative_IoRingKick (re-added) wraps liburing's public io_uring_enter() to ask the kernel to process whatever has been published so far. Called by the managed TrySubmit after releasing s_lock, exactly as before the liburing rewrite. - PortableThreadPool.IoUring.Unix.cs: TrySubmit calls IoRingKick after releasing s_lock again; updated doc comments accordingly. Verified: - All 187 RandomAccess async tests pass with io_uring ON and with DOTNET_USE_IO_URING=0. - Benchmark (concurrency 1/4/16/32/64, corerun, tmpfs regular files, 4KB buffers, ops/sec, two runs): io_uring ON is back to 9,232-9,441 / 29,554-30,793 / 62,591-63,920 / 69,663-72,057 / 79,268-82,329, matching the pre-liburing-rewrite split-lock numbers (15,790/45,977/81,028/88,305/87,962) within run-to-run noise - confirming the regression is fixed and this rewrite no longer trades away the earlier concurrency fix. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Add opportunistic batching in TrySubmit: an Interlocked counter tracks threads currently between 'about to submit' and 'finished submitting'. Only the thread whose decrement observes the counter back at 0 (the last one out of the current wave of concurrent submitters) issues the IoRingKick (io_uring_enter) syscall; others skip it, since their already-published SQEs will be picked up by that same kick. The kick fires whenever the counter reaches 0, regardless of whether this thread's own submission succeeded. Gating it on this thread's own success would be unsound: a successful submitter that isn't last out correctly defers to a later thread, but if that later thread's own submission then fails (e.g. queue momentarily full), nobody would kick at all - and io_uring_wait_cqe does not submit pending SQEs on its own, so the earlier successful submission's completion would never arrive, hanging that operation indefinitely. Since a thread's decrement only ever runs after its own submit-or-fail attempt is fully complete, observing the counter at 0 guarantees every submission in the wave has already been durably published, making the unconditional kick safe. Benchmarked (RandomAccess async ops via corerun, concurrency 1-64): - tmpfs 4KB buffers: up to +30% at concurrency 64 (89K -> 116K ops/sec). - ext4 O_DIRECT 4KB buffers: up to +13% at concurrency 64. - ext4 O_DIRECT 64KB buffers: no measurable change (disk-bandwidth bound at that buffer size, not syscall-count bound). Verified 187/187 RandomAccess async tests pass, both io_uring enabled (default) and with DOTNET_USE_IO_URING=0. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
When the io_uring driver drains multiple completions in one pass, it previously called ThreadPool.UnsafeQueueUserWorkItem once per completion. This adds a batched enqueue path so all ready continuations from a drained batch are queued via a single call: - ThreadPoolWorkQueue: added EnqueueBatch(ReadOnlySpan<IThreadPoolWorkItem>, bool) and internal ThreadPool.UnsafeQueueUserWorkItems(...), mirroring the single-item Enqueue/UnsafeQueueUserWorkItem path but consolidating the logging check and EnsureWorkerRequested() call to once per batch instead of once per item. - IIoUringOperation.CompleteFromIoUring now returns the IThreadPoolWorkItem to queue (or null if not yet ready, e.g. a partial write was resubmitted and remains in flight) instead of enqueuing itself directly. - PortableThreadPool.IoUringThreadPool's driver (DispatchBatch, replacing the old per-completion Dispatch) collects the non-null returned work items from each drained batch (up to 64, one syscall's worth) into a reused scratch array - safe without extra synchronization since only the CAS-elected driver thread ever touches it - then queues them all in one UnsafeQueueUserWorkItems call. - SafeFileHandle.ThreadPoolValueTaskSource.CompleteFromIoUring/ TryContinuePartialWrite updated to return the fallback work item instead of enqueuing it directly at the two partial-write-resubmission-failure call sites. Verified: clean rebuild; RandomAccess async tests (87 tests) pass across 3 runs with io_uring enabled and 1 run with DOTNET_USE_IO_URING=0. Benchmarked (io_uring ON; OFF path unaffected): - tmpfs, 4KB buffers: +2-8% at concurrency 32/64 (91-97K -> 98K, 109-117K -> 117K ops/sec), where completions arrive in the largest batches per driver wakeup. - O_DIRECT, 4KB and 64KB buffers: within normal disk-I/O run-to-run noise in both directions, no regressions - expected, since real disk latency dominates there and reducing enqueue-call count matters far less. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…ecture 3)
Implements TCP/stream socket Accept/Connect/Recv/Send via io_uring on top of
the single shared ring + CAS-elected rotating driver design already used for
RandomAccess file I/O ("Option 3" / io_uring_files), as a companion experiment
to the per-thread-ring port on io_uring_sockets. Goal: measure the maximum
socket throughput achievable with a single shared ring.
- pal_io.h/pal_io.c: new IoRingOp_Accept/Connect/Recv/Send opcodes and
Flags/SockAddr/SockAddrLen IoRingRequest fields, plus the corresponding
io_uring_prep_* cases in IoRingFillSqe.
- Interop.IoRing.cs: managed mirror of the above.
- New public System.Threading.IoUring API (IsSupported, TrySubmitRecv/Send/
Accept/Connect) that submits directly onto the existing shared-ring
IoUringThreadPool. Unlike the per-thread-ring version, completions are
always queued/batched (never run inline), matching this architecture's
constraint that the CAS-elected driver must never execute continuations
itself.
- SocketAsyncContext.IoUring.Unix.cs (new, architecture-agnostic, ported
verbatim) and the SocketAsyncContext.Unix.cs call-site diff wiring
Accept/Connect/Receive/ReceiveFrom/SendTo to attempt the io_uring path
first, falling back to the existing epoll-based SocketAsyncEngine path
unchanged on any submit failure/ineligibility.
- Registered the new files in System.Net.Sockets.csproj and
System.Private.CoreLib.Shared.projitems; added the public IoUring ref
assembly stub to System.Runtime.cs.
- Deliberately not ported: the per-thread-ring-specific
IoUring.TryDriveCurrentThreadRing()/SafeSocketHandle hook - there is no
"current thread's own ring" concept in the shared-ring design, so it does
not apply.
Verification:
- clr.corelib+clr.nativecorelib+libs.pretest -c Release -rc Release builds
clean.
- Core Accept/Connect/Send/Receive functional tests pass (61/61 with
io_uring OFF as a sanity baseline; sampled ON happy-path tests pass too).
- Found and documented (not fixed, out of scope, inherited from the
per-thread-ring branch too): *GetsCanceledByDispose/UDP-cancel tests hang
with io_uring ON, because neither port implements IORING_OP_ASYNC_CANCEL -
closing a socket fd does not interrupt an in-flight io_uring Accept/Recv
the way epoll's non-blocking retry does. Everything else passes.
Benchmark (loopback HTTP echo, Release build, ThreadPool pinned to
Environment.ProcessorCount, corerun):
| Connections | epoll (OFF) rps | shared-ring io_uring (ON) rps | ON/OFF |
|---|---|---|---|
| 1 | 11,257 | 10,917 | 0.97x |
| 4 | 38,374 | 35,722 | 0.93x |
| 16 | 114,659 | 46,767 | 0.41x |
| 50 | 343,532 | 45,640 | 0.13x |
| 100 | 363,928 | 40,956 | 0.11x |
| 200 | 424,061 | 53,823 | 0.13x |
Conclusion: the single-shared-ring architecture caps socket throughput at
roughly the same order of magnitude as the earlier per-thread-ring result at
high connection counts (both fall to ~0.1-0.16x of epoll), but is worse at
low concurrency (0.93-0.97x here vs. the per-thread-ring's previously
measured 1.36x win at 4 connections) - the shared ring's single CAS-elected
driver plus lock-guarded submission is pure added overhead at low
concurrency with no offsetting benefit, since there is no per-thread ring to
amortize it against. Neither single-ring architecture beats epoll for
sockets in this experiment; per-thread rings remain the better of the two
io_uring designs tried so far.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
liburing requires libraring.so to be present on the target machine, which isn't the case in typical benchmarking/deployment scenarios (crank agents, containers without liburing-dev installed). Reverts the liburing rewrite (c5c0425) back to talking to the kernel directly via raw io_uring_setup(2)/io_uring_enter(2) syscalls and manual SQ/CQ/SQE ring mmap'ing, while keeping the submit/kick split and the socket opcodes (Accept/Connect/Recv/Send) added on top of liburing in later commits. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
sqe->off was unconditionally set to -1 (non-positional file offset marker) before the opcode switch, but only the Accept/Connect cases explicitly overwrote it afterwards. For Recv/Send, sqe->off was left at -1 (0xFFFFFFFFFFFFFFFF), which the kernel rejects with EINVAL since that field must stay at its zeroed default for those opcodes. This went unnoticed because liburing-dev wasn't installed, so HAVE_LIBURING_H was undefined and the entire io_uring PAL layer silently compiled to no-op stubs - once liburing-dev was installed and the raw-syscall implementation actually started running, every io_uring-backed socket Recv/Send crashed with SocketException(22). Moved the off assignment into only the Read/Write/ReadV/WriteV cases, where it is actually meaningful. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…bursts Gate the WorkerThread.MaybeAddWorkingWorker() call in TrySubmit with a CAS flag (s_driverWakeRequested), so only the first submitter that notices "no driver active" in a given window pays for the wake. Concurrent submitters piling in behind it during a burst previously each triggered their own wake/park cycle even though only one thread can ever win the s_isDriving CAS - pure wasted work. The flag is reset as soon as any worker visits TryBecomeDriverAndDrive() while an operation is in flight, regardless of whether it wins the CAS, so it can never get stuck. Benchmarked as throughput-neutral (~120K ops/sec ON either way at concurrency=32) under quiesced system conditions; kept as a correctness/robustness improvement that reduces theoretical wake/park churn without measured throughput cost. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…er thread)
Adds IORING_SETUP_SINGLE_ISSUER + IORING_SETUP_DEFER_TASKRUN support to the io_uring
PAL layer, replacing the shared-ring CAS-driver design's submission path with a
single dedicated OS thread ("the issuer") that owns all io_uring_enter calls for
both submission and completion reaping.
- pal_io.c/h: SystemNative_IoRingCreate gains a singleIssuer parameter that sets
IORING_SETUP_SINGLE_ISSUER | IORING_SETUP_DEFER_TASKRUN when enabled.
SystemNative_IoRingWaitForCompletions now always calls io_uring_enter with
IORING_ENTER_GETEVENTS (required to pump DEFER_TASKRUN's deferred completions
even when minComplete is 0).
- Interop.IoRing.cs: mirrors the new singleIssuer parameter.
- PortableThreadPool.IoUring.Unix.cs: removes the shared-ring lock/CAS-driver
design; adds an MPSC ConcurrentQueue + ManualResetEventSlim hand-off from
TrySubmit (always returns true, unbounded queue) to a dedicated issuer thread
running IssuerLoop (submit-then-drain-then-wait). Uses a handshake pattern in
the static constructor to avoid a CLR type-initialization deadlock between the
cctor thread and the issuer thread.
Validation on this branch (io_uring_single_issuer):
- Release build: 0 errors.
- System.IO.FileSystem.Tests: 9761 total, only pre-existing flaky
NoDataIsLostWhenWritingToFile failures (confirmed unrelated, pass in isolation).
- perfcollect trace confirms the kernel-side io_uring_enter contention seen in the
shared-ring architecture (osq_lock/mutex_spin_on_owner ~13% self time) is
eliminated (io_uring/mutex-related symbols <0.5% self time) with SINGLE_ISSUER.
- Throughput is currently on par with or slightly behind epoll across
concurrency 1-64; a new ThreadNative_SpinWait hotspot (~14% self time) in the
issuer thread's poll loop is the next optimization target.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…ption 3) The single-issuer io_uring architecture's IssuerLoop previously used a ManualResetEventSlim to wake the dedicated issuer thread on new submissions, with a bounded PollIntervalMs wait whenever operations were in flight (needed to periodically pump IORING_ENTER_GETEVENTS due to DEFER_TASKRUN semantics). Profiling showed this caused a real, measurable CPU cost (~14% self-time in ThreadNative_SpinWait) from the CLR-level spin-before-block behavior of Wait() being invoked in a tight cycle under load. A prior attempt at a zero-spin-count fix traded that cost for worse low-concurrency latency (a net regression at c=1/c=4). This change replaces the ManualResetEventSlim entirely with a single eventfd, registered on the ring via IORING_REGISTER_EVENTFD. Both the kernel (on CQE post / deferred completion task-work becoming ready) and TrySubmit (via a direct EventFdWrite from any thread) write to this same fd; the issuer thread does a single real (poll(2)-based) kernel-blocking EventFdWait with no CLR-level spin at all. This is the documented intended usage pattern for combining IORING_SETUP_DEFER_TASKRUN with a registered eventfd for single-threaded event loops. Native PAL additions (pal_io.c/pal_io.h, entrypoints.c): - SystemNative_IoRingRegisterEventFd: creates an EFD_NONBLOCK|EFD_CLOEXEC eventfd and registers it on the ring via IORING_REGISTER_EVENTFD. - SystemNative_EventFdWrite: writes 1 to an arbitrary eventfd from any thread (used by TrySubmit to wake the issuer thread). - SystemNative_EventFdWait: poll(2)-based blocking wait with a timeout, draining the counter via read() on success. - SystemNative_IoRingClose now also closes the registered eventfd. Managed changes (Interop.IoRing.cs, PortableThreadPool.IoUring.Unix.cs): - Added LibraryImport declarations for the three new native functions. - The static constructor's ring-creation handshake now also registers the eventfd (still using only captured locals from the dedicated issuer thread, per the existing type-initialization-deadlock-avoidance pattern). - TrySubmit calls EventFdWrite instead of ManualResetEventSlim.Set(). - IssuerLoop calls EventFdWait instead of ManualResetEventSlim.Wait(), using an indefinite wait when idle and a defensive (not routinely relied upon) bounded safety-net timeout when operations are in flight. Benchmark results (concurrency 1-256, EventPipe disabled, 64-byte messages): io_uring (ON) now beats epoll (OFF) at every concurrency level tested, confirmed via a repeated run - a clean win, unlike the prior zero-spin-count attempt which regressed at low concurrency. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
WorkerThread no longer needs to poll for a chance to become the io_uring driver: since the single-issuer architecture moved completion-reaping exclusively onto the dedicated issuer thread, TryBecomeDriverAndDrive was already a permanent no-op (always returning false). Removed the method along with its call site in PortableThreadPool.WorkerThread.cs and the now-stale doc comment references. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…call Previously IssuerLoop's per-iteration DrainAndSubmit always called IoRingKick (a plain io_uring_enter with to_submit=SqEntries, no GETEVENTS) right after filling each batch of SQEs, and the subsequent DrainCompletions always called IoRingWaitForCompletions with to_submit=0 (submit nothing, only reap completions/pump deferred DEFER_TASKRUN task-work). That meant two separate io_uring_enter syscalls per loop iteration even in the common case of a single small batch, even though io_uring_enter can submit and reap completions in one call. SystemNative_IoRingWaitForCompletions now always passes ring->SqEntries (an upper bound, exactly as safe as SystemNative_IoRingKick's existing use of the same value - io_uring_enter never over-consumes or double-processes entries) as to_submit instead of 0, so its one syscall both flushes any already-published-but-not-yet-submitted SQEs and reaps completions. DrainAndSubmit no longer kicks after the last batch it fills - those SQEs are left for the always-immediately-following DrainCompletions call to flush together with reaping completions. It only still kicks between batches when more remains queued (an unusually large burst spanning multiple MaxRequestsPerSubmitBatch-sized batches), to relieve SQ backpressure before continuing to fill more. Net effect: the common single-batch-per-iteration case now performs exactly one io_uring_enter syscall per IssuerLoop iteration instead of two. Verified via smoke tests at concurrency 1/4/64/256 (no hangs, correct results) and a full ON/OFF benchmark matrix (results noisy due to a loaded shared machine, but no regression observed; functionality confirmed correct). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…pt hang SocketAsyncContext.IoUring.Unix.cs: remove the separate DOTNET_SYSTEM_NET_SOCKETS_USE_IO_URING opt-in gate and s_ioUringSocketsEnabled field. All four io_uring code paths (receive/send/accept/connect) now check System.Threading.IoUring.IsSupported directly, which already reflects the single DOTNET_USE_IO_URING switch. This is the only opt-in that should exist. pal_io.c: fix a real bug in SystemNative_IoRingWaitForCompletions that surfaced once sockets actually started using io_uring end-to-end: combining a non-zero to_submit with IORING_ENTER_GETEVENTS in the same io_uring_enter call reliably prevents completions (e.g. IORING_OP_ACCEPT) from ever being posted under IORING_SETUP_DEFER_TASKRUN, even when to_submit is a harmless no-op upper bound. This made Kestrel's listening-socket AcceptAsync hang indefinitely as soon as io_uring was genuinely engaged for sockets. Root-caused with a minimal, dependency-free, single-threaded C repro (raw io_uring_setup/io_uring_enter syscalls, no managed code) that reproduces the hang with the same ring flags, and confirms splitting submission and completion-waiting into two separate io_uring_enter calls (submit-only, then a separate GETEVENTS-only call with to_submit=0) fixes it. Verified: sockets_bench regression-passes, and Kestrel/TechEmpower now handles requests correctly with DOTNET_USE_IO_URING=1 under sustained wrk load (previously hung on every request). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Profiling the TechEmpower JSON benchmark under wrk load showed thousands of individual EventFdWrite syscalls (one per TrySubmit call), even though the issuer thread only needed a small fraction of that many actual wake-ups: many concurrent TrySubmit calls from different Thread Pool worker threads land in the same "the issuer thread is already awake and about to drain the queue anyway" window. Add a s_wakeSignaled coalescing flag: only the thread that wins the 0->1 CAS transition actually writes to the eventfd. IssuerLoop resets the flag immediately before re-checking the queue/waiting, and re-checks the queue right after the reset, so a TrySubmit call racing with the reset is still safely picked up - no wake-up is ever missed. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Profiling showed that at low-to-moderate concurrency, requests tend to arrive in small, closely-spaced bursts rather than a steady stream large enough to always find the issuer thread already awake and mid-drain. A full EventFdWait park-then-wake round trip costs a real scheduling wake-up (potentially many microseconds under load), so before parking, briefly spin (SpinCountBeforeBlocking iterations of Thread.SpinWait(1)) checking s_pendingSubmissions between each iteration, letting an imminent TrySubmit call be observed almost immediately instead. Only spins when something is already in flight (s_inFlightCount > 0), so a fully idle ring never spins. This measurably closed the throughput gap between io_uring and epoll in TechEmpower JSON benchmarking, especially at lower concurrency. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
SystemNative_IoRingWaitForCompletions always issued two separate io_uring_enter syscalls per call: a submit-only flush (to publish any SQEs the issuer thread queued via SystemNative_IoRingSubmit but hasn't yet asked the kernel to consume) followed by a GETEVENTS-only call to reap completions. The submit-only call was unconditional, even when nothing new had been submitted since the previous flush. Under light load (e.g. low wrk concurrency), completions arrive in small trickles, so the issuer thread wakes frequently with nothing new to submit, yet still paid for that first syscall every single time. Track the last flushed SQ tail (IoRing.SqFlushedTail, updated by both SystemNative_IoRingKick and this submit-flush call) and skip the submit-only io_uring_enter entirely when the SQ tail hasn't moved since the last flush. Only ever read/written by the single dedicated issuer thread, same as SqTail itself, so no synchronization is needed. Measured with perf stat -e syscalls:sys_enter_io_uring_enter under the TechEmpower JSON benchmark at wrk concurrency 32: syscalls/request dropped from ~1.27 to ~0.82 (-36%), with throughput up +11.5% (211,837 -> 236,119 req/s). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
This reverts commit 8ceda56. The spin loop added complexity to the issuer thread's wait logic without clear enough evidence of benefit under the TechEmpower-faithful benchmark methodology (primer/warmup/measured-run with correct concurrency levels). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Previously DEFER_TASKRUN was unconditionally requested together with SINGLE_ISSUER. Profiling showed it costing throughput at higher core counts (the issuer thread's io_uring_enter(GETEVENTS) call spends significant time synchronously pumping deferred task-work via io_run_local_work), while still being a net win at low core counts. - pal_io.c/pal_io.h: SystemNative_IoRingCreate now takes a separate deferTaskRun parameter (only honored when singleIssuer is also requested, per kernel requirement). SystemNative_IoRingWaitForCompletions branches per-ring on whether DEFER_TASKRUN was actually enabled: the two-call submit/wait split is only used when required (DEFER_TASKRUN), otherwise submission and completion-waiting are combined into a single io_uring_enter call. - Interop.IoRing.cs: mirrors the new native parameter. - PortableThreadPool.IoUring.Unix.cs: adds GetDeferTaskRunConfig(), backed by DOTNET_IORING_SETUP_DEFER_TASKRUN (default: enabled when Environment.ProcessorCount <= 6, disabled otherwise; caller override always wins). The config value is computed on the static constructor's own thread and captured into the issuer-thread lambda - calling it directly from the issuer thread would deadlock, since it's a static method on the type whose static constructor is still running on the other thread. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
TrySubmitRecv/TrySubmitSend are only called after an optimistic userspace recv(2)/send(2) already returned EWOULDBLOCK, so having io_uring immediately retry the syscall (its default behavior) is redundant work. IORING_RECVSEND_POLL_FIRST (sqe->ioprio) tells the kernel to instead arm poll and wait for readiness before attempting the syscall, matching what's already known to be true at submission time. /json benchmark, wrk -t 8 -d 15s, DOTNET_EnableEventPipe=0, 12-core box: | Concurrency | epoll | io_uring (default, before) | io_uring (default, +POLL_FIRST) | io_uring (DT=1, before) | io_uring (DT=1, +POLL_FIRST) | |---|---|---|---|---|---| | 32 | 261,480 | 208,938 | 246,234 (+17.9%) | 219,459 | 278,252 (+26.8%) | | 64 | 301,823 | 260,668 | 298,664 (+14.6%) | 278,709 | 330,638 (+18.6%) | | 128 | 322,847 | 294,401 | 322,527 (+9.6%) | 324,973 | 355,910 (+9.5%) | | 256 | 330,672 | 311,758 | 324,468 (+4.1%) | 347,894 | 366,050 (+5.2%) | Both io_uring configs now beat or nearly match epoll at every concurrency level tested; DEFER_TASKRUN=1 beats epoll outright at every concurrency level. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Without COOP_TASKRUN, the kernel may send an inter-processor interrupt to force our dedicated issuer thread to stop whatever it's doing in userspace and immediately run deferred task-work (e.g. posting a completion) as soon as it becomes available. Since this thread already deterministically transitions into the kernel on its own on every wakeup (SystemNative_IoRingWaitForCompletions is its entire loop body), that forced preemption buys nothing here and the IPI itself is pure overhead. Requesting COOP_TASKRUN lets the kernel skip sending it and instead let task-work accumulate until the next transition. TASKRUN_FLAG (only meaningful together with COOP_TASKRUN) additionally exposes an IORING_SQ_TASKRUN bit in the SQ ring flags - not consumed by this PAL today, but harmless to request and keeps the door open. /json benchmark, wrk -t 8 -d 15s, DOTNET_EnableEventPipe=0, 12-core box: | Concurrency | epoll | default (before) | default (+COOP_TASKRUN) | DT=1 (before) | DT=1 (+COOP_TASKRUN) | |---|---|---|---|---|---| | 32 | 261,480 | 246,234 | 261,246 (+6.1%) | 278,252 | 273,385 (-1.7%) | | 64 | 301,823 | 298,664 | 320,899 (+7.4%) | 330,638 | 333,413 (+0.8%) | | 128 | 322,847 | 322,527 | 340,215 (+5.5%) | 355,910 | 358,294 (+0.7%) | | 256 | 330,672 | 324,468 | 353,036 (+8.7%) | 366,050 | 366,068 (~0%) | Default config (DEFER_TASKRUN off) now beats epoll at every concurrency level tested, up from only nearly matching it before this change. DEFER_TASKRUN=1 is roughly flat, as expected since it already avoids the IPI-driven preemption COOP_TASKRUN targets. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Replaces the one-shot-per-call Accept submission with a single, persistent IORING_ACCEPT_MULTISHOT submission per listening socket, kept alive across every accepted connection instead of resubmitting for each one. Because a multishot accept's peer-address buffer would be reused/overwritten across every connection it produces (a real correctness hazard given this codebase's batched completion-draining design), the SQE is submitted without an address buffer at all; getpeername(2) is called on the accepted fd instead once a connection is actually handed to a caller. Native/interop: - IoRingRequest gains a Multishot field (pal_io.h/.c, Interop.IoRing.cs); IoRingFillSqe sets IORING_ACCEPT_MULTISHOT on the SQE when requested. - Exposed IORING_CQE_F_MORE as Interop.Sys.IoRingCqeFlagMore. Thread pool: - IIoUringOperation.CompleteFromIoUring now takes the raw CQE flags and returns operationCompleted via an out parameter, since a multishot operation can receive many completions for the same UserData/GCHandle over its lifetime - DispatchBatch only frees the handle and decrements the in-flight count once operationCompleted is true. All existing (one-shot) implementers now set this to true unconditionally. Public API: - Added System.Threading.IoUring.TrySubmitAcceptMultishot, backed by a new MultishotAcceptOperation that resubmits transparently when the kernel benignly stops generating completions for a submission (F_MORE unset, non-negative result), and surfaces a terminal error otherwise. Sockets: - SocketAsyncContext.IoUring.Unix.cs's TryAcceptViaIoUring now starts (once, lazily) a persistent multishot accept submission per listening socket, matching buffered completions against waiting AcceptAsync callers via two queues guarded by a per-socket lock. Benchmarked (TechEmpower /json, wrk -t 8 -d 15s, DOTNET_EnableEventPipe=0, 12-core box), before vs. after this change: Concurrency | Before (default) | After (default) | Before (DEFER_TASKRUN=1) | After (DEFER_TASKRUN=1) 32 | 261,246 | 267,617 | 273,385 | 283,355 64 | 320,899 | 325,492 | 333,413 | 343,494 128 | 340,215 | 344,143 | 358,294 | 362,516 256 | 353,036 | 356,479 | 366,068 | 383,901 Consistent, modest gains (+1% to +5%) across both configurations and every concurrency level tested; no functional regressions observed (clean startup/ shutdown, no EEXIST/exceptions in logs at up to c=256). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
This reverts commit 3cb663b.
This reverts commit 0984794.
…ing (port of dotnet#35330) Ported the epoll fix from dotnet#35330 to the single-issuer io_uring completion hand-off path: instead of the issuer thread resolving every completion in a drained batch itself and moving the whole batch to the Thread Pool via one UnsafeQueueUserWorkItems call (DispatchBatch), it now just copies the raw completions into a ConcurrentQueue and schedules (CAS-gated, at most one at a time) a single CompletionProcessor work item. That work item resolves and runs one completion's continuation inline, reschedules another CompletionProcessor before doing so (to keep growing parallelism on demand, matching SocketAsyncEngine.ScheduleToProcessEvents), and keeps draining the queue itself for up to a 15ms time slice before yielding the thread back to the pool - identical structure/reasoning to SocketAsyncEngine's _eventQueueProcessingRequested/IThreadPoolWorkItem.Execute in dotnet#35330. The older DispatchBatch path is kept (not deleted) and remains reachable via DOTNET_IORING_PARALLELIZED_ENQUEUE=0, so the two can still be A/B compared; the new parallelized-enqueue path is the default. Local single-machine loopback TechEmpower /json benchmark (wrk -t 8 -d 15s, DOTNET_EnableEventPipe=0, 2 runs averaged per config) showed the two paths within noise of each other (~350-365k RPS range on this machine, well below the throughput ceiling where the external dual-machine infra shows the epoll/io_uring gap): | Concurrency | Legacy batch (avg) | Parallelized (avg) | Δ | |---|---|---|---| | 32 | 268,246 | 266,844 | -0.5% | | 64 | 328,903 | 327,052 | -0.6% | | 128 | 351,083 | 350,816 | -0.1% | | 256 | 360,591 | 365,105 | +1.3% | This local loopback setup isn't a good proxy for the effect being tested (thread pool scheduling contention under real network load/higher core counts); needs validation on the external dual-machine infra where the epoll-vs-io_uring gap was originally observed.
Previously all io_uring submission/completion work funneled through one dedicated issuer thread and one ring for the whole process. Trace analysis of an external 8-core-vs-12-core benchmark comparison showed this is a hard, non-scaling bottleneck: the issuer thread's absolute busy time stayed roughly constant regardless of core count (it alone executes every socket's recv/poll syscall work inline as part of io_uring_enter), so throughput regressed as core count/ThreadPool size grew (620k RPS @ 8 cores vs 560k RPS @ 12 cores on the same machine). This replaces the single global ring with a configurable number of independent single-issuer rings (Ring), each with its own ring handle, wake-eventfd, pending-submissions queue, completion queue, and dedicated issuer thread: - New PortableThreadPool.IoUringThreadPool.Ring nested class holds all the per-ring state that used to live in static fields. - GetRingCount() defaults to one ring per 8 cores (max(1, Environment.ProcessorCount / 8)), overridable via DOTNET_IORING_THREAD_COUNT (or the System.Threading.ThreadPool.IoUringThreadCount AppContext switch). - The static constructor eagerly spawns one dedicated issuer thread per ring at startup (same ring-creation handshake as before, repeated N times) - no lazy ring creation, per this being an experiment. - A [ThreadStatic] field gives every calling thread a sticky round-robin assignment to exactly one ring on its first TrySubmit call, so a given thread's requests/completions always stay on the same ring/issuer thread instead of being scattered across all of them. - IssuerLoop, DrainAndSubmit, DrainCompletions, EnqueueCompletions, CompleteOperation, DispatchBatch, and the completion-processing work item are all now parameterized on (or scoped to) a specific Ring. TrySubmit's public contract/signature is unchanged; no other files needed changes. Validated: - clr.corelib+clr.nativecorelib build: 0 errors/warnings. - Smoke-tested the TechEmpower app with the default ring count (1 ring on this 12-core machine) and with DOTNET_IORING_THREAD_COUNT=3 (confirmed via /proc/<pid>/task/*/comm that exactly 3 '.NET IoUring Issuer' threads are created), both serving /json correctly. - Ran a 15s wrk load test (c=64) against the 3-ring configuration with no errors/exceptions/EEXIST in the app log. - Ran the local /json benchmark matrix (c=32/64/128/256) for both configurations for a basic regression check: | Concurrency | 1 ring (default) | 3 rings (DOTNET_IORING_THREAD_COUNT=3) | |---|---|---| | 32 | 348,010 | 244,463 | | 64 | 425,837 | 313,560 | | 128 | 449,309 | 342,135 | | 256 | 419,369 | 378,376 | This local single-machine loopback setup (12 cores shared by client+server+this very CLI session) is not representative of the external multi-machine 8-vs-12-core scaling regression this change targets - it lacks the spare cores for extra issuer threads to pay off and cannot reproduce the effect either way. Real validation requires the external infra. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
An issuer and worker can complete a request before TrySubmit returns. Assigning pin/vector ownership and the SafeHandle reference flag afterward allows completion to observe incomplete state, leaking references and causing subsequent file opens to fail with sharing violations. Initialize that state before publication for reads, writes, scatter/gather operations, and partial gather resubmissions. Clear it when publication fails so local cleanup does not leave stale ownership in the reusable value-task source. Correct the associated completion/lifetime comments. The reproduced sharing failures disappeared with this ordering; all 28 targeted FileStream conformance cases passed in final-series validation. Unrelated timestamp failures also reproduce on the baseline. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Passing ring capacity rather than the pending SQE count can make io_uring_enter return before GETEVENTS processing. This explains the previous apparent need to separate submission from deferred task-work. Compute pending submissions from SQ tail/head instead of a flushed-tail marker. Combine exact-count submission with GETEVENTS and preserve unconsumed entries after short submissions. Pump deferred work after short submissions or submission-side EBUSY/EAGAIN; return retryable EAGAIN when pending entries cannot make progress. On the managed side, reap completions to free a full SQ instead of spinning. Retry transient submission failure with bounded backoff without parking on eventfd or relinquishing request ownership. Surface permanent publication/wait errors rather than silently abandoning requests. Update the single-issuer and completion-dispatch contracts. Add Linux regression coverage for pending accept, bursts beyond ring capacity, blocked callbacks, callback reentrancy, and Unpin-triggered receives. These also protect the subsequent state-reuse changes. The exact-count ACCEPT reproducer was confirmed on local and external kernels. Final-series Checked validation passed all ten focused cases in both dispatch modes with injected EAGAIN; the pre-fix snapshot aborted. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Repeated nonblocking CQ drains otherwise enter the kernel for every completion chunk and the final empty probe. Map SQ flags and remember whether TASKRUN_FLAG was enabled. Skip a nonblocking enter only when there are no pending submissions, deferred task-work notifications, or CQ-overflow notifications. Continue entering for blocking waits and conservatively for rings/build headers without the notification support. This preserves DEFER_TASKRUN progress while allowing already-visible CQEs to be consumed without another syscall. Document the conditional entry contract at the PAL boundary. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Draining either managed queue until empty lets sustained traffic delay the other side of a ring. Publish at most one existing 256-request batch per turn and reap at most 256 completions in 64-entry chunks. Report an exhausted completion budget as more work so the issuer returns to submissions instead of parking. Do not wait to fill a batch. Keep ring ownership and the self-replicating worker completion scheduler unchanged; the pooled-tree scheduling experiment did not improve the external JSON workload. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
The issuer's sleep decision needs to know whether kernel completions are outstanding, not whether worker callbacks have finished. Count requests when the issuer dequeues a batch, before any SQE can be published, and subtract CQEs when they are reaped. Requests still in the managed queue remain covered by the existing queue rechecks. This removes producer/worker atomic updates and their shared cache-line traffic from every operation. Correlation tokens, pins, and SafeHandle references retain their existing worker-side completion lifetimes. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Allocating an ActionIoUringOperation for every pending socket operation adds avoidable steady-state allocation to the shared io_uring bridge. Cache one adapter per thread and clear handle/delegate references on return. Copy completion state to locals and return the adapter before releasing the handle or invoking user code so reentrant submission can safely reuse it. Use finally-based cleanup when submission fails. Keep callbacks on ThreadPool workers, including when legacy dispatch resolves the raw completion on the issuer. Update the prototype API documentation to describe sharded single-issuer rings accurately. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
The pending receive path allocated a closure and delegate for each operation even when repeatedly reading from the same socket. Cache one receive-state object per SocketAsyncContext, including its bound completion delegate. Concurrent receives can allocate independent state instead of sharing an active pin or callback. Capture completion state in locals, clear references, and return the object before disposing the pin or invoking the callback. Both callback reentrancy and a custom MemoryManager.Unpin starting another receive must be safe. Unpin on failed submission, including exception paths. The earlier regression cases exercise both forms of reentrancy and exactly-once pin release. Combined with callback-adapter reuse, this removes nearly all measured managed allocation from pending receives. Cancellation/close integration remains a pre-existing prototype gap. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
After managed wrapper reuse, allocating and freeing a normal GCHandle for every operation still adds correlation bookkeeping overhead. Root operations in a bounded 1024-slot array per ring. Encode the slot index and generation in tagged user data; normal GCHandles have the tag bit clear and remain the fallback when all slots are occupied. Validate the generation on completion and return the slot before invoking operation completion, preserving reentrant submission. Clear strong references on release and roll back either token form if the managed submission queue throws before publishing the request. The existing burst tests exceed slot capacity and the repeated receive tests exercise generation reuse. Final-series validation includes forced compacting collections and Checked EAGAIN injection in both dispatch modes. External JSON results for the complete series, not this change alone: | Configuration | Mean RPS | Mean p99 ms | Mean peak CPU | | --- | ---: | ---: | ---: | | Original io_uring, 7 rings | 1,946,344 | 1.273 | 84.7% | | Final io_uring, 8 rings | 2,100,679 | 0.980 | 86.7% | | Epoll on final binaries | 2,247,628 | 0.685 | 86.7% | Each row uses three load passes sharing one application process. The combined gain includes ring-count tuning; the generic default is unchanged. Epoll parity is not achieved. Existing enabled-suite cancellation and execution-context failures/timeouts also remain; no green full enabled socket suite is claimed. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
The completion processor runs multiple callbacks inside one ThreadPool work item, so normal dispatch cleanup happens only after the entire batch. An AsyncLocal value left by one callback can consequently be observed by the next callback in that batch. Reset ExecutionContext and SynchronizationContext after each completion using the existing ThreadPool helper, then reset mutable worker-thread state after any AsyncLocal change notifications have executed. Preserve the existing queues, batching, and scheduling policy. Add four single-worker receive-burst cases covering one and three rings, with and without AsyncLocal change notifications. Verify callback context, synchronization context, thread name, and final cleanup. All four cases fail on the preserved baseline and pass with the fix in Release and Checked, including both dispatch modes with injected EAGAIN. ThreadPool, TCP transfer, file-sharing, and compacting-GC checks also passed. External throughput impact is not yet measured; the comparison was interrupted and benchmark-network access is currently unavailable. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
The completion processor executes multiple I/O callbacks inside one ThreadPool work item. Counting only the outer work item makes hill climbing measure batches rather than the completion work performed within them. Report progress through the existing ThreadPool helper when continuing to another completion. Leave the last completion to the outer dispatch loop so it is not counted twice. Keep the existing queues, scheduling, issuer ownership, and worker-minimum defaults unchanged. Extend the single-worker receive-burst regression to verify that completed work accounting includes the callbacks. All four cases reported only four completed work items for 128 callbacks before this change and pass with it. Final focused coverage passes 14 cases in Release and Checked, including both dispatch modes with injected EAGAIN. ThreadPool (68 passed, one skipped), TCP transfers (36), and file sharing (28) also pass. External TechEmpower /json, 56 logical CPUs, eight single-issuer rings, 512 connections, 32 load threads, 90-second warmup and 60-second measurement. Each row averages two independent application launches; the confirmation pair reverses candidate/control order. Both use the existing ThreadPool minimum-worker setting at 16 (DOTNET_ThreadPool_ForceMinWorkerThreads=0x10). | Variant | Mean RPS | Mean p99 ms | Mean reported peak CPU | | --- | ---: | ---: | ---: | | Context-isolation fix only | 2,223,767 | 0.611 | 92.0% | | Also report completion progress | 2,255,593 | 0.540 | 90.5% | This is a 1.43% throughput improvement at matched tuned settings. No meaningful throughput improvement was established at the default worker minimum. Lowering that minimum contributes separately to the overall gain; the machine-specific tuning is not hardcoded by this commit. Short-warmup results ramp across passes, so they are not treated as steady state. All accepted measurements have zero response/socket errors and verified deployed runtime hashes. The previously recorded approximately 2.25M RPS epoll result is reused as the goal, not claimed as a new matched control. CPU parity with epoll is not established. Known full enabled socket-suite failures remain outside this performance change. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Add persistent fd-routed accepts with ordered callbacks, targeted cancellation, pressure-driven rearming and owned socket cleanup. Expose Socket.AcceptMultishotAsync with a portable fallback and validate cancellation, disposal and endpoint semantics. Preserve the experimental receive API placeholders for the next implementation step. Focused Release and Checked cases pass. Local accept measurements show lower allocations, but no established timing win; queued bursts can increase syscall counts because peer endpoint capture requires getpeername. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Add ring-owned provided buffers, persistent fd-routed receives, ordered callbacks, cancellation barriers, and pressure-aware pause/rearm handling. Expose Socket.ReceiveMultishotAsync with independent buffer ownership and a portable asynchronous fallback. Fix synchronous-prefix and partial-completion accounting in io_uring sends, required for reliable transport transfers. Cover buffer exhaustion, cancellation races, ownership, reset state, and large sends. Validated the exact source snapshot with 64 focused Release and Checked cases, the disabled socket suite, ThreadPool and RandomAccess guardrails, and forced provided-buffer fallback. Matched Kestrel JSON measurements establish throughput gains but retain a tail-latency regression; this remains an experimental API. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Reduce provided-buffer storage from 4 MiB to 2 MiB per ring. Adjust the existing exhaustion test to retain eight buffers on each of sixteen connections. No builds, benchmarks, or tests were run for this change, as requested. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
This reverts commit efee005. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
This reverts commit ea79f1d. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
This reverts commit 3dca870. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
No description provided.