Skip to content

gateway: run uvicorn on the stdlib loop by default (gateway.loop) - #488

Merged
alex-clickhouse merged 1 commit into
mainfrom
pufit/gateway-asyncio-loop
Oct 6, 2026
Merged

alex-clickhouse merged 1 commit into
mainfrom
pufit/gateway-asyncio-loop

Conversation

@pufit

@pufit pufit commented Oct 5, 2026

Copy link
Copy Markdown
Member

Problem

run_server() calls uvicorn.run() without a loop= argument, so uvicorn's default auto selects uvloop — always, because uvicorn[standard] installs it. uvloop spawns subprocesses with a full fork(), which copies the parent's page tables. Nerve's gateway spawns constantly (every agent CLI session, every gh api call of the GitHub source enrichment, every cron-gate plugin), and once memU is loaded the process is large, so every spawn blocks the event loop and the gateway's own HTTP/WebSocket traffic stalls with it.

Measured on a production deployment (~115 GB RSS, 59 worker slots) over the 10 days since its last restart:

  • main thread spent 28.4 h in the kernel = 11.9 % of wall time — in fork() page-table copies (py-spy: MainThread → create_subprocess_exec ← sources/github.py:_gh_api_get; strace on the same box earlier: ~2.5–2.9 s of 100 % system time per spawn)
  • ~29.5 K APScheduler "run time … was missed" warnings in 32 h; the 20 s wakeup job missed by p50 4 s / p90 12 s / max 26 s
  • web UI / 9.9 s and /api/cron/jobs 8.3 s during spawn bursts vs 0.06 s idle — the panel looks dead while the fleet is fine
  • the same interpreter on the stdlib loop spawns via vfork() (strace-verified) — milliseconds, independent of process size

uvloop's faster I/O buys nothing at the gateway's request rates; the fork cost is what dominates.

Change

  • gateway.loop (new, asyncio | uvloop | auto, default asyncio) — handed to uvicorn.run(loop=…) explicitly, never left to auto. Validated at load (ValueError naming the key), env-reference friendly (${VAR:-default}), blank = default.
  • nerve reload reports a changed gateway.loop as restart-required (_RESTART_ONLY_PATHS + the docs table it is checked against).
  • _env.py docstring: the OpenBLAS atfork collision only exists on the fork() path; the cap stays as defense in depth for the uvloop opt-in.
  • proxy/service.py comments: the kwargs are avoided for uvloop compatibility, which is now an opt-in rather than the daemon's fixed loop.
  • docs (docs/config.md layer table + restart table) and the settings.yaml template.

Tests

tests/test_gateway_loop.py (14 tests): default is the stdlib loop; every uvicorn value accepted; case/whitespace normalised; blank = default; unknown value refused by name; loads from config.yaml and via ${VAR:-default}; run_server passes the configured loop to uvicorn.run verbatim and never falls through to auto; restart_required reports a change and stays quiet otherwise. Full suite: all green locally.

Rollout

Startup-only setting: a running gateway picks it up on its next restart. Deployments that want uvloop back set gateway.loop: uvloop in settings.yaml.

Generated by Nerve

@pufit

pufit commented Oct 6, 2026

Copy link
Copy Markdown
Member Author

Follow-up measurement on the production deployment this PR was written for (~115 GB RSS, 59 worker slots, uvloop), to separate "fork cost" from "blocking sync calls in coroutines" — the two look identical from outside.

Method. 5-minute windows with an external /health probe at 4 Hz, strace -tt -T -e trace=clone,clone3,vfork,fork on the loop thread only, py-spy dump of the loop thread at 2 Hz, per-thread utime/stime/wchan at 2–5 Hz, and py-spy record --gil (samples only the thread holding the GIL). Plus an AST scan of every async def for direct blocking calls (subprocess.run, time.sleep, sync sqlite/HTTP, file I/O) not wrapped in to_thread/run_in_executor.

Results.

  • Static scan: no blocking sync calls on the loop in hot paths (the only two hits are in the Telegram channel's /restart and zip-extract handlers; cron gate plugins do millisecond sqlite reads).
  • clone() on the loop thread measured at 3.2 s each (3 spawns = 9.4 s in a 5-min window; the GitHub source's notification enrichment then chains dozens of gh api spawns → a 7m48s window with /health at 20–57 s).
  • py-spy --gil during such a window: 95.6–99.9 % of all GIL-held samples are the loop thread inside create_subprocess_exec — uvloop holds the GIL for the entire fork+exec handshake, so every other Python thread stops too; the memU thread was observed parked on futex_do_wait (GIL) and lock_mm_and_find_vma (the mmap lock fork() holds while copying page tables).
  • One 4.3-min window: 192 s unresponsive (74 %), of which 96 % was spawning (GitHub source + a cron gate's gh calls + agent CLI starts) and 3.4 % a gc.collect() from the old per-file _release_memory (one 6.5 s stop-the-world pass over the 2M-item memU heap — already replaced on main by _trim_memory/gc.collect(1) in cfa5a95; this deployment predates it).
  • Control: the stdlib loop in the same venv spawns via vfork(); numpy's full-matrix matvec/argpartition in the memU index release the GIL (max GIL gap 1.3–4.9 ms in a controlled run), so they are not a loop blocker.

So the mechanism is a blocking synchronous call inside a coroutine — the spawn itself, in uvloop's subprocess implementation — and loop="asyncio" is what removes it.

Generated by Nerve

@alex-clickhouse alex-clickhouse left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@alex-clickhouse
alex-clickhouse merged commit 9d88ca1 into main Oct 6, 2026
3 checks passed
@alex-clickhouse
alex-clickhouse deleted the pufit/gateway-asyncio-loop branch October 6, 2026 10:52
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants