Repository navigation
fix(rpc): bound every request with a wall-clock budget so slots always come back - #1208
Conversation
…s come back A client-side RPC timeout rejects locally without ever telling the host, so a request stuck waiting on the global state lock kept its in-flight slot after the caller had moved on. With MAX_IN_FLIGHT_RPC=32 and a ~130s client budget, enough queued requests burn through all slots in minutes and every session then fails with HOST_OVERLOADED until the app restarts (issue vastsa#1071). Give each request a wall-clock budget enforced inside serve(): - non-tool requests: 135s — just past the client budget, so the slot is guaranteed to come back within seconds of the client giving up; - tools.execute: the tool's own effective timeout plus a 90s grace window (Bash may legitimately run for hours; a tool with no effective timeout keeps the old unbounded behavior); - a budget expiry answers -32030 HOST_RPC_TIMEOUT and releases the slot. Known boundary: tokio timeouts only fire at await points, so a handler stuck inside one long synchronous call is not interrupted — that class needs the sync work moved off the async workers. What this guarantees is that the dominant queueing case (waiting on the global lock) is cut loose independently per request, degrading a full 32-slot outage to at most one slot held by the wedged handler itself. Deep-dive tests: concurrent lock waiters time out independently and leave the lock acquirable; a fresh request succeeds right after the holder releases; the tools.execute budget matrix covers Bash's 6h ceiling and unbounded tools.
|
Review result: do not merge yet. Issue #1071 is real, but the budget is not applied to the actual The new matrix test passes because its fixtures use the wrong snake_case keys and does not cover the wire shape. The non-tool 135s budget is useful, but this is an incomplete root fix for #1071. Please read the camelCase fields (or deserialize the request before calculating the budget) and add a regression test using The PR is also behind the current |
Business-logic validation (real host-core process, added post-open)Drove the actual patched
The core case is the business guarantee itself: a 137s task runs past the 135s mark where a uniform budget would have killed it with Desktop smoke with Script: drives handshake → |
…budgets ToolsExecuteParams is serde-renamed to camelCase, so the real wire carries toolName / timeoutMs. request_budget_ms read the snake_case spellings and saw an empty tool name with no timeout on every real tools.execute request, leaving them unbounded — the stalled-Glob path from vastsa#1071 stayed open. Read toolName / timeoutMs first with the snake_case spelling as a fallback, and pin the wire shape with a dedicated regression test (review on vastsa#1208).
|
Corrected — thank you for catching this. The review is exactly right, and it also exposed why my own validation missed it twice:
Also refreshed the base onto current |
The tools.execute handler may wait on the user's answer to an ask prompt before the tool runs, and that wait has no timeout of its own. Without headroom, a slow approval plus a full-length default Bash (~120s wait + 60s run) exceeded the tool-timeout-plus-grace budget (150s) and the request would die mid-execution with HOST_RPC_TIMEOUT. Add a 120s permission-wait cap to the tools.execute budget (matching the unattended ask budget order), so the budget is permission cap + tool timeout + grace. Verified against the review matrix and the business check on the real binary; long-tool semantics unchanged.
|
One more hardening pass on my own diff while awaiting re-review — a scenario the current budget would have gotten wrong: Slow approval + full-length tool. The Head |
execute_tool_with_path_access called the synchronous Read, Glob, Grep, Write, and Edit bodies inline on the async Tokio worker that handled the request. The host runs the default multi-thread runtime (one worker per core), the ignore walker never yields, and the rg fast path blocks until its child exits. A few large Globs therefore occupied every worker, and cheap RPCs such as session.list waited seconds behind them; the tools.execute budget from vastsa#1208 cannot help because a timer only fires at an await point. Run those five bodies through tokio::task::spawn_blocking with owned inputs. HashlineStore is an Arc<Mutex<_>> handle, so moving a clone onto the blocking thread is cheap. Bash keeps its existing async path. Tool results, error codes, and the wire contract are unchanged; the existing read and mutation class limits also bound the blocking threads. A panicking body now returns INTERNAL instead of dropping the response. The test-only rg override is thread-local, so it is carried onto the blocking thread for the duration of the body. Refs vastsa#1071
execute_tool_with_path_access called the synchronous Read, Glob, Grep, Write, and Edit bodies inline on the async Tokio worker that handled the request. The host runs the default multi-thread runtime (one worker per core), the ignore walker never yields, and the rg fast path blocks until its child exits. A few large Globs therefore occupied every worker, and cheap RPCs such as session.list waited seconds behind them; the tools.execute budget from #1208 cannot help because a timer only fires at an await point. Run those five bodies through tokio::task::spawn_blocking with owned inputs. HashlineStore is an Arc<Mutex<_>> handle, so moving a clone onto the blocking thread is cheap. Bash keeps its existing async path. Tool results, error codes, and the wire contract are unchanged; the existing read and mutation class limits also bound the blocking threads. A panicking body now returns INTERNAL instead of dropping the response. The test-only rg override is thread-local, so it is carried onto the blocking thread for the duration of the body. Refs #1071
Root cause (issue #1071)
The JS client rejects a host RPC locally after ~130s without ever telling the host, and an in-flight slot is only released when the handler task actually finishes. When one handler wedges the global state lock, every other request queues behind it, times out client-side, and keeps its slot anyway. With
MAX_IN_FLIGHT_RPC = 32, enough queued requests burn through all slots in minutes — and every session then fails withHOST_OVERLOADED/ "API key auth failed" until the app restarts.What changed
serve()now wraps each request task in a wall-clock budget (with_request_budget):tools.executemethodtools.executeeffective_timeout_ms+ 90s graceA budget expiry answers
-32030 HOST_RPC_TIMEOUT, logswarn!(budget_ms), and the slot is released with the task.Known boundary, stated up front: tokio timeouts only fire at await points, so a handler stuck inside one long synchronous call is not interrupted — that class needs the sync work moved off the async workers (follow-up). What this guarantees is that the dominant failure chain in #1071 — requests queueing on the global lock — is cut loose independently per request: a wedged handler can hold at most its own slot instead of starving all 32.
Validation
HOST_RPC_TIMEOUT, and the lock is acquirable immediately after the holder releases;tools.executeBash 60s → 60s+90s grace; Bash 6h ceiling → 6h+90s (never clipped to the fixed budget); a tool without an effective timeout stays unbounded;cargo test -p host-core: 663 passed, 0 failed (658 prior + 5 new/extended).cargo fmt --check: clean.Relationship to #1071
This is the layer-A stopgap from the layered plan proposed in #1071: it removes the "must restart the app" outcome on its own. The deeper items (sync tools onto
spawn_blocking, lock discipline for DB access, slow-RPC observability) remain follow-ups and can land independently.