Skip to content

[sprites 1-3] onConnect: in-memory reattach before sprite resolution - #2005

Merged
2witstudios merged 9 commits into
masterfrom
pu/sprites-1-3-onconnect-reorder
Jul 11, 2026
Merged

2witstudios merged 9 commits into
masterfrom
pu/sprites-1-3-onconnect-reorder

Conversation

@2witstudios

@2witstudios 2witstudios commented Jul 11, 2026 •

Copy link
Copy Markdown
Owner

Why

Tab-back to an open terminal was slow for a reason that had nothing to do with the platform. onConnect ran the entire fused checkAuth chain — resolve the agent-terminal row → getSprite → write a code-execution audit row → reserve a concurrency slot — and only then looked in sessionMap to discover that a live in-memory session was sitting right there, ready to reattach.

docs.sprites.dev/concepts/lifecycle puts a warm wake at 100–500ms and a cold one at 1–2s. That is the platform floor; everything we stacked above it on a tab-back was our own orchestration overhead.

What changed

checkAuth no longer resolves the Sprite (or reserves a slot) eagerly. It returns the cheap, DB-only access verdict (sessionKey, payerId) plus an uncalled resolveSandbox() thunk. onConnect runs the access check, looks up the live session (the key derives without a Sprite, per 1-1), and a pure planConnect({ accessResult, existingSession }) returns a discriminated deny | reattach | create plan carrying the narrowed payload each branch needs. Only the cold path invokes the thunk.

The audit row moved into resolveSandbox with it — it records a PTY actually being launched, and a reattach launches nothing.

  • apps/realtime/src/terminal/agent-terminal-handler.ts — planConnect (pure), split AgentTerminalCheckAuthResult / AgentTerminalSandboxResult, reordered onConnect.
  • apps/realtime/src/terminal/agent-terminal-access.ts — buildAgentTerminalCheckAuth returns the lazy thunk, which owns slot reservation + release.

AgentTerminalCheckAuthDeps is unchanged, so the index.ts wiring needed no edits.

🐛 A severe pre-existing bug this surfaced (please read)

Reviewing the reorder turned up a bug that the reorder would otherwise have shipped on top of. acquireCodeExecutionSlot (packages/lib/src/services/sandbox/quota.ts) is a per-user counter — free: 1, pro: 2, founder: 3, business: 5 — and checkAuth reserved a slot on every call, including calls that start no PTY:

  1. Reattach was impossible at the limit. A free-tier user with one live session already holds the only slot. The tab-back's access check tried to acquire a second, failed, and was denied concurrency_limit. The fast path this PR exists to build was unreachable for the most common tier.
  2. Live sessions were killed after 60s. The re-auth tick also calls checkAuth. At the limit it too failed to acquire a slot — and the tick reads any denial as a revoked authorization, so it called teardownAgentTerminalSession. A free-tier PTY died ~60 seconds after opening.

Both are pre-existing on master; they are masked in production only because CODE_EXECUTION_ENABLED is off. My handler tests missed them because they mock checkAuth — the bug lives in the real composition.

The fix falls out of the same lazy split: a slot bounds how many PTYs a user has running, so it is reserved inside resolveSandbox (the only path that starts one), and releaseSlot is surfaced on the sandbox success result — there is no slot to release unless one was actually reserved. Reattach takes none; re-auth reserves nothing and hands nothing back. Regression tests are at the access layer where the real slot logic lives.

🐛 Second bug: concurrent cold connects (double-mount)

The leaf requires that a double-mount not double-create. The sequential remount was already safe (connect #2 finds the session via getByKey). The genuinely concurrent race was not — and this PR widens its window, because the cold path is exactly where a connect now spends seconds inside resolveSandbox.

My first fix was to let both open a PTY and then discard the loser. An adversarial review pass found that this was destructive, and I replaced it. Recording it here because the reason is subtle and worth a reviewer's attention:

openPtyShell with a non-null sessionId calls sprite.attachSession(id) — it connects to a server-side exec session — and PtyShell.kill() is SIGKILL on the process, not a local socket close. So whenever a streamSessionId is persisted (any warm reattach after a realtime restart), both racers attach to the same remote session, and killing the duplicate SIGKILLs the very process the survivor is attached to. It turned a recoverable orphan into a guaranteed terminal kill, on exactly the path it was meant to protect.

The fix in place: serialize at the key, so a second PTY is never opened at all. A cold connect claims the key before its first await; a concurrent connect for that key awaits the in-flight create and joins its session.

  • TerminalSessionMap gains trackCreate / pendingCreate.
  • The claim is released on every exit from the cold path — success, denial, or throw — so a failed create cannot wedge the key. The joiner loops, in case another create queued behind a failure.
  • No await sits between the join-loop's exit and the claim, so it is atomic w.r.t. the event loop and needs no lock.

This dissolves the class rather than patching it: no duplicate PTY → no kill; no double slot; no double hold; no ambiguous session-id discovery.

🐛 Fourth/fifth/sixth: found by the same adversarial pass

  • Concurrency slot leaked when billing.gate throws. The slot was reserved but no session owned it yet, so a rejection escaped onConnect with the slot still held. activeByUser is a process-lifetime counter — one transient billing/DB blip permanently locked a free-tier user (limit 1) out of terminals on that replica until restart. Now released before rethrowing.
  • discoverNewSessionId took the FIRST new tty session. One Sprite hosts every agent terminal on its machine, so a sibling terminal launching in the same window is indistinguishable from ours — persisting the sibling's id points this terminal's next cold connect at another terminal's PTY. Now abstains unless exactly one appeared (the rule sprites-shell.ts's newTtySessionId already applies).
  • Re-auth checked the CREATOR's userId forever. A session outlives its creator's connection, and any authorized user may reattach. A viewer whose access was revoked mid-session kept receiving output and could keep typing — because the long-gone creator still passed the check. Sessions now carry viewerUserId, set on create and on every reattach, and the tick re-checks whoever is actually driving the PTY.

🐛 Third bug: re-auth stopped noticing a deleted scope (review catch — codex P2)

Making the sandbox resolution lazy also made the (scope, name) existence check lazy — because the fused checkAuth only ever learned "this terminal still exists" as a side effect of resolving its sandbox. The 60s re-auth tick calls only the access half, so after the split it stopped noticing that a terminal's project, branch, or own row had been deleted: the orphaned PTY kept running against a scope that no longer existed.

Restoring the check naively would have re-broken the epic. resolveAgentTerminal fuses two very differently-priced questions:

  • cheap — does the row exist? (resolveScopeKey + store.findByName: a couple of indexed reads, no Sprite)
  • expensive — where does its Sprite live? (machineSandbox.acquire, which for machine/project scope can reconnect or resume a hibernated Sprite)

Calling it from the access half would wake the Sprite on every re-auth tick and every tab-back. So I split the question, not just the call site:

  • New resolveAgentTerminalRow (packages/lib/src/services/machines/agent-terminals.ts) — DB-only. Its tests pass machineSandbox: undefined throughout, which is the proof it cannot wake a Sprite: it is never handed one.
  • The access half runs it eagerly, after the read-only access gates — so a user without edit rights still learns nothing about which terminals exist.
  • The Sprite wake/read stays lazy, so reattach still performs zero Sprite SDK calls.

Requirements

Requirement How it is satisfied
Live detached session + allowed user → agent-terminal:ready (with scrollback) after the access check, ZERO sprite SDK calls Reattach returns before resolveSandbox() is ever called. Test spies on the thunk and every sprite method (listSessions/spawn/createSession/attachSession/updateNetworkPolicy) and asserts none fired — plus that no slot was taken.
No live session → cold path unchanged (resolve sandbox → billing gate → openShell) Create path calls the thunk, then gates, then opens. All pre-existing handler tests (launch command, cwd, session-id discovery/persistence, billing holds, heartbeat settle, re-auth teardown) pass untouched.
Denied user + live session → refuse, not reattach planConnect returns { kind: 'deny' } for a denied verdict regardless of existingSession (pure test). In the shell a denied verdict never reaches the lookup. Handler test: an unauthorized socket gets agent-terminal:error, never :ready, and the victim's session stays bound to its original socket.
Two connects for the same key (double-mount) → no double-create Regression test: the second connect finds the session via getByKey and reattaches — openShell called exactly once. setNew/getByKey semantics preserved, no new locking.

Tests

TDD: planConnect is pure and tested with no mocks (riteway assert({given, should, actual, expected})); the handler is tested with an injected spy sprite client and a real createTerminalSessionMap.

apps/realtime  (all)   19 files / 641 tests pass · branch coverage 98.02% (gate 98%)
packages/lib  (machines) 8 files / 165 tests pass

tsc --noEmit clean, eslint clean. No any.

Notes for reviewers

  • A reattach no longer writes a code-execution audit row. Intended ("no policy writes" in the leaf spec) — the row asserts a PTY was launched, and a reattach launches none. The cold path still audits (test-covered). Flag it if compliance expects a row per connect rather than per launch.
  • Side benefit of the split: the 60s re-auth tick no longer resolves or wakes the Sprite. It calls checkAuth and never touches the thunk.
  • Untouched, as scoped out: the cold-path wake hack (1-4) and frontend keep-alive (4-1).

🤖 Generated with Claude Code

https://claude.ai/code/session_01HnFmynYiN6bmrtibxdzLDJ

…g the Sprite

A tab-back inside the 30-min detached grace window had to wait out a full
sandbox resolution (resolve the agent-terminal row -> getSprite -> audit write)
before onConnect ever looked in sessionMap, even though the live in-memory
session it was about to reattach to made every bit of that work redundant. The
Sprites platform floor for a warm wake is 100-500ms (docs.sprites.dev); this was
seconds of our own orchestration stacked on top of it.

Now that the access decision is DB-only (1-2) and the session key derives from
the (scope, name) target without a Sprite (1-1), checkAuth hands the sprite half
back as an UNCALLED `resolveSandbox` thunk. onConnect runs the cheap access
check, looks the session up, and a new pure `planConnect({accessAllowed,
existingSession})` decides reattach / create / deny. The reattach path returns
without ever invoking the thunk: zero sprite SDK calls, no wake exec, no audit
row. The cold path calls it and behaves exactly as before.

A denied verdict never reaches the session lookup, so losing access can never be
shortcut by a still-live PTY.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HnFmynYiN6bmrtibxdzLDJ
@coderabbitai

coderabbitai Bot commented Jul 11, 2026 •

Copy link
Copy Markdown
Contributor

Warning

Review limit reached

@2witstudios, you've reached your PR review limit, so we couldn't start this review.

Next review available in: 57 minutes

Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available.
You're only billed for reviews past your plan's rate limits ($0.25/file).

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro

Run ID: 063f0c83-3f72-498d-a03f-645dce7ff1dd

📥 Commits

Reviewing files that changed from the base of the PR and between 9138066 and c50c36f.

📒 Files selected for processing (9)
  • apps/realtime/src/index.ts
  • apps/realtime/src/terminal/__tests__/agent-terminal-access.test.ts
  • apps/realtime/src/terminal/__tests__/agent-terminal-handler.test.ts
  • apps/realtime/src/terminal/__tests__/terminal-session-map.test.ts
  • apps/realtime/src/terminal/agent-terminal-access.ts
  • apps/realtime/src/terminal/agent-terminal-handler.ts
  • apps/realtime/src/terminal/terminal-session-map.ts
  • packages/lib/src/services/machines/__tests__/agent-terminals.test.ts
  • packages/lib/src/services/machines/agent-terminals.ts
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch pu/sprites-1-3-onconnect-reorder

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 4f2163fc8e

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread apps/realtime/src/terminal/agent-terminal-handler.ts Outdated
Comment thread apps/realtime/src/terminal/agent-terminal-access.ts
2witstudios and others added 7 commits July 11, 2026 07:47
…he access check

Reviewing the reorder surfaced a severe bug it would otherwise have shipped on
top of. `acquireCodeExecutionSlot` is a per-user counter (free: 1, pro: 2,
founder: 3, business: 5), and checkAuth reserved a slot on EVERY call — including
calls that start no PTY at all:

  - Reattach: a free-tier user with one live session already holds the only slot,
    so the tab-back's access check could not acquire a second and was denied
    `concurrency_limit`. The fast path this PR exists to build was unreachable for
    the most common tier.
  - Re-auth: the 60s tick also calls checkAuth. At the limit it too failed to
    acquire, and the tick reads a denial as a REVOKED authorization — so it tore
    the session down. A free-tier PTY was killed ~60s after it opened.

Both are pre-existing on master; they are masked only because
CODE_EXECUTION_ENABLED is off. Neither path starts a PTY, so neither needs a slot.

The lazy split makes the fix natural: move acquireSlot into `resolveSandbox` (the
only path that starts a PTY) and surface `releaseSlot` on the sandbox SUCCESS
result, so there is no slot to release unless one was actually reserved. The
reattach path now takes no slot and releases none; the re-auth tick reserves
nothing and hands nothing back.

Also makes planConnect's decision load-bearing: it returns a discriminated union
carrying the narrowed access/session payload, so the handler executes the plan
instead of re-deriving the deny via a separate `!ok` check (the previous 'deny'
variant was never read in production).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HnFmynYiN6bmrtibxdzLDJ
…e key

The leaf requires that a double-mount not double-create. The sequential remount
was already safe (the second connect finds the session via getByKey), but the
genuinely CONCURRENT race was not — and this PR widens its window, because the
cold path is exactly where a connect now spends seconds inside resolveSandbox.

Two connects for one key both see an empty sessionMap, both openShell, and
`setNew` silently overwrites: the winner's PTY is orphaned where nothing can
reach it to kill it, and its concurrency slot is stranded for the life of the
process.

- Before setNew, re-check the key. If another cold connect claimed it while we
  were awaiting, discard OUR duplicate (kill the PTY, release the slot and hold)
  and join the winner instead — indistinguishable from a reattach client-side.
- Guard endAgentTerminalSession's deleteByKey with an identity check, so the
  discarded PTY's late onExit cannot evict the live winner from the map.
- Collapse the connect's slot release behind an idempotent wrapper: the slot is a
  bare counter, so a double release silently hands back capacity the connect
  never held, letting a user exceed their tier.
- Extract attachToLiveSession, shared by the tab-back fast path and the race loser.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HnFmynYiN6bmrtibxdzLDJ
…lost race

The discarded duplicate is killed the instant it opened, so it ran no billable
window — its hold belongs back in the pool, exactly like the openShell-throw path.
But killing the shell fires onExit asynchronously, and endAgentTerminalSession
would then settle the very hold just released, double-handling it against a window
that never happened. Clear connectedAt/holdId before the kill so no settle path can
reach it, and release the hold explicitly.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HnFmynYiN6bmrtibxdzLDJ
…-only)

Review catch (codex, P2). Making the sandbox resolution lazy also, unintentionally,
made the (scope, name) EXISTENCE check lazy — because the fused checkAuth only ever
learned "this terminal still exists" as a side effect of resolving its sandbox. The
60s re-auth tick calls only the access half, so after the split it stopped noticing
that a terminal's project, branch, or own row had been deleted: the orphaned PTY
kept running against a scope that no longer existed.

Restoring it naively would have re-broken the epic: resolveAgentTerminal fuses the
cheap question (does the row exist — a couple of indexed reads) with the expensive
one (where does its Sprite live — machineSandbox.acquire, which can RECONNECT OR
RESUME a hibernated Sprite). Calling it from the access half would wake the Sprite
on every re-auth tick and every tab-back.

So split the question, not just the call site:
- New `resolveAgentTerminalRow` in lib — resolveScopeKey + store.findByName, DB-only,
  provably Sprite-free (its tests pass `machineSandbox: undefined`).
- The access half runs it eagerly, AFTER the read-only access gates, so a user
  without edit rights still learns nothing about which terminals exist.
- The Sprite wake/read stays lazy in resolveSandbox.

Re-auth once again tears down a terminal whose scope was deleted — now without
waking anything to find out.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HnFmynYiN6bmrtibxdzLDJ
…the loser

An adversarial review of b873841 found that my own race fix was destructive.

`openPtyShell` with a non-null sessionId calls `sprite.attachSession(id)` — it
connects to a SERVER-SIDE exec session. When a persisted streamSessionId exists,
two racing cold connects both attach to the SAME remote session, and `PtyShell.kill()`
is SIGKILL on the process, not a local socket close. So "discard the duplicate"
SIGKILLed the very process the survivor was attached to: the guard turned a
recoverable orphan into a guaranteed terminal kill, on exactly the warm-reattach
path it was meant to protect.

Serialize at the key instead, so a second PTY is never opened at all:
- TerminalSessionMap gains trackCreate/pendingCreate. A cold connect claims the key
  before its first await; a concurrent connect for that key awaits the in-flight
  create and joins its session. Claim released on every exit (success, deny, throw),
  so a failed create can't wedge the key — and the joiner loops, in case another
  create queued behind a failure.
- This dissolves the class: no duplicate PTY, so no kill; no double slot; no double
  hold; and no ambiguous session-id discovery from concurrent creates.

Also fixed, from the same review:
- Slot LEAKED when billing.gate throws: the slot was reserved but no session owned it
  yet, so a rejection escaped onConnect with the slot still held. activeByUser is a
  process-lifetime counter, so one transient blip permanently locked a free-tier user
  (limit 1) out of terminals on that replica. Now released before rethrowing.
- discoverNewSessionId took the FIRST new tty session. One Sprite hosts every agent
  terminal on its machine, so a sibling terminal launching in the same window is
  indistinguishable from ours — persisting the sibling's id points this terminal's
  next cold connect at another terminal's PTY. Abstain unless exactly one appeared
  (the rule sprites-shell.ts's newTtySessionId already applies).
- Re-auth checked the CREATOR's userId forever. A session outlives its creator, and
  any authorized user may reattach — so a viewer whose access was revoked mid-session
  kept receiving output and could keep typing, since the long-gone creator still
  passed. Sessions now carry viewerUserId, set on create and on every reattach, and
  the tick re-checks the user actually driving the PTY.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HnFmynYiN6bmrtibxdzLDJ
Pure whitespace — the create path gained a try/finally (releasing the session-key
claim) without its body being re-indented.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HnFmynYiN6bmrtibxdzLDJ
The claim was only covered end-to-end through onConnect. Pin the primitive itself,
including the two cases that make it safe rather than merely present: a REJECTED
create still drops its claim (a failed create must never wedge the key forever),
and a stale create settling late must not revoke a newer create's claim.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HnFmynYiN6bmrtibxdzLDJ
@2witstudios

Copy link
Copy Markdown
Owner Author

@coderabbitai review

@coderabbitai

coderabbitai Bot commented Jul 11, 2026 •

Copy link
Copy Markdown
Contributor
✅ Action performed

Review finished.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

… serialization

CI failure was the coverage gate, not a test failure: all tests pass, but the new
branches dropped apps/realtime branch coverage to 97.73% (< 98%). Short by 5 arms.

Cover them with real, reachable behaviours rather than lowering the bar:
- The idempotent slot release (the double-release guard the round-2 review flagged
  as untested): a re-auth teardown releases the slot, and the killed PTY's late
  onExit must be a no-op — asserted to release exactly once.
- The idle reap hands its slot back (the other gap the review named).
- settleAccruedWindow / startSettleHeartbeat error arms that the added denominator
  tipped over: a settle rejecting with a non-Error value, the re-hold gate doing
  the same, a heartbeat firing after the session was removed, and an overlapping
  heartbeat skipped while a settle is still in flight.

Also refresh the endAgentTerminalSession deleteByKey comment: it referenced the
"cold-path racer that lost", a scenario per-key serialization removed. The identity
guard is now purely defensive (one session per key at a time), and the comment says so.

Handler branch coverage 94.4% -> 98.02%; realtime gate green (exit 0). 641 tests pass.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HnFmynYiN6bmrtibxdzLDJ
@2witstudios
2witstudios merged commit f4980cc into master Jul 11, 2026
3 checks passed
@2witstudios
2witstudios deleted the pu/sprites-1-3-onconnect-reorder branch July 11, 2026 16:51
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant