Repository navigation
[sprites 5-2] Checkpoint before destructive agent operations - #2025
Conversation
Checkpoints are ~300ms copy-on-write filesystem snapshots designed
exactly for unattended agent experiments (docs.sprites.dev/concepts/
checkpoints), and we used them nowhere — while sprites.ts's own
spawnWithSelfHealingCwd doc worried an agent that `rm -rf`s /workspace
would otherwise brick the sandbox.
- checkpoint-policy.ts: pure shouldCheckpoint({flagEnabled,
lastCheckpointAt, turnId, lastCheckpointTurnId, now}) — at most once
per agent turn, plus a rapid-batch throttle safety net. Also the
checkpoint-comment formatter, the flag resolver (default ON outside
production, OFF in prod pending PR discussion), and in-process
per-sandbox bookkeeping (mirrors quota.ts's machineActivityByKey —
no persistence, no reaper).
- sprites.ts: SpriteInstanceLike.createCheckpoint (matches the real
@fly/sprites SDK's createCheckpoint(comment?) -> CheckpointStream),
drained via wrap()'s ExecutableSandbox.createCheckpoint.
- machine-host.ts / sprite-machine-host.ts / machine-host-adapter.ts:
plumb createCheckpoint through MachineHandle (optional — a future
non-Sprite backend need not support it) so the agent-bash path,
which goes through MachineHost, gets a real ExecutableSandbox.
- tool-runners.ts: SandboxCheckpointDeps (fully optional seam) +
maybeCheckpointBeforeBatch, called before runCommand in
runBashInSandbox. Fail-open: any failure is logged and swallowed,
never blocks the batch; a failed checkpoint is not recorded, so a
later batch in the same turn gets another attempt.
- apps/web wiring: turnId is lazily stamped once per streamText run
onto ToolExecutionContext (same mutate-in-place pattern as
activeMachine) and threaded into SandboxActorContext.turnId;
buildRealSandboxRunDeps wires the checkpoint dep to the real SDK
call + in-process state.
Restore stays a manual/admin action (future epic) — this leaf only
ever creates checkpoints, never restores one. Terminal PTY path
checkpointing can follow once this proves out.
Test evidence: 680 sandbox unit tests green in packages/lib, 2284
apps/web unit tests green (one unrelated pre-existing failure:
activity-tools.test.ts's DB role, and one unrelated pre-existing
ClickHouse container OOM in analytics-gdpr.integration.test.ts — both
predate this branch). Full monorepo typecheck clean.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018LbMMKiXdkHrcxdtLsZH4Y
|
Warning Review limit reached
Next review available in: 35 minutes Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available. How can I continue?After more reviews become available, a review can be triggered using the To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews. How do review limits work?CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability. For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window. Please refer docs for additional details. Review details⚙️ Run configurationConfiguration used: defaults Review profile: CHILL Plan: Pro Run ID: 📒 Files selected for processing (1)
📝 WalkthroughWalkthroughAdds a Sprite-backed checkpoint API, turn-based checkpoint policy and coalescing, fail-open pre-batch checkpoints for sandbox bash execution, and lazy propagation of stable agent turn IDs. ChangesSandbox checkpointing
Estimated code review effort: 4 (Complex) | ~45 minutes Sequence Diagram(s)sequenceDiagram
participant AgentTool
participant resolveSandboxActorContext
participant runBashInSandbox
participant ExecutableSandbox
participant SpriteCheckpointStream
AgentTool->>resolveSandboxActorContext: resolve shared tool context
resolveSandboxActorContext-->>AgentTool: stable turnId
AgentTool->>runBashInSandbox: run bash batch with turnId
runBashInSandbox->>ExecutableSandbox: createCheckpoint(comment)
ExecutableSandbox->>SpriteCheckpointStream: drain checkpoint stream
SpriteCheckpointStream-->>ExecutableSandbox: completion or error
ExecutableSandbox-->>runBashInSandbox: checkpoint result
runBashInSandbox->>runBashInSandbox: execute bash batch
Possibly related PRs
🚥 Pre-merge checks | ✅ 3 | ❌ 2❌ Failed checks (1 warning, 1 inconclusive)
✅ Passed checks (3 passed)
✨ Finishing Touches🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: bdb12b72c7
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| if (lastCheckpointAt !== null && now.getTime() - lastCheckpointAt.getTime() < CHECKPOINT_MIN_INTERVAL_MS) { | ||
| return false; |
There was a problem hiding this comment.
Checkpoint new turns even inside the interval
With the default CHECKPOINT_MIN_INTERVAL_MS of 30s, this skips checkpointing whenever a different turn starts soon after the previous one. Separate chat turns can easily happen within that window, so if the second turn runs destructive bash, the newest restore point is still from before the previous turn and a restore would discard any legitimate work done in between. The interval throttle should not suppress the first checkpoint for a new turn.
Useful? React with 👍 / 👎.
There was a problem hiding this comment.
Confirmed — good catch. Fixed in 2a1a3eb: removed the CHECKPOINT_MIN_INTERVAL_MS cross-turn gate entirely. shouldCheckpoint is now purely flagEnabled && turnId !== lastCheckpointTurnId — a new turnId always checkpoints regardless of how recently a different turn's checkpoint was taken, since the policy only ever remembers the single most-recent turn and a differing id is by definition never-before-checkpointed.
Added a regression test (both at the pure-policy level and through the full runBashInSandbox path) that pins two different turns checkpointing at the identical instant — the most adversarial case for a time-based throttle. Left the doc comment on shouldCheckpoint explaining why a cross-turn interval gate is unsafe here, so it doesn't get reintroduced by a future "let's rate-limit this" change without solving the underlying tension.
Leaving this thread open for you to verify the fix rather than resolving it myself.
… a new turn's checkpoint Addresses a P2 review finding from chatgpt-codex-connector on PR #2025. shouldCheckpoint's CHECKPOINT_MIN_INTERVAL_MS gate suppressed a NEW turn's first checkpoint if it fell within 30s of a prior, DIFFERENT turn's checkpoint. Since the policy only ever remembers the single most-recent turnId, a turnId that differs from it is by definition a turn that has never been checkpointed — so throttling it on elapsed time silently defeated the safety net: two legitimate turns close together (an ordinary rapid back-and-forth) would leave only the OLDER turn's restore point on record, and a restore after the newer turn's destructive work would discard the newer turn's real work too. Removed the interval gate entirely (and the now-dead envInt/ CHECKPOINT_MIN_INTERVAL_MS/lastCheckpointAt/now plumbing that only existed to support it). shouldCheckpoint is now exactly what the leaf's actual requirement needs: at most once per agent turn via a pure turnId comparison, with a documented warning against reintroducing a cross-turn interval gate without solving this. Added a regression test at both the pure-policy level and the full runBashInSandbox integration level, asserting a new turn checkpoints even when `now` is identical to the prior (different) turn's checkpoint — the most adversarial case for a time-based throttle. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018LbMMKiXdkHrcxdtLsZH4Y
|
@codex review |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 2a1a3eba3a
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| ) { | ||
| return; | ||
| } | ||
| await checkpoint.createCheckpoint({ sandbox, comment: checkpointComment(turnId) }); |
There was a problem hiding this comment.
Bound checkpoint creation so fail-open cannot hang
When checkpointing is enabled and the Sprites checkpoint call returns a stream that stalls instead of rejecting, this await never reaches the catch, so the bash command is never started and the quota/billing slot stays held until the outer request times out. Since the new policy is meant to be fail-open, wrap checkpoint creation/draining in a short timeout (or otherwise abort it) and proceed on timeout just like other checkpoint failures.
Useful? React with 👍 / 👎.
There was a problem hiding this comment.
Confirmed, and good catch — fixed in e56ab7b.
Bounded checkpoint.createCheckpoint(...) with a CHECKPOINT_TIMEOUT_MS (10s) race at the shell layer (maybeCheckpointBeforeBatch, this file) so the promise from the caller's perspective is bounded regardless of what deps.checkpoint happens to be wired to — a timeout is caught by the existing try/catch and treated exactly like any other checkpoint failure (logged, swallowed, batch proceeds). Also bounded it at the driver layer (sprites.ts's createCheckpoint), since the SDK exposes no timeout/abort of its own for either the initial call or the stream it returns; on timeout there it also best-effort calls stream.close() so a stalled read doesn't hold the underlying connection open indefinitely in the background after we've given up waiting on it.
While digging into this I also found (via an internal multi-agent review pass) a related real bug: the check-then-act on the in-process per-turn state isn't atomic, so two bash tool calls dispatched concurrently by the AI SDK in one agent step could both pass the "already checkpointed this turn" check before either recorded — producing two checkpoints for one turn. Fixed that too, in the same commit, with a coalesceCheckpointAttempt helper that registers an in-flight promise synchronously so concurrent callers for the same sandbox share one attempt.
Added regression tests for both (fake-timer timeout tests in sprites.test.ts/tool-runners.test.ts, and a concurrent-dispatch test for the race).
Addresses a P2 finding from chatgpt-codex-connector plus a multi-agent code-review pass on PR #2025 (8 independent finder angles, several converging on the same issues from different directions). Critical fixes: - Bound the checkpoint SDK call so it can never hang the batch. maybeCheckpointBeforeBatch's fail-open contract ("never block agent work on checkpoint availability") was violated by an unbounded await: a stalled checkpoint stream held the bash command's concurrency/billing slot until the outer request timed out. Added CHECKPOINT_TIMEOUT_MS (10s) at the shell layer (tool-runners.ts, bounds the promise regardless of the injected implementation) AND at the driver layer (sprites.ts, also best-effort closes the stalled stream to release its underlying connection — SDK exposes no timeout/abort of its own). - Fix a real concurrency race in "at most once per turn". The AI SDK can execute multiple tool calls from one agent step concurrently (Promise.all); two bash calls in the same turn could both pass shouldCheckpoint's synchronous check before either recorded, producing two checkpoints for one turn. Added coalesceCheckpointAttempt (checkpoint-policy.ts): concurrent attempts for the same sandbox share one in-flight promise, registered synchronously so the race window closes regardless of timing between callers. Simplifications (2 independent review agents converged on the same finding for MachineHandle.createCheckpoint): - MachineHandle.createCheckpoint is now required, not optional — it existed as optional only to guard a hypothetical future non-Sprite backend that doesn't exist; the only implementer (sprite-machine-host) always supplies it. Removed the runtime Promise.reject fallback in machine-host-adapter.ts. - De-duplicated the two return branches of createResolveSandboxActorContext (9 of 10 fields were identical) into one parallel fetch + shared base object, preserving the original findDrive/findUser/getActorInfo concurrency. - Added an opportunistic eviction sweep to the in-process checkpoint state map (mirrors quota.ts's machineActivityByKey pattern) so it doesn't grow unbounded for the life of the process. New/updated tests: checkpoint-policy.test.ts (coalescing, eviction), tool-runners.test.ts (timeout fail-open, concurrent-turn race), sprites.test.ts (stream timeout + close() on hang), machine-host-adapter.test.ts (createCheckpoint delegation), plus createCheckpoint stubs added to every MachineHandle test fixture now that the field is required. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018LbMMKiXdkHrcxdtLsZH4Y
|
@codex review |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: e56ab7b95d
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| const drain = async (): Promise<void> => { | ||
| stream = await sprite.createCheckpoint(comment); | ||
| let streamError: string | undefined; | ||
| await stream.processAll((message) => { |
There was a problem hiding this comment.
Drain checkpoint streams with the Node SDK API
When this runs against the real @fly/sprites Node SDK, the checkpoint stream returned by sprite.createCheckpoint(...) is an async iterable; the official Node example drains it with for await (const msg of stream) on https://sprites.dev/api/sprites/checkpoints, while processAll is shown for the Go client. In environments where checkpointing is enabled, this call will throw TypeError: stream.processAll is not a function, the fail-open wrapper will swallow it, and no pre-agent checkpoint will ever be created or recorded before destructive bash runs.
Useful? React with 👍 / 👎.
There was a problem hiding this comment.
This one I'm going to push back on — I believe it's factually incorrect for this repo's pinned SDK version, though I appreciate the scrutiny given how consequential it would be if true.
Both apps/web/package.json and apps/realtime/package.json pin @fly/sprites to the EXACT version 0.0.1-rc37 (no range), and that's what's actually installed (confirmed via bun.lock's hash). I read the real, compiled, installed source directly rather than relying on the docs site:
node_modules/@fly/sprites/dist/checkpoint.d.ts:
export declare class CheckpointStream {
next(): Promise<StreamMessage | null>;
processAll(handler: (msg: StreamMessage) => void | Promise<void>): Promise<void>;
close(): void;
[Symbol.asyncIterator](): AsyncIterableIterator<StreamMessage>;
}node_modules/@fly/sprites/dist/checkpoint.js (the actual runtime implementation, not just the type declaration):
async processAll(handler) {
try {
let msg;
while ((msg = await this.next()) !== null) {
await handler(msg);
}
}
finally {
this.close();
}
}So processAll is a real, implemented method on CheckpointStream in the Node SDK we actually ship — it will NOT throw TypeError: stream.processAll is not a function. The class supports BOTH consumption styles: for await (const msg of stream) via [Symbol.asyncIterator], and the processAll(handler) convenience method I'm using — both drive the same internal next() loop. The docs page's Node example apparently just didn't happen to show the processAll variant, but that doesn't mean it's Go-only or absent from the Node client.
One thing this DID surface, though: processAll's own finally block already calls this.close() on completion or error — so my own stream.close() call in sprites.ts's timeout branch is redundant on the happy/error path, but still meaningful for the actual timeout case: if next() is mid-await this.reader.read() when our outer timer fires, calling close() from outside cancels that pending read via reader.cancel(), which is the real unstick mechanism. So that part of the fix stands as designed.
If I'm wrong about the pinned version or missed something, happy to be corrected — but I wanted to show my work with the actual installed source rather than either blindly applying the suggestion or dismissing it without evidence.
|
@coderabbitai review |
✅ Action performedReview finished.
|
There was a problem hiding this comment.
Actionable comments posted: 1
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@packages/lib/src/services/sandbox/sandbox-client/__tests__/sprites.test.ts`:
- Around line 1128-1140: Rename the test case around createCheckpoint in
sprites.test.ts to state that it propagates the SDK rejection and lets the
caller decide the fail-open policy. Keep the existing rejecting assertion and
test behavior unchanged.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: defaults
Review profile: CHILL
Plan: Pro
Run ID: 2f69a318-ece0-4d35-9275-59d0ca9d33b7
📒 Files selected for processing (26)
apps/web/src/lib/ai/core/types.tsapps/web/src/lib/ai/tools/__tests__/sandbox-tools-runtime.test.tsapps/web/src/lib/ai/tools/__tests__/sandbox-tools.test.tsapps/web/src/lib/ai/tools/sandbox-tools-runtime.tspackages/lib/package.jsonpackages/lib/src/services/machines/__tests__/agent-terminals.test.tspackages/lib/src/services/machines/__tests__/machine-branches.test.tspackages/lib/src/services/machines/__tests__/machine-projects.test.tspackages/lib/src/services/sandbox/__tests__/checkpoint-policy.test.tspackages/lib/src/services/sandbox/__tests__/git-tool-runners.test.tspackages/lib/src/services/sandbox/__tests__/machine-diff.test.tspackages/lib/src/services/sandbox/__tests__/machine-fs.test.tspackages/lib/src/services/sandbox/__tests__/machine-git-blob.test.tspackages/lib/src/services/sandbox/__tests__/persistent-machine-fs.test.tspackages/lib/src/services/sandbox/__tests__/tool-runners.test.tspackages/lib/src/services/sandbox/checkpoint-policy.tspackages/lib/src/services/sandbox/machine-host.tspackages/lib/src/services/sandbox/sandbox-client/__tests__/machine-host-adapter.test.tspackages/lib/src/services/sandbox/sandbox-client/__tests__/sprite-machine-host.test.tspackages/lib/src/services/sandbox/sandbox-client/__tests__/sprites.test.tspackages/lib/src/services/sandbox/sandbox-client/__tests__/wake-retry.test.tspackages/lib/src/services/sandbox/sandbox-client/machine-host-adapter.tspackages/lib/src/services/sandbox/sandbox-client/sprite-machine-host.tspackages/lib/src/services/sandbox/sandbox-client/sprites.tspackages/lib/src/services/sandbox/sandbox-client/types.tspackages/lib/src/services/sandbox/tool-runners.ts
Addresses a CodeRabbit review nit on PR #2025: "resolves cleanly when the SDK call itself rejects" asserted .rejects.toThrow(...) — the promise rejects, it does not resolve. Renamed to "propagates the SDK's rejection (caller decides fail-open policy)". Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018LbMMKiXdkHrcxdtLsZH4Y
Summary
Checkpoints are ~300ms copy-on-write filesystem snapshots designed exactly for "unattended agent experiments" and destructive changes (docs.sprites.dev/concepts/checkpoints) — and we used them nowhere, while
sprites.ts's ownspawnWithSelfHealingCwddoc worries an agent thatrm -rfs/workspacewould otherwise brick the sandbox. This checkpoints before each agent bash batch, fail-open, at most once per agent turn.Restore is explicitly out of scope (future epic, manual/admin action only) — this leaf only ever creates checkpoints.
Requirements → how satisfied
Given an agent bash tool batch about to execute (flag on), should create a checkpoint tagged with a recognizable comment, at most once per agent turn.
checkpoint-policy.ts's pureshouldCheckpoint({flagEnabled, turnId, lastCheckpointTurnId})— sameturnIdas the last checkpoint → skip; a differentturnId→ always checkpoint.tool-runners.ts'smaybeCheckpointBeforeBatchcalls it beforerunCommandinrunBashInSandbox, tagging withcheckpointComment(turnId)→pagespace-pre-agent-<turnId>.turnIdis lazily stamped once per streamText run ontoToolExecutionContext(same mutate-in-place pattern as the existingactiveMachinefield) and threaded intoSandboxActorContext.turnId. Concurrent tool calls in the same turn (the AI SDK can dispatch several viaPromise.all) are coalesced onto one checkpoint attempt viacoalesceCheckpointAttempt— see "Review fixes" below.Given checkpoint API failure, should proceed with the batch (fail-open) and log — never block agent work on checkpoint availability.
maybeCheckpointBeforeBatchwraps the whole decision+call in try/catch; any failure (including a bounded timeout — see below) callssafeLogWarnand returns normally. A failed checkpoint is not recorded, so a later batch in the same turn gets another attempt.Given repeated tool batches within one turn, should not create additional checkpoints (pure throttle).
shouldCheckpoint'sturnId === lastCheckpointTurnIdcomparison — a pure per-turn dedup, made race-safe under concurrent dispatch bycoalesceCheckpointAttempt.Given the checkpoint list growing, should rely on platform auto-pruning; do not build a custom reaper.
No reaper of the platform's checkpoint list is built — see Findings below. (The in-process throttle bookkeeping map does get an opportunistic eviction sweep — a different, purely local concern, see Review fixes.)
Review fixes
Two rounds of review — the automated Codex reviewer plus an internal 8-angle multi-agent code-review pass — surfaced real issues, all fixed:
P2 (chatgpt-codex-connector): cross-turn interval throttle could suppress a new turn's checkpoint. discussion — The original
shouldCheckpointalso gated on 30s elapsed since the last checkpoint, regardless of turn, so two legitimate turns close together would leave only the OLDER turn's restore point on record. Fixed (2a1a3eba3): removed the interval gate — a differentturnIdis by definition never-before-checkpointed, so it always checkpoints now.P2 (chatgpt-codex-connector): a stalled checkpoint stream could hang the batch forever. discussion — The SDK exposes no timeout for
createCheckpoint/its stream, so anawaiton a hung connection never reached thecatch, holding the concurrency/billing slot until the outer request timed out — the opposite of "fail-open." Fixed (e56ab7b95): bounded withCHECKPOINT_TIMEOUT_MS(10s) at both the shell layer (tool-runners.ts, bounds the promise regardless of the injected implementation) and the driver layer (sprites.ts, also best-effortstream.close()s on timeout to release the connection).(internal review, 2 independent finder agents) Concurrent-turn race in "at most once per turn." The AI SDK can execute multiple tool calls from one agent step concurrently; two bash calls in the same turn could both pass
shouldCheckpoint's synchronous check before either recorded, producing two checkpoints for one turn. Fixed (e56ab7b95):coalesceCheckpointAttemptregisters an in-flight promise synchronously so concurrent callers for the same sandbox share one attempt regardless of exact timing.(internal review, 2 independent finder agents)
MachineHandle.createCheckpointwas optional purely for a hypothetical future non-Sprite backend that doesn't exist. Premature abstraction per project convention. Fixed (e56ab7b95): made it required (matchingExecutableSandbox.createCheckpoint); removed the runtimePromise.rejectfallback inmachine-host-adapter.ts.(internal review) Dead
CheckpointStatefield / duplicated branches. A repo-wide grep confirmed nothing readlastCheckpointAtfor the throttle decision anymore after fix Upload files, agents, dm's, more #1 — re-justified by repurposing it for an opportunistic eviction sweep (below) rather than deleting it outright. Also de-duplicated the two near-identical return branches ofcreateResolveSandboxActorContextinsandbox-tools-runtime.ts(9 of 10 fields were byte-identical) into one parallel fetch + shared base object, preserving the originalfindDrive/findUser/getActorInfoconcurrency.(internal review) In-process checkpoint state map had no bound.
stateBySandboxIdhad no symmetric acquire/release and would grow by one entry per distinct sandbox ever seen for the life of the process. Fixed: added an opportunistic 24h-TTL eviction sweep, mirroringquota.ts'smachineActivityByKeypattern.Findings (named-checkpoint accumulation vs. auto-pruning)
I could not exercise a live Sprite in this environment, so I read the SDK (
@fly/spritessprite.d.ts/checkpoint.d.ts) and the docs instead of hitting the API directly. What I found:auto-ids as "pruned over time" and explicitly call them "a safety net rather than a retention strategy."sprite.createCheckpoint(comment)— the same explicit, user-facing creation path assprite checkpoint create --comment "..."from the CLI, not the platform's own automatic background mechanism. The docs don't say whether explicitly-created checkpoints are also subject to automatic pruning, or whether they accumulate until a human/admin prunes them.GET /v1/sprites/{name}/checkpointsafter some days of agent activity) before this ships broadly. If explicit checkpoints do NOT get pruned, a lightweight follow-up (e.g. keep only the last Npagespace-pre-agent-*checkpoints) would be worth scoping as a separate leaf.Other implementation notes
isCheckpointBeforeAgentBatchEnabled()— explicitSANDBOX_CHECKPOINT_BEFORE_AGENT_BATCH=true|falsealways wins; unset defaults ON outside production, OFF in production. The leaf spec explicitly deferred the production default to PR discussion — please weigh in.SpriteInstanceLike.createCheckpoint(comment?)mirrors the real SDK'sSprite.createCheckpoint(returns aCheckpointStream);wrap()drains it viaprocessAll, bounded byCHECKPOINT_TIMEOUT_MS, surfacing the firsterror-type message as a rejection (purecheckpointStreamErrorMessagehelper, unit tested).Map<sandboxId, {lastCheckpointAt, lastCheckpointTurnId}>+ a coalescingMap<sandboxId, Promise<void>>for in-flight attempts (both mirrorquota.ts's in-process-Map conventions). A process restart just re-checkpoints on the next batch, which is harmless (COW, ~300ms)../services/sandbox/checkpoint-policysubpath export (needed it after hitting the "new subpath modules need an exports entry" gotcha from a prior leaf).Test evidence
Checklist for the orchestrator
SANDBOX_CHECKPOINT_BEFORE_AGENT_BATCH(currently OFF-by-default in prod pending this discussion).auto-) checkpoints are pruned automatically, or need a follow-up reaper leaf.runBashInSandboxsince that's the explicitly named "batch about to execute" surface and the highest-risk one (arbitrary shell); file-write tools already go through path-escape checks and single-file diffs. (An internal review pass also flagged this exact question independently — noted here for visibility, not acted on without a scope decision.)🤖 Generated with Claude Code
https://claude.ai/code/session_018LbMMKiXdkHrcxdtLsZH4Y
Summary by CodeRabbit
New Features
Reliability
Configuration