Skip to content

Agent Code Execution — bash/code in chat via Fly Sprites (epic) - #1487

Merged
2witstudios merged 13 commits into
masterfrom
pu/flash-sandbox
Jun 3, 2026
Merged

2witstudios merged 13 commits into
masterfrom
pu/flash-sandbox

Conversation

@2witstudios

@2witstudios 2witstudios commented Jun 2, 2026 •

Copy link
Copy Markdown
Owner

Agent Code Execution (epic) — bash/code in chat via Fly Sprites

Lets PageSpace AI agents run bash and code in the existing AI chat, executed in isolated Fly Sprites microVMs, behind a default-OFF feature flag. Ships dark — unreachable by agents until the flag is enabled AND the Fly environment is provisioned.

Composed PRs (all merged into pu/flash-sandbox, each individually CI-gated)

Safety posture

microVM isolation + drive/role authz + per-tenant quotas/concurrency + kill-switch + secret-scrubbed env + default-deny egress + resume re-authz + command policy + output truncation + immutable audit. Gated on feature flag + authz (NOT deployment mode).

Validation

Full CI green against master (Lint & TypeScript, Unit Tests, Security Test Suite, CodeQL, Static Security Analysis, Dependency Audit, Secret Scanning). Unit/integration tests use injected fakes — no live Fly calls.

⚠️ Enablement gates before flag-on (tracked in the "Fly Sprites Enablement & Gap Verification" task list)

The @fly/sprites network policy is domain-based only (no IP/CIDR), so it cannot itself block IP-literal egress. G1 (hard gate): empirically verify a Sprite cannot reach the Fly 6PN / *.internal / metadata / Tigris / Postgres by IP — requires the authed Fly env (CLI/token), which is NOT available in this workspace, so it could not be run here. Plus: provision Fly + SPRITES_TOKEN (G8), orphan-sweep reaper (G5), staged per-tenant rollout (G11).

Rollout

Merge is safe while dark. Enable per-tenant only after G1 passes.

🤖 Generated with Claude Code

Summary by CodeRabbit

Release Notes

  • New Features

    • Added sandboxed code execution capability with bash command execution and file read/write operations
    • Implemented code execution quotas and rate limiting with per-tier concurrency controls
  • Security

    • Added comprehensive audit logging for code execution activities
    • Implemented authorization gates and command policy enforcement for sandbox operations
    • Added egress filtering and kill-switch controls for the code execution feature
  • Infrastructure

    • Updated CI/CD workflows to support sandbox-related testing
    • Added environment configuration variables for code execution settings

2witstudios and others added 7 commits June 1, 2026 14:08
The Test Suite and Security workflows only triggered on pull_request to
[main, master, develop], so PRs targeting the pu/flash-sandbox integration
branch (the Agent Code Execution epic) ran no real CI. Add pu/flash-sandbox
to both pull_request branch filters so main's checks gate these PRs too.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
#1465)

* feat(sandbox): add canRunCode authorization gate with kill-switch

Compose getUserDrivePermissions + getAgentAccessLevel behind a single
fail-closed authorization chokepoint for agent code execution. Checks are
ordered cheapest-first: a default-OFF CODE_EXECUTION_ENABLED kill-switch and
cloud-only deployment gate deny before any DB round-trip. Never throws — any
dependency error resolves to a denial. DB-backed helpers are injected (lazily
imported in the default wiring) so the unit tests exercise the composition
with fakes and never touch the database.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(sandbox): add pure resolveExecutionPolicy with safe defaults

Resolve an explicit per-run policy (timeout, vCPU, memory, output cap, region)
instead of inheriting platform defaults. Egress is default-deny (empty
allowlist) and sandboxes are ephemeral (persistent: false) on every profile;
an unknown profile falls back to the most-restrictive safe minimum so a typo
can never widen the blast radius.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(sandbox): add pure buildSandboxEnv allowlist scrub

Build the sandbox environment by allowlist — copying only a fixed set of
explicitly-safe keys from the validated env — so no DB credential, signing
secret, or API key can ever reach untrusted code. Never spreads process.env or
the validated env wholesale; a newly-added secret is excluded by default.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(sandbox): add code execution quota and per-tier concurrency

Add app-level sub-limits so one tenant's runaway agent can't starve or bill
everyone under Vercel's account-wide caps. A per-user in-process semaphore
(ceiling scales by subscription tier, modeled on upload-semaphore) plus a daily
run budget via a new CODE_EXECUTION distributed-rate-limit entry applied
independently to user/drive/tenant scopes. checkCodeExecutionQuota checks
concurrency first so a saturated system rejects without spending any budget.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(sandbox): add code execution audit record + writer

Every run yields an immutable audit record (actor, redacted code, profile,
exit, duration, cost, timestamp) written to the hash-chained activity log via a
new code_execution ActivityOperation. Anomalous runs (timeout, OOM, blocked
command, non-zero exit) additionally raise a security audit event. Builders are
pure (injected timestamp); secrets in captured code are redacted before
persistence. The writer is fire-and-forget — a failing sink never breaks the
run it records.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* refactor(sandbox): bound code length before secret redaction

Truncate submitted code to the audit cap before running redaction regexes so
they only ever execute over bounded input, removing any ReDoS surface on large
submissions. codeTruncated still reflects the raw submitted length.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* refactor(sandbox): freeze resolved execution policies (review F1)

Policies are returned by reference from module constants; readonly only guards
compile time. Freeze the constants and their egress arrays so a downstream
caller cannot mutate the shared default-deny egress baseline (which would widen
egress globally for every subsequent run). Add a regression test asserting the
allowlist cannot be pushed to.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs(sandbox): reword redaction comment to avoid literal key prefixes

Keep secret-scanner-trigger patterns out of source comments.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(sandbox): don't fail env validation on stray kill-switch values (review P2-D)

CODE_EXECUTION_ENABLED was z.enum(['true','false']), so any other value (e.g.
=0 or =TRUE) made the app-wide validateEnv() throw — breaking unrelated startup
and health checks instead of leaving the feature disabled. Accept any string;
isCodeExecutionEnabled() already enables only on the exact value 'true'.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(sandbox): drop deployment-mode gate from canRunCode (review P2-C)

!isCloud() wrongly denied DEPLOYMENT_MODE=tenant, which is a cloud deployment
(isolated image per tenant) with the same feature set. We don't serve on-prem
(a local execution path is future work, not wired here), so any DEPLOYMENT_MODE
/ isOnPrem gate would only add a dead branch. Remove the check entirely; authz,
kill-switch, and quota remain the gates. Drops the isOnPrem dep and the
not_cloud denial reason.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(sandbox): redact secrets across the truncation boundary (review P2-A)

The earlier ReDoS fix truncated to 4 KB before redacting, so a quoted secret
whose closing quote fell past the cap escaped SECRET_ASSIGNMENT and leaked a
prefix into the immutable audit log. Redact over a wider 16 KB scan window,
then truncate the redacted result — a secret starting before the cap is fully
collapsed first. Simplify STANDALONE_TOKEN to a single non-ambiguous run so the
wider scan stays linear. Add a straddle-the-cap regression test.

Also document the invariant that sandbox code is always model-generated
(isAiGenerated, review F3) and that redaction is best-effort audit hygiene, not
the security boundary (review F4).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(sandbox): make quota check a non-incrementing preflight (review P2-B, F1)

checkDistributedRateLimit increments the bucket per call, so checking
user -> drive -> tenant in sequence charged the earlier scopes even when a
later one denied — letting an exhausted drive drain a user's daily budget in
unrelated drives. Switch the default dep to getDistributedRateLimitStatus (the
read-only sibling) so the multi-scope check consumes nothing. Document that the
check is advisory: the single real charge per run and acquireCodeExecutionSlot()
are wired at execution time in PR3, which must handle acquire === false.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs(sandbox): note agent-path drive-as-root coupling (review F5)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(sandbox): require actor authz for agent-origin runs (review P1)

Agent-origin runs only checked the agent page's drive edit access and
skipped the actor's owner/admin drive-role gate entirely, letting a plain
member escalate by triggering an agent that holds drive edit access.

canRunCode now always clears the actor user through authorizeUser first,
then — for agent origin — additionally requires the agent page to hold
drive edit access. Both the human and the agent must be entitled.

Adds two regression tests: agent-origin run denied when the triggering
user is a plain member (insufficient_role) or has no drive membership
(no_drive_access).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(sandbox): resolve unknown profile via own-key check (proto bypass)

resolveExecutionPolicy used `PROFILES[profile] ?? SAFE_MINIMUM_PROFILE`.
Bracket lookup resolves inherited keys, so a profile of '__proto__',
'constructor', 'toString', etc. returned a truthy Object.prototype member
and skipped the safe-minimum fallback — yielding a "policy" with undefined
timeout, vCPU, memory, output cap, and egress allowlist. The signature
accepts an arbitrary string, so an untrusted profile (PR3 wiring) reaches
it. Guard with an own-property check so only real profiles resolve and
everything else falls back to the most-restrictive policy, as documented.

Adds a regression test for the prototype-key inputs.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* refactor(sandbox): simplify audit redaction + dedupe resource id

Two cleanups surfaced by a proactive review pass on the audit builder:

- Drop the 16 KB REDACTION_SCAN_LIMIT window: redact the full code, then
  truncate. The window never added safety (anything past the 4 KB storage
  cap is dropped by the truncate either way) and carried its own straddle
  edge. Redact-before-truncate still fully collapses a secret straddling
  the storage cap; the regexes are linear so full-input scanning is cheap.
- Extract resolveAuditResourceId() so buildActivityLogInput and
  buildSecurityAuditEvent derive the run's resource id from one place —
  forensic correlation across the activity and security logs breaks if the
  two precedence chains ever diverge.

No behavior change to stored output; 51 sandbox tests green.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(sandbox): add adversarial ReDoS payload for secret redaction

Failing test: the SECRET_ASSIGNMENT regex backtracks polynomially on many
repetitions of 'key' (CodeQL js/polynomial-redos). Redaction runs over
agent-supplied code, so this is an attacker-triggerable CPU-exhaustion path.
Asserts redaction completes in linear time while still redacting a real secret.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(sandbox): make secret redaction regex linear-time (CodeQL js/polynomial-redos)

The SECRET_ASSIGNMENT regex matched a key as `[A-Za-z0-9_-]*keyword[A-Za-z0-9_-]*`
— two unbounded runs around a keyword whose characters the runs also match — so
it backtracked polynomially (~O(n^2)) on inputs like "keykeykey…". Redaction
runs over agent-supplied code, making it an attacker-triggerable CPU-exhaustion
path. Replace it with a single linear ASSIGNMENT matcher (`[A-Za-z0-9_-]+` for
the identifier, a distinct `[:=]` separator) and move the secret-keyword test
into the replace callback. Every redaction regex is now a single non-ambiguous
quantifier per class. Adversarial 384 KB payload: ~2 ms (was ~7 s).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The security workflow is path-gated and the Agent Code Execution code lives
in packages/lib/src/services/sandbox/** + the new sandbox_sessions schema,
which matched none of its filters — so the most security-critical code
(session isolation, resume re-authz) merged with only Lint+Unit Tests.
Add the sandbox paths so the Security suite + CodeQL scan this feature.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…tion) (#1469)

* feat(sandbox): derive unguessable conversation session keys

HMAC-SHA256 over (tenant + drive + conversation) keyed by a server-held
secret. Namespaced so distinct conversations never collide onto one
sandbox, and unguessable so an actor who knows a conversation id cannot
reconstruct the sandbox name to probe another session's warm VM. Pure —
the secret is injected by the effect layer.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(sandbox): plan conversation sandbox lifecycle

Pure planner deciding create / resume / idle-teardown / session-end
teardown / deny. Encodes two security invariants: an unauthorized actor
is denied even when a warm session exists (resume re-authz — never hand
back prior-actor state), and session-end always tears down regardless of
authorization (cleanup is unconditional, no orphaned VMs).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(sandbox): map execution policy to sandbox create options

Pure translation of the resolved ExecutionPolicy bounds (timeout, vCPU,
memory, persistent, region) into the option object the effect layer hands
the sandbox client. Explicit caps from policy; never platform defaults.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(sandbox): add sandbox_sessions table for the conversation link

Persists the sandboxId<->conversation link keyed by the opaque session
key (unique), so later turns reconnect to the same warm sandbox. A row is
deleted on teardown; lastActiveAt drives idle reclamation. Generated
migration 0142; exported from schema.ts + db package exports.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(sandbox): session store for the sandboxId<->conversation link

Small interface over sandbox_sessions (find/save/touch/remove) with a
Drizzle-backed impl that upserts on the unique session key. Lazily imports
the db module so callers injecting a fake never load the DB graph; the
orchestrator is unit-tested against an in-memory store.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(sandbox): conversation sandbox lifecycle effects + teardown

acquireConversationSandbox ties the pure pieces to injected IO: derive
key, look up the link, RE-AUTHORIZE the current actor, plan, then execute
create/resume/idle-teardown against the sandbox client + store. Enforces
resume re-authz (deny never reconnects a warm VM) and no-orphans (stop the
new VM if the link can't be persisted). teardownConversationSandbox is
idempotent and never throws — stop is best-effort, the link is always
removed. Adds SANDBOX_SESSION_SECRET (min-32, optional, fail-closed) and
the new module exports.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* refactor(sandbox): type session store queries against real drizzle types

Drop the hand-rolled structural db interface + cast in favour of writing
the Drizzle queries directly in createDbSandboxSessionStore against the
real, lazily-imported db/eq/table. The store interface remains the unit-test
seam (the orchestrator injects an in-memory fake); the concrete queries are
now fully type-checked rather than cast through 'unknown'.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(sandbox): fail closed on store IO errors; best-effort resume touch

Wrap acquireConversationSandbox so an unexpected IO failure (store lookup,
link removal, client.get) denies with reason 'error' instead of throwing —
DB-backed checks must fail closed. Make the resume lastActiveAt update
best-effort so a failed metadata write never denies an authorized,
confirmed-live resume.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* refactor(sandbox): hash session key with SHA3-256 (repo convention)

Switch the session-key HMAC digest from SHA-256 to SHA3-256 to match the
repo's convention for security tokens hashed at rest (auth/token-utils.ts;
CLAUDE.md 'SHA3-256 hashed at rest'). HMAC keying is retained — the key's
inputs are low-entropy, so the server secret is what makes it unguessable.
Output is still 64 hex chars; no behavioural change beyond the primitive.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(sandbox): address Codex review — blank secret + teardown IO guarding

- env-validation: accept a blank SANDBOX_SESSION_SECRET placeholder
  (.or(z.literal('')), mirroring the URL vars) so 'SANDBOX_SESSION_SECRET='
  disables sandbox acquisition (lifecycle fails closed) instead of failing
  app-wide env validation at instrumentation startup. A non-empty value
  still must be >= 32 chars.
- teardownConversationSandbox: guard the store lookup in try/catch and make
  the link removal best-effort (safeRemove), so end/idle/crash/failure
  cleanup never propagates a store error — honouring its documented
  idempotent, never-throws contract.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs(sandbox): tighten teardown contract wording + type guarded lookup

Doc/type-only: correct the teardown doc comments to state that the lookup
is guarded and the stop + link removal are best-effort (a lingering link
self-corrects on next acquire), and give the guarded `existing` lookup an
explicit SandboxSessionRecord|null type instead of an evolving any. No
behaviour change.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…ox (Agent Code Execution) (#1472)

* feat(sandbox): add @vercel/sandbox dep + pure safety primitives

Adds the @vercel/sandbox dependency and the pure, fail-closed building
blocks the PR3 execution path composes inline:

- egress.ts: maps a policy egress allowlist to a @vercel/sandbox network
  policy — `deny-all` by default, and never reachable to the cloud
  metadata endpoint or any RFC1918/CGNAT/link-local range even when
  widened for an external registry (subnet denies take precedence).
- command-policy.ts: size + empty + metadata-IP block, linear regex only
  (no polynomial backtracking), returns allow/block without throwing.
- output-limit.ts: byte-bounded truncation of untrusted stdout/stderr.
- sandbox-paths.ts: confines tool file IO to the sandbox root via the
  shared resolvePathWithinSync validator.
- sandbox-options.ts: carry the policy egress allowlist through to the
  client so provisioning can translate it.

Adds package.json exports for the new modules. Pure functions only;
no execution path is wired yet.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(sandbox): real @vercel/sandbox client adapter

Implements PR2's SandboxClient seam against the actual SDK and extends it
with the execution surface PR3 needs (runCommand / writeFiles /
readFileToBuffer) on an ExecutableSandbox handle.

Provisioning is locked down explicitly, never inheriting platform
defaults: deny-by-default egress from the policy allowlist, allowlisted
env via buildSandboxEnv (fail-safe to empty — never host secrets),
explicit vCPUs/persistence, and a VM lifetime that outlives the per-run
cap and the idle-reclaim window so a conversation's warm sandbox is
reused across turns. The per-run timeout is applied to runCommand, not to
the VM. Outbound credentials (future) are brokered via the network
policy, never injected as raw secrets.

The SDK statics are injected so create-param mapping, exit/stdout/stderr
surfacing, and get->null on a vanished sandbox are unit-tested with a
fake — never against the real Vercel API.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(sandbox): per-run budget charge + bash/writeFile/readFile runners

Adds chargeCodeExecutionBudget — the single real per-run charge that
increments the user/drive/tenant CODE_EXECUTION windows once, sharing the
scope-id construction with the non-incrementing preflight.

Adds the execution orchestration that is the body of each tool's execute,
with the whole safety layer inline and in order: kill-switch re-check;
command/path policy before any VM work (a blocked op never provisions and
is audited); quota preflight -> concurrency reservation -> budget charge
only once a live authorized sandbox is in hand; acquireConversationSandbox
(authz + resume re-authz) + reconnect; run/write/read; output truncation
to the policy cap; audit every executed run and every blocked op; and a
guaranteed concurrency-slot release in finally. The command is passed to
the sandbox as a structured arg array (sh -c), never a host-side shell
string. All IO is injected and unit-tested with fakes.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(sandbox): bash/writeFile/readFile AI SDK tool wrappers (unregistered)

Thin AI SDK tool() wrappers over the conversation sandbox: each execute
reads the chat context, resolves the actor (drive from the active
location, tenant = drive owner, concurrency tier = acting user), and
delegates to the @pagespace/lib runner where the safety layer lives. The
runner deps and context resolver are injected so the wrappers are
unit-tested with fakes (no DB, no real Vercel API).

NOT REGISTERED: these are not spread into pageSpaceTools and are not
tool_search-discoverable. Exposure to agents (registration + default-OFF
feature flag) is PR4 — until then this module is unreachable from chat.

Adds the optional VERCEL_TOKEN / VERCEL_TEAM_ID / VERCEL_PROJECT_ID env
vars to serverEnvSchema (the SDK falls back to OIDC when absent; a partial
triad never half-authenticates).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…s in agent chat (#1477)

* feat(sandbox): add call-time tool gate (kill-switch + authz + quota), default-OFF

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(sandbox): register flag-gated bash/writeFile/readFile tools in agent chat

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…ode Execution) (#1481)

* feat(sandbox): swap execution driver from @vercel/sandbox to @fly/sprites

Replace the Vercel Sandbox client with a Fly Sprites driver behind the
existing provider-neutral SandboxClient seam. The entire safety layer
(can-run-code, quota, audit, command-policy, output-limit, sandbox-paths,
session-key, lifecycle, tool-gate) is unchanged.

- Add sandbox-client/{types,sprites}.ts implementing ExecSandboxClient over
  @fly/sprites (getOrCreate resumes/creates by session key; stop DESTROYS;
  per-command timeout enforced in-driver with guaranteed teardown).
- Egress lockdown is reshaped to the Fly L3 NetworkPolicy: default-deny
  catch-all, with explicit internal-Fly denies (*.internal, _api.internal,
  *.flycast, Tigris) placed BEFORE any allow as SSRF defence-in-depth.
- Fresh Sprites are destroyed if the egress policy can't be applied — never
  handed back with open egress.
- env: drop Vercel OIDC triad, add SPRITES_API_TOKEN (blank → fail-closed).
- SANDBOX_ROOT → /workspace; region iad1 → iad.
- Pin @fly/sprites to 0.0.1-rc37 (the published 0.0.1 release regressed and
  dropped the network-policy + filesystem APIs); remove @vercel/sandbox.
- DELETE vercel-sandbox-client.ts + its test — hard cutover, no compat shim.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* refactor(sandbox): run registered bash/writeFile/readFile tools on Sprites

Reconcile the merged PR4 registration with the Sprites driver. Split the
flag-gated, default-OFF tools so the provider SDK never loads in the factory's
tests:

- sandbox-tools.ts is now the provider-agnostic factory only (schemas +
  context resolution + call-time gate + delegation); no DB, no backing-provider
  SDK import.
- sandbox-tools-runtime.ts holds the production wiring (DB-backed session store,
  Fly Sprites driver, quota, audit, actor resolver) and the gate wiring, and
  exports buildSandboxTools.
- ai-tools.ts imports buildSandboxTools from the runtime module; registration
  stays flag-gated and default-OFF (PR4 behaviour preserved, just on Sprites).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(sandbox): lazy-load Sprites driver off the default path; robust ExecError shape

Addresses Codex review on #1481.

P1 — @fly/sprites is ESM-only and requires Node >=24, but the web images run
Node 22. The static import chain (ai-tools → sandbox-tools-runtime → sprites →
@fly/sprites) pulled the SDK into the module graph on every chat request,
including the default code-execution-OFF path. Replace it with a dynamic
import inside getSandboxClient() so the SDK is loaded only when a sandbox tool
actually runs (kill-switch ON). getSandboxClient is now async; acquire/reconnect
await it. The off-path no longer evaluates the unsupported SDK.

P2 — Make the ExecError duck-type accept both the nested `.result` shape and
the flattened `exitCode/stdout/stderr` shape the SDK also exposes, so a version
skew can't turn a real non-zero exit into a transport failure. Add a test for
the flat shape.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@vercel

vercel Bot commented Jun 2, 2026 •

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated (UTC)
pagespace-marketing Ready Ready Preview, Comment Jun 2, 2026 10:20pm

@coderabbitai

coderabbitai Bot commented Jun 2, 2026 •

Copy link
Copy Markdown
Contributor

Review Change Stack

Warning

Review limit reached

@2witstudios, we couldn't start this review because you've reached your PR review rate limit.

More reviews will be available in 47 minutes and 58 seconds. Learn how PR review limits work.

Your organization has run out of usage credits. Purchase more in the billing tab.

⌛ How to resolve this issue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

We recommend that you space out your commits to avoid hitting the rate limit.

🚦 How do rate limits work?

CodeRabbit enforces hourly rate limits for each developer per organization.

Our paid plans include higher PR review limits than trial, open-source, and free plans. In all cases, reviews become available again over time. During sustained high-volume PR review activity, CodeRabbit may temporarily slow when the next review becomes available.

Please see our Fair Usage Limits Policy for further information.

ℹ️ Review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro

Run ID: 20f1b83d-fa91-4c72-893b-3098d10ecdfa

📥 Commits

Reviewing files that changed from the base of the PR and between a697221 and a614f9f.

⛔ Files ignored due to path filters (1)
  • bun.lock is excluded by !**/*.lock
📒 Files selected for processing (8)
  • .github/workflows/security.yml
  • packages/db/drizzle/0145_flawless_living_mummy.sql
  • packages/db/drizzle/meta/0145_snapshot.json
  • packages/db/drizzle/meta/_journal.json
  • packages/db/package.json
  • packages/db/src/schema.ts
  • packages/lib/package.json
  • packages/lib/src/services/sandbox/egress.ts
📝 Walkthrough

Walkthrough

This PR adds a sandboxed code-execution platform: DB-backed session tracking, Sprites-based executable sandboxes, authorization and quota gating, lifecycle planning, audit/redaction pipelines, tool runners (bash/write/read), web integration, CI workflow triggers, and packaging/type exports.

Changes

Sandboxed Code Execution Platform

Layer / File(s) Summary
CI, env, and package wiring
.github/workflows/security.yml, .github/workflows/test.yml, packages/lib/package.json, packages/lib/src/config/env-validation.ts, packages/lib/src/config/__tests__/env-validation.test.ts, packages/lib/src/monitoring/activity-logger.ts, packages/lib/src/security/distributed-rate-limit.ts
Workflow triggers now include sandbox-related paths and pu/flash-sandbox branch; packages/lib adds ./services/sandbox/* exports, typesVersions, and @fly/sprites dependency; serverEnvSchema accepts CODE_EXECUTION_ENABLED, SANDBOX_SESSION_SECRET, SPRITES_API_TOKEN; activity op includes code_execution; distributed limits add CODE_EXECUTION.
Database schema and schema exports
packages/db/drizzle/0144_zippy_kid_colt.sql, packages/db/drizzle/meta/_journal.json, packages/db/package.json, packages/db/src/schema.ts, packages/db/src/schema/sandbox-sessions.ts
Adds sandbox_sessions migration, journal entry, package export for ./schema/sandbox-sessions, Drizzle schema module and indexes for conversation/lastActiveAt.
Policies & pure utilities
packages/lib/src/services/sandbox/execution-policy.ts, packages/lib/src/services/sandbox/command-policy.ts, packages/lib/src/services/sandbox/egress.ts, packages/lib/src/services/sandbox/output-limit.ts, packages/lib/src/services/sandbox/sandbox-paths.ts, packages/lib/src/services/sandbox/sandbox-env.ts, packages/lib/src/services/sandbox/sandbox-options.ts
Defines execution profiles (default/minimal), command pre-checks (empty/size/metadata IP block), egress allowlist sanitizer and Sprites policy builder, UTF‑8-safe truncation, sandbox path confinement under /workspace, environment allowlist (NODE_ENV), and mapping from policy→sandbox options.
Sandbox client & Sprites driver
packages/lib/src/services/sandbox/sandbox-client/types.ts, packages/lib/src/services/sandbox/sandbox-client/sprites.ts, packages/lib/src/services/sandbox/sandbox-client/__tests__/sprites.test.ts
Introduces ExecutableSandbox and ExecSandboxClient contracts; Sprites client implements run/write/read with output caps, timeouts (SIGKILL), egress lockdown, and vanished-sprite detection.
Session key & store
packages/lib/src/services/sandbox/session-key.ts, packages/lib/src/services/sandbox/session-store.ts
Pure HMAC-SHA3-256 session key derivation; DB-backed SandboxSessionStore with findBySessionKey, save (upsert), touch, and remove.
Lifecycle & session manager
packages/lib/src/services/sandbox/lifecycle.ts, packages/lib/src/services/sandbox/session-manager.ts, packages/lib/src/services/sandbox/__tests__/session-manager.test.ts
Pure lifecycle planner and orchestrator implementing create/resume/teardown/deny/noop behaviors, idle expiration handling, best-effort cleanup, and idempotent non-throwing teardown.
Quota & concurrency
packages/lib/src/services/sandbox/quota.ts, packages/lib/src/services/sandbox/__tests__/quota.test.ts, packages/lib/src/security/distributed-rate-limit.ts
Per-user concurrency slots scaled by subscription tier, advisory preflight reading distributed rate-limit status across user/drive/tenant scopes, and budget charging that increments per-scope windows once per run.
Authorization & gate
packages/lib/src/services/sandbox/can-run-code.ts, packages/lib/src/services/sandbox/tool-gate.ts, tests
Kill-switch isCodeExecutionEnabled, canRunCode authorization with explicit denial reasons, and gateSandboxToolCall that enforces kill-switch → authorize → quota in sequence.
Tool runners (bash/write/read)
packages/lib/src/services/sandbox/tool-runners.ts, packages/lib/src/services/sandbox/__tests__/tool-runners.test.ts
Implements runBashInSandbox, writeSandboxFile, readSandboxFile: policy validation, quota preflight + slot reservation, session open, sandbox delegation, output truncation, anomaly auditing, and guaranteed slot release.
Audit pipeline
packages/lib/src/services/sandbox/audit.ts, packages/lib/src/services/sandbox/__tests__/audit.test.ts
Redaction and truncation of code, immutable audit record construction, activity log input mapping to code_execution, conditional security events for anomalies, and resilient sink writes via Promise.allSettled.
Web integration and tests
apps/web/src/lib/ai/tools/sandbox-tools.ts, apps/web/src/lib/ai/tools/sandbox-tools-runtime.ts, apps/web/src/lib/ai/core/ai-tools.ts, tests
createSandboxTools() factory and runtime wiring (buildRealSandboxRunDeps, resolveSandboxActorContext) and buildPageSpaceTools() conditional registration of code-execution tools gated by CODE_EXECUTION_ENABLED.

🎯 4 (Complex) | ⏱️ ~60 minutes

  • Possibly related issue:

    • #1491: Node>=24 runtime assertion and dynamic import for the Sprites driver are directly related to runtime wiring in sandbox-tools-runtime.ts.
  • Possibly related PRs:

"I nibble logs and hop on keys,
I guard the code with fuzzy knees,
Redact the secrets, cap the bytes,
Spin up VMs on moonlit nights,
A rabbit cheers: safe sandbox, please!"

✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch pu/flash-sandbox

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 179a140807

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread packages/lib/src/services/sandbox/sandbox-client/sprites.ts Outdated
Comment thread apps/web/src/lib/ai/tools/sandbox-tools-runtime.ts Outdated
…0142 collision)

Master added 0142_sparkling_maverick + 0143_yielding_praxagora; our branch
added 0142_dizzy_ares (sandbox_sessions). Took master's migration meta as base,
dropped the colliding 0142, kept both schema exports (sandbox-sessions + credits),
and regenerated a fresh migration (0144_zippy_kid_colt) that only adds
sandbox_sessions.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…, full tests (#1489)

* refactor(sandbox): reconcile egress to SDK include:'defaults' preset

Replace the hand-rolled internal-deny domain list (*.internal, *.flycast,
*.tigris.dev, _api.internal) with the SDK's maintained { include: 'defaults' }
PolicyRule preset — the only lever the Sprites network-policy API exposes for
the internal surface. It is prepended before any allow whenever the allowlist
is deliberately widened, so a later misconfiguration that allowed * still could
not reach the internal targets. The v1 empty-allowlist case stays a pure
deny-all and leans on no preset semantics.

Document the known limitation: domain rules cannot block IP-literal egress, so
the empirical 6PN/metadata isolation (gate G1) is a deployment concern in the
enablement checklist, not encoded here.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* refactor(sandbox): reconcile driver to real @fly/sprites SDK surface

Verified the SDK surface against @fly/sprites@0.0.1-rc37 type defs and aligned
the driver to it:

- Hard wall-clock timeout via spawn + kill('SIGKILL'). The SDK's promise-based
  exec/execFile expose no timeout and no abort handle, so the run is driven
  through spawn (same structured file+args[] form, no host shell string) which
  returns a SpriteCommand we can SIGKILL on a timer. We replicate the SDK's own
  execFile stream collection (stdout/stderr data listeners, exit event) and kill
  the command — not the Sprite — so the warm session survives a single slow run.
- maxBuffer: cap buffered stdout+stderr at the policy output cap; an output
  flood SIGKILLs the command and fails the run (host-memory DoS guard).
- Explicit storage cap: add storageGb to the policy and map it onto SpriteConfig
  (ramMB/cpus/storageGB/region) so every Sprite gets explicit caps, not the
  quota default.

A non-zero exit now resolves as a result (spawn's wait surfaces the code) rather
than being recovered from a thrown ExecError, dropping the ExecError duck-typing.
Tested with a non-terminating command (SIGKILL + timeout) and an over-cap output.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* refactor(sandbox): purity — inject validated env + forward policy output cap

buildSandboxEnv was reading getValidatedEnv() in a default param, the one IO in
an otherwise pure module. Make env a required injected argument so the allowlist
construction is fully pure and deterministic; move the getValidatedEnv() read
into defaultBuildEnv (the effect seam in tool-runners), matching the DI pattern
used across the sandbox layer.

Forward the policy output cap to the driver as maxBytes (mapped onto the SDK's
maxBuffer) alongside the existing wall-clock timeout, so the runner bounds both
the run duration and the buffered output from a single place.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* refactor(sandbox): stop retaining command output after the run settles

Guard the spawn output collector so a not-yet-dead command's late stdout/stderr
chunks are dropped once the run has settled (overflow / timeout / exit). Bounds
host memory in the brief window between SIGKILL and the process actually dying
on an untrusted output flood.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(sandbox): enforce output cap before retaining the chunk

The maxBuffer collector pushed each chunk before testing the cap, so a single
oversized stdout/stderr frame was retained in memory before the overflow was
detected. Compute the projected length first and, on overflow, SIGKILL and fail
WITHOUT retaining the offending chunk — buffered memory now never exceeds the
cap. Found in self-review.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 17

🧹 Nitpick comments (5)
packages/lib/src/config/__tests__/env-validation.test.ts (1)

49-76: ⚡ Quick win

Add matching coverage for SPRITES_API_TOKEN.

env-validation.ts gives SPRITES_API_TOKEN the same empty-string/invalid-value contract, and resolveSpritesToken() depends on that fail-closed behavior. Right now this suite only locks down the session-secret half of the change.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@packages/lib/src/config/__tests__/env-validation.test.ts` around lines 49 -
76, Add tests mirroring the SANDBOX_SESSION_SECRET cases for SPRITES_API_TOKEN
in env-validation.test.ts: create one test that supplies SPRITES_API_TOKEN as an
empty string and asserts serverEnvSchema.safeParse returns success, and another
that supplies a too-short non-empty SPRITES_API_TOKEN and asserts safeParse
fails and the validation error issues include 'SPRITES_API_TOKEN'; this ensures
the schema behavior matches resolveSpritesToken and the empty-string fail-closed
contract.
apps/web/src/lib/ai/core/__tests__/ai-tools.test.ts (1)

226-257: ⚡ Quick win

Don’t erase the sandbox tool contract with as never.

Those casts let this test keep compiling even if buildPageSpaceTools or the sandbox tool shape changes, which weakens the registration guard you just added. Prefer a typed stub helper or satisfies against the actual factory return type here. As per coding guidelines, "Always use proper TypeScript types. Never use any types".

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@apps/web/src/lib/ai/core/__tests__/ai-tools.test.ts` around lines 226 - 257,
The tests use unsafe `as never` casts in the sandboxToolsFactory stubs which
breaks the type contract; update the two tests using buildPageSpaceTools to
return a properly typed stub instead of `as never` by either creating a small
typed helper object that matches the sandbox factory's return type (the expected
sandbox tools shape used by buildPageSpaceTools) or by using TypeScript's
`satisfies` operator against that factory return type for the literal objects
(e.g., replace `... }) as never` with `} satisfies ReturnType<typeof
sandboxToolsFactory>` or an explicit typed variable), and remove the `as never`
casts so the compiler will enforce the sandbox tool shape for `bash`,
`writeFile`, and `readFile`.
packages/lib/src/services/sandbox/execution-policy.ts (1)

16-35: ⚡ Quick win

Make ExecutionPolicy readonly to match the frozen runtime contract.

These objects are shared singletons and frozen before being returned, but the exported interface still permits policy.timeoutMs = ... at compile time. Marking the fields readonly keeps consumers from writing code that only fails at runtime. As per coding guidelines, "Always use proper TypeScript types. Never use any types".

♻️ Suggested typing change
 export interface ExecutionPolicy {
   /** Profile this policy represents (echoed back for audit/logging). */
-  profile: ExecutionProfile;
+  readonly profile: ExecutionProfile;
   /** Hard wall-clock cap for a single run, in milliseconds. */
-  timeoutMs: number;
+  readonly timeoutMs: number;
   /** vCPU allocation. */
-  vcpus: number;
+  readonly vcpus: number;
   /** Memory allocation, in megabytes. */
-  memoryMb: number;
+  readonly memoryMb: number;
   /** Disk allocation, in gigabytes. An explicit per-sprite cap, not the quota default. */
-  storageGb: number;
+  readonly storageGb: number;
   /** Maximum stdout/stderr bytes retained before truncation. */
-  maxOutputBytes: number;
+  readonly maxOutputBytes: number;
   /** Egress firewall allowlist. Empty means default-deny (no outbound). */
-  egressAllowlist: readonly string[];
+  readonly egressAllowlist: readonly string[];
   /** Whether the sandbox survives between runs. Always false in v1. */
-  persistent: boolean;
+  readonly persistent: boolean;
   /** Explicit deployment region. */
-  region: string;
+  readonly region: string;
 }
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@packages/lib/src/services/sandbox/execution-policy.ts` around lines 16 - 35,
The ExecutionPolicy interface currently allows mutation at compile time; update
the ExecutionPolicy declaration so every property is readonly (e.g., readonly
profile, readonly timeoutMs, readonly vcpus, readonly memoryMb, readonly
storageGb, readonly maxOutputBytes, readonly egressAllowlist, readonly
persistent, readonly region) to match the runtime frozen singleton contract used
by the sandbox execution policy; keep the existing readonly on egressAllowlist
if present and ensure the interface only exposes immutable fields so consumers
cannot assign to properties like policy.timeoutMs.
packages/lib/src/services/sandbox/__tests__/audit.test.ts (1)

90-103: ⚡ Quick win

Avoid a wall-clock SLA in this unit test.

performance.now() on shared CI runners measures scheduler noise as much as regex complexity, so this can flap even when the matcher stays linear. Keep the adversarial payload coverage, but move the strict timing bound to a benchmark/softer perf smoke test.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@packages/lib/src/services/sandbox/__tests__/audit.test.ts` around lines 90 -
103, The test uses a strict wall-clock assertion with performance.now() and
expect(elapsedMs).toBeLessThan(1000), which flakes on CI; remove the timing
measurement and assertion in the adversarial-key test in audit.test.ts (keep the
payload, secret, buildAuditRecord call and the assertion that record.code does
not contain the secret) and move any strict performance check into a separate
benchmark or softer perf smoke test; update the test that references
buildAuditRecord and record.code accordingly so it only asserts correct
redaction, not elapsed milliseconds.
packages/lib/src/services/sandbox/sandbox-env.ts (1)

31-41: ⚡ Quick win

The return type promises keys this function may omit.

result starts empty and only gets a property when env[key] is a string, but the signature says Record<AllowlistedKey, string>. For partial inputs this can return {} while callers are told every allowlisted key exists. Return Partial<Record<AllowlistedKey, string>> instead, or materialize defaults for all allowlisted keys. As per coding guidelines, "Always use proper TypeScript types. Never use any types".

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@packages/lib/src/services/sandbox/sandbox-env.ts` around lines 31 - 41, The
function buildSandboxEnv currently returns Record<AllowlistedKey, string> but
only populates entries when env[key] is a string, which can violate the return
type; change its signature to return Partial<Record<AllowlistedKey, string>>
(or, if you prefer to keep the full Record type, ensure you materialize defaults
for every key in SANDBOX_ENV_ALLOWLIST) and update the variable declaration of
result and any callers to expect the partial type; locate buildSandboxEnv, the
result variable, SANDBOX_ENV_ALLOWLIST, AllowlistedKey and ServerEnv to apply
the type change (or implement default population) so the TypeScript types
accurately reflect possible missing keys.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In @.github/workflows/security.yml:
- Around line 29-30: The workflow currently lists pu/flash-sandbox only under
pull_request which leaves direct pushes/merges into that branch ungated; update
the workflow triggers so pu/flash-sandbox is included for push events as well
(or remove the exclusion that omits pu/flash-sandbox on push) so Security/CodeQL
runs on direct pushes and merges into pu/flash-sandbox; look for the
pull_request and push sections and ensure the branch name pu/flash-sandbox
appears under push.branches (or eliminate the exclusion in the push trigger) to
enforce the gate.
- Around line 6-9: Update the paths array in .github/workflows/security.yml to
include the web-layer registration files and generated migration artifacts so
the security workflow covers the full sandbox boundary: add patterns for
apps/web/src/lib/ai/tools/sandbox-tools.ts,
apps/web/src/lib/ai/tools/sandbox-tools-runtime.ts,
apps/web/src/lib/ai/core/ai-tools.ts and packages/db/drizzle/** (and mirror the
same additions where lines 31-34 repeat) so PRs changing tool exposure or
emitted migrations trigger the workflow.

In @.github/workflows/test.yml:
- Line 7: The workflow defines the pu/flash-sandbox branch under the
pull_request trigger but not under the push trigger, so direct pushes to that
branch skip the checks; update the .github/workflows/test.yml workflow to add
pu/flash-sandbox to the branches list under the push trigger (mirror the same
branches array used for pull_request) so both push and pull_request triggers
include pu/flash-sandbox and enforce the same gating.

In `@apps/web/src/lib/ai/tools/sandbox-tools-runtime.ts`:
- Around line 44-62: The memoized sandboxClientPromise in getSandboxClient
currently preserves a rejected promise forever; modify getSandboxClient so that
the dynamic import chain (the Promise assigned to sandboxClientPromise) has a
.catch handler that clears sandboxClientPromise = null before rethrowing the
error, ensuring subsequent calls retry; target sandboxClientPromise and
getSandboxClient (and the createSpritesSandboxClient import chain) when adding
the catch-and-rethrow logic.

In `@packages/lib/package.json`:
- Around line 190-279: The package.json's typesVersions mapping is missing
entries for several sandbox subpath exports listed under "exports" (e.g.,
"./services/sandbox/command-policy", "./services/sandbox/egress",
"./services/sandbox/output-limit", "./services/sandbox/sandbox-paths",
"./services/sandbox/sandbox-client/types",
"./services/sandbox/sandbox-client/sprites", "./services/sandbox/tool-runners",
"./services/sandbox/tool-gate"); update the "typesVersions" block to include
each of these subpaths mapping to their respective "./dist/... .d.ts" files
(matching the pattern used for the other sandbox entries) so TypeScript resolves
the correct declaration files for those exports. Ensure keys match the export
names and values point to the corresponding dist/*.d.ts paths used elsewhere in
the file.

In `@packages/lib/src/security/distributed-rate-limit.ts`:
- Around line 585-588: Update the comment block that begins "Agent code
execution daily run budget" to reference the correct provider: replace the
mention of "Vercel account-level Spend Management" with the sandbox driver's
provider ("Fly Sprites") and the appropriate control-plane term (e.g., "Fly
Sprites control plane" or "Fly Sprites cost management"). Locate this comment in
distributed-rate-limit.ts (the block preceding the daily run budget logic) and
update the wording so breadcrumbs point to the Fly Sprites control plane instead
of Vercel.

In `@packages/lib/src/services/sandbox/egress.ts`:
- Around line 45-55: The egress allowlist currently turns any non-empty string
into a Sprites domain allow rule, which allows '*' or non-host values to
short-circuit the deny-all; update buildSpriteNetworkPolicy to validate and
canonicalize egressAllowlist entries before mapping to rules: trim and lowercase
each entry, reject wildcard '*' and any value that is not a literal hostname
(reject IP literals, values containing ':' '/', or URL schemes/paths), and only
map the remaining valid hostnames to { domain, action: 'allow' } (drop invalid
entries and consider emitting a debug/log message for dropped values).

In `@packages/lib/src/services/sandbox/lifecycle.ts`:
- Around line 71-83: Move the idle-expired session reclamation ahead of the
unconditional deny: if an existingSession exists and (now.getTime() -
existingSession.lastActiveAt.getTime()) >= idleTimeoutMs then return the
teardown action (sandboxId and reason 'idle') regardless of authorization,
before checking authorization.ok; keep the original deny return for non-expired
sessions (authorization.ok false) so unauthorized actors still don't receive a
live session. Update the logic around authorization, existingSession,
idleTimeoutMs, idleFor, now and the return branches to reflect this reordered
check.

In `@packages/lib/src/services/sandbox/output-limit.ts`:
- Around line 30-34: truncateToBytes currently slices bytes to maxBytes then
decodes leniently, but the replacement char can produce more bytes than
maxBytes; update truncateToBytes so the returned text's UTF-8 byte length never
exceeds maxBytes by either trimming the byte buffer to a valid UTF-8 boundary
before decoding or by decoding and then re-checking Buffer.byteLength(decoded,
'utf8') and iteratively removing the last Unicode codepoint from decoded until
byte length <= maxBytes; reference the existing symbols cut, decoded, maxBytes,
and originalBytes and ensure truncated=true when trimmed.

In `@packages/lib/src/services/sandbox/quota.ts`:
- Around line 166-185: The current chargeCodeExecutionBudget function uses
Promise.all over budgetScopeIds with deps.charge and can partially apply charges
if some promises fail; fix by making multi-scope charging atomic: either add an
atomic bulk method to ChargeBudgetDeps (e.g., deps.bulkCharge(scopeIds, runId))
and call that from chargeCodeExecutionBudget, or implement sequential charging
with compensating rollback—charge scopes one-by-one, and on any failure call
deps.refund (or deps.uncharge) for all successfully charged scopes and surface
the original error; ensure the ChargeBudgetDeps interface is updated to include
the new bulkCharge or refund method and include an idempotency/run identifier so
retries are safe.

In `@packages/lib/src/services/sandbox/sandbox-client/sprites.ts`:
- Around line 313-320: The getOrCreate function currently returns wrap(await
sdk.getSprite(name)) without reapplying policy options; update getOrCreate to,
when sdk.getSprite(name) succeeds, compare the resumed sprite's effective
policy/caps with the incoming options and then either reapply
lockdownFreshSprite (or call a new helper like lockdownSpriteOnResume) to
enforce mutable limits or invalidate/rotate the session key when immutable caps
differ; use the existing lockdownFreshSprite signature and sdk/session-rotation
logic to ensure resumed sprites are re-locked to the provided options or a new
session is issued when policy versioning changes.
- Around line 316-331: The current catch blocks around sdk.getSprite in
createSpritesSandboxClient.getOrCreate and createSpritesSandboxClient.get
swallow all errors; change each catch to capture the error (e.g., catch (err))
and only handle the not-found condition (inspect err.code / err.status / use the
SDK's NotFoundError predicate) — on 404/not-found proceed to createSprite or
return null, otherwise rethrow the error so auth, rate-limit, and control-plane
errors surface. Ensure you update both the getOrCreate wrapper that calls
sdk.getSprite(name) and the get wrapper that calls sdk.getSprite(sandboxId).

In `@packages/lib/src/services/sandbox/session-key.ts`:
- Around line 44-53: deriveSessionKey currently accepts an empty secret which
yields a predictable HMAC; update deriveSessionKey to validate the secret input
(e.g., ensure it's a non-empty string) and throw a clear error when secret is
missing/empty before calling createHmac, so the function fails closed rather
than producing a predictable digest; reference the deriveSessionKey function and
the createHmac call when adding this guard.

In `@packages/lib/src/services/sandbox/session-manager.ts`:
- Around line 84-102: safeStop currently swallows failures which lets callers
(like provisionFresh and teardownConversationSandbox) remove the DB row even if
SandboxClient.stop failed; change safeStop to surface success/failure (e.g.,
return a boolean or throw a specific error) instead of always swallowing, and
update callers to only call SandboxSessionStore.remove (or safeRemove) when
safeStop indicates a confirmed stop; for failed stops hand the sandboxId to a
durable retry/reaper path rather than deleting the session row. Update all uses
(safeStop and safeRemove spots including the other occurrences referenced around
the provisionFresh and teardownConversationSandbox call sites) to follow this
pattern.

In `@packages/lib/src/services/sandbox/session-store.ts`:
- Around line 88-90: The conflict update on sandboxSessions using
onConflictDoUpdate currently sets { sandboxId, lastActiveAt: now, updatedAt: now
} but omits refreshing userId, so when a row is reused via sessionKey the
actor's userId remains stale; update the onConflictDoUpdate set map to also
assign userId (e.g., userId: newUserId or the function parameter name used where
sandboxId is coming from) so the row's user identity is updated when a session
row is reused, referencing sandboxSessions, sessionKey, userId, and the
onConflictDoUpdate call to locate the change.

In `@packages/lib/src/services/sandbox/tool-runners.ts`:
- Around line 284-289: The branch in the tool runner that resolves cwd (use of
resolvedCwd, resolveSandboxPath and SANDBOX_ROOT) returns fail('path_escape')
when resolveSandboxPath returns falsy but does not create an audit entry, unlike
writeSandboxFile/readSandboxFile which log 'blocked_command'; add an audit/log
call for the blocked sandbox-escape attempt (using the same 'blocked_command'
audit semantics) immediately before returning fail('path_escape') so all denied
cwd escapes are consistently audited.
- Around line 427-457: The current flow reads the entire file into a Buffer via
session.sandbox.readFileToBuffer and only then applies truncateToBytes, which
can OOM; change the sandbox read to enforce policy.maxOutputBytes at the API
boundary (e.g. add/use a readFileToLimitedBuffer or a maxBytes option on
session.sandbox.readFileToBuffer) so the client never returns more than
policy.maxOutputBytes, and ensure decoding respects UTF‑8 boundaries
(truncateToBytes or equivalent should operate on the limited Buffer, not after
full decode); keep the existing safeAudit calls (profile, code, exitCode,
durationMs, anomaly) and error paths but replace the raw full-buffer read in
this block with the bounded read call and then decode/truncate only that bounded
data.

---

Nitpick comments:
In `@apps/web/src/lib/ai/core/__tests__/ai-tools.test.ts`:
- Around line 226-257: The tests use unsafe `as never` casts in the
sandboxToolsFactory stubs which breaks the type contract; update the two tests
using buildPageSpaceTools to return a properly typed stub instead of `as never`
by either creating a small typed helper object that matches the sandbox
factory's return type (the expected sandbox tools shape used by
buildPageSpaceTools) or by using TypeScript's `satisfies` operator against that
factory return type for the literal objects (e.g., replace `... }) as never`
with `} satisfies ReturnType<typeof sandboxToolsFactory>` or an explicit typed
variable), and remove the `as never` casts so the compiler will enforce the
sandbox tool shape for `bash`, `writeFile`, and `readFile`.

In `@packages/lib/src/config/__tests__/env-validation.test.ts`:
- Around line 49-76: Add tests mirroring the SANDBOX_SESSION_SECRET cases for
SPRITES_API_TOKEN in env-validation.test.ts: create one test that supplies
SPRITES_API_TOKEN as an empty string and asserts serverEnvSchema.safeParse
returns success, and another that supplies a too-short non-empty
SPRITES_API_TOKEN and asserts safeParse fails and the validation error issues
include 'SPRITES_API_TOKEN'; this ensures the schema behavior matches
resolveSpritesToken and the empty-string fail-closed contract.

In `@packages/lib/src/services/sandbox/__tests__/audit.test.ts`:
- Around line 90-103: The test uses a strict wall-clock assertion with
performance.now() and expect(elapsedMs).toBeLessThan(1000), which flakes on CI;
remove the timing measurement and assertion in the adversarial-key test in
audit.test.ts (keep the payload, secret, buildAuditRecord call and the assertion
that record.code does not contain the secret) and move any strict performance
check into a separate benchmark or softer perf smoke test; update the test that
references buildAuditRecord and record.code accordingly so it only asserts
correct redaction, not elapsed milliseconds.

In `@packages/lib/src/services/sandbox/execution-policy.ts`:
- Around line 16-35: The ExecutionPolicy interface currently allows mutation at
compile time; update the ExecutionPolicy declaration so every property is
readonly (e.g., readonly profile, readonly timeoutMs, readonly vcpus, readonly
memoryMb, readonly storageGb, readonly maxOutputBytes, readonly egressAllowlist,
readonly persistent, readonly region) to match the runtime frozen singleton
contract used by the sandbox execution policy; keep the existing readonly on
egressAllowlist if present and ensure the interface only exposes immutable
fields so consumers cannot assign to properties like policy.timeoutMs.

In `@packages/lib/src/services/sandbox/sandbox-env.ts`:
- Around line 31-41: The function buildSandboxEnv currently returns
Record<AllowlistedKey, string> but only populates entries when env[key] is a
string, which can violate the return type; change its signature to return
Partial<Record<AllowlistedKey, string>> (or, if you prefer to keep the full
Record type, ensure you materialize defaults for every key in
SANDBOX_ENV_ALLOWLIST) and update the variable declaration of result and any
callers to expect the partial type; locate buildSandboxEnv, the result variable,
SANDBOX_ENV_ALLOWLIST, AllowlistedKey and ServerEnv to apply the type change (or
implement default population) so the TypeScript types accurately reflect
possible missing keys.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro

Run ID: 19100e3b-95f1-4f20-9109-4dbdc5869dbb

📥 Commits

Reviewing files that changed from the base of the PR and between 6a823cc and 50ac066.

⛔ Files ignored due to path filters (1)
  • bun.lock is excluded by !**/*.lock
📒 Files selected for processing (52)
  • .github/workflows/security.yml
  • .github/workflows/test.yml
  • apps/web/src/lib/ai/core/__tests__/ai-tools.test.ts
  • apps/web/src/lib/ai/core/ai-tools.ts
  • apps/web/src/lib/ai/tools/__tests__/sandbox-tools.test.ts
  • apps/web/src/lib/ai/tools/sandbox-tools-runtime.ts
  • apps/web/src/lib/ai/tools/sandbox-tools.ts
  • packages/db/drizzle/0144_zippy_kid_colt.sql
  • packages/db/drizzle/meta/0144_snapshot.json
  • packages/db/drizzle/meta/_journal.json
  • packages/db/package.json
  • packages/db/src/schema.ts
  • packages/db/src/schema/sandbox-sessions.ts
  • packages/lib/package.json
  • packages/lib/src/config/__tests__/env-validation.test.ts
  • packages/lib/src/config/env-validation.ts
  • packages/lib/src/monitoring/activity-logger.ts
  • packages/lib/src/security/distributed-rate-limit.ts
  • packages/lib/src/services/sandbox/__tests__/audit.test.ts
  • packages/lib/src/services/sandbox/__tests__/can-run-code.test.ts
  • packages/lib/src/services/sandbox/__tests__/command-policy.test.ts
  • packages/lib/src/services/sandbox/__tests__/egress.test.ts
  • packages/lib/src/services/sandbox/__tests__/execution-policy.test.ts
  • packages/lib/src/services/sandbox/__tests__/lifecycle.test.ts
  • packages/lib/src/services/sandbox/__tests__/output-limit.test.ts
  • packages/lib/src/services/sandbox/__tests__/quota.test.ts
  • packages/lib/src/services/sandbox/__tests__/sandbox-env.test.ts
  • packages/lib/src/services/sandbox/__tests__/sandbox-options.test.ts
  • packages/lib/src/services/sandbox/__tests__/sandbox-paths.test.ts
  • packages/lib/src/services/sandbox/__tests__/session-key.test.ts
  • packages/lib/src/services/sandbox/__tests__/session-manager.test.ts
  • packages/lib/src/services/sandbox/__tests__/tool-gate.test.ts
  • packages/lib/src/services/sandbox/__tests__/tool-runners.test.ts
  • packages/lib/src/services/sandbox/audit.ts
  • packages/lib/src/services/sandbox/can-run-code.ts
  • packages/lib/src/services/sandbox/command-policy.ts
  • packages/lib/src/services/sandbox/egress.ts
  • packages/lib/src/services/sandbox/execution-policy.ts
  • packages/lib/src/services/sandbox/lifecycle.ts
  • packages/lib/src/services/sandbox/output-limit.ts
  • packages/lib/src/services/sandbox/quota.ts
  • packages/lib/src/services/sandbox/sandbox-client/__tests__/sprites.test.ts
  • packages/lib/src/services/sandbox/sandbox-client/sprites.ts
  • packages/lib/src/services/sandbox/sandbox-client/types.ts
  • packages/lib/src/services/sandbox/sandbox-env.ts
  • packages/lib/src/services/sandbox/sandbox-options.ts
  • packages/lib/src/services/sandbox/sandbox-paths.ts
  • packages/lib/src/services/sandbox/session-key.ts
  • packages/lib/src/services/sandbox/session-manager.ts
  • packages/lib/src/services/sandbox/session-store.ts
  • packages/lib/src/services/sandbox/tool-gate.ts
  • packages/lib/src/services/sandbox/tool-runners.ts

Comment thread .github/workflows/security.yml
Comment thread .github/workflows/security.yml
Comment thread .github/workflows/test.yml
Comment thread apps/web/src/lib/ai/tools/sandbox-tools-runtime.ts
Comment thread packages/lib/package.json
Comment thread packages/lib/src/services/sandbox/session-key.ts
Comment thread packages/lib/src/services/sandbox/session-manager.ts Outdated
Comment thread packages/lib/src/services/sandbox/session-store.ts Outdated
Comment thread packages/lib/src/services/sandbox/tool-runners.ts Outdated
Comment thread packages/lib/src/services/sandbox/tool-runners.ts
2witstudios added a commit that referenced this pull request Jun 2, 2026
…TODO

Backs the interim fail-closed guard with a tracked follow-up (relocate the
@fly/sprites driver to a Node>=24 runtime) per the Codex review on #1487.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…xecution

Hardens the dark-shipped code-execution feature against the review round on
PR #1487 (2 Codex P1 + 17 CodeRabbit threads):

Sprites driver (sprites.ts):
- Re-apply the deny-default egress lockdown on EVERY hand-back (fresh OR
  resumed), closing the crash-window where a Sprite created before lockdown
  could run the next command with open egress; a resumed lockdown failure
  rejects without destroying the warm session.
- Narrow the getSprite fallback to genuine not-found only (new
  isSpriteNotFoundError): auth/rate-limit/outage errors now surface instead of
  masquerading as a vanished Sprite (which could spawn a duplicate / drop a
  healthy session).

Pure-fn correctness:
- egress: sanitizeEgressAllowlist rejects '*', IP literals, and non-host
  strings so a wildcard can't short-circuit the terminating deny.
- output-limit: truncateToBytes is now a HARD byte cap — trims the trailing
  U+FFFD replacement char so the result never exceeds maxBytes.

Boundaries / fail-closed:
- session-key: reject an empty HMAC secret (guessable-name guard, defence in
  depth on top of upstream env validation).
- session-store: refresh userId on the conflict upsert so audit metadata tracks
  the live sandbox's creator after re-provisioning.
- session-manager: safeStop now reports confirmation; teardown removes the
  session link ONLY after a confirmed stop, keeping it on an unconfirmed stop so
  a retry / the idle reaper reclaims the VM instead of orphaning it.
- tool-runners: audit a blocked bash `cwd` path escape (parity with
  writeFile/readFile path-escape auditing).

Web runtime (sandbox-tools-runtime.ts):
- Reset the cached client promise on a lazy-load failure (no poisoned rejection
  until restart).
- Fail closed with an actionable message if the SDK is loaded on Node < 24
  (the @fly/sprites runtime gate), so flipping the flag on a Node 22 image
  surfaces the deployment requirement instead of a cryptic SDK crash.

CI / packaging:
- security.yml + test.yml: gate direct pushes to pu/flash-sandbox and extend the
  security path filters to the web-layer sandbox registration files and
  migrations.
- package.json: complete typesVersions for all 18 sandbox subpath exports.

Documented (in-code) as fail-safe / enablement-gate items rather than changed:
multi-scope budget charge is non-atomic but fail-safe (over-counts, never
under), idle-session reclaim for denied actors is the reaper's job, and a
read-side host-memory cap needs a bounded read at the SDK boundary.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…'*' unreachable)

The allowlist sanitizer now drops '*', so the prior 'even if a later rule
allowed *' example is moot; reframe the preset-first ordering as defence in
depth on top of sanitization.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

♻️ Duplicate comments (1)
packages/lib/src/services/sandbox/quota.ts (1)

166-195: ⚠️ Potential issue | 🟠 Major | 🏗️ Heavy lift

Restore atomic-looking charging instead of documenting partial charges as acceptable.

Promise.all still lets earlier scope increments stick when a later scope rejects, so a run can return an error after consuming user/drive/tenant budget anyway. That breaks the “denied run never charges” contract in this function and can unfairly exhaust quotas until the window rolls over.

Compensating-rollback sketch
 export interface ChargeBudgetDeps {
   /** Increment the daily budget for one scoped identifier (consumes the window). */
   charge: (id: string) => Promise<void>;
+  /** Undo one previously applied charge. */
+  refund: (id: string) => Promise<void>;
 }
@@
 export async function chargeCodeExecutionBudget({
   userId,
   driveId,
   tenantId,
   deps = defaultChargeDeps,
 }: {
@@
 }): Promise<void> {
-  await Promise.all(budgetScopeIds({ userId, driveId, tenantId }).map((id) => deps.charge(id)));
+  const charged: string[] = [];
+  try {
+    for (const id of budgetScopeIds({ userId, driveId, tenantId })) {
+      await deps.charge(id);
+      charged.push(id);
+    }
+  } catch (error) {
+    await Promise.allSettled(charged.map((id) => deps.refund(id)));
+    throw error;
+  }
 }
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@packages/lib/src/services/sandbox/quota.ts` around lines 166 - 195,
chargeCodeExecutionBudget currently uses Promise.all which can leave earlier
scope charges applied if a later charge fails; change it to perform sequential
charges using budgetScopeIds({ userId, driveId, tenantId }) and deps.charge so
you can track successful ids, and on any failure run a compensating rollback
calling a new/extended refund/uncharge method on ChargeBudgetDeps (e.g.,
deps.refund or deps.uncharge) for each already-charged id; ensure rollback is
best-effort (log rollback errors but do not mask the original charge error) and
rethrow the original error so callers see the denial, and update
ChargeBudgetDeps/defaultChargeDeps accordingly to expose the refund/uncharge
operation.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In @.github/workflows/security.yml:
- Around line 10-13: security.yml currently gates sandbox-related paths but
doesn't execute the sandbox Vitest tests; add a job/step in the workflow to
install deps and run the sandbox test suite (e.g., run vitest or the repo's test
script) limited to src/services/sandbox/** tests so changes to
packages/lib/src/services/sandbox/** will actually run
src/services/sandbox/__tests__/egress.test.ts and
src/services/sandbox/__tests__/session-manager.test.ts; ensure the step runs
after checkout and setup-node (or equivalent), uses the repo package manager
(pnpm/npm/yarn) to install, and invokes the test command targeting the sandbox
test patterns.

---

Duplicate comments:
In `@packages/lib/src/services/sandbox/quota.ts`:
- Around line 166-195: chargeCodeExecutionBudget currently uses Promise.all
which can leave earlier scope charges applied if a later charge fails; change it
to perform sequential charges using budgetScopeIds({ userId, driveId, tenantId
}) and deps.charge so you can track successful ids, and on any failure run a
compensating rollback calling a new/extended refund/uncharge method on
ChargeBudgetDeps (e.g., deps.refund or deps.uncharge) for each already-charged
id; ensure rollback is best-effort (log rollback errors but do not mask the
original charge error) and rethrow the original error so callers see the denial,
and update ChargeBudgetDeps/defaultChargeDeps accordingly to expose the
refund/uncharge operation.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro

Run ID: cc1b5b9c-5512-4790-8cac-e6895978d479

📥 Commits

Reviewing files that changed from the base of the PR and between 50ac066 and a697221.

📒 Files selected for processing (20)
  • .github/workflows/security.yml
  • .github/workflows/test.yml
  • apps/web/src/lib/ai/tools/sandbox-tools-runtime.ts
  • packages/lib/package.json
  • packages/lib/src/security/distributed-rate-limit.ts
  • packages/lib/src/services/sandbox/__tests__/egress.test.ts
  • packages/lib/src/services/sandbox/__tests__/output-limit.test.ts
  • packages/lib/src/services/sandbox/__tests__/session-key.test.ts
  • packages/lib/src/services/sandbox/__tests__/session-manager.test.ts
  • packages/lib/src/services/sandbox/__tests__/tool-runners.test.ts
  • packages/lib/src/services/sandbox/egress.ts
  • packages/lib/src/services/sandbox/lifecycle.ts
  • packages/lib/src/services/sandbox/output-limit.ts
  • packages/lib/src/services/sandbox/quota.ts
  • packages/lib/src/services/sandbox/sandbox-client/__tests__/sprites.test.ts
  • packages/lib/src/services/sandbox/sandbox-client/sprites.ts
  • packages/lib/src/services/sandbox/session-key.ts
  • packages/lib/src/services/sandbox/session-manager.ts
  • packages/lib/src/services/sandbox/session-store.ts
  • packages/lib/src/services/sandbox/tool-runners.ts
🚧 Files skipped from review as they are similar to previous changes (15)
  • .github/workflows/test.yml
  • packages/lib/src/services/sandbox/tests/session-key.test.ts
  • packages/lib/src/security/distributed-rate-limit.ts
  • packages/lib/src/services/sandbox/output-limit.ts
  • packages/lib/src/services/sandbox/session-key.ts
  • packages/lib/package.json
  • packages/lib/src/services/sandbox/sandbox-client/tests/sprites.test.ts
  • packages/lib/src/services/sandbox/tests/output-limit.test.ts
  • apps/web/src/lib/ai/tools/sandbox-tools-runtime.ts
  • packages/lib/src/services/sandbox/lifecycle.ts
  • packages/lib/src/services/sandbox/session-store.ts
  • packages/lib/src/services/sandbox/tests/tool-runners.test.ts
  • packages/lib/src/services/sandbox/session-manager.ts
  • packages/lib/src/services/sandbox/sandbox-client/sprites.ts
  • packages/lib/src/services/sandbox/tool-runners.ts

Comment thread .github/workflows/security.yml
security.yml now gates sandbox paths but its job only ran src/security,
src/auth, and named utils — so a sandbox-only change triggered the gate without
executing any sandbox test. Add a step running the full src/services/sandbox
suite (fake-injected, no live Fly/DB) so the security gate actually validates the
code-execution boundary it guards. (CodeRabbit thread.)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2witstudios added a commit that referenced this pull request Jun 2, 2026
Bumps the apps/web Docker image (deps/builder/runner stages) from
node:22.17.0-alpine to node:24.16.0-alpine.

Why: the agent code-execution driver (`@fly/sprites`, behind the default-OFF
flag in PR #1487) is Node 24+/ESM-only. The web process loads it in-process, so
enabling code execution requires the web runtime on Node 24. This is the
prerequisite runtime bump, landed separately from the (dark) feature.

Validated: full production image build on node:24.16.0-alpine succeeds —
bun install, @pagespace/db + @pagespace/lib build, and `next build`
(208/208 static pages) all pass; image runs Node v24.16.0; no engine warnings.
Scoped to web only — processor (TensorFlow native) stays on Node 22.

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
# Conflicts:
#	packages/db/drizzle/meta/0144_snapshot.json
#	packages/db/drizzle/meta/_journal.json
#	packages/db/src/schema.ts
@2witstudios
2witstudios merged commit 8ec55d5 into master Jun 3, 2026
20 checks passed
2witstudios added a commit that referenced this pull request Jun 3, 2026
…llision)

Merging master brought Agent Code Execution's 0145_flawless_living_mummy (#1487),
colliding with our regenerated 0145. Took master's 0145 as canonical and regenerated
the credits delta as 0146_furry_abomination via db:generate. Verified mechanically:
0146.prevId == master 0145.id, 0145.prevId == master 0144.id, journal idx + when
timestamps strictly increasing — chain points at the correct parent. Re-added the
pendingMillicents range CHECK (drizzle-kit doesn't emit CHECK constraints).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2witstudios added a commit that referenced this pull request Jun 3, 2026
…econcile, accounting + e2e (#1484)

* fix(billing): meter all AI provider calls and bill resolved model names

Close revenue leaks where AI provider calls bypassed credit metering or
recorded $0 cost. Billing only happens via AIMonitoring.trackUsage →
consumeCredits with a real userId and a model present in AI_PRICING.

Unmetered call sites — add trackUsage (real userId was already in scope):
- ask_agent (agent-communication-tools): the largest leak — a user-triggerable
  tool loop of up to stepCountIs(20) round-trips, never billed; also reached by
  every channel @mention of an agent. Returns full ProviderResult from
  getConfiguredModel and meters response.totalUsage so every round-trip counts.
- Memory discovery/integration/compaction: run per active user on memory cron,
  on the expensive pro/glm-5 tier; discovery fires 3 passes/run.
- Zoom extract-action-items and generate-summary: per webhook.

Mis-metered ($0) call sites — track the resolved providerResult.modelName
instead of the raw stored model (PageSpace tier aliases 'standard'/'pro' and
the unpriced default 'glm-4.5-air' all hashed to AI_PRICING.default = $0):
- /api/v1/chat/completions (was page.aiModel ?? 'unknown')
- page-agents/consult (was agent.aiModel || 'glm-4.5-air'); also switch to
  result.totalUsage since it is a stepCountIs(100) tool loop
- /api/ai/chat (was raw currentModel)

Catalog↔pricing drift:
- Add 'glm-4.5-air' to AI_PRICING (0.35/1.55, matching z-ai/glm-4.5-air); it was
  selectable via the glm provider but unpriced, so it metered at $0.

Correctly-metered paths (global assistant, pulse generate/cron, workflow
executor) already used providerResult.modelName and are unchanged.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(billing): lock in metering for ask_agent and glm-4.5-air pricing

Regression guards for the leak fixes in this PR:
- ai-monitoring: assert PageSpace-tier backend models (glm-4.5-air, glm-4.7,
  glm-5) all price above $0, and that glm-4.5-air bills at its published rate.
  Catches future catalog↔pricing drift that would meter at $0.
- agent-communication-tools: assert ask_agent bills the requesting user against
  the resolved model name (glm-5) using totalUsage (all tool-loop round-trips),
  proving the previously-unmetered path is now metered.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(billing): make AI-usage persistence durable and stop $0 ledger churn

Two metering-correctness gaps left open by #1475 (audit-leaks), found in the
prepaid AI-credits audit. Both are pure correctness/safety; no billing-policy
change.

1. Durability (was fire-and-forget). `trackAIUsage` detached the
   writeAiUsage -> consumeCredits chain with `.then()` and returned before it
   settled. Callers `await` trackAIUsage from a stream onFinish / post-response
   handler, but a serverless freeze could drop the detached promise — losing
   BOTH the usage log AND the charge. With no aiUsageLogs row, the reconcile
   cron's orphan sweep has nothing to recover from, so the charge is gone for
   good. Now the chain is awaited, so the write is durable before the request
   returns. Still never throws into the AI request.

2. $0 ledger churn. A free/local model — or a tool-only analytics log with no
   tokens (trackAIToolUsage) — produced amountCents 0, yet consumeCredits still
   opened a balance transaction, took the row lock, and ran a $0 decrement,
   writing a misleading "applied/monthly" ledger row per call. In an agent tool
   loop that serialized N no-op locks on the user's balance row. Now a
   zero-charge call settles the claimed row as 'skipped' without the balance
   transaction. The claim row still exists, so the orphan sweep stays idempotent
   and never re-processes it.

Tests: +1 durability test (asserts consume runs before the awaited trackAIUsage
resolves, no setTimeout flush) and +1 zero-charge test (no transaction, row
marked 'skipped'). credit-consume + ai-monitoring suites: 76 passed.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(billing): extend usage-persistence durability to the tool-call path

Addresses Codex review (P2) on #1476: trackAIToolUsage called trackAIUsage
without returning/awaiting it, so a caller that `await`s trackToolUsage resolved
immediately — the durability guarantee didn't reach tool-analytics logs, and the
writeAiUsage / zero-charge ledger settlement could still be dropped on a
serverless freeze after onFinish.

trackAIToolUsage now RETURNS the trackAIUsage promise (no longer an async wrapper
that discards it), so awaiting it waits for the log to persist. +1 test asserting
the returned promise stays pending until writeAiUsage settles.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(billing): schedule reconcile cron, drain backfill, bill errored-but-real AI spend

The "billed exactly once across crashes/deploys" guarantee had three holes:

W4a — reconcile cron unscheduled. /api/cron/reconcile-credits existed and was
HMAC-protected but absent from docker/cron/crontab, so nothing ever ran the
reconciliation. Added a signed GET every 10 minutes using the same cron-curl
pattern as the other ~16 jobs.

W4b — backfill never drained. backfillCredits() did a single LIMIT 200 pending
sweep + single LIMIT 200 orphan sweep; a backlog >200 silently left the rest.
It now loops until a pass returns fewer than BATCH from both sweeps (settled
rows drop out of the next query, so re-querying makes forward progress),
bounded by MAX_PASSES=50 as an unbounded-run backstop. GRACE_MS cutoff and the
isBillingEnabled() guard are unchanged; returns cumulative {retried, orphans}.

R1 — errored-but-real spend was dropped (deliberate billing-policy change).
Tokens consumed before a mid-stream error/abort are real provider cost, but
trackAIUsage only billed when success===true and the orphan sweep filtered
success=true, so an errored generation that produced tokens was logged with a
real cost and billed by neither path. Now:
  - trackAIUsage bills when aiUsageLogId && (success || totalTokens > 0). A
    token-less pre-generation failure still carries 0 tokens and is skipped;
    consumeCredits still settles a zero-charge call as 'skipped' (no $0 churn).
  - the orphan sweep reconciles success:false rows carrying cost > 0, and now
    filters gt(cost, 0) so no/zero-cost rows stay excluded.
The base PR intentionally left failed calls unbilled; the audit owner has
decided errored-but-real spend MUST be billed.

Tests: backfill drains a >200 backlog across passes, stops at the safety cap,
bills a success:false orphan with cost, and asserts the sweep no longer gates
on success; trackAIUsage bills an errored call with tokens but not a token-less
failure. @pagespace/lib typecheck clean; billing + ai-monitoring suites green.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(billing): fund prepaid balances from Stripe (invoice.paid refill + credit-pack top-up)

The pure routing/arithmetic in credit-core (classifyStripeEvent,
computeMonthlyRefill, applyTopup) had zero callers and there was no funding
shell, so a paid invoice or credit-pack purchase never became spendable credit.
The webhook only logged invoice.paid and ignored credit-pack checkouts.

Add credit-funding.ts — an imperative shell that:
  - invoice.paid -> resets the monthly bucket to the tier allowance, rolls the
    billing window forward from the invoice period, and records a monthly_grant
    ledger row keyed on the invoice id.
  - checkout.session.completed (mode=payment, kind=credit_pack) -> adds the pack
    to the never-expiring top-up bucket via applyTopup, recording a topup_purchase
    ledger row keyed on the session id.

Exactly-once: each funding ledger insert uses onConflictDoNothing against the
partial unique index credit_ledger_stripe_ref_unique (predicate restated as the
arbiter), and the balance mutation only runs when that insert actually inserted —
so a redelivered Stripe event credits the balance exactly once. Ledger insert and
balance write share one transaction. Funding never throws into the webhook: a
failure is logged and swallowed; Stripe retry / the reconcile cron re-delivers.
No-op when billing is disabled (tenant/onprem) and for tier_change events (tier
persistence stays in handleSubscriptionChange; the next invoice.paid refills at
the new allowance).

Wire applyStripeFunding into the webhook for invoice.paid and
checkout.session.completed, additively — existing logging and subscription-tier
behavior are unchanged.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(billing): enforce prepaid gate at AI entry points + log usage when metadata missing

Wire canConsumeAI() into every user-facing AI generation entry point so an
out-of-credits user is blocked (HTTP 402) before the model runs, instead of
the platform silently fronting the overage. Also fix the global-messages route
to always write an aiUsageLogs row (R4) so the orphan-sweep can recover/bill
calls where the provider returned no usage metadata.

Entry points gated (402 out_of_credits when !gate.allowed):
- api/ai/chat
- api/ai/global/[id]/messages
- api/v1/chat/completions
- api/ai/page-agents/consult
- api/pulse/generate (on-demand; cron path intentionally not gated)

R4: api/ai/global/[id]/messages now always calls AIMonitoring.trackUsage in
onFinish (0/undefined tokens are fine — $0 cost, but the log row exists).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(billing): make credit-pack top-up funding race-safe; correct funding-failure docs

Self-review of the funding shell surfaced a lost-update on the top-up money path:
two concurrent first-time credit-pack purchases (distinct session ids, so their
ledger inserts don't serialize each other) would both SELECT ... FOR UPDATE on a
not-yet-existent balance row — locking nothing — both read 0, and the second
write would overwrite the first instead of adding to it, silently dropping a paid
top-up.

Fix: inside the funding transaction, ensure the balance row exists first
(INSERT ... ON CONFLICT DO NOTHING), then SELECT ... FOR UPDATE always locks a
real row, making the read-add-write atomic. applyTopup is still the source of the
new value; concurrent purchases now serialize on the row lock and both increments
apply. Add a first-time-buyer regression test alongside the existing add-to-
existing-balance test.

Also corrected the module/function docs: funding swallows its own errors (never
500s) and the webhook's coarse stripeEvents guard blocks same-event reprocessing,
so a failed funding event is not auto-recovered by Stripe retry. The previous
comment overclaimed retry/cron recovery; it now states failures are surfaced via
logs for operator follow-up.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(billing): non-zero reserve floor, shortfall-as-debt, sub-cent remainder, gate-driven monthly reset

Closes four correctness gaps in the prepaid credit core:

R2a — Reserve floor default 0 -> 25¢, bounding the single in-flight call
that can overshoot zero (cost is only known post-stream). Env override kept.

R2b — Shortfall is no longer silently discarded. decrementAndSettle records
appliedCents (what actually left the balance) on the usage row and, when the
balance can't cover the charge, writes the uncovered remainder as a terminal
'adjustment' (debt) row in the same txn — visible, queryable by aiUsageLogId,
recoverable. Balances stay >= 0 (DB CHECK); debt lives in the ledger. The
usage-log unique index is scoped to entryType='usage' so the debt row can
share the call's aiUsageLogId.

R3 — Sub-cent costs no longer round to $0. Charges accrue in millicents into a
per-user pendingMillicents carry; each settle debits floor(pending/1000) whole
cents and banks the remainder. New pure core: chargeMillicents / accruePending
/ accrueCharge. No float ever reaches stored state.

W3-free — canConsumeAI now stamps a monthlyPeriod{Start,End} on lazy-init and,
when the window has expired, resets the monthly bucket to the tier allowance and
rolls the window forward — giving free/no-subscription users a monthly reset
without a cron. The reset UPDATE re-checks expiry in its WHERE so a racing
invoice.paid refill naturally wins.

Schema: + credit_balances.pendingMillicents, + credit_ledger.appliedCents,
+ credit_ledger.chargeMillicents; usage-log unique index scoped to 'usage'.
Migration 0144 generated (not hand-written).

Out of scope (tracked separately): per-user in-flight concurrency cap /
reservation, which requires threading a reservation id through routes+monitoring.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(billing): mark funding ledger rows settled so the backfill cron can't claw them back

Critical: monthly_grant and topup_purchase rows inherited creditLedger's default
consumeStatus 'pending'. backfillCredits() sweeps EVERY pending ledger row through
settlePendingLedgerRow() -> decrementAndSettle(), which SUBTRACTS abs(amountCents)
from the balance (it exists to settle unsettled *usage* charges). A funding row has
a positive amountCents, so after the 5-minute grace period the cron would reverse
every grant/top-up — clawing back exactly the credit funding just added.

Funding applies its balance change in the same transaction as the ledger insert, so
the row is already settled the moment it is written. Insert funding rows with
consumeStatus 'applied' so the pending sweep skips them. Add assertions to the
monthly-refill and top-up tests pinning consumeStatus 'applied'.

Reported by Codex review (P1).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(billing): end-to-end credits-flow integration tests

Drive the real billing shells (applyStripeFunding, canConsumeAI,
consumeCredits, backfillCredits) wired together over one shared
in-memory DB, proving the prepaid money path fund→gate→consume→reconcile
works as a single system. Covers happy path, idempotency (aiUsageLogId +
stripeRef), crash recovery (pending settle, orphan sweep, success:false
billing, >BATCH multi-pass drain), monthly reset, sub-cent accrual,
shortfall/debt, and billing-disabled no-op.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(billing): align gate tests with #1473 (mock createAdminRestrictedResponse + isAdminOnlyProvider)

* fix(billing): await ask_agent metering so the sub-agent charge is durable

Addresses Codex P2: trackAIUsage now persists+debits inside its returned
promise, so the ask_agent call must await it (matching the chat/v1 handlers)
or the sub-agent usage log and credit debit can be dropped on a serverless
return.

* fix(billing): make funding failures retryable via Stripe redelivery

Codex P1 (re-raised): the webhook commits its stripeEvents idempotency marker
before processing, so a swallowed funding failure was lost forever — Stripe's
redelivery short-circuits as "already processed" and the backfill cron reconciles
only usage rows, not funding. A transient DB error during invoice.paid or a
credit-pack checkout could leave a paying customer permanently unfunded.

applyStripeFunding now logs and RE-THROWS genuine failures (non-actionable cases —
billing disabled, ignored events, unknown customer, missing ids — still return
quietly). The webhook wraps funding in fundOrLetStripeRetry: on a funding failure
it deletes the stripeEvents marker and rethrows, so the route returns 500 and
Stripe redelivers, reprocessing the event. Funding is idempotent on
creditLedger.stripeRef, so the balance is still credited exactly once.

Safe to reprocess: the handlers that run before funding on these events are
log-only / no-op (handleInvoicePaid only logs; handleCheckoutCompleted acts only
for mode 'subscription', whereas a credit-pack top-up is mode 'payment').

Tests: rethrow-on-failure assertion replaces the old swallow test; added a
non-actionable-cases test asserting no throw for unknown customer / ignored /
billing-disabled. 55 billing tests pass.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(lib): export ./billing/credit-funding so the web build can resolve it

The Stripe webhook route imports @pagespace/lib/billing/credit-funding, but the
new module was missing from package.json exports + typesVersions. tsconfig-path
typecheck resolved it locally, but the Next build resolves via package exports
and failed collecting /api/stripe/webhook. Verified: web build now exits 0.

* fix(billing): make the whole funding-relevant webhook path retryable, not just funding

Codex P1: the marker-cleanup only wrapped the funding call, but for invoice.paid
handleInvoicePaid runs first and does a DB user lookup. A transient failure THERE
threw before the cleanup, leaving the stripeEvents marker in place — the outer
catch returns 500, and Stripe's redelivery short-circuits at the marker conflict
and never reaches applyStripeFunding, so the paid monthly credit is still lost.

Replace fundOrLetStripeRetry(event) with withFundingRetry(eventId, run): it wraps
the WHOLE funding-relevant case body (pre-funding handler + applyStripeFunding) and
deletes the marker on ANY failure before rethrowing, so Stripe redelivers and the
whole path reruns. Safe to reprocess: handleInvoicePaid only logs;
handleCheckoutCompleted's sole throwable is an idempotent customer-link upsert (and
it acts only for mode 'subscription' — a credit-pack top-up is mode 'payment'; its
provisioning POST swallows its own errors); funding is idempotent on stripeRef.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(billing): gate self-heals a NULL monthly period so top-up-first users get the free allowance

If a credit-pack purchase creates the first credit_balances row (topup funding
inserts { userId } only — monthly 0, monthlyPeriodEnd NULL) before the user's
first AI request, the gate previously skipped both the reset (period was NULL,
not expired) and lazy-init (a row existed), so the free monthly allowance was
never granted and never reset. The gate now resets on (period IS NULL OR
expired), stamping a window and granting the tier allowance. Race-safe: the
UPDATE re-checks the same predicate. +1 test; integration fake-DB engine gains
an 'or' operator.

* fix(billing): address CodeRabbit review (gate ordering, paid-user reset, migration CHECK, test rigor)

- chat route: run the prepaid gate BEFORE persisting the user message, so a 402
  no longer leaves an orphaned/duplicate prompt in chat history on retry.
- credit-gate: restrict gate-driven monthly reset to FREE/non-subscription users.
  Paid tiers refill authoritatively via invoice.paid; gate-resetting them would
  over-grant when a renewal invoice is late or retried. +tests (paid user blocked).
- migration 0144: add the credit_balances_pending_millicents_range CHECK so DBs
  upgraded through this migration enforce the 0<=pending<1000 carry invariant.
- global R4 test: prove trackUsage is AWAITED via a never-resolving deferred
  (a synchronous mock + 'was called' could not catch a fire-and-forget regression).
- integration test: derive the billing window from Date.now() so the funded
  period is always active (a hardcoded past window let the gate refill mid-test).
- consult test: drop the file-wide no-explicit-any disable; type the mock helpers.

* fix(billing): bound drain loop on no-progress, not just MAX_PASSES

Review hardening for the backfill drain loop:

- No-forward-progress break. `decrementAndSettle` (credit-consume.ts) leaves a
  ledger row 'pending' when the user has no balance row yet, so such rows never
  drop out of the pending sweep. The drain loop would therefore re-fetch and
  re-attempt the same unprocessable batch every pass up to MAX_PASSES (50× the
  work, every 10-min cron run, for a balance-less backlog). Now each pass
  fingerprints the fetched rows (order-independent); two identical consecutive
  passes mean nothing settled, so the loop stops instead of churning. A truly
  stuck full batch now ends in 2 passes, not 50. Partial progress still drains
  normally via the short-pass break.

- Hoisted the cap-exhaustion warning out of the loop body. It now fires exactly
  when the loop ran to MAX_PASSES (a real remaining backlog), not on a narrow
  last-iteration batch-size coincidence, and is now covered by tests.

Tests: +1 (stuck full batch stops early, no MAX_PASSES warning) and the cap test
now asserts the warning fires. @pagespace/lib typecheck clean; full lib suite
178 files / 4361 tests green.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
(cherry picked from commit 2e993ec)

* fix(billing): normalize appliedCents to avoid storing -0 on sub-cent settles

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
(cherry picked from commit 34a6cd2)

* fix(billing): clamp month rollover in gate reset to avoid month-end overflow

CodeRabbit P2: setUTCMonth(+1) turns Jan 31 into Mar 3, making the 'monthly'
reset window longer than a month and delaying the next allowance refill for
users initialized/reset near month end. addOneMonth now clamps to the last
valid day of the target month (Jan 31 -> Feb 28/29). Exported + unit-tested
across mid-month, month-end, leap-year, and year-rollover cases.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
(cherry picked from commit 7e69435)

* fix(billing): address credits-remediation review findings (gate ordering, funding tier/customer, expired paid monthly, webhook test)

Five follow-ups from the /aidd:review of pu/credits-remediation:

- [P1] Mock applyStripeFunding in the webhook unit tests. The new funding call
  loaded the real billing module against a mock db that can't satisfy its query
  chain, turning every funding-relevant event into a 500. Mock the module and
  add wiring assertions (funding invoked once per event; 500 is retryable).

- [P1] Resolve a credit-pack buyer from trusted session metadata.userId before
  the customer link. A first-time payment-mode checkout doesn't link the Stripe
  customer to a user, so the customer lookup missed and the top-up was silently
  dropped. resolveTopupUser prefers metadata.userId, falls back to the customer.

- [P2] Run the global-assistant credit gate BEFORE persisting the user message,
  matching the page-chat route. A denied request no longer leaves an orphaned
  prompt that duplicates on top-up + retry.

- [P2] Exclude a paid user's expired monthly bucket from the gate decision. Once
  monthlyPeriodEnd has passed (renewal delayed), leftover monthly allowance no
  longer funds calls — only the never-expiring top-up does (blocked-until-renewal).

- [P2] Base the monthly refill on the tier derived from the PAID invoice line,
  not the stored users.subscriptionTier, so an invoice.paid that races ahead of
  the subscription webhook still grants the correct allowance. applyStripeFunding
  takes an optional { tier }; the webhook derives it via getTierFromPrice.

Tests: +4 lib billing (95 pass), +4 web route (webhook/global gate). lib + web
typecheck clean.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs(marketing): move pricing/FAQ/terms copy to metered AI credits (#1495)

Rewrite every "N AI calls per day" surface in apps/marketing to the
prepaid metered AI-credits story: each tier includes a monthly $ AI-credit
allowance that meters usage, you can buy more anytime via top-up packs,
unused monthly credits reset each billing period, and model access still
differs by tier (free = standard models; paid = standard + Pro models).

Numbers are sourced from packages/lib billing/credit-pricing.ts via a new
single-source-of-truth module (apps/marketing/src/lib/credits.ts) so public
copy can't drift from what the app actually meters.

Surfaces updated: pricing page (cards + comparison table), FAQ (incl. new
"how credits work" entry + reworded out-of-credits answer), Terms (plan
list + usage-limits section), getting-started + features/ai docs, privacy,
schema.org offers, search index, and the BYOK blog post. Storage/file-size
copy left unchanged.

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(billing): AI unit-economics observability (margin queries + admin view) (#1494)

* feat(billing): AI unit-economics observability (margin queries + admin view)

Add margin aggregation queries joining aiUsageLogs with creditLedger usage
rows to surface real provider cost vs charged credits, gross margin %, and
uncovered debt per period, model/provider, and user.

- monitoring-queries.ts: computeMarginPct + getUnitEconomicsSummary,
  getMarginByPeriod, getMarginByModel, getTopSpendersByMargin,
  getOutstandingDebtByUser. Magnitudes via ABS(); debt summed from
  'adjustment' rows by ledger createdAt (no join) so retention purges
  can't under-report. Granularity is a bound param, not interpolated.
- GET /api/admin/unit-economics (withAdminAuth): JSON snapshot + CSV export.
- /admin/unit-economics admin view: summary cards, margin by model, top
  spenders, outstanding debt, margin-over-time; linked from /admin nav.
- Unit tests for margin logic, filters, and entryType scoping (14 tests).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(billing): sum precise sub-cent fields in unit-economics aggregates

Summing per-row-rounded realCostCents/amountCents rounded high-volume
sub-cent traffic to $0 and reported bogus margin. Aggregate the precise
fields and round once: charged from SUM(chargeMillicents)/1000, real cost
from SUM(aiUsageLogs.cost)*100. appliedCents stays exact (whole-cent debit,
remainder banked in pendingMillicents). Debt keeps summing amountCents since
'adjustment' rows carry only the whole-cent shortfall.

Addresses CodeRabbit P2 on PR #1494.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(email): announce metered AI credits to users (template + broadcast script) (#1496)

* feat(email): announce metered AI credits to users (template + broadcast script)

Workstream F of the metered-AI-credits cutover: a one-time announcement
email telling users their AI usage is moving from daily call limits to a
monthly pool of prepaid AI credits (with buy-more top-ups).

- CreditsChangeEmail React Email template matching the existing template
  visual style; states the per-tier monthly allowance + top-up packs and
  reassures that documents/tasks/channels/collaboration are unaffected.
- credits-change-content helper derives all per-tier dollar figures straight
  from billing/credit-pricing (TIER_MONTHLY_ALLOWANCE_CENTS + CREDIT_PACKS)
  so the email can never quote a number the gate doesn't actually grant.
- render-email helper wraps @react-email/components render so repo-root
  scripts can produce email HTML without depending on it directly.
- send-credits-change-notifications.ts broadcast script mirrors
  send-tos-notifications.ts: queries all users with a valid email, sends via
  the shared rate-limited sendEmail, and is idempotent/resumable via a local
  JSONL ledger (re-runs skip already-sent recipients; failures retry).
  Supports --dry-run, --verified-only, --limit, --delay-ms, --log.

Verified end-to-end with --dry-run against a seeded DB: per-tier numbers
render correctly (free $5 / pro $15 / founder $50 / business $100), invalid
emails skip, ledger entries skip, and --verified-only/--limit behave. No
real emails sent. lint, typecheck, build, and lib tests green.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(email): address Codex review on credits broadcast script

- Refuse a live (non --dry-run) send when the resolved app base URL points at
  localhost, so the broadcast can never email a broken CTA. Resolve the URL
  from NEXT_PUBLIC_APP_URL then WEB_APP_URL, preferring the first non-localhost
  value (handles a setup where only the server-side WEB_APP_URL is production).
- Make the idempotency ledger crash-safe: open + validate writability before
  the first send, fsync each record, and treat a ledger-write failure after a
  successful send as fatal — abort and name the unrecorded recipient so a
  re-run can never silently double-send it.
- Drop the unused emailVerified column from the user select.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(billing): credit reservation holds + free in-flight cap; retire daily AI rate limits (#1497)

* fix(billing): credit reservation holds + free in-flight cap; retire daily AI rate limits

Workstream A — reservation/hold + free-tier in-flight cap:
- New credit_holds table (id, userId fk cascade, estCents, aiUsageLogId,
  createdAt, expiresAt; indexed on userId + expiresAt) via db:generate.
- credit-core (pure): reservationCents(), holdExpiresAt(), and evaluateGate
  extended to subtract reservedCents + estCostCents from spendable and to deny
  with a new too_many_in_flight reason when inFlightCount >= maxInFlight.
- credit-pricing: CREDIT_HOLD_ESTIMATE_CENTS (default = reserve floor),
  CREDIT_HOLD_TTL_SECONDS (900), MAX_FREE_INFLIGHT (2), all env-tunable.
- credit-gate canConsumeAI: authoritative decision now runs in one transaction
  that locks the balance row, sums & counts non-expired holds, denies the
  free-tier in-flight cap and out-of-credits, else inserts a hold and returns
  { allowed, holdId }. GateResult gains holdId.
- credit-consume: consumeCredits({…, holdId?}) releases the hold inside the
  settle transaction (and on the zero-charge path); new releaseHold() frees a
  reservation for token-less failures that never bill.
- credit-backfill reconcile: sweeps holds past expiresAt so a crashed stream's
  reservation can't permanently shrink spendable (BackfillResult.expiredHolds).
- holdId threaded gate -> route -> billing: AIUsageData/trackUsage ->
  consumeCredits, across all 5 AI routes (+ agent-communication-tools note).
  Shared credit-gate-response helper maps out_of_credits -> 402,
  too_many_in_flight -> 429.

Workstream B — retire daily AI rate limits:
- chat + global-assistant routes: removed the getCurrentUsage ->
  createRateLimitResponse (429) blocks and the incrementUsage/broadcastUsageEvent
  calls in onFinish. Model-tier gating (requiresProSubscription / admin-only
  providers) preserved. usage-service / rate-limit-cache / rate_limit_buckets /
  sweep-expired left intact — confirmed they back auth/login/integration limits.

Tests: extended billing unit + integration suites (hold accounting, in-flight
cap, hold release on settle, expiry sweep, full gate->consume hold lifecycle)
and added route 429 + helper coverage. lib 4426 + web routes green; typecheck,
lint, web build all pass.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(billing): address Codex review — settle-time monthly expiry + release holds on early route exits

P2 (credit-gate.ts): use-it-or-lose-it now holds at SETTLE too. decrementAndSettle
excludes an expired monthly window (paid user past monthlyPeriodEnd) so allocateSpend
no longer silently draws the forfeited monthly allowance — it spends top-up only and
drops the stale monthly, matching the gate's exclusion.

P2 (global/[id]/messages + chat + consult + pulse routes): release the credit hold on
pre-generation early returns/throws. A holdHandedOff flag + finally frees the reservation
whenever the request exits after the gate but before the stream/billing takes ownership
(auth/permission/provider/save failures), instead of stranding it against the user's
balance + in-flight cap until the reconcile sweep. v1 unchanged (no explicit early return
after its gate; a throw falls through to the reconcile backstop like any crashed stream).

Also: export @pagespace/lib/billing/credit-consume (exports + typesVersions) so routes can
import releaseHold; add settle-time expired-monthly unit test.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(db): regenerate credits migration onto master's 0144 (resolve migration-number collision)

Merging master brought in Canvas Publishing's 0144_flawless_ma_gnuci, colliding with
our hand-numbered 0144_big_valkyrie/0145_brave_exodus. Took master's 0144 as canonical
and regenerated a single 0145 capturing the credits schema delta (pendingMillicents,
appliedCents, chargeMillicents, credit_holds, usage-log index rescope) on top of it.
Re-added the pendingMillicents range CHECK by hand — drizzle-kit in this repo doesn't
emit/track CHECK constraints, so the regenerate would otherwise silently drop the
[0,1000) money-path invariant the original migration enforced.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(billing): credit balance UI, buy-credits checkout, and out-of-credits UX (#1499)

* feat(billing): credit balance UI, buy-credits checkout, and out-of-credits UX

Workstream C + the in-app (apps/web) copy of Workstream D for the metered
AI-credits cutover. Gives users the surfaces that make the hard cutover humane:
see their balance, get correct 402/429 messages with a CTA, and buy more credits.

- Balance API: GET /api/credits → { monthly, topup, spendable, reserved } with a
  read-only getCreditBalance() that mirrors the gate's window semantics for display
  (free-tier lapsed → full allowance; paid lapsed → 0 monthly; spendable nets holds).
  SWR hook useCreditBalance with live socket updates.
- Live updates: replace the retired daily-quota usage:updated socket event with
  credits:updated (broadcastCreditsEvent + emitCreditsUpdated), emitted after a call
  settles (the two interactive AI routes) and after funding (the Stripe webhook).
- Widget: replace UsageCounter with a CreditBalance header widget (remaining + low
  warning + Buy credits) and a CreditBalanceCard on settings/billing. settings/plan
  + settings/billing now show credits, not aiCalls/day.
- Buy-credits checkout: POST /api/stripe/create-credit-topup mirrors create-subscription
  (mode:'payment', inline price_data from CREDIT_PACKS, metadata.kind='credit_pack'),
  so the existing webhook funds the top-up bucket. BuyCreditsButton in settings + in
  the out-of-credits chat error states.
- Error UX: classifyAIError distinguishes out_of_credits (402) and too_many_in_flight
  (429) with distinct copy; SidebarChatTab + ChatInputArea show a Buy-credits CTA.
- Copy: plans.ts limits move from aiCalls/pro/day to a monthly credit allowance +
  proModels capability, sourced from credit-pricing via a new web credits.ts helper
  (mirrors apps/marketing/src/lib/credits.ts). Retire the orphaned /api/subscriptions/usage.

Tests: balance logic, GET /api/credits, the top-up route, error classification, and
the renamed socket event. Lint + web typecheck + web build green.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(billing): hide BuyCreditsButton on iOS / billing-disabled (Codex P2)

The out-of-credits error CTAs (ChatInputArea, SidebarChatTab) rendered
BuyCreditsButton unconditionally, exposing a Stripe checkout on iOS Capacitor
builds where billing UI must be hidden for App Store compliance. Make
BuyCreditsButton self-hide via useBillingVisibility so every call site —
including the error states — is compliant.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(ai): dedupe users import in v1 completions after merging master

master #1500 (server-side tool execution) and our credit gate both added
`import { users }` to the completions route; the clean text-merge left a
duplicate identifier (TS2300). Removed the redundant import.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(billing): address CodeRabbit review on metered-credits PR

- v1 completions: release the credit hold on stream failure (was leaking the
  reservation until TTL/reconcile, leaving users artificially short).
- chat route: create the conversation row AFTER the credit gate so a denied
  first prompt leaves no orphaned conversation.
- monitoring-queries: anchor the unit-economics window on creditLedger.createdAt
  and LEFT JOIN aiUsageLogs (was inner-join + usage-log timestamp, which dropped
  charged credits/margin once a usage log was retention-purged); bucket
  purged-log rows under 'unknown' model/provider. Test updated to match.
- admin CSV export: neutralize spreadsheet formula injection (=,+,-,@) in
  attacker-controlled name/email cells.
- error classifier: tighten to exact codes/phrases so "context window limit
  exceeded" / generic "ai credits" no longer misroute to rate-limit/buy-credits.
- admin unit-economics page: render period buckets from the server string
  (no Date reparsing) to avoid timezone-shifted / mislabeled month buckets.
- credits route: validate subscriptionTier at runtime instead of casting.
- create-credit-topup: reject malformed/non-object JSON with 400, not 500.
- AiUsageMonitor: only compare ids this monitor is scoped to (page-agent mode
  has no conversationId, so it was dropping every credits:updated event).
- send-credits-change-notifications: count ATTEMPTS against --limit so a
  provider outage can't blow past a canary cap.

Not changed: CodeRabbit's "decouple apps/marketing from @pagespace/lib" — master
already couples it (contact route imports @pagespace/lib/security), so the
standalone premise is stale; the import is the no-drift source of truth. Replied
on-thread.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(billing): CREDITS_ENFORCEMENT_ENABLED kill-switch (dark-launch the gate)

The metered-credits cutover otherwise hard-enforces (live 402/429) the moment it
deploys, on placeholder allowances. Add an env flag so the gate can be dark-launched:
deploy the code in observe-only mode (meter + record real cost/charged credits for the
unit-economics view) and flip blocking on deliberately once the numbers are validated.

- credit-pricing.ts: envBool + isCreditsEnforcementEnabled() (default FALSE), read at
  call time so it toggles via env+redeploy and is settable per-test.
- credit-gate.ts: the gate still does ALL bookkeeping (lazy-init, monthly reset, balance
  read, hold on the allow path); when enforcement is OFF it only overrides a would-be
  denial (out_of_credits / too_many_in_flight) to allowed:'enforcement_disabled'. A
  credit-having user is unchanged (normal allow + hold). consumeCredits is untouched, so
  metering/observability run regardless.
- credit-core.ts: add 'enforcement_disabled' to GateReason (an allowed reason;
  credit-gate-response only maps deny reasons, so no HTTP change).
- tests: the two suites that exercise real enforcement set CREDITS_ENFORCEMENT_ENABLED=true;
  new dark-launch cases assert denials are suppressed while bookkeeping still runs.

To enforce in production: set CREDITS_ENFORCEMENT_ENABLED=true and redeploy.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(db): rebase credits migration onto master's 0145 (resolve 0145 collision)

Merging master brought Agent Code Execution's 0145_flawless_living_mummy (#1487),
colliding with our regenerated 0145. Took master's 0145 as canonical and regenerated
the credits delta as 0146_furry_abomination via db:generate. Verified mechanically:
0146.prevId == master 0145.id, 0145.prevId == master 0144.id, journal idx + when
timestamps strictly increasing — chain points at the correct parent. Re-added the
pendingMillicents range CHECK (drizzle-kit doesn't emit CHECK constraints).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(billing): exempt numeric cells from CSV spreadsheet-injection guard

sanitizeSpreadsheetCell ran on every cell, so legit negative exports (marginUSD
"-0.37", marginPct "-12.50") got quote-prefixed and landed as text in Excel/Sheets.
Exempt plain numbers (/^-?\d+(\.\d+)?$/); keep quoting only non-numeric text that
starts with a formula trigger (=,+,-,@), i.e. the name/email columns.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@2witstudios
2witstudios deleted the pu/flash-sandbox branch June 5, 2026 18:49

This branch was previously deployed

1 inactive deployment
Preview — a614f9f7 Deployed Jun 2, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant