Repository navigation
ci(preview): PR-branch lifecycle overhaul — in-place recycle, Electric hardening, orphan sweep - #255
Merged
Merged
Conversation
…c hardening, orphan sweep Extracts the CI/CD work from feat/slack-agent (slack changes excluded): PlanetScale (scripts/planetscale-pr-branch.ts): - Reuse the pr-<n> branch across deploys with an in-place SQL reset (packages/db/scripts/reset-preview-branch.ts) instead of the ~9-minute delete → recreate; fall back to delete → recreate if the reset fails - Ownership normalization (packages/db/scripts/normalize-preview-ownership.ts, run after migrate + grant): REASSIGN OWNED BY CURRENT_USER TO postgres — pscale_api_* roles are never grantable, so the next run's role can only own the prior run's objects through its inherited postgres membership - Two-role split: the main CI role stays non-replication; Electric gets its own --with-replication role (MAPLE_PG_ELECTRIC_URL) - Electric-required cluster parameters via branch resize, with ALTER SYSTEM fallback when the token lacks resize permission; 15-min ready-poll budget on top of the CLI's own create --wait timeout - `up` refuses to provision when the PR is already closed Electric Cloud (scripts/electric-pr-branch.ts): - Reuse the pr-<n> environment (deletion soft-reserves the name), reset its services each deploy; suffix fallback when the base name is stuck - Own activation polling (CLI --wait caps at 300s and discards state), replication-attribute + logical-slot preflight probes with actionable errors - Mirror the synced-table list into the Cloud source's own publication (manual table publishing) after reassigning synced tables to postgres - Reassign the cloud publication to postgres so the next in-place reset can drop it - sslmode=require postgresql:// URL shaping for Electric's validation Sweep (new .github/workflows/cleanup-preview-orphans.yml): - `sweep` subcommand on all three branch scripts (PlanetScale, Electric, Tinybird) deleting pr-* resources whose PR is closed; scheduled workflow as the safety net for close-event teardowns GitHub never runs (conflicted PRs, stale workflow versions, post-close redeploys) Verified on the feat/slack-agent preview: run 30055008127 recycled the branch in 9 seconds (previously every deploy took the ~10-minute recreate path). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Contributor
|
Your Pullfrog Router balance is exhausted. You have a card on file but auto-reload is disabled, so runs paused once your balance went past the overdraft buffer. Top up balance → · Enable auto-reload →
|
From the PR #255 code + security reviews: - reset-preview-branch.ts: refuse to run outside CI without RESET_PREVIEW_CONFIRM=1 — the script empties whatever DATABASE_URL points at and the connected role inherits postgres, so a stray prod/stg URL in a developer shell must not be enough to gut a real database - reset-preview-branch.ts: drop migration-created extra schemas (owned by postgres post-normalization; vendor/system schemas untouched) so a replayed CREATE SCHEMA migration can't wedge every subsequent deploy with no recreate fallback - reset-preview-branch.ts: run the inactive-replication-slot sweep over the replication-role connection (REPLICATION_DATABASE_URL) — the main role deliberately lacks the REPLICATION attribute, so the old sweep could never actually drop a slot; warn loudly on failure (a stale slot pins WAL) - planetscale-pr-branch.ts: mint the Electric replication credential before the reset and pass it through; mask the full exported connection URLs, not just the raw passwords (encodeURIComponent can defeat raw-password masking) - planetscale-pr-branch.ts: drop the dead ALTER SYSTEM fallback (equally permission-denied on PlanetScale) and downgrade the params-convergence timeout from a deploy-failing error to a warning — Electric activates without the params (run 30055008127) - electric-pr-branch.ts: append sslmode=require when the URL never carried an sslmode param (previously only replaced an existing one); fail fast when `services create` returns no service id instead of dying later with a misleading get-secret error - all three sweeps: hard-fail when GITHUB_REPOSITORY/GITHUB_TOKEN are absent — a token-less sweep resolves every PR to "unknown", skips everything, and green-no-ops forever, which is the exact failure class the safety net exists to catch - comment fixes: nonexistent mintReplicationCredential reference, stale MAPLE_PG_URL mention in the Electric workflow step, single-repo naming caveat on the sweep workflow Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
JeremyFunk
added a commit
that referenced
this pull request
Jul 24, 2026
From the PR #255 code + security reviews: - reset-preview-branch.ts: refuse to run outside CI without RESET_PREVIEW_CONFIRM=1 — the script empties whatever DATABASE_URL points at and the connected role inherits postgres, so a stray prod/stg URL in a developer shell must not be enough to gut a real database - reset-preview-branch.ts: drop migration-created extra schemas (owned by postgres post-normalization; vendor/system schemas untouched) so a replayed CREATE SCHEMA migration can't wedge every subsequent deploy with no recreate fallback - reset-preview-branch.ts: run the inactive-replication-slot sweep over the replication-role connection (REPLICATION_DATABASE_URL) — the main role deliberately lacks the REPLICATION attribute, so the old sweep could never actually drop a slot; warn loudly on failure (a stale slot pins WAL) - planetscale-pr-branch.ts: mint the Electric replication credential before the reset and pass it through; mask the full exported connection URLs, not just the raw passwords (encodeURIComponent can defeat raw-password masking) - planetscale-pr-branch.ts: drop the dead ALTER SYSTEM fallback (equally permission-denied on PlanetScale) and downgrade the params-convergence timeout from a deploy-failing error to a warning — Electric activates without the params (run 30055008127) - electric-pr-branch.ts: append sslmode=require when the URL never carried an sslmode param (previously only replaced an existing one); fail fast when `services create` returns no service id instead of dying later with a misleading get-secret error - all three sweeps: hard-fail when GITHUB_REPOSITORY/GITHUB_TOKEN are absent — a token-less sweep resolves every PR to "unknown", skips everything, and green-no-ops forever, which is the exact failure class the safety net exists to catch - comment fixes: nonexistent mintReplicationCredential reference, stale MAPLE_PG_URL mention in the Electric workflow step, single-repo naming caveat on the sweep workflow Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The scheduled orphan-sweep safety net now lives on its own branch (ci/preview-orphan-sweep, stacked on this one) so this PR stays focused on the branch recycle + Electric hardening. Moved out: - .github/workflows/cleanup-preview-orphans.yml (the scheduled workflow) - the `sweep` subcommand on scripts/planetscale-pr-branch.ts (sweepOrphanBranches + parseArgs/usage/dispatch + header paragraph) - the `sweep` subcommand on scripts/electric-pr-branch.ts (sweep + its sweep-only fetchPrState helper + parseArgs/usage/dispatch + header paragraph) - the entire sweep addition to scripts/tinybird-pr-branch.ts, which was that file's only change — it is now byte-identical to main and drops out of this PR fetchPrState STAYS in scripts/planetscale-pr-branch.ts: the `up` closed-PR guard uses it to refuse provisioning a branch for an already-closed PR. Incidental comments that used "sweep" as a generic verb (reset-preview-branch.ts drop passes, the inactive-slot cleanup notes) are reworded so nothing here reads as a reference to the removed workflow. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
JeremyFunk
added a commit
that referenced
this pull request
Jul 24, 2026
Electric Cloud now refuses `environments delete --force` while the
environment still holds services ("Cannot delete environment with existing
services. Delete all services first.", run 30081184677 — PR #255's close
teardown). deleteEnvironment() drops the services first, fixing both the
close-event `down` and the orphan sweep, which would have hit the same
error on every env that still had a source attached.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
JeremyFunk
added a commit
that referenced
this pull request
Jul 24, 2026
- normalize-preview-ownership.ts gets the same local-misuse tripwire as the reset script (refuse unless CI or RESET_PREVIEW_CONFIRM=1) — bun auto-loads .env and mise loads .env.local, so a stray real DATABASE_URL in a dev shell must not silently rewrite ownership. - reset-preview-branch.ts emptiness verification now also counts routines and standalone enum/domain types: a survivor of those classes previously sailed past the check and hit duplicate_object at migrate replay, where no recreate fallback exists. - RESET_PRESERVE_SCHEMAS (comma-separated) exempts named schemas from the extra-schema drop — the ops lever for a future postgres-owned PlanetScale vendor schema. - electric-pr-branch.ts: resolve sourceId via `?? fail(...)` so it is typed string outright, and drop the dead `if (sourceId)` wrapper and re-checks that guarded already-fatal states. Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
JeremyFunk
added a commit
that referenced
this pull request
Jul 24, 2026
* ci(preview): scheduled orphan sweep for PR-preview resources Re-adds the orphan-sweep safety net split out of the branch-lifecycle PR: the `sweep` subcommand on the PlanetScale/Electric/Tinybird lifecycle scripts and the scheduled cleanup-preview-orphans workflow that runs them. Content is identical to the implementation as reviewed on the lifecycle branch (2017c6c), including the hard fail when GITHUB_TOKEN is missing. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * ci(preview): sweep orphaned Hyperdrive PR configs The account caps Hyperdrive configs at 25, and leaked maple-db-pr-* configs from closed PRs eventually block every new PR preview ("This account has reached its limit for how many Hyperdrives it can have (25)", run 30077724144). Unlike the other preview resources, Hyperdrive teardown is owned solely by `alchemy destroy` in the close-event run — which executes the PR branch's OWN workflow version, so branches predating the placeholder-MAPLE_PG_URL destroy fix fail with "Missing required deployment env: MAPLE_PG_URL" on every rerun, forever (runs 29737841130 et al., attempt 2). Those orphans are unreachable by rerunning; only a sweep from current main code can collect them. scripts/hyperdrive-orphan-sweep.ts mirrors the sibling sweeps' double gate: exact `maple-db-pr-<digits>` name match (maple-prd / maple-db-stg / maple-db-dev-* can never match) AND the GitHub API affirming the PR is closed; unknown/open → keep; missing token → hard fail. Wired into cleanup-preview-orphans.yml gated on CLOUDFLARE_API_TOKEN. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * ci(preview): delete Electric services before the environment Electric Cloud now refuses `environments delete --force` while the environment still holds services ("Cannot delete environment with existing services. Delete all services first.", run 30081184677 — PR #255's close teardown). deleteEnvironment() drops the services first, fixing both the close-event `down` and the orphan sweep, which would have hit the same error on every env that still had a source attached. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * ci(preview): route the Electric sweep through the services-first delete Adversarial review of the post-review commits caught that the sweep path deleted environments via a raw `environments delete --force`, bypassing the services-first fix — so the canonical orphan (an env leaked from a failed close-teardown with its source still attached) would have failed with "Cannot delete environment with existing services" on every scheduled tick. Drop the services in the sweep loop too. Also two hardening nits from the same review: fail loudly when the Hyperdrive list response shape is unexpected instead of green-no-op'ing the safety net, and gate the Hyperdrive sweep step on CLOUDFLARE_ACCOUNT_ID as well as the token so a half-configured Infisical env skips cleanly. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * ci(preview): surface the failing statement when the in-place reset dies The reset hit a nondeterministic server-side "malformed array literal: ''" for "unnamed portal parameter $2" (run 30083630959) — no query this script sends carries two parameters, and the same schema reset cleanly minutes earlier, so the suspect is PlanetScale-side DDL interception. The server's error report doesn't include the statement; postgres.js errors do. Print query + parameters on the way out so the next occurrence identifies itself. The delete → recreate fallback already absorbs the failure, so this only costs log lines. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * ci(preview): exercise the recycle path with the reset diagnostic in place Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * ci(preview): inline the pg_% pattern — bind params trip the reset on PlanetScale The recurring in-place-reset failure (malformed array literal: "" for "unnamed portal parameter $2", 3 consecutive runs on pr-257) is parameter-related: the diagnostic shows the failing catalog query was sent with parameters ["pg\_%", ""] — a phantom empty second parameter on a query that interpolates exactly one value. It does not reproduce against plain Postgres (3/3 clean resets locally with identical schema), only against PlanetScale over the direct branch connection. Sidestep whatever mangles extended-protocol parameters there by inlining the LIKE pattern as a SQL literal — the owners and extra-schemas queries now carry zero bind parameters. Verified locally end-to-end. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
JeremyFunk
added a commit
that referenced
this pull request
Jul 28, 2026
…ed Slack agent (#246) * feat(slack-agent): eve-based Slack agent with Railway self-deploy Adds apps/slack-agent, a standalone eve-framework Slack agent that deploys itself to Railway (Docker) rather than Cloudflare Workers: - Workers AI (REST) as the model backend, self-managed Slack app (bot token + signing secret) instead of a Connect integration - @workflow/world-postgres for durable runs; EVE_WORKFLOW_WORLD is baked in at build time via the Dockerfile - excluded from the bun workspace ("!apps/slack-agent") so its own bun.lock and eve toolchain resolve independently Also adds scripts/ingest-dummy.ts for pushing dummy OTLP traces/logs at the local ingest gateway, drops turbo concurrency to 15, and gitignores the .eve model-catalog cache. * fix(slack-agent): use a model that streams structured tool calls The agent was posting raw tool-call JSON into Slack: {"type": "function", "name": "ask_question", "parameters": {…}} Not a formatting bug — a leaked tool call. @cf/meta/llama-3.3-70b-instruct-fp8-fast only parses tool calls on non-streaming requests; eve's harness always streams, and in streaming mode Workers AI returns that model's raw tool-call JSON as ordinary `response` text deltas, which eve has no reason to treat as anything but assistant text. (Even non-streaming it stringifies non-string args: "allowFreeform": "true".) workers-ai-provider has a salvage path for leaked tool calls, but it's gated on a forced tool choice — eve uses auto, so it never engages. Switch to @cf/zai-org/glm-5.2, which streams OpenAI-shaped incremental delta.tool_calls (name + id first, argument fragments keyed by index after) ending in finish_reason: "tool_calls", and emits chain-of-thought on reasoning_content so it maps to reasoning parts instead of message text. Bump the declared context window to its 256K. Verified through a live eve session: modelId workersai/@cf/zai-org/glm-5.2, actions.requested → action.result, correct Tokyo time in prose. Also documents the streaming-tool-call constraint, a curl to check any replacement model's SSE shape, and the price trade-off vs gpt-oss-120b (also verified good, and ~5x cheaper on output if spend beats capability). Includes the pending PORT 3000 -> 8080 alignment with the documented Railway setup. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * Implement basic slack agent * chore: include alerting worker in dev:essentials Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Refactor * fix(mcp): emit object-typed inputSchema for no-param tools An empty Schema.Struct({}) compiles to `{ anyOf: [{type:object},{type:array}] }` in effect 4.0.0-beta.93, with no top-level `type: "object"`. Strict MCP clients (the Vercel AI SDK used by the eve Slack agent) validate each tool's inputSchema.type against z.literal("object"), so list_source_repositories (tool #40) fails to parse and aborts the ENTIRE tools/list response — dropping every Maple tool from the connection. Normalize a no-property struct to an explicit empty object schema centrally in toInputSchema so all current and future no-param tools stay MCP-compliant. Add a test asserting every registered tool emits inputSchema.type === "object". Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(web): purge foreign Clerk instance cookies to stop cross-env 403s All deployed stages share the maple.dev registrable domain, but PR previews/staging run the dev Clerk instance while prod runs the production one. Both write __client_uat cookies on Domain=maple.dev, so browsing multiple environments makes the instances overwrite each other's session hints and ClerkJS handshakes against the wrong frontend API — surfacing as transient 403s on app-pr-N/api-pr-N and prod. Purge foreign-instance parent-domain __client_uat* cookies (suffix derived from the publishable key, mirroring @clerk/shared) before ClerkJS initializes. Mitigation until preview/staging move to their own registrable domain, per Clerk's same-domain limitation. Also document the channels:history Slack bot scope in the slack-agent README. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(electric-sync): 503 on half-configured Electric Cloud creds The shape proxy only treated a missing ELECTRIC_URL as "not configured". A deploy with ELECTRIC_URL + ELECTRIC_SOURCE_ID but no ELECTRIC_SECRET (e.g. a PR preview inheriting shared Infisical Electric config while its per-PR source step is skipped) forwarded an unauthenticated request to Electric Cloud, which returns 401 MISSING_SECRET — surfacing as a hard- broken shape stream in the browser instead of graceful degradation. Treat incoherent Cloud credentials (exactly one of source_id/secret) as not-configured and take the existing 503 path, and log the misconfig for telemetry. Adds isElectricConfigCoherent + unit tests. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * Refactors & improvements * Parity with chat-flue * Fix deployment * Fix slack deployment * ci(preview): reuse PlanetScale PR branch with in-place SQL reset A PS-DEV Postgres branch takes ~9 min to provision, and the preview deploy paid that on every push by deleting + recreating pr-<n>. The branch now lives for the PR's lifetime: each deploy resets it in SQL (assume prior owner roles, drop publications, drop the drizzle schema, empty public per-object, sweep inactive replication slots) and falls back to delete -> recreate if the reset fails. Also treat `pscale branch create --wait`'s own ~10-min timeout as non-fatal (the branch keeps provisioning server-side) and keep polling until ready. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * ci(preview): raise PlanetScale ready-poll budget to 15 min Provisioning has been observed to exceed the 10-minute budget; the create path's total allowance is now ~10 min (CLI --wait) + 15 min (our branch-show poll). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * ci(preview): match electric-pr-branch.ts to the real @electric-sql/cli 0.0.10 The Electric step ran for the first time (token landed in Infisical dev) and died on `environments create --json returned no environment id`: the script was written against an assumed CLI. Verified the actual interface from the CLI's bundled source and fixed: - `environments create --json` returns { environmentId }, not { id } — the root cause of the CI failure. - `--publication <name>` does not exist; the real flag is `--manual-table-publishing` (Electric's default publication name is the migration-owned `electric_publication_default`). Pass it by default (prod parity), opt out via ELECTRIC_MANUAL_TABLE_PUBLISHING=false; ELECTRIC_PUBLICATION is gone. - Add `--wait` so the source is active before the alchemy deploy binds it. - The postgres service id IS the shape-API source_id; secret comes from `sourceSecret` (fallback: `services get-secret` → { secret }). - Not-found/conflict detection now uses the CLI's typed exit codes (3/5) with text matching as fallback; JSON errors arrive on stderr. - Drop the unscoped `environments list` fallback (--project is required) and pin the CLI to @electric-sql/cli@0.0.10 (interface-verified). Redaction, up/down contract, and the GITHUB_ENV exports are unchanged. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * ci(preview): retry Electric env create while the deleted name is still reserved Environment deletion is async server-side: the env leaves `environments list` immediately, but create then transiently fails with VALIDATION_ERROR "an environment with this name already exists" (seen on run 30030020767). Retry name conflicts with the existing 5s/2min backoff budget instead of failing. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * ci(preview): reuse the Electric PR environment; reset services instead Electric soft-deletes environments: a deleted env leaves `environments list` immediately but its name stays reserved (observed >13 min on runs 30030020767/30031075727), so the delete-then-recreate reset loops forever on "name already exists". Rework the lifecycle: - `up` reuses the existing `pr-<n>` environment and resets only its services (they are what point at the dropped tables/rotated creds after the PlanetScale in-place reset), creating a fresh postgres source each deploy. - Fresh creates keep a short name-conflict retry, then fall back to a unique `pr-<n>-r<run-id>` name when the base is soft-delete-reserved (the current state of pr-246). Matching is exact-name or `pr-<n>-` prefix everywhere. - `down` deletes all matching environments (base + suffix fallbacks). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * ci(preview): pass Electric a postgresql:// scheme database URL Run 30032252411 got the environment created (suffix fallback worked) but `services create postgres` failed server-side with a generic "Input validation failed". MAPLE_PG_URL uses the short `postgres://` scheme; all Electric examples use `postgresql://` — normalize before passing. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * ci(preview): isolate the Electric create-source validation failure in one run The Cloud API rejects `services create postgres` with a field-less "Input validation failed" (runs 30032252411/30033111516; scheme normalization did not help). Attempt the full-fidelity create first, then degrade the two suspects — the manual-table-publishing option and the ?sslmode query — one at a time, logging which variant is accepted. Non-validation failures still fail immediately. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * ci(preview): hand Electric an sslmode=require connection string The isolation ladder (run 30034103225) showed sslmode=verify-full is rejected at input validation while a bare URL passes input validation but fails the server's database probe ("Database validation failed" — TLS-less connect to PlanetScale). Electric's PlanetScale guide prescribes exactly `postgresql://...?sslmode=require` — normalize to that. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * ci(preview): log the PR branch's replication state before the Electric create sslmode=require got past input validation (run 30035178700) but every variant now fails Electric's database probe. Temporarily print wal_level, slot/sender limits, role replication privilege, and the publication's existence from the branch itself so the next run tells us which prerequisite the PS-DEV branch is missing. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * ci(preview): grant the CI role REPLICATION before creating the Electric source Diagnostic (run 30036098689) pinpointed it: wal_level=logical, slots, and the publication are all in place, but rolreplication=false. The credential role inherits `postgres` via membership, and REPLICATION is a role attribute — attributes are never inherited — so Electric's database probe kept answering "Database validation failed". Grant the attribute in-place (direct ALTER, falling back to SET LOCAL ROLE postgres on a pinned connection), verify it stuck, and fail with a precise message if not. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * ci(preview): mint the PlanetScale CI role with REPLICATION In-place granting is impossible on PlanetScale — both `ALTER ROLE CURRENT_USER WITH REPLICATION` and the SET ROLE postgres fallback get "permission denied to alter role" (run 30036963759). The supported path is minting the role with the attribute: `pscale role create --with-replication` (requires --inherited-roles postgres, already passed). The electric script now only verifies rolreplication and fails with a precise message on regression. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * ci(preview): own the Electric source activation wait Run 30037847577 passed validation (REPLICATION role fixed it) but the CLI's --wait timed out at its hard 300s cap and discarded the service state. Create without --wait, capture id+secret immediately, and poll `services get` with a 10-min budget, logging status transitions and dumping the non-secret service state on error/timeout. The input-validation isolation ladder is gone — the accepted input shape is settled (sslmode=require + manual-table-publishing). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * ci(preview): surface failover-slot prerequisites + slot state around activation Run 30038967625: the source sits in status=pending for the full 10-min budget with no error detail. Electric always creates FAILOVER-enabled slots, which PlanetScale only accepts with sync_replication_slots=on and hot_standby_feedback=on — settings the diagnostic never checked. Log those (plus max_connections) pre-create, and dump pg_replication_slots on activation error/timeout so the step log shows whether Electric's slot ever landed and why it is stuck. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * ci(preview): set Electric-required cluster parameters on the PR branch Run 30040575049 diagnostics: the branch runs sync_replication_slots=off, hot_standby_feedback=off, max_connections=25 — Electric creates its slot but never activates (it needs failover-capable slots and headroom for its 20-connection pool), leaving the source pending forever. Apply pgconf.sync_replication_slots=on, hot_standby_feedback=on, and max_connections=100 via `pscale branch resize --parameters --wait` during branch up, skipping when the live settings already satisfy them, and poll the live settings afterwards (the change can restart the cluster). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * ci(preview): degrade gracefully when the CI token cannot resize the branch `pscale branch resize` is denied for the CI service token (run 30042253492: "User does not have permission to perform this action"). Keep the resize attempt (self-heals once the token gains branch-change access) but fall back to ALTER SYSTEM as `postgres` for the SIGHUP-reloadable settings (sync_replication_slots, hot_standby_feedback) on a pinned single-connection session + pg_reload_conf(), and shrink the Electric source pool to --db-pool-size 5 (ELECTRIC_DB_POOL_SIZE override) so it fits under the branch's max_connections=25, which only a resize can raise. Parameter failures warn instead of failing the deploy. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * ci(preview): probe logical-slot creation before the Electric create The freshly recreated branch draws "Database validation failed" again (run 30042533336) — the old branch had wal_level=logical, a fresh one may not. Log wal_level + publication in the checklist and dry-run pg_create_logical_replication_slot: its error text is the real reason behind Electric's opaque rejection. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * ci(preview): split the Electric credential off the main CI role Every deploy since --with-replication landed lost the in-place branch reset ("permission denied to grant role" on the previous run's role — PlanetScale replication-attribute roles are not grantable) and paid the ~9-min delete → recreate path. Mint the main role WITHOUT replication again (restores role assumption / fast resets) and a second `-repl`-suffixed role WITH replication used only by the Electric source, exported as MAPLE_PG_ELECTRIC_URL (electric-pr-branch.ts prefers it, falling back to MAPLE_PG_URL). Note: the next deploy still recreates once (the current branch owner is a replication role from the previous scheme); resets are fast again after. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Fix MCP list for slack agent * ci(preview): publish synced tables to Electric Cloud's own publication Electric Cloud ignores electric_publication_default: with --manual-table-publishing each Cloud source creates its own publication (cloud_electric_pub_svc_<name>) and refuses to add tables to it, so every shape request 400'd with "missing from the publication". After the source activates, mirror the migration-owned table list into the cloud publication(s) via an idempotent server-side DO block. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Render slack-agent charts in the Maple dark theme Replace the placeholder light-blue palette in agent/lib/chart.ts with the Maple dark-theme tokens (tokens.css .dark, oklch converted to hex for resvg): card surface + hairline border canvas, border/50 grid, muted axis ink, gradient area fills, and per-unit semantic series colors (latency amber, throughput purple, error-rate red, bytes teal, counts blue). Type is now Geist Mono: the Dockerfile decompresses the @fontsource woff2 faces to TTF (resvg reads TTF only) into the system font dir. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Fix blank chart text: pass explicit fontDirs to resvg resvg-js 2.x's loadSystemFonts discovers fonts on Linux via fontconfig's /etc/fonts/fonts.conf; the slim production image has no fontconfig package, so font discovery silently found nothing and every glyph (title, axis labels, latest-value label) was dropped from the PNG — charts shipped as bare lines on a card. Point resvg directly at the font dirs the Dockerfile populates (fontDirs needs no fontconfig), keep loadSystemFonts for local dev, and set the Geist Mono default. Verified in a node:24-slim container against resvg-js 2.6.2. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * ci(preview): assume table-owner roles before publishing to the cloud publication The Electric role owns the cloud publication but not the synced tables (owned by the ephemeral migrate-step pscale role), so ALTER PUBLICATION ADD TABLE failed with "must be owner of table api_keys". Assume each table-owner role via GRANT <owner> TO CURRENT_USER first — the same trick reset-preview-branch.ts uses, backed by the inherited postgres grant's admin over the ephemeral roles. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * ci(preview): reassign synced tables to postgres before cloud-publication mirror GRANT <table-owner> TO the Electric role is a dead end — replication roles can't receive role grants on PlanetScale, so run 30050755374 still failed with "must be owner of table actors" (and the DO block's RAISE NOTICE diagnostics never reach the client). Instead the main CI role now reassigns the synced tables' ownership to postgres (allowed: it owns them and is a member of postgres), which the Electric role also inherits — making it a table owner for the ALTER PUBLICATION mirror. All statements run client-side with per-statement error logging. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Add sweep * ci(preview): normalize branch ownership to postgres so the in-place reset works The reuse path never actually ran: reset-preview-branch.ts tried to GRANT the prior run's pscale_api_* roles to CURRENT_USER, but PlanetScale refuses to grant those roles at all ("permission denied to grant role", run 30051727304), so DROP PUBLICATION electric_publication_default failed with "must be owner" and every deploy fell back to the ~9-minute delete → recreate path. Every pscale role inherits `postgres`, so hand ownership over at the END of each deploy instead: - new packages/db/scripts/normalize-preview-ownership.ts (db:normalize-preview, run after migrate + grant in the workflow) — REASSIGN OWNED BY CURRENT_USER TO postgres as the main role, covering the drizzle schema, all of public, and electric_publication_default - electric-pr-branch.ts reassigns the Cloud source's cloud_electric_pub_svc_* publication to postgres after activation (its owner is the replication role) The next run's role then owns everything via its own postgres membership and the reset's drops just work. Both steps are non-fatal: an un-normalized branch only costs the next deploy the recreate fallback, never correctness. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * ci(preview): empty commit to exercise the in-place branch recycle path Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * ci(preview): address review findings on the branch-lifecycle scripts From the PR #255 code + security reviews: - reset-preview-branch.ts: refuse to run outside CI without RESET_PREVIEW_CONFIRM=1 — the script empties whatever DATABASE_URL points at and the connected role inherits postgres, so a stray prod/stg URL in a developer shell must not be enough to gut a real database - reset-preview-branch.ts: drop migration-created extra schemas (owned by postgres post-normalization; vendor/system schemas untouched) so a replayed CREATE SCHEMA migration can't wedge every subsequent deploy with no recreate fallback - reset-preview-branch.ts: run the inactive-replication-slot sweep over the replication-role connection (REPLICATION_DATABASE_URL) — the main role deliberately lacks the REPLICATION attribute, so the old sweep could never actually drop a slot; warn loudly on failure (a stale slot pins WAL) - planetscale-pr-branch.ts: mint the Electric replication credential before the reset and pass it through; mask the full exported connection URLs, not just the raw passwords (encodeURIComponent can defeat raw-password masking) - planetscale-pr-branch.ts: drop the dead ALTER SYSTEM fallback (equally permission-denied on PlanetScale) and downgrade the params-convergence timeout from a deploy-failing error to a warning — Electric activates without the params (run 30055008127) - electric-pr-branch.ts: append sslmode=require when the URL never carried an sslmode param (previously only replaced an existing one); fail fast when `services create` returns no service id instead of dying later with a misleading get-secret error - all three sweeps: hard-fail when GITHUB_REPOSITORY/GITHUB_TOKEN are absent — a token-less sweep resolves every PR to "unknown", skips everything, and green-no-ops forever, which is the exact failure class the safety net exists to catch - comment fixes: nonexistent mintReplicationCredential reference, stale MAPLE_PG_URL mention in the Electric workflow step, single-repo naming caveat on the sweep workflow Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(alerts): redesign Slack alert messages per Block Kit guidance - Wrap slack-bot blocks in a colored attachment so the severity bar renders (parity with the webhook destination) - Truncate the default header to Slack's 150-char limit (long rule names previously failed delivery with invalid_blocks) - Escape mrkdwn control chars in rule names, group keys, incident ids - Replace the five-field dump with a lead summary sentence + compact fields (emoji-paired capitalized severity; group only when set) - Add a context footer with incident ref and <!date^…> local-time stamp - Full one-line notification fallback text instead of "rule: Triggered" - Button styles per Slack guidance: primary on Open in Maple, no danger on navigation links Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(alerts): align default notification templates with the redesigned format A rule that customizes only the title (or only the body) has its other field filled from the built-in defaults, which still produced the old five-line field dump — so partially-templated rules kept sending the old-style "basic" message. Update DEFAULT_BODY_TEMPLATE and the rule editor placeholder to the lead-sentence format the designed blocks use. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(alerts): stop Slack rendering the fallback text as a duplicate line With attachments and no top-level blocks, chat.postMessage renders the top-level text in-channel above the color bar (confirmed on the PR-246 preview). Move the notification-preview one-liner into the attachment's fallback field, which Slack uses only for push/desktop previews. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Refactor * fix(tinybird): drop completed service_operations_minutely forward query * Add bot follow-up without metnion * Fixes * Remvoe unnecessary tool * Refactors * Format * feat(slack-agent): strip __maple_ui payloads from MCP tool results Port of chat-flue's splitToolResult change from main: Maple's MCP server emits a second content entry per tool result carrying the web chat's structured UI payload (createDualContent, tagged __maple_ui). chat-flue now splits it off client-side; eve has no result-transform hook, so the Slack agent's model was receiving the raw UI JSON duplicated next to the text report on every Maple tool call. Extend the vendored eve patch to drop those entries in McpConnectionClient.executeTool, with a canary test alongside the botToken one. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * merge migrations * Fixes * fix(slack): apply review fixes; move Slack event intake to the Railway bot Review fixes (multi-agent effect-review-v4 + general review): - SlackIntegrationService: probeAndRevokeIfDead no longer counts a failed DB revoke as revoked (logs and returns false); uninstall's revocation update is now a compare-and-set with one retry so a concurrent reinstall can't be clobbered while its fresh API key survives unrevoked; Service.of() construction; Arr.map/Arr.filter over native methods - slack-bot-token: SlackBotTokenResolver.of() construction - integrations.http.test: complete the 8-method die-stub - slack-agent thread-follow-up: bound conversations.replies with oldest/latest so the engagement window is correct for >100-reply threads; single ~2s promotion deadline inside Slack's ack budget - slack-agent maple.ts: type-check resolve payload fields; render_chart: reject non-finite point values Also lands the events-intake refactor: SlackEventsRouter is removed from the API (Slack allows one Events URL per app, already pointed at the bot); the bot detects app_uninstalled/tokens_revoked (uninstall-detection.ts, bot-token-scoped) and calls POST /internal/slack/workspaces/:teamId/revoke, with the 6-hourly reconcile cron as backstop. SLACK_SIGNING_SECRET leaves the API env. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(api): provide SlackIntegrationServiceStubLayer in setup-audit harness AllV2GroupLayersLive now includes HttpV2SlackIntegrationsLive, so every harness composing it must stub SlackIntegrationService; setup-audit was the one file missed. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Increase timeout * docs: move the after-hours compute limit out of the shared CLAUDE.md Machine-local preference, not a project rule — belongs in the owner's CLAUDE.local.md or ~/.claude/CLAUDE.md instead. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(slack): surface remote disconnects on the integration status When the Maple app is removed (or its tokens revoked) from Slack's own UI, the integration card silently reverted to the never-installed state, which read as a Maple bug. Persist why a workspace row was revoked (revoked_reason: uninstalled / superseded / app_uninstalled / tokens_revoked / reconciliation), expose the remote reasons on the v2 status endpoint (disconnected_reason / _team_name / _at), and show a warning banner on the card explaining the disconnect came from Slack's side. Dashboard uninstalls and replacement installs still read as plain not-installed. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(slack): neutral copy for the remote-disconnect banner Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
This branch had an error being deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.

Extracts all CI/CD work from
feat/slack-agent(Slack-agent changes excluded). Seven files, all touched exclusively by theci(preview):commit series.PlanetScale PR branches — in-place recycle
scripts/planetscale-pr-branch.tsreuses thepr-<n>branch across deploys and resets it in SQL (packages/db/scripts/reset-preview-branch.ts: drops the drizzle schema, allpublicobjects, publications, inactive replication slots) instead of the ~9-minute delete → recreate. Reset failure falls back to delete → recreate, so it can only cost time, never correctness.packages/db/scripts/normalize-preview-ownership.ts, new workflow step after migrate + grant):REASSIGN OWNED BY CURRENT_USER TO postgres. PlanetScale never makespscale_api_*roles grantable, so the next run's ephemeral role can only own the prior run's objects through its inheritedpostgresmembership — this is what makes the in-place reset actually work.--with-replicationrole (MAPLE_PG_ELECTRIC_URL).branch resizewith anALTER SYSTEMfallback; 15-min ready-poll budget;uprefuses to provision for already-closed PRs.Electric Cloud PR sources
pr-<n>environment (deletion soft-reserves the name — recreate conflicts), reset its services each deploy; suffixed-name fallback when the base name is stuck.--waitcaps at 300s and discards state), REPLICATION-attribute + logical-slot preflight probes with actionable error text.cloud_electric_pub_svc_*publication after reassigning synced tables topostgres; reassign that publication topostgrestoo so the next reset can drop it.Orphan sweep → stacked PR
The scheduled orphan-sweep safety net (the
sweepsubcommand on the three lifecycle scripts + thecleanup-preview-orphans.ymlworkflow) moved out to the stacked #257 to keep this PR focused; retarget it tomainafter this merges. The Tinybird script's only change here was the sweep, so it dropped out of this PR entirely.Verification
On the
feat/slack-agentpreview, run 30055008127 recycled the branch in 9 seconds (every prior deploy took the ~10-minute recreate path after failing onmust be owner of publication electric_publication_default).🤖 Generated with Claude Code