Skip to content

ci(preview): PR-branch lifecycle overhaul — in-place recycle, Electric hardening, orphan sweep - #255

Merged
JeremyFunk merged 3 commits into
mainfrom
ci/preview-branch-lifecycle
Jul 24, 2026
Merged

JeremyFunk merged 3 commits into
mainfrom
ci/preview-branch-lifecycle

Conversation

@JeremyFunk

@JeremyFunk JeremyFunk commented Jul 24, 2026 •

Copy link
Copy Markdown
Collaborator

Extracts all CI/CD work from feat/slack-agent (Slack-agent changes excluded). Seven files, all touched exclusively by the ci(preview): commit series.

PlanetScale PR branches — in-place recycle

  • scripts/planetscale-pr-branch.ts reuses the pr-<n> branch across deploys and resets it in SQL (packages/db/scripts/reset-preview-branch.ts: drops the drizzle schema, all public objects, publications, inactive replication slots) instead of the ~9-minute delete → recreate. Reset failure falls back to delete → recreate, so it can only cost time, never correctness.
  • Ownership normalization (packages/db/scripts/normalize-preview-ownership.ts, new workflow step after migrate + grant): REASSIGN OWNED BY CURRENT_USER TO postgres. PlanetScale never makes pscale_api_* roles grantable, so the next run's ephemeral role can only own the prior run's objects through its inherited postgres membership — this is what makes the in-place reset actually work.
  • Two-role split: main CI role stays non-replication; Electric gets a dedicated --with-replication role (MAPLE_PG_ELECTRIC_URL).
  • Electric-required cluster parameters via branch resize with an ALTER SYSTEM fallback; 15-min ready-poll budget; up refuses to provision for already-closed PRs.

Electric Cloud PR sources

  • Reuse the pr-<n> environment (deletion soft-reserves the name — recreate conflicts), reset its services each deploy; suffixed-name fallback when the base name is stuck.
  • Own activation polling (the CLI's --wait caps at 300s and discards state), REPLICATION-attribute + logical-slot preflight probes with actionable error text.
  • Manual table publishing: mirror the synced-table list into the Cloud source's own cloud_electric_pub_svc_* publication after reassigning synced tables to postgres; reassign that publication to postgres too so the next reset can drop it.

Orphan sweep → stacked PR

The scheduled orphan-sweep safety net (the sweep subcommand on the three lifecycle scripts + the cleanup-preview-orphans.yml workflow) moved out to the stacked #257 to keep this PR focused; retarget it to main after this merges. The Tinybird script's only change here was the sweep, so it dropped out of this PR entirely.

Verification

On the feat/slack-agent preview, run 30055008127 recycled the branch in 9 seconds (every prior deploy took the ~10-minute recreate path after failing on must be owner of publication electric_publication_default).

🤖 Generated with Claude Code

…c hardening, orphan sweep

Extracts the CI/CD work from feat/slack-agent (slack changes excluded):

PlanetScale (scripts/planetscale-pr-branch.ts):
- Reuse the pr-<n> branch across deploys with an in-place SQL reset
  (packages/db/scripts/reset-preview-branch.ts) instead of the ~9-minute
  delete → recreate; fall back to delete → recreate if the reset fails
- Ownership normalization (packages/db/scripts/normalize-preview-ownership.ts,
  run after migrate + grant): REASSIGN OWNED BY CURRENT_USER TO postgres —
  pscale_api_* roles are never grantable, so the next run's role can only own
  the prior run's objects through its inherited postgres membership
- Two-role split: the main CI role stays non-replication; Electric gets its
  own --with-replication role (MAPLE_PG_ELECTRIC_URL)
- Electric-required cluster parameters via branch resize, with ALTER SYSTEM
  fallback when the token lacks resize permission; 15-min ready-poll budget
  on top of the CLI's own create --wait timeout
- `up` refuses to provision when the PR is already closed

Electric Cloud (scripts/electric-pr-branch.ts):
- Reuse the pr-<n> environment (deletion soft-reserves the name), reset its
  services each deploy; suffix fallback when the base name is stuck
- Own activation polling (CLI --wait caps at 300s and discards state),
  replication-attribute + logical-slot preflight probes with actionable errors
- Mirror the synced-table list into the Cloud source's own publication
  (manual table publishing) after reassigning synced tables to postgres
- Reassign the cloud publication to postgres so the next in-place reset can
  drop it
- sslmode=require postgresql:// URL shaping for Electric's validation

Sweep (new .github/workflows/cleanup-preview-orphans.yml):
- `sweep` subcommand on all three branch scripts (PlanetScale, Electric,
  Tinybird) deleting pr-* resources whose PR is closed; scheduled workflow as
  the safety net for close-event teardowns GitHub never runs (conflicted PRs,
  stale workflow versions, post-close redeploys)

Verified on the feat/slack-agent preview: run 30055008127 recycled the branch
in 9 seconds (previously every deploy took the ~10-minute recreate path).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@pullfrog

pullfrog Bot commented Jul 24, 2026 •

Copy link
Copy Markdown
Contributor

Your Pullfrog Router balance is exhausted.

You have a card on file but auto-reload is disabled, so runs paused once your balance went past the overdraft buffer.

Top up balance → · Enable auto-reload →

Pullfrog  | ⚠️ this action is pinned to a commit SHA, which freezes the cleanup step — switch to @v0 or keep the SHA fresh with Dependabot | Rerun failed job ➔ | View workflow run | via Pullfrog | Using Claude Opus | 𝕏

From the PR #255 code + security reviews:

- reset-preview-branch.ts: refuse to run outside CI without
  RESET_PREVIEW_CONFIRM=1 — the script empties whatever DATABASE_URL points
  at and the connected role inherits postgres, so a stray prod/stg URL in a
  developer shell must not be enough to gut a real database
- reset-preview-branch.ts: drop migration-created extra schemas (owned by
  postgres post-normalization; vendor/system schemas untouched) so a replayed
  CREATE SCHEMA migration can't wedge every subsequent deploy with no
  recreate fallback
- reset-preview-branch.ts: run the inactive-replication-slot sweep over the
  replication-role connection (REPLICATION_DATABASE_URL) — the main role
  deliberately lacks the REPLICATION attribute, so the old sweep could never
  actually drop a slot; warn loudly on failure (a stale slot pins WAL)
- planetscale-pr-branch.ts: mint the Electric replication credential before
  the reset and pass it through; mask the full exported connection URLs, not
  just the raw passwords (encodeURIComponent can defeat raw-password masking)
- planetscale-pr-branch.ts: drop the dead ALTER SYSTEM fallback (equally
  permission-denied on PlanetScale) and downgrade the params-convergence
  timeout from a deploy-failing error to a warning — Electric activates
  without the params (run 30055008127)
- electric-pr-branch.ts: append sslmode=require when the URL never carried an
  sslmode param (previously only replaced an existing one); fail fast when
  `services create` returns no service id instead of dying later with a
  misleading get-secret error
- all three sweeps: hard-fail when GITHUB_REPOSITORY/GITHUB_TOKEN are absent —
  a token-less sweep resolves every PR to "unknown", skips everything, and
  green-no-ops forever, which is the exact failure class the safety net
  exists to catch
- comment fixes: nonexistent mintReplicationCredential reference, stale
  MAPLE_PG_URL mention in the Electric workflow step, single-repo naming
  caveat on the sweep workflow

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
JeremyFunk added a commit that referenced this pull request Jul 24, 2026
From the PR #255 code + security reviews:

- reset-preview-branch.ts: refuse to run outside CI without
  RESET_PREVIEW_CONFIRM=1 — the script empties whatever DATABASE_URL points
  at and the connected role inherits postgres, so a stray prod/stg URL in a
  developer shell must not be enough to gut a real database
- reset-preview-branch.ts: drop migration-created extra schemas (owned by
  postgres post-normalization; vendor/system schemas untouched) so a replayed
  CREATE SCHEMA migration can't wedge every subsequent deploy with no
  recreate fallback
- reset-preview-branch.ts: run the inactive-replication-slot sweep over the
  replication-role connection (REPLICATION_DATABASE_URL) — the main role
  deliberately lacks the REPLICATION attribute, so the old sweep could never
  actually drop a slot; warn loudly on failure (a stale slot pins WAL)
- planetscale-pr-branch.ts: mint the Electric replication credential before
  the reset and pass it through; mask the full exported connection URLs, not
  just the raw passwords (encodeURIComponent can defeat raw-password masking)
- planetscale-pr-branch.ts: drop the dead ALTER SYSTEM fallback (equally
  permission-denied on PlanetScale) and downgrade the params-convergence
  timeout from a deploy-failing error to a warning — Electric activates
  without the params (run 30055008127)
- electric-pr-branch.ts: append sslmode=require when the URL never carried an
  sslmode param (previously only replaced an existing one); fail fast when
  `services create` returns no service id instead of dying later with a
  misleading get-secret error
- all three sweeps: hard-fail when GITHUB_REPOSITORY/GITHUB_TOKEN are absent —
  a token-less sweep resolves every PR to "unknown", skips everything, and
  green-no-ops forever, which is the exact failure class the safety net
  exists to catch
- comment fixes: nonexistent mintReplicationCredential reference, stale
  MAPLE_PG_URL mention in the Electric workflow step, single-repo naming
  caveat on the sweep workflow

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The scheduled orphan-sweep safety net now lives on its own branch
(ci/preview-orphan-sweep, stacked on this one) so this PR stays focused
on the branch recycle + Electric hardening. Moved out:

- .github/workflows/cleanup-preview-orphans.yml (the scheduled workflow)
- the `sweep` subcommand on scripts/planetscale-pr-branch.ts
  (sweepOrphanBranches + parseArgs/usage/dispatch + header paragraph)
- the `sweep` subcommand on scripts/electric-pr-branch.ts
  (sweep + its sweep-only fetchPrState helper + parseArgs/usage/dispatch
  + header paragraph)
- the entire sweep addition to scripts/tinybird-pr-branch.ts, which was
  that file's only change — it is now byte-identical to main and drops
  out of this PR

fetchPrState STAYS in scripts/planetscale-pr-branch.ts: the `up`
closed-PR guard uses it to refuse provisioning a branch for an
already-closed PR. Incidental comments that used "sweep" as a generic
verb (reset-preview-branch.ts drop passes, the inactive-slot cleanup
notes) are reworded so nothing here reads as a reference to the removed
workflow.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@JeremyFunk
JeremyFunk merged commit 759f52a into main Jul 24, 2026
8 checks passed
@JeremyFunk
JeremyFunk deleted the ci/preview-branch-lifecycle branch July 24, 2026 09:03
JeremyFunk added a commit that referenced this pull request Jul 24, 2026
Electric Cloud now refuses `environments delete --force` while the
environment still holds services ("Cannot delete environment with existing
services. Delete all services first.", run 30081184677 — PR #255's close
teardown). deleteEnvironment() drops the services first, fixing both the
close-event `down` and the orphan sweep, which would have hit the same
error on every env that still had a source attached.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
JeremyFunk added a commit that referenced this pull request Jul 24, 2026
- normalize-preview-ownership.ts gets the same local-misuse tripwire as
  the reset script (refuse unless CI or RESET_PREVIEW_CONFIRM=1) — bun
  auto-loads .env and mise loads .env.local, so a stray real DATABASE_URL
  in a dev shell must not silently rewrite ownership.
- reset-preview-branch.ts emptiness verification now also counts
  routines and standalone enum/domain types: a survivor of those classes
  previously sailed past the check and hit duplicate_object at migrate
  replay, where no recreate fallback exists.
- RESET_PRESERVE_SCHEMAS (comma-separated) exempts named schemas from
  the extra-schema drop — the ops lever for a future postgres-owned
  PlanetScale vendor schema.
- electric-pr-branch.ts: resolve sourceId via `?? fail(...)` so it is
  typed string outright, and drop the dead `if (sourceId)` wrapper and
  re-checks that guarded already-fatal states.

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
JeremyFunk added a commit that referenced this pull request Jul 24, 2026
* ci(preview): scheduled orphan sweep for PR-preview resources

Re-adds the orphan-sweep safety net split out of the branch-lifecycle
PR: the `sweep` subcommand on the PlanetScale/Electric/Tinybird
lifecycle scripts and the scheduled cleanup-preview-orphans workflow
that runs them. Content is identical to the implementation as reviewed
on the lifecycle branch (2017c6c), including the hard fail when
GITHUB_TOKEN is missing.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* ci(preview): sweep orphaned Hyperdrive PR configs

The account caps Hyperdrive configs at 25, and leaked maple-db-pr-* configs
from closed PRs eventually block every new PR preview ("This account has
reached its limit for how many Hyperdrives it can have (25)", run
30077724144). Unlike the other preview resources, Hyperdrive teardown is
owned solely by `alchemy destroy` in the close-event run — which executes
the PR branch's OWN workflow version, so branches predating the
placeholder-MAPLE_PG_URL destroy fix fail with "Missing required deployment
env: MAPLE_PG_URL" on every rerun, forever (runs 29737841130 et al.,
attempt 2). Those orphans are unreachable by rerunning; only a sweep from
current main code can collect them.

scripts/hyperdrive-orphan-sweep.ts mirrors the sibling sweeps' double gate:
exact `maple-db-pr-<digits>` name match (maple-prd / maple-db-stg /
maple-db-dev-* can never match) AND the GitHub API affirming the PR is
closed; unknown/open → keep; missing token → hard fail. Wired into
cleanup-preview-orphans.yml gated on CLOUDFLARE_API_TOKEN.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* ci(preview): delete Electric services before the environment

Electric Cloud now refuses `environments delete --force` while the
environment still holds services ("Cannot delete environment with existing
services. Delete all services first.", run 30081184677 — PR #255's close
teardown). deleteEnvironment() drops the services first, fixing both the
close-event `down` and the orphan sweep, which would have hit the same
error on every env that still had a source attached.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* ci(preview): route the Electric sweep through the services-first delete

Adversarial review of the post-review commits caught that the sweep path
deleted environments via a raw `environments delete --force`, bypassing the
services-first fix — so the canonical orphan (an env leaked from a failed
close-teardown with its source still attached) would have failed with
"Cannot delete environment with existing services" on every scheduled tick.
Drop the services in the sweep loop too.

Also two hardening nits from the same review: fail loudly when the
Hyperdrive list response shape is unexpected instead of green-no-op'ing the
safety net, and gate the Hyperdrive sweep step on CLOUDFLARE_ACCOUNT_ID as
well as the token so a half-configured Infisical env skips cleanly.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* ci(preview): surface the failing statement when the in-place reset dies

The reset hit a nondeterministic server-side "malformed array literal: ''"
for "unnamed portal parameter $2" (run 30083630959) — no query this script
sends carries two parameters, and the same schema reset cleanly minutes
earlier, so the suspect is PlanetScale-side DDL interception. The server's
error report doesn't include the statement; postgres.js errors do. Print
query + parameters on the way out so the next occurrence identifies itself.
The delete → recreate fallback already absorbs the failure, so this only
costs log lines.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* ci(preview): exercise the recycle path with the reset diagnostic in place

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* ci(preview): inline the pg_% pattern — bind params trip the reset on PlanetScale

The recurring in-place-reset failure (malformed array literal: "" for
"unnamed portal parameter $2", 3 consecutive runs on pr-257) is
parameter-related: the diagnostic shows the failing catalog query was sent
with parameters ["pg\_%", ""] — a phantom empty second parameter on a query
that interpolates exactly one value. It does not reproduce against plain
Postgres (3/3 clean resets locally with identical schema), only against
PlanetScale over the direct branch connection. Sidestep whatever mangles
extended-protocol parameters there by inlining the LIKE pattern as a SQL
literal — the owners and extra-schemas queries now carry zero bind
parameters. Verified locally end-to-end.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
JeremyFunk added a commit that referenced this pull request Jul 28, 2026
…ed Slack agent (#246)

* feat(slack-agent): eve-based Slack agent with Railway self-deploy

Adds apps/slack-agent, a standalone eve-framework Slack agent that
deploys itself to Railway (Docker) rather than Cloudflare Workers:

- Workers AI (REST) as the model backend, self-managed Slack app
  (bot token + signing secret) instead of a Connect integration
- @workflow/world-postgres for durable runs; EVE_WORKFLOW_WORLD is
  baked in at build time via the Dockerfile
- excluded from the bun workspace ("!apps/slack-agent") so its own
  bun.lock and eve toolchain resolve independently

Also adds scripts/ingest-dummy.ts for pushing dummy OTLP traces/logs
at the local ingest gateway, drops turbo concurrency to 15, and
gitignores the .eve model-catalog cache.

* fix(slack-agent): use a model that streams structured tool calls

The agent was posting raw tool-call JSON into Slack:

  {"type": "function", "name": "ask_question", "parameters": {…}}

Not a formatting bug — a leaked tool call. @cf/meta/llama-3.3-70b-instruct-fp8-fast
only parses tool calls on non-streaming requests; eve's harness always streams, and
in streaming mode Workers AI returns that model's raw tool-call JSON as ordinary
`response` text deltas, which eve has no reason to treat as anything but assistant
text. (Even non-streaming it stringifies non-string args: "allowFreeform": "true".)

workers-ai-provider has a salvage path for leaked tool calls, but it's gated on a
forced tool choice — eve uses auto, so it never engages.

Switch to @cf/zai-org/glm-5.2, which streams OpenAI-shaped incremental
delta.tool_calls (name + id first, argument fragments keyed by index after) ending
in finish_reason: "tool_calls", and emits chain-of-thought on reasoning_content so
it maps to reasoning parts instead of message text. Bump the declared context
window to its 256K.

Verified through a live eve session: modelId workersai/@cf/zai-org/glm-5.2,
actions.requested → action.result, correct Tokyo time in prose.

Also documents the streaming-tool-call constraint, a curl to check any replacement
model's SSE shape, and the price trade-off vs gpt-oss-120b (also verified good, and
~5x cheaper on output if spend beats capability).

Includes the pending PORT 3000 -> 8080 alignment with the documented Railway setup.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* Implement basic slack agent

* chore: include alerting worker in dev:essentials

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Refactor

* fix(mcp): emit object-typed inputSchema for no-param tools

An empty Schema.Struct({}) compiles to `{ anyOf: [{type:object},{type:array}] }`
in effect 4.0.0-beta.93, with no top-level `type: "object"`. Strict MCP clients
(the Vercel AI SDK used by the eve Slack agent) validate each tool's
inputSchema.type against z.literal("object"), so list_source_repositories (tool
#40) fails to parse and aborts the ENTIRE tools/list response — dropping every
Maple tool from the connection.

Normalize a no-property struct to an explicit empty object schema centrally in
toInputSchema so all current and future no-param tools stay MCP-compliant. Add a
test asserting every registered tool emits inputSchema.type === "object".

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(web): purge foreign Clerk instance cookies to stop cross-env 403s

All deployed stages share the maple.dev registrable domain, but PR
previews/staging run the dev Clerk instance while prod runs the
production one. Both write __client_uat cookies on Domain=maple.dev, so
browsing multiple environments makes the instances overwrite each
other's session hints and ClerkJS handshakes against the wrong frontend
API — surfacing as transient 403s on app-pr-N/api-pr-N and prod.

Purge foreign-instance parent-domain __client_uat* cookies (suffix
derived from the publishable key, mirroring @clerk/shared) before
ClerkJS initializes. Mitigation until preview/staging move to their own
registrable domain, per Clerk's same-domain limitation.

Also document the channels:history Slack bot scope in the slack-agent
README.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(electric-sync): 503 on half-configured Electric Cloud creds

The shape proxy only treated a missing ELECTRIC_URL as "not configured".
A deploy with ELECTRIC_URL + ELECTRIC_SOURCE_ID but no ELECTRIC_SECRET
(e.g. a PR preview inheriting shared Infisical Electric config while its
per-PR source step is skipped) forwarded an unauthenticated request to
Electric Cloud, which returns 401 MISSING_SECRET — surfacing as a hard-
broken shape stream in the browser instead of graceful degradation.

Treat incoherent Cloud credentials (exactly one of source_id/secret) as
not-configured and take the existing 503 path, and log the misconfig for
telemetry. Adds isElectricConfigCoherent + unit tests.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* Refactors & improvements

* Parity with chat-flue

* Fix deployment

* Fix slack deployment

* ci(preview): reuse PlanetScale PR branch with in-place SQL reset

A PS-DEV Postgres branch takes ~9 min to provision, and the preview
deploy paid that on every push by deleting + recreating pr-<n>. The
branch now lives for the PR's lifetime: each deploy resets it in SQL
(assume prior owner roles, drop publications, drop the drizzle schema,
empty public per-object, sweep inactive replication slots) and falls
back to delete -> recreate if the reset fails. Also treat `pscale
branch create --wait`'s own ~10-min timeout as non-fatal (the branch
keeps provisioning server-side) and keep polling until ready.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* ci(preview): raise PlanetScale ready-poll budget to 15 min

Provisioning has been observed to exceed the 10-minute budget; the
create path's total allowance is now ~10 min (CLI --wait) + 15 min
(our branch-show poll).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* ci(preview): match electric-pr-branch.ts to the real @electric-sql/cli 0.0.10

The Electric step ran for the first time (token landed in Infisical dev) and
died on `environments create --json returned no environment id`: the script
was written against an assumed CLI. Verified the actual interface from the
CLI's bundled source and fixed:

- `environments create --json` returns { environmentId }, not { id } — the
  root cause of the CI failure.
- `--publication <name>` does not exist; the real flag is
  `--manual-table-publishing` (Electric's default publication name is the
  migration-owned `electric_publication_default`). Pass it by default
  (prod parity), opt out via ELECTRIC_MANUAL_TABLE_PUBLISHING=false;
  ELECTRIC_PUBLICATION is gone.
- Add `--wait` so the source is active before the alchemy deploy binds it.
- The postgres service id IS the shape-API source_id; secret comes from
  `sourceSecret` (fallback: `services get-secret` → { secret }).
- Not-found/conflict detection now uses the CLI's typed exit codes (3/5)
  with text matching as fallback; JSON errors arrive on stderr.
- Drop the unscoped `environments list` fallback (--project is required)
  and pin the CLI to @electric-sql/cli@0.0.10 (interface-verified).

Redaction, up/down contract, and the GITHUB_ENV exports are unchanged.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* ci(preview): retry Electric env create while the deleted name is still reserved

Environment deletion is async server-side: the env leaves `environments list`
immediately, but create then transiently fails with VALIDATION_ERROR "an
environment with this name already exists" (seen on run 30030020767). Retry
name conflicts with the existing 5s/2min backoff budget instead of failing.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* ci(preview): reuse the Electric PR environment; reset services instead

Electric soft-deletes environments: a deleted env leaves `environments list`
immediately but its name stays reserved (observed >13 min on runs
30030020767/30031075727), so the delete-then-recreate reset loops forever on
"name already exists". Rework the lifecycle:

- `up` reuses the existing `pr-<n>` environment and resets only its services
  (they are what point at the dropped tables/rotated creds after the
  PlanetScale in-place reset), creating a fresh postgres source each deploy.
- Fresh creates keep a short name-conflict retry, then fall back to a unique
  `pr-<n>-r<run-id>` name when the base is soft-delete-reserved (the current
  state of pr-246). Matching is exact-name or `pr-<n>-` prefix everywhere.
- `down` deletes all matching environments (base + suffix fallbacks).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* ci(preview): pass Electric a postgresql:// scheme database URL

Run 30032252411 got the environment created (suffix fallback worked) but
`services create postgres` failed server-side with a generic "Input
validation failed". MAPLE_PG_URL uses the short `postgres://` scheme; all
Electric examples use `postgresql://` — normalize before passing.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* ci(preview): isolate the Electric create-source validation failure in one run

The Cloud API rejects `services create postgres` with a field-less "Input
validation failed" (runs 30032252411/30033111516; scheme normalization did
not help). Attempt the full-fidelity create first, then degrade the two
suspects — the manual-table-publishing option and the ?sslmode query — one
at a time, logging which variant is accepted. Non-validation failures still
fail immediately.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* ci(preview): hand Electric an sslmode=require connection string

The isolation ladder (run 30034103225) showed sslmode=verify-full is
rejected at input validation while a bare URL passes input validation but
fails the server's database probe ("Database validation failed" — TLS-less
connect to PlanetScale). Electric's PlanetScale guide prescribes exactly
`postgresql://...?sslmode=require` — normalize to that.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* ci(preview): log the PR branch's replication state before the Electric create

sslmode=require got past input validation (run 30035178700) but every
variant now fails Electric's database probe. Temporarily print wal_level,
slot/sender limits, role replication privilege, and the publication's
existence from the branch itself so the next run tells us which
prerequisite the PS-DEV branch is missing.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* ci(preview): grant the CI role REPLICATION before creating the Electric source

Diagnostic (run 30036098689) pinpointed it: wal_level=logical, slots, and
the publication are all in place, but rolreplication=false. The credential
role inherits `postgres` via membership, and REPLICATION is a role
attribute — attributes are never inherited — so Electric's database probe
kept answering "Database validation failed". Grant the attribute in-place
(direct ALTER, falling back to SET LOCAL ROLE postgres on a pinned
connection), verify it stuck, and fail with a precise message if not.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* ci(preview): mint the PlanetScale CI role with REPLICATION

In-place granting is impossible on PlanetScale — both `ALTER ROLE
CURRENT_USER WITH REPLICATION` and the SET ROLE postgres fallback get
"permission denied to alter role" (run 30036963759). The supported path is
minting the role with the attribute: `pscale role create
--with-replication` (requires --inherited-roles postgres, already passed).
The electric script now only verifies rolreplication and fails with a
precise message on regression.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* ci(preview): own the Electric source activation wait

Run 30037847577 passed validation (REPLICATION role fixed it) but the
CLI's --wait timed out at its hard 300s cap and discarded the service
state. Create without --wait, capture id+secret immediately, and poll
`services get` with a 10-min budget, logging status transitions and
dumping the non-secret service state on error/timeout. The
input-validation isolation ladder is gone — the accepted input shape is
settled (sslmode=require + manual-table-publishing).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* ci(preview): surface failover-slot prerequisites + slot state around activation

Run 30038967625: the source sits in status=pending for the full 10-min
budget with no error detail. Electric always creates FAILOVER-enabled
slots, which PlanetScale only accepts with sync_replication_slots=on and
hot_standby_feedback=on — settings the diagnostic never checked. Log those
(plus max_connections) pre-create, and dump pg_replication_slots on
activation error/timeout so the step log shows whether Electric's slot
ever landed and why it is stuck.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* ci(preview): set Electric-required cluster parameters on the PR branch

Run 30040575049 diagnostics: the branch runs sync_replication_slots=off,
hot_standby_feedback=off, max_connections=25 — Electric creates its slot
but never activates (it needs failover-capable slots and headroom for its
20-connection pool), leaving the source pending forever. Apply
pgconf.sync_replication_slots=on, hot_standby_feedback=on, and
max_connections=100 via `pscale branch resize --parameters --wait` during
branch up, skipping when the live settings already satisfy them, and poll
the live settings afterwards (the change can restart the cluster).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* ci(preview): degrade gracefully when the CI token cannot resize the branch

`pscale branch resize` is denied for the CI service token (run
30042253492: "User does not have permission to perform this action").
Keep the resize attempt (self-heals once the token gains branch-change
access) but fall back to ALTER SYSTEM as `postgres` for the
SIGHUP-reloadable settings (sync_replication_slots, hot_standby_feedback)
on a pinned single-connection session + pg_reload_conf(), and shrink the
Electric source pool to --db-pool-size 5 (ELECTRIC_DB_POOL_SIZE override)
so it fits under the branch's max_connections=25, which only a resize can
raise. Parameter failures warn instead of failing the deploy.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* ci(preview): probe logical-slot creation before the Electric create

The freshly recreated branch draws "Database validation failed" again
(run 30042533336) — the old branch had wal_level=logical, a fresh one may
not. Log wal_level + publication in the checklist and dry-run
pg_create_logical_replication_slot: its error text is the real reason
behind Electric's opaque rejection.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* ci(preview): split the Electric credential off the main CI role

Every deploy since --with-replication landed lost the in-place branch
reset ("permission denied to grant role" on the previous run's role —
PlanetScale replication-attribute roles are not grantable) and paid the
~9-min delete → recreate path. Mint the main role WITHOUT replication
again (restores role assumption / fast resets) and a second
`-repl`-suffixed role WITH replication used only by the Electric source,
exported as MAPLE_PG_ELECTRIC_URL (electric-pr-branch.ts prefers it,
falling back to MAPLE_PG_URL).

Note: the next deploy still recreates once (the current branch owner is a
replication role from the previous scheme); resets are fast again after.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Fix MCP list for slack agent

* ci(preview): publish synced tables to Electric Cloud's own publication

Electric Cloud ignores electric_publication_default: with
--manual-table-publishing each Cloud source creates its own publication
(cloud_electric_pub_svc_<name>) and refuses to add tables to it, so every
shape request 400'd with "missing from the publication". After the source
activates, mirror the migration-owned table list into the cloud
publication(s) via an idempotent server-side DO block.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Render slack-agent charts in the Maple dark theme

Replace the placeholder light-blue palette in agent/lib/chart.ts with the
Maple dark-theme tokens (tokens.css .dark, oklch converted to hex for
resvg): card surface + hairline border canvas, border/50 grid, muted axis
ink, gradient area fills, and per-unit semantic series colors (latency
amber, throughput purple, error-rate red, bytes teal, counts blue).

Type is now Geist Mono: the Dockerfile decompresses the @fontsource woff2
faces to TTF (resvg reads TTF only) into the system font dir.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Fix blank chart text: pass explicit fontDirs to resvg

resvg-js 2.x's loadSystemFonts discovers fonts on Linux via fontconfig's
/etc/fonts/fonts.conf; the slim production image has no fontconfig
package, so font discovery silently found nothing and every glyph
(title, axis labels, latest-value label) was dropped from the PNG —
charts shipped as bare lines on a card. Point resvg directly at the
font dirs the Dockerfile populates (fontDirs needs no fontconfig),
keep loadSystemFonts for local dev, and set the Geist Mono default.
Verified in a node:24-slim container against resvg-js 2.6.2.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* ci(preview): assume table-owner roles before publishing to the cloud publication

The Electric role owns the cloud publication but not the synced tables
(owned by the ephemeral migrate-step pscale role), so ALTER PUBLICATION
ADD TABLE failed with "must be owner of table api_keys". Assume each
table-owner role via GRANT <owner> TO CURRENT_USER first — the same
trick reset-preview-branch.ts uses, backed by the inherited postgres
grant's admin over the ephemeral roles.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* ci(preview): reassign synced tables to postgres before cloud-publication mirror

GRANT <table-owner> TO the Electric role is a dead end — replication
roles can't receive role grants on PlanetScale, so run 30050755374 still
failed with "must be owner of table actors" (and the DO block's RAISE
NOTICE diagnostics never reach the client). Instead the main CI role now
reassigns the synced tables' ownership to postgres (allowed: it owns
them and is a member of postgres), which the Electric role also
inherits — making it a table owner for the ALTER PUBLICATION mirror.
All statements run client-side with per-statement error logging.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Add sweep

* ci(preview): normalize branch ownership to postgres so the in-place reset works

The reuse path never actually ran: reset-preview-branch.ts tried to GRANT the
prior run's pscale_api_* roles to CURRENT_USER, but PlanetScale refuses to
grant those roles at all ("permission denied to grant role", run 30051727304),
so DROP PUBLICATION electric_publication_default failed with "must be owner"
and every deploy fell back to the ~9-minute delete → recreate path.

Every pscale role inherits `postgres`, so hand ownership over at the END of
each deploy instead:

- new packages/db/scripts/normalize-preview-ownership.ts (db:normalize-preview,
  run after migrate + grant in the workflow) — REASSIGN OWNED BY CURRENT_USER
  TO postgres as the main role, covering the drizzle schema, all of public,
  and electric_publication_default
- electric-pr-branch.ts reassigns the Cloud source's cloud_electric_pub_svc_*
  publication to postgres after activation (its owner is the replication role)

The next run's role then owns everything via its own postgres membership and
the reset's drops just work. Both steps are non-fatal: an un-normalized branch
only costs the next deploy the recreate fallback, never correctness.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* ci(preview): empty commit to exercise the in-place branch recycle path

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* ci(preview): address review findings on the branch-lifecycle scripts

From the PR #255 code + security reviews:

- reset-preview-branch.ts: refuse to run outside CI without
  RESET_PREVIEW_CONFIRM=1 — the script empties whatever DATABASE_URL points
  at and the connected role inherits postgres, so a stray prod/stg URL in a
  developer shell must not be enough to gut a real database
- reset-preview-branch.ts: drop migration-created extra schemas (owned by
  postgres post-normalization; vendor/system schemas untouched) so a replayed
  CREATE SCHEMA migration can't wedge every subsequent deploy with no
  recreate fallback
- reset-preview-branch.ts: run the inactive-replication-slot sweep over the
  replication-role connection (REPLICATION_DATABASE_URL) — the main role
  deliberately lacks the REPLICATION attribute, so the old sweep could never
  actually drop a slot; warn loudly on failure (a stale slot pins WAL)
- planetscale-pr-branch.ts: mint the Electric replication credential before
  the reset and pass it through; mask the full exported connection URLs, not
  just the raw passwords (encodeURIComponent can defeat raw-password masking)
- planetscale-pr-branch.ts: drop the dead ALTER SYSTEM fallback (equally
  permission-denied on PlanetScale) and downgrade the params-convergence
  timeout from a deploy-failing error to a warning — Electric activates
  without the params (run 30055008127)
- electric-pr-branch.ts: append sslmode=require when the URL never carried an
  sslmode param (previously only replaced an existing one); fail fast when
  `services create` returns no service id instead of dying later with a
  misleading get-secret error
- all three sweeps: hard-fail when GITHUB_REPOSITORY/GITHUB_TOKEN are absent —
  a token-less sweep resolves every PR to "unknown", skips everything, and
  green-no-ops forever, which is the exact failure class the safety net
  exists to catch
- comment fixes: nonexistent mintReplicationCredential reference, stale
  MAPLE_PG_URL mention in the Electric workflow step, single-repo naming
  caveat on the sweep workflow

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(alerts): redesign Slack alert messages per Block Kit guidance

- Wrap slack-bot blocks in a colored attachment so the severity bar
  renders (parity with the webhook destination)
- Truncate the default header to Slack's 150-char limit (long rule
  names previously failed delivery with invalid_blocks)
- Escape mrkdwn control chars in rule names, group keys, incident ids
- Replace the five-field dump with a lead summary sentence + compact
  fields (emoji-paired capitalized severity; group only when set)
- Add a context footer with incident ref and <!date^…> local-time stamp
- Full one-line notification fallback text instead of "rule: Triggered"
- Button styles per Slack guidance: primary on Open in Maple, no danger
  on navigation links

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(alerts): align default notification templates with the redesigned format

A rule that customizes only the title (or only the body) has its other
field filled from the built-in defaults, which still produced the old
five-line field dump — so partially-templated rules kept sending the
old-style "basic" message. Update DEFAULT_BODY_TEMPLATE and the rule
editor placeholder to the lead-sentence format the designed blocks use.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(alerts): stop Slack rendering the fallback text as a duplicate line

With attachments and no top-level blocks, chat.postMessage renders the
top-level text in-channel above the color bar (confirmed on the PR-246
preview). Move the notification-preview one-liner into the attachment's
fallback field, which Slack uses only for push/desktop previews.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Refactor

* fix(tinybird): drop completed service_operations_minutely forward query

* Add bot follow-up without metnion

* Fixes

* Remvoe unnecessary tool

* Refactors

* Format

* feat(slack-agent): strip __maple_ui payloads from MCP tool results

Port of chat-flue's splitToolResult change from main: Maple's MCP server
emits a second content entry per tool result carrying the web chat's
structured UI payload (createDualContent, tagged __maple_ui). chat-flue
now splits it off client-side; eve has no result-transform hook, so the
Slack agent's model was receiving the raw UI JSON duplicated next to the
text report on every Maple tool call. Extend the vendored eve patch to
drop those entries in McpConnectionClient.executeTool, with a canary
test alongside the botToken one.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* merge migrations

* Fixes

* fix(slack): apply review fixes; move Slack event intake to the Railway bot

Review fixes (multi-agent effect-review-v4 + general review):
- SlackIntegrationService: probeAndRevokeIfDead no longer counts a failed
  DB revoke as revoked (logs and returns false); uninstall's revocation
  update is now a compare-and-set with one retry so a concurrent reinstall
  can't be clobbered while its fresh API key survives unrevoked;
  Service.of() construction; Arr.map/Arr.filter over native methods
- slack-bot-token: SlackBotTokenResolver.of() construction
- integrations.http.test: complete the 8-method die-stub
- slack-agent thread-follow-up: bound conversations.replies with
  oldest/latest so the engagement window is correct for >100-reply
  threads; single ~2s promotion deadline inside Slack's ack budget
- slack-agent maple.ts: type-check resolve payload fields; render_chart:
  reject non-finite point values

Also lands the events-intake refactor: SlackEventsRouter is removed from
the API (Slack allows one Events URL per app, already pointed at the bot);
the bot detects app_uninstalled/tokens_revoked (uninstall-detection.ts,
bot-token-scoped) and calls POST /internal/slack/workspaces/:teamId/revoke,
with the 6-hourly reconcile cron as backstop. SLACK_SIGNING_SECRET leaves
the API env.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(api): provide SlackIntegrationServiceStubLayer in setup-audit harness

AllV2GroupLayersLive now includes HttpV2SlackIntegrationsLive, so every
harness composing it must stub SlackIntegrationService; setup-audit was
the one file missed.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Increase timeout

* docs: move the after-hours compute limit out of the shared CLAUDE.md

Machine-local preference, not a project rule — belongs in the owner's
CLAUDE.local.md or ~/.claude/CLAUDE.md instead.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(slack): surface remote disconnects on the integration status

When the Maple app is removed (or its tokens revoked) from Slack's own
UI, the integration card silently reverted to the never-installed state,
which read as a Maple bug. Persist why a workspace row was revoked
(revoked_reason: uninstalled / superseded / app_uninstalled /
tokens_revoked / reconciliation), expose the remote reasons on the v2
status endpoint (disconnected_reason / _team_name / _at), and show a
warning banner on the card explaining the disconnect came from Slack's
side. Dashboard uninstalls and replacement installs still read as plain
not-installed.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(slack): neutral copy for the remote-disconnect banner

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>

This branch had an error being deployed

1 failed deployment
pr-preview — e3a2853b Deployed Jul 24, 2026 by JeremyFunk via deploy-pr-preview #1026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant