Skip to content

feat(api): run apps/api on Cloudflare Workers - #25

Merged
Makisuo merged 20 commits into
mainfrom
cf-migration-api
Apr 16, 2026
Merged

Makisuo merged 20 commits into
mainfrom
cf-migration-api

Conversation

@Makisuo

@Makisuo Makisuo commented Apr 14, 2026

Copy link
Copy Markdown
Collaborator

Summary

Moves apps/api from a Bun HTTP server to a Cloudflare Worker so it can
be deployed independently, without waiting on the in-progress
alchemy-effect infra rewrite on cf-migration.

  • Refactors apps/api into a Workers fetch handler (src/worker.ts +
    shared src/app.ts), with DatabaseD1Live using the D1 binding and a
    new EdgeCacheService backed by Workers caches.default (memory
    fallback for tests/Node).
  • Splits DatabaseLive into an abstract Database Context.Service with
    two live implementations: DatabaseLibsqlLive (tests, apps/alerting)
    and DatabaseD1Live (worker).
  • Wires drizzle D1 migrations natively: wrangler.jsonc sets
    migrations_dir: "../../packages/db/drizzle" and alchemy.run.ts
    passes migrationsDir on the D1Database resource. Drops the one-shot
    SQL bundle approach (couldn't re-run incrementally).
  • Adds apps/api/alchemy.run.ts (upstream alchemy package, mirroring
    apps/chat-agent) with stage-based naming and domain mapping, plus
    deploy:stack / destroy:stack scripts.

What this PR does not include

Everything below stays on cf-migration until it's ready:

  • .context/alchemy-effect subtree bump
  • apps/{web,landing,chat-agent,alerting}/alchemy.run.ts updates
  • packages/infra/src/railway rewrite

Notes

  • apps/api/src/services/WorkerEnvironment.ts is a local Effect Context
    tag replacing what was originally imported from
    alchemy/Cloudflare/Workers. With customConditions: ["bun"], tsc
    resolves alchemy/* to raw TS source, which fails typecheck against
    the workspace's effect beta version. Keeping a local tag means src/**
    has zero alchemy imports and the deploy-time alchemy.run.ts (outside
    src/**) doesn't pollute the tsc graph.
  • First-time deploy steps for the reviewer:
    1. bun install
    2. bun --filter=@maple/api deploy:stack (alchemy provisions D1,
      applies drizzle migrations via migrationsDir, deploys Worker)
    3. Set secrets via env vars before running: TINYBIRD_TOKEN,
      MAPLE_INGEST_KEY_ENCRYPTION_KEY,
      MAPLE_INGEST_KEY_LOOKUP_HMAC_KEY, etc. (see alchemy.run.ts for
      the full list)

Test plan

  • bun --filter=@maple/api typecheck
  • bun --filter=@maple/alerting typecheck
  • bun --filter=@maple/db typecheck
  • cd apps/api && bun test — 160/160 passing
  • wrangler d1 migrations apply MAPLE_DB --local applies all 16
    drizzle migrations; re-run is a no-op
  • Smoke test bun run dev (wrangler dev + local D1)
  • Staging deploy via bun run deploy:stack
  • Hit /health on the deployed Worker and verify one authenticated
    route against D1

🤖 Generated with Claude Code

Makisuo and others added 2 commits April 15, 2026 00:57
Refactors apps/api from a Bun HTTP server into a Workers fetch handler
so we can deploy via `wrangler deploy` without waiting on the
alchemy-effect infra rewrite.

- split DatabaseLive into an abstract Context.Service with D1 and libsql
  live layers; tests and apps/alerting use the libsql variant
- add EdgeCacheService backed by Workers caches.default with an in-memory
  fallback for Node/test runtimes; QueryEngineService uses it in place of
  Effect.Cache
- add packages/db/client factories, D1 migration bundle, and the
  scripts/prepare-local-d1.ts dev helper
- local WorkerEnvironment Context tag so we don't pull alchemy into the
  tsc graph (alchemy 2.0.0-beta.3 fails type-check against workspace
  effect 4.0.0-beta.48)

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Replaces the one-shot SQL bundle with drizzle's native migration folder,
which both wrangler and alchemy read directly. The bundle approach
couldn't run incrementally — it would fail on the second apply because
tables already existed.

- wrangler.jsonc: add migrations_dir pointing at packages/db/drizzle so
  `wrangler d1 migrations apply` tracks applied migrations in the
  standard d1_migrations table
- alchemy.run.ts: declare the D1Database with migrationsDir and the
  Worker with D1 binding + env/secret wiring, mirroring the pattern in
  apps/chat-agent/alchemy.run.ts; domains per stage (api.maple.dev /
  api-staging.maple.dev)
- drop the bundle scripts (create-d1-migration-bundle, export-app-state,
  render-d1-import-sql) and the .generated SQL artifact
- rewrite scripts/prepare-local-d1.ts to delegate to
  `wrangler d1 migrations apply --local` — idempotent, picks up new
  migrations without wiping
- add alchemy dep (same pkg.pr.new version as web/landing/chat-agent) and
  deploy:stack / destroy:stack scripts

Verified: all 16 drizzle migrations apply cleanly via
`wrangler d1 migrations apply MAPLE_DB --local` and re-running is a
no-op. Typecheck and tests still pass.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
@pullfrog

pullfrog Bot commented Apr 14, 2026 •

Copy link
Copy Markdown
Contributor

no API key found. Pullfrog requires at least one LLM provider API key.

to fix this, add the required secret to your GitHub repository:

  1. go to: https://github.com/Makisuo/maple/settings/secrets/actions
  2. click "New repository secret"
  3. set the name to your provider's key (e.g., ANTHROPIC_API_KEY, OPENAI_API_KEY, GEMINI_API_KEY)
  4. set the value to your API key
  5. click "Add secret"

configure your model at https://pullfrog.com/console/Makisuo/maple

for full setup instructions, see https://docs.pullfrog.com/keys

Pullfrog  | Rerun failed job ➔ | View workflow run | Triggered by Pullfrog | Using Claude Opus | 𝕏

CI deploy-prd failed with `Secret cannot be undefined` because
`alchemy.secret(optionalEnv(...))` was invoked for env vars that weren't
set in CI (MAPLE_ROOT_PASSWORD, CLERK_*, etc.). The D1 database and
migrations did apply successfully before the Worker step died.

Introduces `optionalSecret` / `optionalString` helpers that omit the
binding entirely when the env var is unset. The Worker runtime already
handles missing optional env via `Option` in `services/Env.ts`, so
omitting the binding is the right shape.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
@railway-app
railway-app Bot temporarily deployed to maple / pr-25 April 14, 2026 23:16 Inactive
@railway-app
railway-app Bot temporarily deployed to maple / pr-25 April 15, 2026 00:49 Inactive
@railway-app
railway-app Bot temporarily deployed to maple / pr-25 April 15, 2026 10:43 Inactive
@railway-app
railway-app Bot temporarily deployed to maple / pr-25 April 15, 2026 11:29 Inactive
@railway-app
railway-app Bot temporarily deployed to maple / pr-25 April 15, 2026 11:34 Inactive
@railway-app
railway-app Bot temporarily deployed to maple / pr-25 April 15, 2026 22:15 Inactive
@railway-app
railway-app Bot temporarily deployed to maple / pr-25 April 15, 2026 22:31 Inactive
@railway-app
railway-app Bot temporarily deployed to maple / pr-25 April 15, 2026 22:33 Inactive
@railway-app
railway-app Bot temporarily deployed to maple / pr-25 April 15, 2026 22:43 Inactive
@railway-app
railway-app Bot temporarily deployed to maple / pr-25 April 15, 2026 23:11 Inactive
@railway-app
railway-app Bot temporarily deployed to maple / pr-25 April 15, 2026 23:12 Inactive
@railway-app
railway-app Bot temporarily deployed to maple / pr-25 April 15, 2026 23:14 Inactive
@Makisuo
Makisuo merged commit 5ee13f0 into main Apr 16, 2026
3 checks passed
@Makisuo
Makisuo deleted the cf-migration-api branch April 16, 2026 00:03
JeremyFunk added a commit that referenced this pull request Sep 28, 2026
… on Anthropic only

Follow-up to the uncacheable-prompt fix (#25). The never-written skip
read a zero cache-write bucket as "nothing was cacheable", but OpenAI-
model emitters (CrewAI, Haystack, LiteLLM, Mastra, Microsoft Agent
Framework, Strands, Vercel AI SDK) stamp a zero write bucket on every
call, so an OpenAI session with long prompts and a real 0% hit rate was
skipped instead of warned. The rule now applies only when every
reporting call is an Anthropic call. The separate "too few calls" early
return is dropped: the cacheable-call gate below it covers it.
JeremyFunk added a commit that referenced this pull request Sep 28, 2026
…mpt-cache check

Follow-up to the uncacheable-prompt fix (#25). The never-written rule was
session-wide and keyed on provider `anthropic` with a write bucket of 0,
so it missed Claude through OpenRouter (flue: provider `openrouter`,
prompts 1,043-1,392 tokens under Haiku 4.5's 4,096 minimum) and
OpenRouter Broadcast (provider `anthropic`, write key absent, prompts
1.2K-2.6K), and a mixed session fell back to judging every call.

A call is now left out on its own when it is Claude (provider
`anthropic` or a `claude` model) and neither wrote nor read the cache,
like the 1024-token cut; the rest of the session is still judged.
JeremyFunk added a commit that referenced this pull request Sep 28, 2026
…sses every Claude minimum

Follow-up to the uncacheable-prompt fix (#25). Leaving out every Claude
call that neither wrote nor read the cache also hid real misses when the
emitter does not report writes: 8 Claude calls with 10K-token prompts
and 4 misses read "passed 90% over 3 calls" instead of "51% over 7".
Such a call is now left out only when its prompt is also under 4,096
tokens, the largest Claude minimum; above it, reading nothing is a miss.
The cold/warm cache fixture is back on a Claude model with no write
bucket and 10K prompts, and must warn.
JeremyFunk added a commit that referenced this pull request Sep 29, 2026
* fix(agent-sessions): stop the prompt-cache check warning on uncacheable prompts

Symptom: the Prompt cache check warned "Cache hit rate 0% over N calls; N
missed the cache" on sessions whose prompts were all too short to cache.
Seen on DSPy (prompts peak 870 tokens, OpenAI minimum 1024), Google ADK,
LangChain, Haystack, CrewAI and Claude Agent SDK (Haiku 4.5 prompts of
1.2K-2.3K tokens against its 4,096-token minimum) test sessions.

Cause: promptCacheCheck judged every call after the first, whatever its
prompt size.

Fix: calls under 1024 prompt tokens (the smallest minimum OpenAI,
Anthropic and Gemini cache) are left out of the judgement, and a session
whose calls report a cache-write bucket but never wrote or read the cache
is skipped: nothing reached the model's minimum, or caching is off. The
"too few calls" message now counts the calls actually judged.

* fix(agent-sessions): count model calls once in the provider-errors headline

Symptom: the Provider errors check read "All 48 model calls were answered
first time" on an OpenRouter Broadcast session of 16 calls ("All 30" for
10), and "All 16 model calls" on a Google ADK session with llmCalls 8.

Cause: buildSessionChecks passed the count of every inference-classified
span (session-checks.ts providerCheck call), so OpenRouter's `provider
attempt N` and `generation` children and ADK's `call_llm` wrapper were
counted beside the call they belong to.

Fix: the headline uses summary.work.llmCalls, the netted count the
session header already shows.

* fix(agent-sessions): judge the prompt cache once per model call

Symptom: on Google ADK sessions the Prompt cache check read "0% over 15
calls" (a1) and "over 19 calls" (b) while the session had 8 and 10 model
calls: the checks did not net ADK's duplicate reporter pair.

Cause: promptCacheCheck took every inference-classified span, so ADK's
`call_llm` wrapper and the `generate_content` span beneath it, both
stamped with the same usage, were judged as two calls each.

Fix: session-summary exports sessionLlmCalls, the netted call list behind
work.llmCalls, and the cache check reads it. The provider headline half
of the same symptom is fixed in the previous commit.

* fix(agent-sessions): count a nested model call's time once in agentTime

Symptom: vitals read 204.8 s of agent time in a 125.5 s OpenRouter
Broadcast session with no parallel work.

Cause: computeAgentTime charged every inference span its full duration,
and Broadcast nests a `provider attempt N` and a `generation` span (both
op `chat`) under each `LLM Generation`, so one call was charged up to
three times. Google ADK's `call_llm` over `generate_content` doubles the
same way.

Fix: an inference span inside another inference span is the same call
observed twice; only the outermost is charged. A TTFT that only a nested
level reported still splits the call.

* fix(agent-sessions): label a turn from its model calls before framework spans

Symptom: every turn of a checkpointed LangGraph thread (OpenInference
LangChain dual-write) was labeled with turn 1's prompt, e.g. all a1 turns
1-6 read "Hi! Briefly introduce yourself.".

Cause: turnLabel (session-turns.ts) fell back to the first span in start
order with any user message. The LangGraph `model` node CHAIN span starts
before its ChatOpenAI call and carries only the thread's first message.

Fix: after the anchor, model calls are asked first, as the transcript's
userRows already does; other spans remain the last resort.

* fix(agent-sessions): read past a "New task:" lead-in in turn labels

Symptom: every smolagents turn label and the session title read
"New task:".

Cause: proseLine (session-turns.ts) labels a turn with the first
non-empty line of the user message, and smolagents sends every task as
"New task:\n<task>".

Fix: a first line of at most three words ending in a colon is a lead-in;
the line after it labels the turn. A longer sentence ending in a colon,
or a lead-in with nothing after it, is kept as before.

* fix(agent-sessions): fold a turn that opened no work into the next one

Symptom: a Microsoft Agent Framework workflow session showed its
one-span `workflow.build` trace as Turn 1 (no label, no agent, 1 span),
the real run as Turn 2, and get_agent_session returned no title.

Cause: `workflow.build` carries the session's id in its own trace and is
read as an agent root by its name, so buildSessionTurns opened a turn on
it; the summary title is turn 1's label.

Fix: on the agent-root and per-trace rules, where the boundary is a
heuristic, an anchor whose spans hold no model or tool call, no user
message and nothing failed joins the next turn, as spans before the
first anchor join turn 1. Conversation-id turns are explicit and kept,
and a session with no work anywhere keeps its anchors.

* fix(agent-sessions): stop at the second line when reading a turn label

A captured prompt can be tens of kilobytes; the label reads at most two
non-empty lines, so it stops collapsing whitespace after them.

* fix(agent-sessions): leave out the first cacheable call, not the first call, in the prompt-cache check

Follow-up to the uncacheable-prompt fix: the size filter ran after
dropping the first call, so a short opening call (a title or router
call) was dropped instead of the first long call, which writes the cache
and then read as a miss.

* fix(agent-sessions): match only smolagents' "New task:" lead-in in turn labels

Follow-up to the lead-in fix: any first line of three words ending in a
colon was skipped, so a user's "Fix this:" over pasted code labeled the
turn with the code's first line. The rule now matches the framework's
literal heading.

* fix(agent-sessions): read a never-written prompt cache as uncacheable on Anthropic only

Follow-up to the uncacheable-prompt fix (#25). The never-written skip
read a zero cache-write bucket as "nothing was cacheable", but OpenAI-
model emitters (CrewAI, Haystack, LiteLLM, Mastra, Microsoft Agent
Framework, Strands, Vercel AI SDK) stamp a zero write bucket on every
call, so an OpenAI session with long prompts and a real 0% hit rate was
skipped instead of warned. The rule now applies only when every
reporting call is an Anthropic call. The separate "too few calls" early
return is dropped: the cacheable-call gate below it covers it.

* fix(agent-sessions): say why nested inference time is charged once, not that it is one call

Follow-up to the agentTime fix (#39): an inference span inside another is
only the same model call under direct nesting; the rule holds because the
nested span's time is inside the outer span's. Comment only.

* fix(agent-sessions): match finish reasons without case or separators

Symptom: Strands TS replies cut off at the output limit were never
flagged; the Reply length check passed.

Cause: Strands TS writes its finish reasons camelCase (`maxTokens`), and
the truncation and refusal matchers (session-findings.ts
TRUNCATION_FINISH_REASONS, session-summary.ts REFUSAL_FINISH_REASONS)
only lowercased before comparing against `max_tokens`/`content_filter`.

Fix: a shared finishReasonsIn helper matches on the reason lowercased
with separators removed, so `maxTokens`, `MAX_TOKENS` and `max_tokens`
are one reason.

* fix(agent-sessions): fold a workless anchor only when it leads straight into the next turn

Follow-up to the empty-turn fix (#41), from review. A workless agent
invocation followed by a pause was folded into the next turn, which then
read the pause as a stall inside it. The fold now needs the next turn to
start within 5 s of the workless anchor ending (setup such as MAF's
`workflow.build`), and merging moves the accumulated bucket forward in
place instead of re-copying it for each workless anchor in a row.

* fix(agent-sessions): charge a model call made under a tool or agent in agentTime

Follow-up to the agentTime fix (#39), from review. The walk to the
outermost inference span crossed agent and tool spans, so a model call a
tool or sub-agent made inside another call was charged nothing. The walk
now stops at an agent or tool: only inference spans nested directly (or
through non-work spans) share the outer span's time.

* fix(agent-sessions): judge a netted call's cache off the observation that reported it

Follow-up to the per-call cache fix (#40), from review. When an app span
and its gateway's mirror share a response id, the netted call list can
keep the app span, which may carry no cache fields, and the cache check
then skipped a session the mirror measured. For such a call the check now
reads the same response's observation that reported cache usage.

* fix(agent-sessions): leave out each uncached Claude call from the prompt-cache check

Follow-up to the uncacheable-prompt fix (#25). The never-written rule was
session-wide and keyed on provider `anthropic` with a write bucket of 0,
so it missed Claude through OpenRouter (flue: provider `openrouter`,
prompts 1,043-1,392 tokens under Haiku 4.5's 4,096 minimum) and
OpenRouter Broadcast (provider `anthropic`, write key absent, prompts
1.2K-2.6K), and a mixed session fell back to judging every call.

A call is now left out on its own when it is Claude (provider
`anthropic` or a `claude` model) and neither wrote nor read the cache,
like the 1024-token cut; the rest of the session is still judged.

* fix(agent-sessions): cover Vercel AI SDK's content-filter refusal in the finish-reason match

Follow-up to the finish-reason fix. The case it fixes today is Vercel AI
SDK's kebab-case ai.response.finishReason: a `content-filter` refusal
read as passed and now warns. Strands TS camelCase reasons live in
output-message parts and reach the checks only once those are read.

* fix(agent-sessions): judge an uncached Claude call once its prompt passes every Claude minimum

Follow-up to the uncacheable-prompt fix (#25). Leaving out every Claude
call that neither wrote nor read the cache also hid real misses when the
emitter does not report writes: 8 Claude calls with 10K-token prompts
and 4 misses read "passed 90% over 3 calls" instead of "51% over 7".
Such a call is now left out only when its prompt is also under 4,096
tokens, the largest Claude minimum; above it, reading nothing is a miss.
The cold/warm cache fixture is back on a Claude model with no write
bucket and 10K prompts, and must warn.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant