Repository navigation
fix(api,web,db): optimistic-locking for dashboards + service-map nested aggregate - #36
Merged
Merged
Conversation
…ed aggregate DeepSec section 4 (dashboard concurrency) + section 6 (service-map SQL). Dashboards previously used read-modify-write with no compare-and-swap, so two concurrent MCP calls or browser edits would silently lose one update. `recordVersion` also computed `latest+1` without a unique index, allowing two concurrent saves to stamp the same `versionNumber`. - Add `dashboards.version` column + unique index on `dashboard_versions(org_id, dashboard_id, version_number)` via migration 0009_dashboard_versioning. - `DashboardPersistenceService.upsert` now reads current version and runs `UPDATE ... WHERE id = ? AND version = ?`. Zero rows affected → typed `DashboardConcurrencyError` (HTTP 409). New `mutate(transform)` API runs the read-CAS-write loop with up to 5 retries before surfacing the conflict; used by all 5 widget MCP tools via `withDashboardMutation`. - `update-dashboard` MCP tool: preserve `existing.variables` in both the metadata-only and full-replacement branches (was wiping it). - `reorder-dashboard-widgets` MCP tool: validate layout geometry server-side (`x>=0, y>=0, 1<=w<=12, h>=1, x+w<=12`). - `useDashboardStore`: per-dashboard FIFO mutation queue so back-to-back edits can't race; on `DashboardConcurrencyError` roll back the optimistic update, surface a refetch banner, and trigger `useAtomRefresh`. - 3 vitest cases in `apps/api/src/mcp/tools/__tests__/dashboard-concurrency.test.ts` exercise concurrent `mutate` (both succeed via retry), concurrent `upsert` (loser surfaces typed conflict), and the refetch+retry recovery flow. Service map (`serviceDependenciesSQL` + `serviceDbEdgesSQL`) was returning "Aggregate function sum(callCount) AS callCount is found inside another aggregate function in query." ClickHouse's UNION-ALL+GROUP-BY optimizer was pushing the outer `sum(callCount)` into each branch where the inner already aliased its aggregate as `callCount`, producing the rejected `sum(sum(...))`. - Inner branches now expose distinct `bucket*` aliases so the outer aggregate has no name to fold against. - Outer average is now properly weighted: `sum(bucketDurationSumMs) / nullIf(sum(bucketCallCount), 0)` instead of `avg(avgDurationMs)` (which averaged averages, ignoring relative call counts per branch). - MV branch reads `Hour < toStartOfHour(endTime)` (complete buckets only), raw branch reads `Timestamp >= toStartOfHour(endTime)` to `endTime` (the in-progress hour). Previously both branches overlapped on the trailing hour and double-counted spans. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Contributor
|
This run croaked 😵 The workflow encountered an error before any progress could be reported. Please check the link below for details. |
…eanup filter
The `cleanupStaleTinybirdDeployments` source already restricts deletion to
terminal `failed`/`error` statuses (this is the plan section 3 fix that
landed on main earlier), but the test still asserted the old behaviour
("delete deploying + data_ready"). Align the test fixture with the actual
filter so CI on this branch passes.
This matches the uncommitted fix already present on the user's main worktree.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
`AlertDestinationCreateRequest` is a `Schema.Union` of `Schema.Class`
constructors, so `Schema.encodeUnknownSync` rejects plain objects on the
input side — they aren't instances of the union members. The test was
passing plain objects directly, which broke when the configs were migrated
to `Schema.Class`.
Construct the inputs with `new SlackAlertDestinationConfig({...})` etc. so
encoding has a class instance to project to the plain wire shape. Same
assertions, same wire-format expectations — only the input side changes.
Pre-existing breakage on main, not introduced by this PR; fixing here so
the branch's CI is green.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
`packages/clickhouse-cli` has no test files but the `test` script ran `bun test`, which exits 1 when no files match the test glob. Turbo then fails the whole pipeline on this package even though there's nothing to test. Remove the script so turbo skips this package's `test` task instead. Pre-existing breakage on main, surfaced once the unrelated domain test failures stopped masking it. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
JeremyFunk
added a commit
that referenced
this pull request
Sep 28, 2026
) Seeds a reused Strands agent over an event-loop span (two turns) and an OpenRouter request whose generation and provider attempt both failed, under their own org, and runs the real compiled page query: 595 tokens and 2 calls for the Strands session, 1 call for the failed request. The unit tests only compare the netting SQL's text. The suite runs in CI once the ClickHouse E2E job lists this file.
JeremyFunk
added a commit
that referenced
this pull request
Sep 29, 2026
* fix(agent-sessions): stop zero-usage spans from absorbing their children's tokens
Symptom: Vercel AI SDK sessions showed exactly twice their tokens on the
session page (docs-verify-vercel-ai-sdk a2: 680 for two chat calls totalling
340; a1 3930 for 1965).
Cause: with `usage: true` the SDK stamps
`ai.usage.outputTokenDetails.reasoningTokens="0"` on the `step` span between
`invoke_agent` and `chat`. That decodes to a reasoning bucket of 0, so
`spanTokenBuckets` returned a zero total and `countableUsageSpans` treated the
step as a reporter. `chargeToNearestReporter` then charged each chat to its
step (netting to nothing) and the agent span kept its whole roll-up.
Fix: a span whose buckets total zero is not a reporter, which is already the
list SQL's rule (`usageReportersExpr` admits `Tokens > 0`).
Seen in: Vercel AI SDK 7 capture docs_vercel-ai-sdk_a.
* fix(agent-sessions): net list usage through spans that report nothing
Symptom: the sessions list showed exactly twice the tokens of the session
page for Strands (docs-verify-strands a1 4712/262 vs 2356/131, b 6146/1246
vs 3073/623) and for the Vercel AI SDK (a2 680 vs 340).
Cause: the list SQL charged each reporter to its direct parent
(`childClaimsExpr` keyed on ParentSpanId, ai-span-columns.ts:121). Strands
puts `execute_event_loop_cycle` between `invoke_agent` and `chat`, the
Vercel AI SDK a `step` span, smolagents a `Step N` chain; the claims were
keyed on those spans, no reporter picked them up, and the agent span kept
its full roll-up. The session page walks the whole ancestry.
Fix: each trace also collects its links (span id -> itself when it
reported usage, else its parent). The session level climbs up to 8 links
from each reporter's parent to the nearest ancestor that reported and
carries it as a twelfth tuple element, which the claims and the "any
ancestor reported" test for unreported calls key on. Verified against the
EU replay: list totals now equal the session page for every Strands, Vercel
AI SDK and smolagents session. A week of the production org nets to the
same totals as before at the same read time (~370ms either way).
Limit: a parent outside the index (a span without a vendor stamp, e.g. the
Strands TypeScript loop span under a custom service name) still stops the
climb.
Seen in: captures docs_strands_{a,b,g}, docs_vercel-ai-sdk_{a,b}.
* fix(agent-sessions): count a failed gateway request once, not once per attempt
Symptom: an OpenRouter Broadcast request whose every provider attempt failed
counted as 2 LLM calls on the list and the session page
(trace:83d675eb59e436078591a758e28cb09a: `LLM Generation` plus
`provider attempt 1: OpenAI`, both op `chat`, both Error, no usage).
Cause: a model call that reported no usage counted unless an ancestor
reported usage (session-summary.ts `countedLlmCalls`, ai-span-columns.ts
`nettedReportersExpr`). The attempt netting only worked when the generation
above it reported usage, which a failed request never does.
Fix: a model call that reported no usage does not count when its parent is
a model call. On the list the session's reporter ids now carry every
reporter (usage and model calls), so the check is one `has` on the direct
parent next to the existing one on the charged ancestor. Verified on the
EU replay (the trace now counts 1; other sessions unchanged) and on a week
of the production org (378 unreported calls counted before and after).
Seen in: OpenRouter Broadcast capture (openrouter guide, failed request).
* fix(agent-sessions): net cost the same way on the list and the session page
Symptom: nested agents that each stamp `gen_ai.usage.cost` summed to
different session costs on the list and the session page (LiteLLM guide,
orchestrator delegating to workers through tool spans): the page subtracted
the sub-agents' cost from the orchestrator's roll-up, the list did not.
Cause: two rules differed. The list charged a claim to its direct parent,
which for a sub-agent is its tool span, so nothing was netted; that half is
fixed by the full-ancestry netting in the list commit before this one. The
page also took a span stamped with a zero cost as a cost reporter
(`costBySpan`, `cost >= 0`, session-summary.ts), where the list only
charges `Cost > 0`: a zero-cost wrapper absorbed its calls' cost and left
the agent above it keeping its whole roll-up, the #28 shape for cost.
Fix: a zero cost is recorded (so the session still reads "free", not
"unmeasured") but is not a reporter claims are charged to.
Seen in: LiteLLM capture docs_litellm_b (nested agents), reconstructed as a
fixture; the replayed session already agrees because the guide now prices
only the outermost agent.
* fix(agent-sessions): count a tool call paused for a human once
Symptom: a Strands session with one human-approved `delete_file` call
showed it twice in the session page's tool ledger and tool-call count
(docs-verify-strands a1: 4 tool calls, `delete_file` x2, for 3 calls).
Cause: the interrupted call ends its `execute_tool` span Ok with no result,
and the resumed turn's trace opens a second span under the same
`gen_ai.tool.call.id` (call_ltoPrLQIg3ZHWkyFnBIOE65u, traces df191678… and
18645746…). `buildSessionSummary` counted every tool span
(session-summary.ts `work.toolCalls`, `toolUsage`).
Fix: `countedToolCalls` takes the session's tool spans with the ones
sharing a call id collapsed to the last to start (the resumed copy, which
carries the result); spans without an id count as before. Both the count
and the ledger read it.
The list's tool-call count is not changed: `ai_trace_index` carries no
call id, so deduping it there needs a new index column and a recreated
materialized view. Across the EU replay (every framework) this is the only
shared call id among tool spans, and a week of the production org has
none.
Seen in: capture docs_strands_a (HITL resume).
* fix(agent-sessions): climb list claims per measure, as the session page charges them
Review follow-up to the #7 list netting. The list climbed one chain for
both measures (nearest ancestor with tokens or cost), while the session page
charges tokens to the nearest token reporter and cost to the nearest cost
reporter. A wrapper that priced but did not count (or the reverse) between
an agent and its call stopped the other measure's claim, and the agent kept
its roll-up of it on the list only.
Each trace now carries two links maps (`tokenLinks`, `costLinks`); each
reporter carries both ancestors (elements 12 and 13), and the child-claims
sumMap enters each reporter twice, its tokens under one and its cost under
the other. The climb is 4 links (past up to three non-reporting spans; the
deepest shape seen is two) to keep the added map lookups down: a
production week netted in full reads ~360-440ms against ~220-370ms before
the netting change, same totals; EU replay totals unchanged.
* fix(agent-sessions): merge only a paused tool call into its resumed copy
Review follow-up to #21. Keying the tool calls on `gen_ai.tool.call.id`
alone merged any two calls sharing an id, and ids are not unique across a
session for every emitter (parallel lanes, providers that number calls per
turn). Only a span that recorded no result and did not fail is now dropped,
and only when a later span with the same id carries a result — the shape a
human-approval pause leaves. A session captured without payloads keeps
every span.
* fix(agent-sessions): count nothing of an agent span whose calls reported (cumulative reporters)
#7, cumulative reporters. Symptom: agents that live across turns report the
conversation so far on their agent span, and both pages counted that
excess over the agent's own calls again.
- smolagents `run(reset=False)` (capture cap_a_v1, session a1): run spans
report 1067, 2243, 4857, 7811 cumulatively over calls summing to 7811;
both pages showed 15978 (2.05x).
- Strands with one Agent reused across requests (capture strands_user,
scenario a): 18044 shown against 4541 billed (~4x).
Cause: a wrapper kept whatever it reported above the reporters beneath it
(`countableUsageSpans` / `costBySpan`, `nettedReportersExpr`), which is
right for a model-call wrapper whose child call reported nothing and wrong
for an agent whose excess is earlier turns.
Fix: a reporter that is not a model call (an agent, a workflow, any wrapper
the index does not flag `IsLlmCall`) claims nothing of a measure once a
reporter of that measure is charged to it; model-call wrappers (an SDK's
`generateText` over `doGenerate`, a gateway generation over its attempts)
keep their excess. Same rule on the session page (`keptClaim`) and the
list (`if(r.6 = 0 AND charged > 0, 0, …)` per measure). Verified with the
compiled page query over literal index rows on ClickHouse 25.8 (smolagents
7811, Strands 1031 for three turns); the EU replay and a production week
have no agent span with an excess, so their totals are unchanged.
* fix(agent-sessions): say which frameworks the paused tool-call merge covers (#21)
OpenAI Agents stamps no gen_ai.tool.call.id on its tool spans, so its
human-approval copies still count twice; the comment no longer lists it.
* test(agent-sessions): prove the list netting on real index rows (#7, #36)
Seeds a reused Strands agent over an event-loop span (two turns) and an
OpenRouter request whose generation and provider attempt both failed, under
their own org, and runs the real compiled page query: 595 tokens and 2 calls
for the Strands session, 1 call for the failed request. The unit tests only
compare the netting SQL's text. The suite runs in CI once the ClickHouse E2E
job lists this file.
* fix(agent-sessions): state which paused tool-call copies the #21 merge covers
Covered: Strands, which records the resumed call's result on span
attributes. Not covered, still counted twice: Google ADK (outcome only in
gcp.vertex.agent.tool_response, and the paused copy's confirmation request
reads as a result), Strands versions that record results in span events,
and OpenAI Agents (no gen_ai.tool.call.id).
This branch was previously deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.

Summary
Closes Section 4 (dashboard concurrency / lost updates) and Section 6 (service-map nested-aggregate SQL error) of the DeepSec triage in
.claude/plans/we-created-an-depsec-cuddly-tarjan.md.dashboards.versioncolumn + CAS upsert. Two concurrent MCP calls or browser edits can no longer silently lose one update; the loser surfaces a typedDashboardConcurrencyError(HTTP 409) and the web hook auto-refetches.sum(sum(...)); aliases are renamed and the average is now properly weighted across MV + raw branches.What changed
Schema (held until human review)
packages/db/drizzle/0009_dashboard_versioning.sql— addsdashboards.version INTEGER NOT NULL DEFAULT 0and a unique index ondashboard_versions(org_id, dashboard_id, version_number).packages/db/src/schema/dashboards.ts,packages/db/drizzle/meta/0009_snapshot.json,_journal.jsonupdated to match.Domain
DashboardConcurrencyError(HTTP 409) added to@maple/domain/http. Wired into theupsert,restoreVersion,create, andinstantiateTemplateHTTP endpoints' error unions.Persistence
DashboardPersistenceService.upsertnow reads current version and runsUPDATE ... WHERE id = ? AND version = ?. Zero rows affected → conflict.DashboardPersistenceService.mutate(orgId, userId, dashboardId, transform)runs the read-CAS-write loop with up to 5 retries before surfacing the typed conflict. The MCPwithDashboardMutationhelper now delegates to it, so every widget tool inherits retry-on-conflict.MCP tools
update-dashboard: preservesexisting.variablesin both the metadata-only and full-replacement branches (the full-replacement path was silently wiping it).reorder-dashboard-widgets: server-side geometry validation (x>=0, y>=0, 1<=w<=12, h>=1, x+w<=12).Web client
useDashboardStorenow serializes mutations per dashboard with a FIFO queue (Map<id, Promise>), so back-to-back edits can't race against each other.DashboardConcurrencyErrorthe optimistic update rolls back, a banner surfaces the conflict, anduseAtomRefreshtriggers a refetch — the existing list-result effect clears the banner once fresh state lands.Service map (separate root-cause, same PR)
serviceDependenciesSQLandserviceDbEdgesSQLexpose distinctbucket*aliases so the outersum(...) AS callCounthas no inner name to fold against.avg(avgDurationMs)(averaging averages) tosum(bucketDurationSumMs) / nullIf(sum(bucketCallCount), 0)(properly weighted across branches).Hour < toStartOfHour(endTime)only; raw branch readsTimestamp >= toStartOfHour(endTime)toendTime. Previously both overlapped on the trailing hour and double-counted spans.Tests
apps/api/src/mcp/tools/__tests__/dashboard-concurrency.test.ts— 3 vitest cases:mutatecalls both land via retry — no lost update.upsertcalls — loser surfacesDashboardConcurrencyErrorinstead of silently overwriting.Whole api test suite (279 tests across 20 files) green.
Test plan
bun typecheckclean on@maple/api,@maple/web,@maple/db,@maple/domain,@maple/query-engine(only pre-existing unrelatedworker.tsimport error remains)bun testfromapps/api— 279 pass, 0 fail0009_dashboard_versioning.sqlapplies cleanly viawrangler d1 migrations apply --localversion=1, audit rowcreated), added a Bar Chart widget (version=2, audit rowwidget_added), fired two parallelPUT /api/dashboards/:idrequests (both committed against fresh state,versionadvanced 2→4 — local D1 serializes so neither write actually overlapped, but both succeeded which is the correct CAS behavior; the actual race-window assertion is covered by the unit tests)sum(sum(...))patterns, singleAS callCount(outer only) in both queriesReviewer notes
0009_dashboard_versioning.sqlbefore running it through the release flow. SQLite-friendly (ADD COLUMN ... DEFAULT 0 NOT NULL), no backfill required.upsertdoes single-attempt CAS (no retry) so the client gets a 409 to refetch from. MCPmutateretries internally because the transform is re-applied on top of fresh state — that's safe; HTTP retry would silently overwrite.recordVersionis still wrapped inEffect.ignore(best-effort audit). The new unique index protects against duplicateversion_numberrows under racyrecordVersioncalls; the worst case is a missed audit row, which is acceptable given the dashboard-table CAS guarantees the actual state is consistent.🤖 Generated with Claude Code
Need help on this PR? Tag
@codesmithwith what you need.