Repository navigation
fix(agent-sessions): normalise GenAI usage, cost, TTFT and agent name at ingest - #1123
JeremyFunk wants to merge 19 commits into
Conversation
…ive whatever provider it names Symptom: a session of Claude calls forwarded by OpenRouter Broadcast read 32,573 tokens where OpenRouter billed 28,325 (session f0f992b0): every cache read was counted twice. Cause: Broadcast stamps the upstream as `gen_ai.provider.name` (`anthropic` for Claude models) while reporting OpenRouter's own OpenAI-shaped usage, whose prompt figure already contains the cached tokens. `openrouter` was only in GENAI_PROVIDER_USAGE_CONVENTIONS, so the vendor lookup missed and the provider lookup applied Anthropic's excludes-cache rule. Fix: `openrouter` joins GENAI_VENDOR_USAGE_CONVENTIONS as NESTED, so the vendor decides before the provider does, on the detail page and in the index expression. Seen in: OpenRouter Broadcast capture (`openrouter`), 153 spans stamped `gen_ai.provider.name=anthropic`.
…ude Code, not on the provider Symptom: an Anthropic session traced by `opentelemetry-instrumentation-genai-anthropic` read 78,144 tokens (input 40,092 + cacheRead 30,324 + cacheWrite 7,581 + output 147) where the instrumentation reported 40,239: every cache read and write was counted twice, on the list and on the detail page. Cause: GENAI_PROVIDER_USAGE_CONVENTIONS (packages/domain/src/gen-ai.ts) applied the raw Messages API rule (`input_tokens` excludes both cache buckets) to every span whose provider is `anthropic`. The semconv's Anthropic mapping has the instrumentation add the cache buckets to `gen_ai.usage.input_tokens`, and spec-conformant emitters do (the genai-anthropic instrumentation; Pydantic AI, whose `input_tokens` is documented as cache-inclusive; Vercel AI SDK spans that land in `unknown:genai`). The only emitter in the captures that passes the raw figures through is Claude Code, whose `claude_code.llm_request` usage the gateway restates verbatim. Fix: `anthropic` leaves the provider table (it takes the nesting default), and `claude_agent_sdk` joins GENAI_VENDOR_USAGE_CONVENTIONS with the excludes-cache rule, so the emitter decides. Seen in: docs_provider-sdks_anthropic (genai-anthropic 1.2b0, chat span input_tokens=7911, cache_write.input_tokens=7581); Claude Agent SDK captures keep their current totals.
…lings of the reasoning and cache-write buckets Symptom: reasoning tokens from Mastra and Pydantic AI, and OpenRouter Broadcast's cache writes, never reached a session's usage; a cache write read as plain prompt tokens. Cause: the usage alias lists (GENAI_LEGACY_ALIASES in ai-integrations.ts, GENAI_USAGE_KEYS in gen-ai-columns.ts) did not carry `gen_ai.usage.reasoning_tokens` (Mastra), `gen_ai.usage.details. reasoning_tokens` (Pydantic AI) or `gen_ai.usage.input_tokens.cache_write` (OpenRouter Broadcast, documented beside the `.cached` read key Maple already maps). Fix: all three join the default integration's aliases and the index's bucket key lists, after the canonical keys, so a span carrying both spellings still reads the canonical one. Seen in: docs_mastra_a/b, mastra_* (`gen_ai.usage.reasoning_tokens`); docs_pydantic-ai_a/b/lf, pydantic_ai_* (`gen_ai.usage.details. reasoning_tokens`); OpenRouter Broadcast OTel collector docs (`gen_ai.usage.input_tokens.cache_write`).
… vendor, as the list does Symptom: an Agno session was priced at $0.0042 on the Agent Sessions list while its detail page (and `get_agent_session`) showed no cost. Cause: the list's `Cost` column reads `llm.cost.total` on every span (GENAI_COST_KEYS), but the detail page mapped it only through the OpenInference integration (ai-vendors.ts), which is registered for `openinference-openai` and `unknown:openinference` alone. Agno, DSPy, smolagents, CrewAI and the OpenAI Agents SDK emit OpenInference under their own vendor stamps. Fix: `llm.cost.total` moves into the default integration's cost aliases, so every vendor decodes it and the two pages read the same keys. Seen in: docs_agno_a/b (`llm.cost.total` on `OpenRouter.invoke` / `OpenRouter.ainvoke` spans, vendor `agno`).
Symptom: LiteLLM and Pydantic AI sessions showed as unpriced on the list and on the detail page although every model call carried a price. Cause: neither cost key was in any alias list: LiteLLM writes `litellm.cost.total` (beside `litellm.cost.input`/`.output`/margins) and Pydantic AI writes Logfire's `operation.cost`. Maple read only `gen_ai.usage.cost`, `gen_ai.usage.total_cost` and `llm.cost.total`. Fix: both join the default integration's cost aliases and the index's GENAI_COST_KEYS, after the existing keys. Both are the instrumentation's own price for the call from its price table (LiteLLM's model cost map with any proxy margin, Pydantic AI's genai-prices), the same kind of figure OpenLLMetry's `gen_ai.usage.cost` and OpenInference's `llm.cost.total` already are, so they read into the same field. Seen in: docs_litellm_a/b/proxy (`litellm.cost.total` 2.895e-05 on `chat openai/gpt-4o-mini`); docs_pydantic-ai_a/b/lf and pydantic_ai_* (`operation.cost` 4.755e-05 on `chat openai/gpt-4o-mini`).
…outer Broadcast and Strands Symptom: no TTFT on the detail page's waterfall and vitals for Pydantic AI, OpenRouter Broadcast and Strands model calls, all of which report it. Cause: `responseTimeToFirstChunk` read `gen_ai.response.time_to_first_ chunk` for every vendor and `gen_ai.client.operation.time_to_first_chunk` for the Vercel AI SDK alone (ai-vendors.ts). Pydantic AI writes the latter too; OpenRouter Broadcast writes `trace.metadata.openrouter.first_token_ms` and Strands writes `gen_ai.server.time_to_first_token`, both in milliseconds (a Strands chat span of 1,511 ms reports 1127). Fix: `gen_ai.client.operation.time_to_first_chunk` (seconds) moves to the default integration's aliases. `openrouter` and `strands` get integrations whose refine lifts their millisecond key into the field as seconds, only when no seconds-valued key already set it. The index carries no TTFT, so no warehouse change is needed. Seen in: docs_pydantic-ai_a/b/lf, docs_vercel-ai-sdk_a/b, eve_slack (`gen_ai.client.operation.time_to_first_chunk`); openrouter (`first_token_ms` 3784 on `LLM Generation`); docs_strands_a/b/g, strands_* (`gen_ai.server.time_to_first_token`).
Symptom: a LangChain agent traced through LangSmith's OTel export showed no agent name on the list (no Agent facet value) or on the detail page. Cause: LangSmith names the agent in `langsmith.metadata.lc_agent_name` on every span of the run and writes no `gen_ai.agent.name`; neither the default integration nor GENAI_AGENT_NAME_KEYS read it. Fix: the key joins the default integration's `agentName` aliases and the index's agent-name key list, after the canonical key. Seen in: docs_langchain_ls (50 spans with `langsmith.metadata.lc_agent_name=assistant`, none with `gen_ai.agent.name`). The second half of the report (LangSmith `GraphInterrupt` read as a failure) is a failure-classification change and is not part of this batch.
…dual-write is off Symptom: a CrewAI crew traced with OpenInference's defaults showed no agent names: the list's Agent facet and the detail page's agent scopes were empty for every sub-agent. Cause: without `enable_genai_semconv`, CrewAI's OpenInference agent spans carry the agent's role only in `graph.node.id`, which no agent-name key list read. Agno's OpenInference spans use the same key for an opaque node id (`c8bddb16e7b7e3cc`), so a plain alias would name Agno agents by hex ids. Fix: a `crewai` integration whose refine lifts `graph.node.id` into `agentName` when nothing else named the agent, and the index's agent-name expression reads the same key under the same vendor gate (CREWAI_AGENT_NAME_KEY, shared by both). Seen in: crewai_agents / crewai_user (`graph.node.id` = orchestrator, weather_worker, ...; no `gen_ai.agent.name`); agno_agents (`graph.node.id` opaque, `graph.node.name` the name).
… usage convention and the added keys The list reads `Tokens`, `Cost` and `AgentName` off `ai_trace_index`, whose materialized view compiles the expressions in gen-ai-columns.ts at creation. The preceding fixes changed those expressions (usage convention keyed on the emitter; Mastra, Pydantic AI and OpenRouter usage spellings; LiteLLM and Pydantic AI cost keys; LangSmith and CrewAI agent names), so the view is recreated for the list to agree with the detail page. Migration 0035 drops and recreates `ai_trace_index_mv`; no column is added. Nothing is backfilled, as in 0026/0029/0031/0032: rows materialized before it keep their old values until raw `traces`' 30-day TTL ages them out. The generated schema, the Tinybird manifest and the local chDB schema (v25 -> v26) are regenerated; the local step only drops and rebuilds the view. The materialization e2e suite seeds one span per emitter under its own org and asserts every changed column off the real view.
…ion summary too The turn summary read off trace_detail_spans (summaryMeasures_ in ai-sessions.ts) reads agentName through the integrations' source keys, which leave out the refine-only CrewAI role key, so its agentNames stayed empty for a CrewAI crew with the GenAI dual-write off while the index and the detail page named the role. It now reads graph.node.id under the same crewai vendor gate.
The vendor table now also holds an emitter that passes a provider's raw figures through (Claude Code) and one that names a provider it did not take its figures from (OpenRouter Broadcast), and the default integration's alias table now holds cross-dialect keys; the comments say so. Drops a requiredForIngest assertion the version test already makes.
|
Note A newer push replaced |
|
Navigate logical layers of code changes, visualize relationships, and explore their blast radius. No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: defaults Review profile: CHILL Plan: Advanced Run ID: ⛔ Files ignored due to path filters (2)
📒 Files selected for processing (27)
Included review availability: This review used your included allowance. Your plan provides up to 4 included reviews per hour; 2 remain after this review. 📝 WalkthroughWalkthroughThe change expands GenAI telemetry attribute handling across integrations, ClickHouse materialization, and session and trace queries. It adds vendor-specific usage conventions and mappings, updates the AI trace index migration and local schema to version 26, and adds migration and materialization coverage. ChangesGenAI telemetry
Priority: ⚪ Not assessed Estimated code review effort: 3 (Moderate) | ~25 minutes Change: Bug fix Sequence Diagram(s)sequenceDiagram
participant Emitter
participant QueryIntegration
participant ClickHouse
participant TraceIndex
Emitter->>QueryIntegration: Provide span attributes
QueryIntegration->>QueryIntegration: Map aliases and vendor-specific values
Emitter->>ClickHouse: Insert spans into traces
ClickHouse->>TraceIndex: Materialize index values from the updated view
Suggested reviewers: Merge Risk: ⚪ Minimal · up to No merge-blocking issue is established. Complete the normal checks and planned Tinybird schema deployment. Security Architecture ReviewSecurity architecture risk: 🟡 Moderate · up to The new calculations do not show an expanded authorization boundary, but a failed view update can leave new sessions absent from the index even when the schema update reports success. Older indexed values also retain their previous meaning until they expire. Retained concerns
Security review detailsSecurity Blast Radius
Trust Boundaries and Controls
Resilience and Maintainability Implications
Hardening Proposals
🚥 Pre-merge checks | ✅ 4 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (4 passed)
Full details: Docstring CoverageExplanation Docstring coverage is 25.00% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 4 functions across 22 files. (4 skipped: 4 unsupported.)
✨ Finishing Touches 💡 1📝 Generate docstrings 💡
🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
There was a problem hiding this comment.
Devin Review found 1 potential issue.
1 flag not posted on this PR by your GitHub settings — view it in Devin Review. (Configure)
| export const GENAI_AGENT_NAME_KEYS = [ | ||
| "gen_ai.agent.name", | ||
| "ai.telemetry.functionId", | ||
| "langsmith.metadata.lc_agent_name", | ||
| ] as const |
There was a problem hiding this comment.
🟡 Conflicting agent names across session views
When both agent-name keys appear, genAiAgentNameExpr chooses ai.telemetry.functionId before LangSmith's name. mapAiSpan chooses LangSmith's name first, so the session list and details disagree.
Learn more
The materialized index uses GENAI_AGENT_NAME_KEYS in order to select one AgentName for each span. The detail mapper builds its key order from genAiSources, then appends vendor keys through mergeSources. This puts LangSmith's key before the Vercel SDK function ID on the detail path, unlike the index. Both values can coexist when the Vercel SDK runs inside a LangSmith-traced agent.
Example: A span contains ai.telemetry.functionId=generateText and langsmith.metadata.lc_agent_name=assistant. The list labels it generateText; its details label it assistant.
Recommended fix: Put the two aliases in the same priority order in GENAI_AGENT_NAME_KEYS and the mapper, then add a test supplying both keys for vercel_ai_sdk.
Was this helpful? React with 👍 or 👎 to provide feedback.
There was a problem hiding this comment.
Fixed in b003a48: GENAI_AGENT_NAME_KEYS now follows the detail page's order (gen_ai.agent.name, lc_agent_name, then functionId), pinned in ai-span-columns.test.ts; 0035 regenerated.
…nction id in the index, as the detail page does The detail page reads the default integration's agentName keys (gen_ai.agent.name, then langsmith.metadata.lc_agent_name) before a vendor's own (ai.telemetry.functionId for vercel_ai_sdk), while the index read the function id second: a span carrying both would be named differently on the list and the detail page. GENAI_AGENT_NAME_KEYS takes the detail page's order, pinned by ai-span-columns.test.ts, and migration 0035 and the local v26 schema are regenerated with it.
Maple reviewConfidence 4/5 · likely safe to merge Re-keys Agent Sessions' usage convention on the emitter (
What was checked
|
The suite that checks every ai_trace_index column against the real view (ai-trace-index-materialization.clickhouse.e2e.test.ts) was never in the ClickHouse E2E job, so its assertions only ran where someone started a local ClickHouse. It now runs beside the trace facets rollup check.
…cache_write_input_tokens (#11) Older Strands releases write both cache buckets as gen_ai.usage.cache_read_input_tokens / cache_write_input_tokens (captures strands_user, strands_agents), which no alias list read. Both join the default integration's cache aliases and the index's bucket key lists; migration 0035 and the local v26 schema are regenerated with them, and the materialization e2e seeds a Strands span carrying them.
GENAI_LEGACY_ALIASES now also holds other dialects' keys any vendor's span can carry, so it becomes GENAI_DEFAULT_ALIASES and its generated test titles stop calling every key deprecated.
Maple reviewConfidence 4/5 · likely safe to merge
What was checked
|
…age-keys # Conflicts: # packages/backend/src/services/warehouse/ai-trace-index-materialization.clickhouse.e2e.test.ts # packages/query-engine-integrations/src/__sql_baseline__/integrations.sql # packages/query-engine-integrations/src/ai/ai-sessions.ts # packages/query-engine-integrations/src/ai/ai-vendors.ts
Maple reviewConfidence 4/5 · likely safe to merge Moves the gen_ai usage convention from the provider to the emitter (anthropic out, claude_agent_sdk and openrouter in), adds the Mastra, Pydantic AI, LiteLLM and LangSmith usage/cost/agent-name keys, and recreates ai_trace_index_mv in migration 0035 so the list and detail page agree. Safe to merge.
What was checked
|
…e canonical key each Frameworks spell these facts their own way (ai.usage.*, llm.token_count.*, litellm.cost.total, a TTFT in milliseconds), and Claude Code's and Gemini's usage figures leave the cache or reasoning buckets out where the semconv folds them in. Every reader re-derived that per vendor and provider. The gateway now settles it once per stamped span (ai_session/canonical.rs): the first parseable spelling of each bucket, the cost, the TTFT and the agent name is written under its gen_ai.* key, Claude Code's input gains both cache buckets and a Gemini provider's output gains the reasoning, so gen_ai.usage.* always carries the semconv meaning. A value already canonical and unfolded is left as the emitter wrote it; dialect keys are kept.
… one key each The ingest gateway now restates every spelling and the usage convention, so the read side drops them: the vendor/provider usage-convention tables, the usage/cost/TTFT/agent-name aliases in the default, Vercel AI SDK and OpenInference integrations, the CrewAI and millisecond-TTFT refines, and the matching key lists and convention branches behind ai_trace_index. Token buckets are carved out under the one semconv convention on both the list and the detail page. Spans ingested before the gateway change keep their old spellings and read as unpriced / without usage on the detail page; that data is not migrated.
…, cost and agent-name keys Migration 0035 and the local chDB v26 step now carry the view compiled from the simplified expressions: one key per fact, no dialect coalesces, no vendor or provider branches. Still a drop-and-recreate of the view only; nothing is backfilled.
Maple reviewConfidence 3/5 · needs attention Ingest now restates every AI span's usage, cost, TTFT and agent name under one canonical
FindingsWarning · F1 · Detail page under-counts tokens for spans ingested before the deploycorrectness ·
Note · F2 ·
|
| input: convention.inputIncludesCache | ||
| ? Math.max(0, reportedInput - cacheRead - cacheWrite) | ||
| : reportedInput, | ||
| input: Math.max(0, reportedInput - cacheRead - cacheWrite), |
There was a problem hiding this comment.
Detail page under-counts tokens for spans ingested before the deploy
F1 · Warning · correctness
spanTokenBuckets now carves both cache buckets out of the prompt figure unconditionally, which is only correct once ingest has folded them in. Spans already in traces were stamped by the old gateway, so a Claude Code (claude_agent_sdk) span still carries Anthropic's exclusive gen_ai.usage.input_tokens: for raw input_tokens: 5000 with cache_read: 1000 the detail page and get_agent_session show 5000 instead of 6000, because the 1000 is subtracted from a figure that never held it. The index rows keep their v25 Tokens, so for the 30-day window the list and the detail page disagree — the opposite of what the change sets out to fix.
The pre-deploy rows are indistinguishable from restated ones at read time, so either state the dip in migration 0035 and the `local-schema-history` v26 entry (they only mention the index's stale values), or have `canonical::normalize` leave a marker the readers can test before carving.
Prompt for an AI agent
In `packages/agent-sessions/src/session-summary.ts:466`: Detail page under-counts tokens for spans ingested before the deploy.
`spanTokenBuckets` now carves both cache buckets out of the prompt figure unconditionally, which is only correct once ingest has folded them in. Spans already in `traces` were stamped by the old gateway, so a Claude Code (`claude_agent_sdk`) span still carries Anthropic's exclusive `gen_ai.usage.input_tokens`: for raw `input_tokens: 5000` with `cache_read: 1000` the detail page and `get_agent_session` show 5000 instead of 6000, because the 1000 is subtracted from a figure that never held it. The index rows keep their v25 `Tokens`, so for the 30-day window the list and the detail page disagree — the opposite of what the change sets out to fix.
Suggested fix: The pre-deploy rows are indistinguishable from restated ones at read time, so either state the dip in migration 0035 and the `local-schema-history` v26 entry (they only mention the index's stale values), or have `canonical::normalize` leave a marker the readers can test before carving.
Verify the problem exists at that location before changing it, and keep the fix to those lines.
Agent Sessions read the wrong usage convention for some emitters, and missed usage, cost, TTFT and agent-name keys that frameworks actually write (bugs #8, #10, #11, #12, #23a, #24, #31, #35 from the agent-tracing docs verification, each from a real OTLP capture).
Design: normalise on the write side. The ingest gateway restates these facts once per stamped span; every reader (the
ai_trace_indexview, the session page, the MCP tools) reads onegen_ai.*key per fact with one meaning. The read-side alias lists, vendor refines and vendor/provider usage-convention tables are deleted.Ingest:
apps/ingest/src/ai_session/canonical.rs(070e5f9e2)Runs on every vendor-stamped span, after the Claude Code restatement.
gen_ai.usage.{input_tokens, cache_read.input_tokens, cache_creation.input_tokens, output_tokens, reasoning.output_tokens}. Coversgen_ai.usage.prompt_tokens/completion_tokens, OpenRouter (input_tokens.cached,input_tokens.cache_write,output_tokens.reasoning), older Strands (cache_*_input_tokens), Mastra and Pydantic AI reasoning, the Vercelai.usage.*and OpenInferencellm.token_count.*dialects, and the registrycache_write.input_tokensspelling.claude_agent_sdk): cache buckets are folded into input (fix: detect wrong-kind ClickHouse objects, restrict Tinybird stale-deployment cleanup #35/fix-native-edit-shortcuts #8: Anthropic's raw figure excludes them).gen_ai.usage.cost←gen_ai.usage.total_cost,llm.cost.total,litellm.cost.total,operation.cost(Enterprise tinybird selfhosted #10, chore(ci): SHA-pin actions, gate publish/deploy with environments, fix RCE in tag input #31).gen_ai.response.time_to_first_chunk(s) ←gen_ai.client.operation.time_to_first_chunk; OpenRoutertrace.metadata.openrouter.first_token_msand Strandsgen_ai.server.time_to_first_token÷ 1000 (Redesign #12).gen_ai.agent.name←langsmith.metadata.lc_agent_name,ai.telemetry.functionId, CrewAI'sgraph.node.id(crewai only) (#23a, Cf migration #24).canonical.rsunit tests (one per rule). The Claude Code restatement test now expects the folded input (2 + 114,514 + 3,549).Read side (
cb1b2e764)gen-ai.ts: usage-convention tables removed.ai-integrations.ts/ai-vendors.ts: usage/cost/TTFT/agent-name aliases, the CrewAI refine and the ms-TTFT refines removed.gen-ai-columns.ts: one key per usage bucket, cost and agent name.byConventionand the provider-key lists are gone. Tokens = the disjoint buckets summed under the one convention.spanTokenBucketsalways carves the cache out of input and the reasoning out of output.ai-span-columns.test.tspins that the list and the page read the same single key.Warehouse: migration 0035 (
8db109d70)ai_trace_index_mvwith the simplified expressions. No column is added and nothing is backfilled.Old data (accepted)
get_agent_sessionno longer read those spellings, so such spans show no usage/cost/TTFT/agent name.Deploy
tinybird:deploy(prod schema).Overlap
ai_session.rs. Whichever lands second rebases.