Skip to content

feat(agent): integrate Pi V2 with recovery, charts and object tools - #2906

Merged
openai0229 merged 134 commits into
release/pi-v2from
feat/agent-v2-contracts
Sep 23, 2026
Merged

openai0229 merged 134 commits into
release/pi-v2from
feat/agent-v2-contracts

Conversation

@openai0229

@openai0229 openai0229 commented Sep 11, 2026 •

Copy link
Copy Markdown
Contributor

Related issue

N/A — the Pi V2 runtime and its follow-up work were requested by the maintainer.

Summary

Pi V2 adds a second, local agent runtime to Community next to the existing Spring AI chat: persistent conversations with datasource context, tool execution with approvals, user questions, saved output files, charts, and scoped database object discovery. Web and desktop share one typed Pi client and one backend operation registry, so a new operation no longer needs a separate route and binding per transport.

Scope: 497 files, +33,090/−296 against main (37f0e126f), which this branch has merged.

  • Runtime and transport — 30 typed operations in chat2db-community-client/src/service/pi/contract.ts. Both transports post the same versioned envelope to /api/v3/ai/pi/invoke, where PiOperationRegistry is the single allowlist, payload validator and serializer, and calls the existing controllers, so authorization, ownership and run semantics stay shared with REST.
  • Sessions and storage — V2 storage keeps one directory per session with separate files for the session, runs, events, approvals and artifacts, and names event files by their sequence. Event pages are range reads by sequence in both directions: opening a long conversation reads its newest page, scrolling up costs one range read per request, and polling a running turn reads only the page after the cursor. Streamed text deltas are merged into a few durable events before they are written. The mixed-version session list keeps V1 conversations visible and routable.
  • Run orchestration and recovery — a run is accepted with a frozen model, skill set and context; tool chronology and timings are persisted; suspended runs stay resumable; completion is evidence-based; startup races, stale errors and orphaned handles are handled, and deleting a session is serialised with starting a run.
  • Pi runtime — on-demand installation from official release archives with a pinned manifest and checksum verification, bounded download and decompression, and task-center integration; process supervision with diagnostics; a strict JSONL RPC client with a silence watchdog; the packaged runtime launch path.
  • Model access — Pi native model protocols (OpenAI chat-completions and responses, Anthropic messages, Google generative-ai) with gateway proxying; the model ticket is passed by configuration and is never exported to the child environment.
  • Tools and approvals — database tools (db_search_datasources, db_search_databases, db_search_schemas, db_search_objects, db_describe_objects, db_query), render_chart, file tools (read, grep, ls, find), and Bash/PowerShell behind a beta gate with per-command approval. Approvals and questions have their own storage and in-conversation cards, and state-changing agent operations are restricted to the desktop bridge or a loopback caller.
  • Query capture bounds — the SPI reads large JDBC values under one shared per-invocation UTF-8 budget (chat2db.agent.v2.outputs.max-capture-bytes, 32 MB default) without materialising them first, marks truncated values with a reason, and the DM and SQL Server executors use the same reader.
  • Outputs and charts — oversized results are saved as conversation-scoped artifacts with paged reads, search and download; charts render from saved query results, including grouped, stacked and horizontal layouts, with a bundled chart skill.
  • Skills — bundled chart and skill-manager skills are installed under the product storage root and replaced in place only when the packaged content changes; user skills are read where they live (~/.chat2db-skills/) and stay editable while a session has them loaded; bundled skills stay read-only; installing a different skill under an existing name asks the user to overwrite or rename.
  • External MCP servers — a user connects a server by asking for it; there is no settings page. A bundled mcp-manager skill explains the workflow and six mcp_* tools change the configuration, while the tools a server advertises join the session catalogue as mcp__<server>__<tool>. Management tools are registered but inactive until a model asks for them through a new tool_search loader, and the catalogue is re-read at the end of every run, so a server added during a conversation is callable in that same conversation. Approvals carry an answer instead of a boolean — allow once, always allow this tool, always allow this server, or deny — and an external tool asks until the user remembers it; a stdio server additionally needs its exact launch line approved, and changing the command, URL, arguments, environment variable names or header names asks again. Servers live in one plain-text storage/agent-v2/mcp/servers.json whose reported path is what the model quotes, endpoints must use HTTPS unless they point at localhost, and link-local and cloud-metadata addresses are refused. Connections use the MCP Java SDK the project already ships (stdio and streamable HTTP), start lazily, and are dropped when a server changes, is removed, or a call fails. A runtime whose packaged extension changed is restarted before the next run instead of failing on a tool its process never had.
  • Frontend — the AI panel gains the V2 session view, runtime selection, session routing, @ object context, skill slash completion, in-conversation model switching, decision and trace rendering, output and chart cards, and beta confirmation. Pi components never call platform or bridge APIs directly.
  • Desktop and packaging — native directory selection, local attachment parsing and output saving stay in host adapters; the desktop layout marker, bootstrap, updater and runtime resources are generated together; a JDK URLStreamHandlerProvider loads nested packaged resources; deleting a session also removes its Pi session files.
  • Supporting surfaces — agent keys are present in every JCEF and start language bundle; script/check-jcef-boundaries.py guards the renderer and agent module boundaries; spec/code/pi-adapter-contract.md documents the transport contract; the agent chat suites run in the frontend prebuild chain and in CI.

Affected surfaces

  • Frontend / Web
  • Backend / API / Storage
  • Database plugin / Driver
  • JCEF / Desktop packaging
  • CI / Build / Release
  • Documentation only

Verification

Every Maven command below ran with tests explicitly enabled (-Dmaven.test.skip=false -DskipTests=false -Dsurefire.failIfNoSpecifiedTests=false -Dmaven.test.failure.ignore=false).

  • Commands and results:
    • mvn -f chat2db-community-server/pom.xml -pl :chat2db-community-domain-core,:chat2db-community-storage,:chat2db-community-tools,:chat2db-community-agent,:chat2db-community-web,:chat2db-community-jcef,:chat2db-community-updater,:chat2db-community-start -am '-Dsurefire.includes=**/*Test.java' test — every module of the reactor reports SUCCESS (all 53 modules built, 4,059 tests, 0 failures, 0 errors, 28 skipped). This is the post-merge run: merging main (37f0e126f) brought 29 commits, including the legacy in-app updater removal and the Excel/JSON/SQL import readers.
    • yarn test:agent-chat, yarn test:i18n, yarn test:mcp-lifecycle, yarn test:import-preview, yarn test:task-center, yarn test:community-boundary, yarn lint:eslint — all passed on the merged tree.
    • mvn -pl :chat2db-community-storage,:chat2db-community-domain-core,:chat2db-community-web,:chat2db-community-start -am -Dtest='McpServerStorageImplTest,McpServerServiceImplTest,LocalAgentV2StorageTest,AiAgentSkillServiceImplTest,AiAgentFileAccessServiceImplTest,AiAgentChartServiceImplTest,AgentApprovalServiceImplTest,AgentToolGatewayServiceTest,AgentNativeToolApprovalTest,AgentMcpToolRegistryTest,SdkMcpToolDiscoveryTest,PiTransportContractTest,AgentControllerTest,AgentServiceImplTest,AiSessionFacadeServiceImplTest,AgentRunCoordinatorTest,AgentRuntimeHandleRegistryTest,AgentRuntimeLifecycleContractTest,AgentSkillResourcesTest,LocalAgentRuntimeConditionTest' test — storage 14, domain-core 52, web 18, start 3 tests; 0 failures, 0 errors. SdkMcpToolDiscoveryTest starts real MCP servers: a stdio fixture that requires the client to offer the 2025-06-18 revision, and a streamable-HTTP fixture whose echo and big tools verify discovery, calls and a result larger than the spooling threshold.
    • node --experimental-vm-modules --test chat2db-community-server/chat2db-community-agent/src/test/js/*.test.mjs — 3 files passed; they cover managed output routing, per-run replay isolation, tool_search activation, a tool that appears later in the catalogue, and the ticket refresh handshake.
    • yarn test:agent-chat — passed; it covers the Pi transport and adapter contract, event-stream tail and paging, transcript building, charts, outputs, questions and context.
    • npx eslint src/blocks/AI/index.tsx src/blocks/AI/agentEventStream.ts src/blocks/AI/agentEventStream.test.ts src/service/pi/contract.ts --max-warnings=0 — clean.
    • npx tsc --noEmit — reports the repository's 210 baseline diagnostics; the files this PR touches add none.
    • python3 script/check-jcef-boundaries.py — passes.
  • Manual verification:
    • Live JCEF desktop window against the rebuilt Community backend and a real 27,484-event conversation: opening the session 28 ms, newest page 26–65 ms, one page further back 39 ms, forward page 20 ms warm (191 ms on the first read, which lists the event directory once); the same page cost 2,940 ms through the previous full-scan implementation, and paging shows no jump or stall in the window.
    • Session deletion removes the private Pi runtime directories while other sessions and shared skill files stay untouched; cancelling a deletion sends no request; the packaged bootstrap loads nested resources.
    • MCP, in the running Community backend with a real model: the model reads the mcp-manager skill, calls tool_search, lists servers (configPath reported), and adds a stdio server; the approval card shows the exact command line; the server connects and its tool is listed with callName; the same conversation then calls mcp__<server>__<tool>, which asks once and, after "always allow this tool", runs without asking again (the answer is persisted and the second call is verified to skip the card). The same flow was repeated over streamable HTTP with a local server, and a 128 KB result came back as a truncated preview plus an output reference (artifactId, path) through the ordinary spooling path.
  • Not run: the full yarn build:web:community, native installers and desktop GUI on Windows and Linux, and a real Pi runtime installation end to end. A Web/desktop smoke of the in-place skill install path (bundled skill replaced on upgrade, same-name user skill prompting overwrite or rename) is still outstanding.
  • UI evidence: N/A — the Pi panel was exercised in the running desktop window; no before/after screenshots are attached.

Risk and compatibility

  • Public API or stored data: V2 storage stays in its own ai-chat-history-v2 directory while V1 keeps its ai-chat-history layout and REST routes, and the mixed-version session list merges both. Event pages are read by sequence, so a missing event file inside a requested window is reported as a broken sequence, while a read that has no watermark yet still verifies the whole sequence from the file names.
  • Database or driver compatibility: the SPI change applies only to V2 agent query capture under an explicit budget; V1 query responses are unchanged. The DM and SQL Server executors were adapted to the same reader.
  • Network, privacy, or security: state-changing agent operations (shell/Pi enablement, tool switches, working directory, approval decisions, native output saving) require the desktop bridge or a loopback caller, so a remote caller can no longer grant itself shell access and approve it. Native host operations require the real desktop bridge context, outputs stay conversation-scoped, and shell tools stay out of the active tool set until enabled, with every command still requiring approval.
  • Community / Local / Pro boundary: Community implementation only; no product migration.
  • MCP servers: the configuration is one plain-text file in the agent data directory (storage/agent-v2/mcp/servers.json, owner-only permissions where the platform supports them); secret values are stored there as text and are never returned by a tool, an event or a log line, and the skill tells the model not to repeat them. Approvals are bound to the loopback tool ticket, so a remote caller cannot configure a server or approve one. Servers needing an interactive OAuth login, and the MCP resource, prompt, sampling and task surfaces, are not supported yet; the client offers every protocol revision the shipped SDK implements and requests 2025-06-18.
  • Backward compatibility: nothing has to read state written by an earlier build of Pi V2, and no such compatibility remains; V1 chat compatibility is kept because that surface is released. Packaged resources use the current main bootstrap layout and must be delivered together.

Reviewer map

  • Start here: chat2db-community-client/src/service/pi/contract.ts, PiOperationRegistry, AgentRunCoordinator, LocalAgentEventStorage.list/listBefore, AgentSkillConfiguration with AiAgentSkillServiceImpl, AiAgentFileAccessServiceImpl (a loaded user skill stays editable while a bundled skill stays read-only), PiProcessSupervisor, BoundedJdbcValueReader, RuntimeUrlStreamHandlerProvider, agentEventStream.ts, AgentMcpToolRegistry with AgentToolGatewayService and McpServerServiceImpl, SdkMcpToolDiscovery, AgentRuntimeSessionHandleImpl.needsRestart, skills/mcp-manager/, spec/code/pi-adapter-contract.md and spec/code/mcp-integration.md.
  • Failure condition: a broken event sequence fails the page that spans it; skill installation fails when the packaged content cannot be published or does not match the extracted directory; live Pi processes or filesystem errors block session deletion before product history is removed.
  • Rollback or disable path: disable Pi through its existing feature setting, and keep a prior full installer/runtime bundle when evaluating the desktop layout.

Contributor declaration

  • I linked the Issue that defines this change.
  • I tested the affected behavior and reported the actual results above.
  • I did not include credentials, private data, or generated build output.
  • I disclosed substantial AI assistance below, or this PR contains no substantial AI-generated code.

AI assistance: implementation, review and verification performed with Codex.

Scrolling up re-ran the whole session load, so the panel flashed and stayed in its loading
state. The earlier stretch is now fetched and merged into the already loaded events:

- no session reload, no panel loading state, no polling restart while scrolling
- a small inline loading label on the load-earlier control instead of the whole-list state
- the trigger stays locked until the user scrolls again, and the reading position is restored
  after the older turns are inserted above it
Asking the agent to list the packaged skill root failed with "Skill resource path is not
loaded", and listing a loaded packaged skill failed the user-directory check.

- a read-only tool may browse the packaged skill root and the directories of loaded packaged
  skills; unloaded packaged content stays unreadable and hidden
- mutations in the packaged area stay blocked for every tool
Loading an earlier stretch while a run was streaming rebuilt the message list under the live
view, so scrolling up snapped back. The earlier-history path now runs only while the session
is idle, and its control is hidden while a run is in flight.
The earlier-history path rebuilt the whole round list and recomputed the transcript, so the
list jumped and the scroll trigger could stay latched after a load that changed nothing.

- prepend only the older rounds, so streaming state, rounds and reading position stay intact
- restore the reading position from a height/top anchor on a dedicated prepend signal, not on
  every message change
- release the scroll trigger in the loader itself, and allow loading while a run streams
Pi sessions render their rounds through AgentV2Session, so the control placed in the other
render path never appeared and scrolling to the top had no visible effect. It now lives inside
the scrolling list, above whichever view renders the rounds.
Loading depended on the live run operation, which an idle session clears, so both the button
and the scroll trigger silently did nothing. History loading now owns its abort signal, created
when the session opens and cancelled when another session is opened.
The manual control is gone: scrolling to the top of the loaded window loads the previous
stretch by itself, and the existing top loading bar shows that work instead of a new widget.
Loading earlier turns read the whole session: the store listed its event
directory, parsed every event file and validated the full sequence before
returning one page. A 27k-event conversation therefore spent seconds on a
single page request, and no amount of client-side throttling could hide it.

Add a backwards range read. Event files are named by their sequence, so the
newest events before a sequence are read by walking down from it and stopping
at the first missing file; unrelated or corrupt files outside the page are
never touched.

- storage/domain/web: `listBefore`/`listEventsBefore`, and `events.list` routes
  to it when the payload carries `beforeSequence`. The controller keeps one
  DTO mapping for REST and for the Pi transport, so both stay identical.
- client: a conversation opens on one tail page (`beforeSequence =
  lastEventSequence + 1`) extended at most once to a turn start, and reading
  further back fetches one page per request. `readAgentHistory` and its
  forward-window replay are gone.
- the scroll zone at the top of the loaded window is the only trigger, so one
  gesture costs one request; a wheel that can no longer move the list asks for
  the previous page too.

Verified on the 27k-event session: opening 28 ms, tail page 65 ms, second page
39 ms, page near the start 5 ms, against 2940 ms for the same page through the
old forward path. 52 backend tests and `yarn test:agent-chat` pass.
Polling a run read the whole session for every page: the store listed its
event directory, parsed every file and validated the sequence before
returning the events after the cursor. A 27k-event conversation paid about
2.9 s for every page while a run polls every 400 ms, so continuing a
conversation in an old session stayed slow.

The watermark already says where the history ends, so a forward page is now
a range read upwards from the cursor. A missing file inside the window is
still a broken sequence, and on a cold cache the watermark is verified
against the file names, which keeps the restart gap check and the recovery
scan working.

Storage tests cover forward paging, the gap inside the window and unreadable
files beyond the requested page. Measured on the 27k-event session: first
page 191 ms (one directory listing), warm page 20 ms, caught-up poll 4 ms,
against 2940 ms before.
Pi V2 has never shipped, so nothing has to read state written by an earlier
build of it. Two shims existed only for that:

- skills redirected a resource path recorded by a run that used the removed
  digest snapshot directories; paths now resolve where the skill lives.
- charts kept a constructor, and a test, for JSON written before groupBy and
  stack existed.

Both are gone, together with a storage test that only proved run and event
JSON from an earlier build still loads. The V1 chat keeps its compatibility
surface because that code is shipped: the V1/V2 session summary merge, the
version-based session routing and the context fallback for requests without
agent context.
A user can now connect an external MCP server by asking for it, without a
settings page: a built-in `mcp-manager` skill explains the workflow, six
`mcp_*` tools change the configuration, and the servers' own tools join the
session catalogue as `mcp__<server>__<tool>`.

- Configuration lives in one plain-text `mcp.json` next to the other user
  settings, written with owner-only permissions. Its absolute path differs per
  product and environment, so the tools report it (`configPath`) and the skill
  forbids stating a path from memory.
- The tools are implemented in the backend and registered but inactive; the
  runtime asks for them through a new `tool_search`, so a session that never
  touches MCP carries one small schema instead of six. The active set follows
  the backend catalogue (`group`/`defaultActive`), which also keeps the file and
  shell tools working when the catalogue is unreachable.
- Approvals now carry an answer instead of a boolean: allow once, always allow
  this tool, always allow this server, or deny. External tool calls ask until
  the user remembers them; a stdio server additionally needs its exact launch
  line approved once, and changing the command, URL, arguments, environment
  variable names or header names asks again.
- HTTP endpoints must use HTTPS unless they point at localhost, and link-local
  and cloud metadata addresses are refused before the backend connects.
- Connections use the MCP Java SDK the project already ships (stdio and
  streamable HTTP), start lazily, and are dropped when a server changes, is
  removed, or a call fails.

Backend: 13 storage, 7 domain-core and 16 web tests cover the new code and pass;
the extension's JS tests cover tool activation and the catalogue refresh.
The stdio fixture server answers initialize, tools/list and tools/call over
JSON-RPC, so the SDK wiring is verified end to end instead of only mocked. It
requires the client to offer the 2025-06-18 revision.

That requirement exposed a real defect: both SDK transports inherit a default
protocol list of 2024-11-05 only, and the client requests the last entry of the
list it is given, so every current server rejected the handshake. The client now
offers all three revisions the SDK implements, ordered oldest to newest, so it
asks for 2025-06-18 and still accepts a server that answers with an older one.

`spec/code/mcp-integration.md` records the surfaces, the configuration file, the
approval answers, the connection lifecycle and the limits of this first version.
Driving the feature through the real web backend and a real stdio server found
three defects that unit tests had not covered:

- `mcp__<server>__<tool>` calls fell through to the database registry and
  answered "Unknown V2 database tool", because only the management tools were
  recognised on the gateway's dispatch path. Names with the MCP prefix now route
  to the MCP branch, and a name whose server is disabled or removed fails with a
  clear message instead of asking for an approval first.
- `mcp_update_server` and `mcp_set_policy` built their schema twice, so the
  provider rejected the whole request with "Invalid schema for function ...
  \"object\" is not of types \"boolean\", \"object\"". The registry test now
  asserts that every property is a schema object and that no schema nests a
  second one.
- Adding a server tried to record its approved launch line before the server
  existed, so `commandApproved` stayed false and the next test call asked again.
  The line is recorded after the change succeeded.

The extension also stops trusting `setActiveTools`: it filters the catalogue to
the tools the runtime actually registered and swallows failures, because an
unknown name inside a promise chain ended the Pi process with exit code 1 right
after the skills were loaded.

Verified end to end in web mode against a real stdio MCP server: skill read,
`tool_search` activation, `mcp_list_servers`, `mcp_add_server` with an approval
card showing the exact command line, discovery of the server's tool, then an
external call approved with "always allow this tool" that runs without asking
again and returns the server's answer.
…hat do not apply

An MCP approval card was labelled "Bash" with a terminal icon, because the card
translated every tool that was not `db_query` or `powershell` into "Bash". The
label now falls back to the tool's own name for anything else, and the icon
follows it: database for SQL, terminal for the shell tools, a plug for MCP.

The card also offered "always allow this tool" and "always allow this server" on
approvals that can never be remembered — managing MCP servers always asks, and a
shell command has its own launch-line approval. Those answers are now offered
only for external MCP tools, which are the only calls that can be remembered.
Three things kept a server that had just been configured from being callable:

- The runtime only knew the tool definitions written when its process started, so
  a server added afterwards appeared in the active names but had no definition to
  register. The gateway now serves the catalogue at `/definitions` and the
  extension re-reads it (with the active names) at session start, at the end of
  every run and on the model refresh command, registering tools that appeared
  meanwhile.
- The management tools reported the server's own tool label, so a model called
  `sequentialthinking` and the runtime answered "Tool not found". They now report
  `callName` (`mcp__<server>__<tool>`) and the skill tells the model to call that
  exact name.
- A running process kept the extension it was started with. The handle now
  compares the packaged extension and its own copy with the digest it started
  with, and the coordinator closes such a handle before the next run so the run
  starts a fresh process instead of failing on a tool the process never had.

The configuration file also moved out of the cache directory to
`<env base path>/storage/agent-v2/mcp/servers.json`, next to the skill directory
it belongs to.

Verified in one session against a real stdio server: add with an approval card,
then call `mcp__fixture__echo` in the same conversation, approve it once with
"always allow this tool", and receive the server's answer.
A second fixture server speaks MCP over streamable HTTP and offers an echo tool
plus a tool that returns a requested number of kilobytes. The test discovers its
tools, calls the small one and calls the large one, so the HTTP transport is
covered by a real server instead of only by the endpoint policy tests.

The same server drove a local acceptance run: adding it over HTTP connected and
listed both tools, and a 128 KB result came back to the session as a truncated
preview plus an output reference (artifactId, path and byte count), which is the
ordinary spooling path the database tools already use.
Conflicts were five JCEF message bundles and the frontend i18n source hashes.
Both sides' keys are kept in the bundles, and the hashes are regenerated from
the merged sources.

main removed the legacy in-app updater (d522066), so those files stay deleted
here, and main's import refactor (Excel/JSON/SQL readers) is taken as-is.
Three tests still assumed the previous shapes and only the module-wide run
surfaced them:

- AgentDatabaseServiceImplTest answers approvals through a dynamic proxy, so the
  change from a boolean result to AgentApprovalDecision compiled but failed with
  a ClassCastException at runtime. The fake now returns ALLOW_ONCE or DENY.
- AiAgentSkillServiceImplTest lists the installed skills and now has to expect
  mcp-manager next to chart and skill-manager.
- DesktopAgentBridgeTest posted {"approved":false} to the approvals route; the
  request body carries a decision value now.
return List.of();
}
Path directory = paths.resourceDirectory(sessionId, "events");
if (!Files.exists(directory, LinkOption.NOFOLLOW_LINKS)) {
if (cached != null) {
// A session directory may have been deleted and recreated with the same id.
// Do not carry the old in-memory watermark into the new lifecycle.
if (cached == 0 || Files.exists(paths.eventFile(sessionId, cached), LinkOption.NOFOLLOW_LINKS)) {
The panel lists Pi's built-in tools, so labels for other tool names are unused.
Drop them, the two unused group labels, and the dead keys they left in every
locale; the tests now require a label for each Pi tool in all locales.
@openai0229
openai0229 force-pushed the feat/agent-v2-contracts branch from 1b5175c to bb2a630 Compare September 23, 2026 01:36
Conflicts were the frontend prebuild test chain, the i18n source hashes and the
About page update check. The chain keeps main's update-check-schedule step and
still ends with the agent chat suite, the About page takes main's guarded manual
check, and the hashes are regenerated from the merged sources.
@openai0229
openai0229 changed the base branch from main to release/pi-v2 September 23, 2026 02:18
@openai0229
openai0229 marked this pull request as ready for review September 23, 2026 02:24
@openai0229
openai0229 requested review from a team, Aias00 and auenger as code owners September 23, 2026 02:24
@openai0229
openai0229 merged commit 7b13327 into release/pi-v2 Sep 23, 2026
19 of 24 checks passed
@openai0229
openai0229 deleted the feat/agent-v2-contracts branch September 23, 2026 02:24
@openai0229 openai0229 moved this from In Progress to Done in Chat2DB Community Sep 23, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Status: Done

Development

Successfully merging this pull request may close these issues.

2 participants