Skip to content

fix(runtime): deliver MCP and plugin tool images as image blocks - #1367

Closed
yexisu wants to merge 1 commit into
vastsa:mainfrom
yexisu:fix/tool-image-shapes
Closed

yexisu wants to merge 1 commit into
vastsa:mainfrom
yexisu:fix/tool-image-shapes

Conversation

@yexisu

@yexisu yexisu commented Oct 4, 2026

Copy link
Copy Markdown
Contributor

Problem

The host tool bridge only recognized a top-level images: [{ data, mimeType }] array. MCP results shaped as content: [{ type: "image", ... }] and plugin tools returning a bare content-block array ([{ type: "text" }, { type: "image" }]) fell into the JSON.stringify branch, so vision models never received an image block — the user believed the model had seen the picture (#1360).

Fix

Normalize the two missing shapes in the host tool bridge (runtime.ts):

  • Bare block array (plugin tools): render text blocks as text, collect well-formed image blocks (vision enabled), set imageCount in details.
  • content: [...] (MCP): same treatment; details keep the rest of the payload plus imageCount.
  • Shared collectImageBlocks helper accepts only well-formed entries ({ type: "image", data, mimeType }); malformed ones are skipped, not fatal.
  • toolResultFromUi (persisted-row restore) gains the same bare-array branch so a restart does not flatten the image back to JSON.

The existing images: [...] shape is untouched.

Validation

  • toolResultFromUi tests 3/3 in runtime.test.ts (incl. new bare-array regression)
  • agent-runtime tsc --noEmit: remaining errors reproduce on pristine main (baseline noise)
  • Live MCP round-trip on a real vision model not run locally (needs a configured MCP image server); blind verification of the shape handling was done by the issue reporter

Fixes #1360

The host tool bridge only recognized a top-level `images` array, so MCP
results shaped as `content: [{ type: "image" }]` and plugin tools
returning a bare content-block array were flattened into JSON text;
vision models never saw the picture (vastsa#1360). Normalize both shapes:
render text blocks as text, emit well-formed image blocks when vision is
enabled, and carry imageCount in details. Apply the same restoration to
persisted bare arrays in toolResultFromUi so a restart keeps the image.

Validated: toolResultFromUi tests 3/3 in runtime.test.ts; remaining
agent-runtime tsc errors reproduce on pristine main (baseline noise).
@vastsa

vastsa commented Oct 4, 2026 •

Copy link
Copy Markdown
Owner

Thanks for fixing MCP and plugin image normalization. The issue was real. I added a direct runtime-bridge regression for both content objects and bare plugin block arrays, plus the bridge contract, in #1372; that follow-up has merged. Closing this duplicate as superseded.

@vastsa vastsa closed this Oct 4, 2026

This branch was previously deployed

1 inactive deployment
Preview — 94dbfcce Deployed Oct 4, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug] MCP / 插件工具返回的图片被拍平成 JSON 文本,视觉模型收不到图像块

2 participants