Skip to content

fix(host-core): return image files from Read as inline image blocks - #1296

Merged
vastsa merged 2 commits into
vastsa:mainfrom
yexisu:fix/read-image-blocks
Oct 2, 2026
Merged

vastsa merged 2 commits into
vastsa:mainfrom
yexisu:fix/read-image-blocks

Conversation

@yexisu

@yexisu yexisu commented Oct 2, 2026

Copy link
Copy Markdown
Contributor

Fixes #1073, fixes #1241.

Problem

Read refused every image file as binary content (TOOL_BINARY_CONTENT), so an agent with a vision-capable model could never look at a picture the user pointed at by path. Desktop attachments and pastes already inline images into the message; only the tool path was missing (the original pi tool returns type: "image" blocks, which our desktop Read never re-exposed).

Fix

tool_read in host-core now recognizes png, jpeg, gif and webp extensions, verifies the file signature before encoding, and returns a structured result:

{ "path": "...", "text": "Image file ...", "images": [{ "data": "<base64>", "mimeType": "image/png" }] }

The agent runtime already converts a result images array into pi-ai ImageContent blocks for vision models (the same path plugin image results use), so host-core is the only change needed. Non-vision models still see the text note.

Preserved behavior:

  • files above the shared 10 MB inline-image budget (MAX_INLINE_IMAGE_BYTES) still refuse, with a size explanation
  • an image extension with non-image bytes still refuses with TOOL_BINARY_CONTENT
  • text reads, binary refusals and denylist handling are untouched

Validation

  • cargo test -p host-core --locked — not run locally: no Rust toolchain on this machine; new tests added for the image-block path, signature mismatch, and oversize refusal
  • CI requested to run host-core tests/clippy/fmt

Spec/docs impact: none — this restores behavior the original pi agent already had; no persisted format or protocol change.

Read refused every image file as binary content (issue vastsa#1073, vastsa#1241), so
the agent could never look at a picture the user pointed it at by path,
even with a vision-capable model. Desktop attachments and pastes already
inline images; only the tool path was missing.

Read now recognizes png, jpeg, gif and webp files by extension, verifies
the file signature before encoding, and returns a structured image block
alongside a short text note. The agent runtime already converts a result
`images` array into pi-ai image blocks for vision models, so host-core is
the only change needed. Files above the shared 10 MB inline-image budget
and image extensions with non-image bytes keep refusing with their
previous error codes.
@vastsa

vastsa commented Oct 2, 2026

Copy link
Copy Markdown
Owner

I verified the runtime boundary and this does not yet reach the model. read_image_block returns { path, text, images: [{ data, mimeType }] }, but packages/agent-runtime/src/runtime.ts::toolResultFromUi only restores image blocks from raw.content[] entries with type: "image"; it never reads the top-level images field. So the serialized result becomes text-only (or JSON text) before the next model turn. Please thread the image as the actual supported tool-result content block and add a runtime-level regression test proving the Read output reaches the model as image content. CI also currently fails cargo fmt --check.

The maintainer review of vastsa#1296 found the runtime boundary dropped the
picture: toolResultFromUi only read content[] blocks with type "image"
and serialized the host Read's top-level images array as plain JSON
text, so the model never saw it across a persisted restore. The
restorer now expands a top-level images array into real image content
blocks (dropping malformed entries), with runtime-level regression
tests proving the Read output reaches the model as image content.
Also applies cargo fmt to the host-core image-block code.
@yexisu

yexisu commented Oct 2, 2026

Copy link
Copy Markdown
Contributor Author

已在 4bf5687cc 按复核意见修复:

  1. 图片现在真正到达模型:toolResultFromUi(packages/agent-runtime/src/runtime.ts)原来只恢复 content[] 里 type: "image" 的块,把 host Read 返回的顶层 images 数组当 JSON 文本序列化。现在恢复路径会识别顶层 images:先输出去掉 images 的文本块,再把每个合法 { data, mimeType } 展开为真实 image 内容块;畸形条目静默丢弃,不影响恢复。
  2. 运行时级回归测试(runtime.test.ts):Read 图片结果经 toolResultFromUi 恢复为 [text, image] 内容块(钉住「到达模型」这一步);畸形 images 条目只留文本、不抛错。
  3. cargo fmt:已在本机装好 stable 工具链并 cargo fmt 修正两处格式,cargo fmt --check 现在通过。

验证:vitest runtime.test.ts 新用例 2/2 通过;cargo fmt --check ✅。cargo test -p host-core 首次编译在本机超时未跑完(逻辑未变,仅格式化),host-core 行为依赖既有测试与 CI 门禁。

@vastsa

vastsa commented Oct 2, 2026

Copy link
Copy Markdown
Owner

复核新 head 4bf5687 后,阻塞合入的 TS 编译错误可以在 runtime.ts:1503–1505 复现:raw 的联合类型没有 images 属性,解构也不是已收窄的对象类型;agent-runtime tsc --noEmit 报 4 个 TS2339/TS2700。请修完后重跑 JS build/typecheck。

另外,host 端先用 metadata.len() 判断 10 MB,再无上限 std::fs::read,文件若在两步间被替换/增长,上限会被绕过且仍可能一次性分配超大缓冲。建议对同一打开的文件句柄做有界读取(最多上限 + 1 字节)再校验签名。

我在最新 origin/main(011fff902)上验证合并无文本冲突;但当前 PR head 不包含该 main(本地 merge-base --is-ancestor origin/main HEAD 未通过),需要刷新基线后再跑门禁。现在不能合入。

@vastsa
vastsa merged commit 0924014 into vastsa:main Oct 2, 2026
2 of 3 checks passed
@vastsa

vastsa commented Oct 2, 2026

Copy link
Copy Markdown
Owner

已由维护者集成 PR #1305 修复并合入:保留了本 PR 的原始提交,同时修复 TS narrowing,并用对打开文件句柄的有界读取堵住 size metadata 竞态。#1305 的 Rust/JS 门禁均通过。此 PR 作为 superseded 关闭。

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug] 没有办法读取用文本路径的图片, [BUG] read 工具无法读取图片

2 participants