fix(host-core): return image files from Read as inline image blocks - #1296
Conversation
Read refused every image file as binary content (issue vastsa#1073, vastsa#1241), so the agent could never look at a picture the user pointed it at by path, even with a vision-capable model. Desktop attachments and pastes already inline images; only the tool path was missing. Read now recognizes png, jpeg, gif and webp files by extension, verifies the file signature before encoding, and returns a structured image block alongside a short text note. The agent runtime already converts a result `images` array into pi-ai image blocks for vision models, so host-core is the only change needed. Files above the shared 10 MB inline-image budget and image extensions with non-image bytes keep refusing with their previous error codes.
|
I verified the runtime boundary and this does not yet reach the model. |
The maintainer review of vastsa#1296 found the runtime boundary dropped the picture: toolResultFromUi only read content[] blocks with type "image" and serialized the host Read's top-level images array as plain JSON text, so the model never saw it across a persisted restore. The restorer now expands a top-level images array into real image content blocks (dropping malformed entries), with runtime-level regression tests proving the Read output reaches the model as image content. Also applies cargo fmt to the host-core image-block code.
|
已在
验证:vitest |
|
复核新 head 另外,host 端先用 我在最新 |
Fixes #1073, fixes #1241.
Problem
Readrefused every image file as binary content (TOOL_BINARY_CONTENT), so an agent with a vision-capable model could never look at a picture the user pointed at by path. Desktop attachments and pastes already inline images into the message; only the tool path was missing (the original pi tool returnstype: "image"blocks, which our desktop Read never re-exposed).Fix
tool_readin host-core now recognizespng,jpeg,gifandwebpextensions, verifies the file signature before encoding, and returns a structured result:{ "path": "...", "text": "Image file ...", "images": [{ "data": "<base64>", "mimeType": "image/png" }] }The agent runtime already converts a result
imagesarray into pi-aiImageContentblocks for vision models (the same path plugin image results use), so host-core is the only change needed. Non-vision models still see the text note.Preserved behavior:
MAX_INLINE_IMAGE_BYTES) still refuse, with a size explanationTOOL_BINARY_CONTENTValidation
cargo test -p host-core --locked— not run locally: no Rust toolchain on this machine; new tests added for the image-block path, signature mismatch, and oversize refusalSpec/docs impact: none — this restores behavior the original pi agent already had; no persisted format or protocol change.