Skip to content

fix(runtime): extract text from PDF in web_fetch and cap chunked responses (#1271) - #1281

Open
Adityakk9031 wants to merge 1 commit into
RightNow-AI:mainfrom
Adityakk9031:fix/web-fetch-pdf-extraction-1271
Open

Adityakk9031 wants to merge 1 commit into
RightNow-AI:mainfrom
Adityakk9031:fix/web-fetch-pdf-extraction-1271

Conversation

@Adityakk9031

@Adityakk9031 Adityakk9031 commented Aug 31, 2026 •

Copy link
Copy Markdown

Description

Fixes #1271.

\web_fetch\ previously decoded every HTTP response directly into string text via
esp.text(). When fetching PDFs, this resulted in raw FlateDecode-compressed binary byte strings entering the agent context (often hundreds of thousands of characters of binary tokens), overflowing context windows and poisoning conversation history. Additionally, chunked responses lacking a \Content-Length\ header bypassed the response size guard and were read unbounded into memory.

Changes

  1. PDF Text Extraction:
    • Added \pdf-extract\ (v0.8) and \lopdf\ (v0.34) workspace dependencies.
    • Added \is_pdf()\ detection matching \pplication/pdf, \pplication/x-pdf, and %PDF\ magic byte headers.
    • Added \extract_pdf_content()\ to extract plain text from PDF bytes in memory, providing structured fallback descriptions if the PDF contains scanned images without OCR or is encrypted/corrupted.
  2. Chunked Streaming & Bounded Response Guard:
    • Replaced unbounded
      esp.text()\ with chunked stream reader
      ead_bounded_body()\ that enforces \max_response_bytes\ dynamically across all chunks, protecting against memory exhaustion even without \Content-Length\ headers.
  3. Binary Content Sniffing:
    • Added \is_binary_content()\ to sniff binary MIME types (\image/, \�ideo/, \�udio/*, \�pplication/zip, \�pplication/octet-stream) and sample null bytes in the first 1024 bytes.
    • Returns structured metadata notes for binary responses instead of dumping raw binary byte strings into context.
  4. Lossy UTF-8 Fallback:
    • Handled non-UTF8 text decoding with lossy UTF-8 conversion (\String::from_utf8_lossy).
  5. Applied to both Engine and Legacy Tool:
    • Updated both \WebFetchEngine::fetch\ and \ ool_web_fetch_legacy\ in \ ool_runner.rs.
  6. Unit Tests:
    • Added unit tests for PDF detection, binary detection, valid PDF text extraction, and corrupted PDF fallback handling.

Testing

  • Ran unit tests in \crates/openfang-runtime/src/web_fetch.rs\ (21 passed, 0 failed).
  • Verified shellcheck on \scripts/install.sh.

@Adityakk9031

Copy link
Copy Markdown
Author

@jaberjaber23 have a look

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

web_fetch injects raw PDF binary into agent context instead of extracting text

1 participant