Skip to content

feat(ai): web_fetch tool — read direct URLs as markdown - #1386

Merged
2witstudios merged 7 commits into
masterfrom
pu/fetch
May 20, 2026
Merged

2witstudios merged 7 commits into
masterfrom
pu/fetch

Conversation

@2witstudios

@2witstudios 2witstudios commented May 20, 2026 •

Copy link
Copy Markdown
Owner

Summary

Adds `web_fetch` AI tool that fetches any public URL and returns its full content as clean markdown, controlled by the same toggle as `web_search`.

What changed

  • New tool: `web_fetch` — fetches a URL, converts HTML→markdown via TurndownService (already a project dep), returns up to 20k chars by default
  • Toggle: `web_fetch` joins `WEB_SEARCH_TOOLS` — one "Web" toggle gates both tools
  • Label: "Web Search" → "Web" in ToolsPopover

Security hardening

  • SSRF protection: `validatePublicHost()` resolves hostname via `dns.lookup({ all: true })` and checks every returned IP — blocks literal IPs and catches public hostnames that resolve to private ranges (DNS rebinding mitigation). Covers RFC1918, loopback, link-local, AWS metadata (169.254.x), and IPv6 private (fe80::/10, fc00::/7, ::ffff: mapped). Fail-closed on DNS failure.
  • https-only: non-HTTPS protocol rejected before any network I/O
  • Content-Type guard: rejects binary responses (PDFs, images, ZIPs) before attempting HTML→markdown
  • Streaming body with 5 MB cap: replaces `response.text()` with a `ReadableStream` reader that hard-caps byte consumption; also checks `Content-Length` header for an early reject
  • URL redaction in logs: `redactUrl()` strips query string and fragment before logging to prevent token/secret leakage

Test plan

  • Enable Web toggle, ask AI to fetch a public URL → readable markdown
  • Disable Web toggle → both `web_search` and `web_fetch` unavailable
  • Toggle label reads "Web" in ToolsPopover
  • Fetch `https://192.168.1.1\` → SSRF block error
  • Fetch `https://169.254.169.254/latest/meta-data/\` → SSRF block error
  • Fetch a PDF URL → content-type error
  • Fetch a paywalled URL → graceful HTTP error response

🤖 Generated with Claude Code

Adds a `web_fetch` AI tool alongside `web_search` that fetches the full
content of a specific URL and returns it as clean markdown via TurndownService
(already a project dependency). Gated by the same "Web" toggle as web_search —
both tools are now grouped under WEB_SEARCH_TOOLS.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
@coderabbitai

coderabbitai Bot commented May 20, 2026 •

Copy link
Copy Markdown
Contributor

Warning

Rate limit exceeded

@2witstudios has exceeded the limit for the number of commits that can be reviewed per hour. Please wait 4 minutes and 38 seconds before requesting another review.

You’ve run out of usage credits. Purchase more in the billing tab.

⌛ How to resolve this issue?

After the wait time has elapsed, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

We recommend that you space out your commits to avoid hitting the rate limit.

🚦 How do rate limits work?

CodeRabbit enforces hourly rate limits for each developer per organization.

Our paid plans have higher rate limits than the trial, open-source and free plans. In all cases, we re-allow further reviews after a brief timeout.

Please see our FAQ for further information.

ℹ️ Review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro

Run ID: 37b12d62-0477-48f7-baed-82742a83beab

📥 Commits

Reviewing files that changed from the base of the PR and between e115fa8 and f5ac25d.

📒 Files selected for processing (2)
  • apps/web/src/lib/ai/tools/__tests__/web-search-tools.test.ts
  • apps/web/src/lib/ai/tools/web-search-tools.ts
📝 Walkthrough

Walkthrough

This PR adds a new web_fetch tool that fetches and converts URLs to Markdown with SSRF protection, integrates it into the tool filtering system alongside web_search, and updates the web-search toggle label from "Web Search" to "Web".

Changes

Web Fetch Tool Addition

Layer / File(s) Summary
SSRF protection helper and dependencies
apps/web/src/lib/ai/tools/web-search-tools.ts
TurndownService is imported and isPrivateHost() helper blocks requests to loopback, RFC1918 private ranges, link-local, and cloud metadata hosts to prevent SSRF attacks.
Web_fetch tool implementation
apps/web/src/lib/ai/tools/web-search-tools.ts
New web_fetch tool requires authenticated userId, validates HTTPS URLs, blocks private hostnames, fetches with 15s timeout, sanitizes HTML by removing script/style tags, converts to fenced Markdown via TurndownService, truncates output, and returns structured success/error payloads with next steps.
Tool filtering and allowlist integration
apps/web/src/lib/ai/core/tool-filtering.ts
WEB_SEARCH_TOOLS set now includes both web_search and web_fetch; getToolsSummary list includes web_fetch so both tools are reported together when web search is enabled or disabled.
UI toggle label update
apps/web/src/components/ui/floating-input/ToolsPopover.tsx
Web-search toggle label shortened from "Web Search" to "Web".

Sequence Diagram

sequenceDiagram
  participant Client
  participant WebFetch as web_fetch Tool
  participant Validator as URL & SSRF Validator
  participant Fetcher as HTTP Fetch
  participant Converter as HTML→Markdown
  participant Response as Result

  Client->>WebFetch: URL, userId
  WebFetch->>Validator: Validate HTTPS
  Validator->>Validator: Check isPrivateHost
  Validator-->>WebFetch: Valid or Error
  WebFetch->>Fetcher: Fetch with 15s timeout
  Fetcher->>Converter: HTML (scripts/styles removed)
  Converter->>Converter: Convert via TurndownService
  Converter->>Converter: Truncate to maxLength
  Converter-->>Response: Success/Error payload
  Response-->>Client: Structured result
Loading

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~25 minutes

Possibly related PRs

  • 2witstudios/PageSpace#1316: Both PRs modify the AI tool allowlist and web-search toggle handling; this PR extends the same WEB_SEARCH_TOOLS set to include the new web_fetch tool alongside the filtering changes in that PR.
  • 2witstudios/PageSpace#98: Directly modifies ToolsPopover.tsx to shorten the web-search toggle label, affecting the same UI component being updated in this PR.

Poem

🐰 A fetch tool hops in with caution so tight,
SSRF guards keep the requests in sight.
HTML to Markdown, with scripts stripped clean,
Web content flows—the safest we've seen! 🔒✨

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Title check ✅ Passed The title 'feat(ai): web_fetch tool — read direct URLs as markdown' directly and specifically describes the main change: adding a new web_fetch tool that reads URLs and converts them to markdown.
Docstring Coverage ✅ Passed Docstring coverage is 100.00% which is sufficient. The required threshold is 80.00%.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch pu/fetch

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 49008499c3

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment on lines +229 to +232
const response = await fetch(url, {
headers: { 'User-Agent': 'Mozilla/5.0 (compatible; PageSpace/1.0)' },
signal: AbortSignal.timeout(15000),
});

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Block internal targets before fetching user-provided URLs

Calling fetch(url) directly on model/user-controlled input allows SSRF against services reachable only from the server network (for example http://127.0.0.1, RFC1918 ranges, or cloud metadata endpoints), which can leak sensitive internal data through tool output. This commit introduces web_fetch as a general URL reader but does not enforce a public-host allowlist/denylist or private-IP resolution checks before requesting the URL, so a prompt can intentionally or indirectly pivot this tool into internal infrastructure access.

Useful? React with 👍 / 👎.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in the current implementation. validatePublicHost() now runs before any network I/O: it calls dns.promises.lookup(hostname, { all: true }) to resolve every IP the hostname can return, and passes each through isPrivateIp() which covers RFC1918, loopback (127.x, ::1), link-local (169.254.x, fe80::/10), and IPv6-mapped private ranges. DNS failure is fail-closed (throws). This catches both literal private IPs and public hostnames that DNS-rebind to private space.

Adds isPrivateHost() guard that rejects loopback (127.x, ::1, localhost),
RFC1918 (10.x, 172.16-31.x, 192.168.x), link-local/cloud-metadata
(169.254.x), and http:// before issuing fetch(). Addresses P1 SSRF
finding from Codex review.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
@2witstudios

Copy link
Copy Markdown
Owner Author

SSRF fix (e115fa8): Added isPrivateHost() before calling fetch() — blocks loopback (127.x, ::1, localhost), RFC1918 (10.x, 172.16–31.x, 192.168.x), link-local + AWS metadata (169.254.x), and enforces https-only. Returns a structured error to the model rather than throwing, so it gets a clear explanation without leaking internal network shape. Addresses the P1 finding from Codex review.

- Reject non-HTML/text responses before running turndown to avoid binary
  garbage in tool output (e.g. PDFs, images, ZIPs)
- Fix truncated flag: compare fullMarkdown.length > maxLength instead of
  markdown.length === maxLength to avoid false positives on exact-length content

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
@2witstudios

Copy link
Copy Markdown
Owner Author

Quality fixes (4871f29):

  • Added Content-Type guard before calling response.text() — binary responses (PDF, image, ZIP) now return a structured error instead of passing garbage through TurndownService
  • Fixed truncated flag: now correctly compares fullMarkdown.length > maxLength rather than markdown.length === maxLength (the old check produced a false positive when content was exactly maxLength chars)

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 3

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@apps/web/src/lib/ai/tools/web-search-tools.ts`:
- Around line 253-268: Before calling response.text(), validate
response.headers: ensure Content-Type is text/html (or startsWith 'text/') and
if Content-Length exists reject if it exceeds a byteCap (e.g., 1MB). If
Content-Length is absent, read response.body as a stream (use
response.body.getReader()) and accumulate bytes up to the byteCap, aborting and
throwing if the cap is exceeded; then decode the accumulated bytes to a string
and pass that to TurndownService (instead of calling response.text()). Apply
these checks in the same block that currently creates `response` and where
`html`, `cleaned`, `td`, and `markdown` are produced, and throw clear errors for
non-text content or oversized responses.
- Line 251: Several logging calls (e.g., the webSearchLogger.debug at "Fetching
URL" and the similar calls around lines 270-275 and 290-292) are logging full
user-supplied URLs including query strings and fragments; change those to log a
redacted URL that strips search and hash. For each place where you call
webSearchLogger.debug/info with the url variable (e.g., the "Fetching URL"
call), construct a redactedUrl by parsing the input with the URL constructor in
a try/catch and using only urlObj.origin + urlObj.pathname (or fallback to the
original host/path if parsing fails), then pass redactedUrl in the log payload
(keep maskIdentifier(userId) as-is); apply the same change to all other logging
sites in this file that log url so no query or fragment is sent to logs.
- Around line 121-135: The current isPrivateHost(hostname) only checks the
literal hostname string; you must resolve the hostname to IP addresses before
allowing the fetch and reject requests whose resolved IPs are in
private/loopback/link-local/metadata or IPv6-local ranges. Update the request
path to perform a DNS resolution (e.g., using dns.promises.lookup or resolve
with all addresses for both A and AAAA) for the parsed.hostname (handling
bracketed IPv6), then create/replace with an isPrivateIp(ip: string) helper that
checks numeric ranges (IPv4: 10.0.0.0/8, 127.0.0.0/8, 169.254.0.0/16,
172.16.0.0/12, 192.168.0.0/16, etc.; IPv6: ::1, fe80::/10, fc00::/7, link-local,
etc.) and reject if any resolved address is private; if DNS resolution fails or
returns no addresses, treat as disallowed (fail closed). Ensure this change is
applied where parsed.hostname was previously passed to isPrivateHost (including
the other locations noted).
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro

Run ID: 4441f936-2c7d-40c2-8e89-9946e2f8a843

📥 Commits

Reviewing files that changed from the base of the PR and between a456f28 and e115fa8.

📒 Files selected for processing (3)
  • apps/web/src/components/ui/floating-input/ToolsPopover.tsx
  • apps/web/src/lib/ai/core/tool-filtering.ts
  • apps/web/src/lib/ai/tools/web-search-tools.ts

Comment thread apps/web/src/lib/ai/tools/web-search-tools.ts Outdated
Comment thread apps/web/src/lib/ai/tools/web-search-tools.ts Outdated
Comment thread apps/web/src/lib/ai/tools/web-search-tools.ts Outdated
… cap

- Replace isPrivateHost() string check with validatePublicHost() which
  resolves hostname via dns.lookup and checks all returned IPs — catches
  public hostnames that point to private IPs (DNS rebinding mitigation);
  also extends IPv6 private coverage (fe80::/10, fc00::/7, ::ffff: mapped)
- Add redactUrl() helper that strips query and fragment before logging to
  prevent signed token or secret leakage in server logs
- Add Content-Length early-reject and stream body with 5 MB hard byte cap
  using ReadableStream reader instead of response.text() — prevents
  unbounded memory buffering of large responses

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
@2witstudios

Copy link
Copy Markdown
Owner Author

Security hardening (b26e352) — addressing CodeRabbit review:

Critical: DNS-based SSRF validation (replacing string-only isPrivateHost)

  • Added async validatePublicHost() that resolves the hostname via dns.lookup({ all: true }) and checks every returned IP against isPrivateIp() — catches public hostnames that point to private IPs at resolution time (DNS rebinding mitigation)
  • Extended IPv6 private range coverage: fe80::/10 (link-local), fc00::/7 (unique-local), ::ffff:-mapped IPv4
  • Fail-closed: DNS resolution failure → request blocked

Major: URL redaction in logs

  • Added redactUrl() that strips query string and fragment before logging — prevents signed tokens or document secrets appearing in server logs

Major: Streaming body with byte cap

  • Replaced response.text() with a ReadableStream reader that hard-caps at 5 MB
  • Added Content-Length early-reject (if header present and > 5 MB, return structured error before fetching any body)

2witstudios and others added 3 commits May 19, 2026 21:43
…happy path

- SSRF: http rejection, private IPv4 literals, localhost, DNS rebinding
  (public hostname resolving to private IP), and DNS failure (fail-closed)
- Content: non-HTML content-type rejection and Content-Length size limit
- Happy path: markdown conversion, truncation flag accuracy

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
…ckLookupFn

vi.mock('dns') factory needs both default and named exports for ESM interop.
Variable must start with 'mock' to be hoisted before vi.mock factory runs.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
…itest 2.x

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
@2witstudios
2witstudios merged commit ad978ca into master May 20, 2026
3 checks passed
@2witstudios
2witstudios deleted the pu/fetch branch May 27, 2026 01:16
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant