Skip to content

New Integration Proposal: MrScraper Components and Agent Tools #3928

Description

@ai-mrscraper

Summary and motivation

MrScraper is a managed web-data platform for retrieving JavaScript-rendered pages, extracting structured data with natural-language instructions, crawling websites, querying Google search results, and running reusable scraping workflows.

Haystack applications frequently need current information that is not available in a model's training data or an existing document store. This integration gives Haystack Pipelines and Agents a supported way to acquire that information without requiring users to build and operate browser automation, proxy rotation, anti-bot handling, extraction prompts, and result-storage infrastructure themselves.

The integration supports use cases including:

  • Supplying Haystack Agents with current web information.
  • Building RAG ingestion pipelines from JavaScript-rendered or protected websites.
  • Extracting products, jobs, articles, property listings, reviews, and other repeated records.
  • Monitoring prices, availability, content changes, and competitors.
  • Discovering URLs across a website before processing selected pages.
  • Starting research workflows from Google SERP results.
  • Reusing saved AI or manually configured scrapers across single URLs or batches.
  • Retrieving stored extraction results for downstream processing.

The proposed integration is implemented in PR #3898.

Adoption signals

  • GitHub stars of the main repository: MrScraper is primarily a commercial hosted service, so its core platform is not represented by one open-source main repository. The official Python SDK repository currently has 0 stars. The MrScraper GitHub organization maintains nine public repositories covering its Python and Node.js SDKs, LangChain integration, n8n node, CLI, MCP server, and agent plugins. These repositories are new, so their current star counts are not yet representative of adoption of the hosted platform.

  • PyPI downloads in the last 30 days: The official mrscraper-sdk received 4,784 downloads in the last 30 days as of September 8, 2026 (live monthly-download badge, PyPI Stats).

  • Release activity: The latest mrscraper-sdk release is 0.3.0, published September 8, 2026. Six versions have been published since its initial release on March 9, 2026: 0.1.0, 0.1.1, 0.1.2, 0.2.0, 0.2.1, and 0.3.0. Releases have accompanied API and SDK development, with updates approximately every one to three months after the initial release.

  • Maintenance: MrScraper is actively maintained by a dedicated commercial company. The Python SDK received commits on September 8, 2026, and the official MCP server received commits on September 7, 2026. The company also maintains product documentation, SDKs, integrations, and support for the hosted API. MrScraper's LinkedIn page lists approximately 5,900 followers and a company size of 51–200 employees.

  • Haystack community demand: We have not found an earlier Haystack Discord thread or issue requesting the integration. This proposal and PR #3898 were initiated by the MrScraper team to make the existing demand for real-time web access and structured web extraction available natively to Haystack users.

  • Anything else: MrScraper already provides integrations for LangChain, n8n, Zapier, and MCP-compatible agents, as well as official Python and Node.js SDKs and a CLI. The n8n package received approximately 694 npm downloads during the latest reported 30-day period. These integrations demonstrate demand for using MrScraper in agent and automation workflows comparable to the proposed Haystack integration.

The public SDK and integration repositories are relatively new. We expect adoption to grow because AI Agents increasingly need reliable access to current, structured web data, while operating browser infrastructure, proxies, retries, and site-specific extraction logic remains expensive for individual application teams.

If helpful for evaluating adoption, the MrScraper team can also provide non-public aggregate metrics such as active accounts, monthly API requests, extracted records, customer growth, or enterprise usage directly to the Haystack maintainers.

Detailed design

The integration is distributed as the independent mrscraper-haystack package under the haystack_integrations namespace. It depends on haystack-ai and httpx and communicates with the documented MrScraper API.

It exposes 15 Haystack components organized into six capability groups:

  • Account

    • MrScraperGetAccountInfo
  • Discovery

    • MrScraperCrawlWebsiteUrls
    • MrScraperSearchGoogleSerp
  • Extraction

    • MrScraperExtractPageByPrompt
    • MrScraperExtractListings
    • MrScraperExtractStructuredData
    • MrScraperFetchRenderedHtml
  • Results

    • MrScraperGetResults
    • MrScraperGetLatestResults
    • MrScraperGetResultDetail
  • Scraper creation

    • MrScraperCreatePromptScraper
    • MrScraperCreateListingScraper
    • MrScraperCreateWebsiteCrawlScraper
  • Saved scraper execution

    • MrScraperRunExistingScraper
    • MrScraperRunExistingScraperBatch

Every component provides both run() and true asynchronous run_async() methods. The asynchronous methods use asynchronous HTTP operations rather than wrapping blocking calls.

All components expose one stable Haystack output:

{"result": value}

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    P3new integrationDiscuss the creation of a new integration in Core

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions