Summary and motivation
MrScraper is a managed web-data platform for retrieving JavaScript-rendered pages, extracting structured data with natural-language instructions, crawling websites, querying Google search results, and running reusable scraping workflows.
Haystack applications frequently need current information that is not available in a model's training data or an existing document store. This integration gives Haystack Pipelines and Agents a supported way to acquire that information without requiring users to build and operate browser automation, proxy rotation, anti-bot handling, extraction prompts, and result-storage infrastructure themselves.
The integration supports use cases including:
- Supplying Haystack Agents with current web information.
- Building RAG ingestion pipelines from JavaScript-rendered or protected websites.
- Extracting products, jobs, articles, property listings, reviews, and other repeated records.
- Monitoring prices, availability, content changes, and competitors.
- Discovering URLs across a website before processing selected pages.
- Starting research workflows from Google SERP results.
- Reusing saved AI or manually configured scrapers across single URLs or batches.
- Retrieving stored extraction results for downstream processing.
The proposed integration is implemented in PR #3898.
Adoption signals
-
GitHub stars of the main repository: MrScraper is primarily a commercial hosted service, so its core platform is not represented by one open-source main repository. The official Python SDK repository currently has 0 stars. The MrScraper GitHub organization maintains nine public repositories covering its Python and Node.js SDKs, LangChain integration, n8n node, CLI, MCP server, and agent plugins. These repositories are new, so their current star counts are not yet representative of adoption of the hosted platform.
-
PyPI downloads in the last 30 days: The official mrscraper-sdk received 4,784 downloads in the last 30 days as of September 8, 2026 (live monthly-download badge, PyPI Stats).
-
Release activity: The latest mrscraper-sdk release is 0.3.0, published September 8, 2026. Six versions have been published since its initial release on March 9, 2026: 0.1.0, 0.1.1, 0.1.2, 0.2.0, 0.2.1, and 0.3.0. Releases have accompanied API and SDK development, with updates approximately every one to three months after the initial release.
-
Maintenance: MrScraper is actively maintained by a dedicated commercial company. The Python SDK received commits on September 8, 2026, and the official MCP server received commits on September 7, 2026. The company also maintains product documentation, SDKs, integrations, and support for the hosted API. MrScraper's LinkedIn page lists approximately 5,900 followers and a company size of 51–200 employees.
-
Haystack community demand: We have not found an earlier Haystack Discord thread or issue requesting the integration. This proposal and PR #3898 were initiated by the MrScraper team to make the existing demand for real-time web access and structured web extraction available natively to Haystack users.
-
Anything else: MrScraper already provides integrations for LangChain, n8n, Zapier, and MCP-compatible agents, as well as official Python and Node.js SDKs and a CLI. The n8n package received approximately 694 npm downloads during the latest reported 30-day period. These integrations demonstrate demand for using MrScraper in agent and automation workflows comparable to the proposed Haystack integration.
The public SDK and integration repositories are relatively new. We expect adoption to grow because AI Agents increasingly need reliable access to current, structured web data, while operating browser infrastructure, proxies, retries, and site-specific extraction logic remains expensive for individual application teams.
If helpful for evaluating adoption, the MrScraper team can also provide non-public aggregate metrics such as active accounts, monthly API requests, extracted records, customer growth, or enterprise usage directly to the Haystack maintainers.
Detailed design
The integration is distributed as the independent mrscraper-haystack package under the haystack_integrations namespace. It depends on haystack-ai and httpx and communicates with the documented MrScraper API.
It exposes 15 Haystack components organized into six capability groups:
-
Account
-
Discovery
MrScraperCrawlWebsiteUrls
MrScraperSearchGoogleSerp
-
Extraction
MrScraperExtractPageByPrompt
MrScraperExtractListings
MrScraperExtractStructuredData
MrScraperFetchRenderedHtml
-
Results
MrScraperGetResults
MrScraperGetLatestResults
MrScraperGetResultDetail
-
Scraper creation
MrScraperCreatePromptScraper
MrScraperCreateListingScraper
MrScraperCreateWebsiteCrawlScraper
-
Saved scraper execution
MrScraperRunExistingScraper
MrScraperRunExistingScraperBatch
Every component provides both run() and true asynchronous run_async() methods. The asynchronous methods use asynchronous HTTP operations rather than wrapping blocking calls.
All components expose one stable Haystack output:
Summary and motivation
MrScraper is a managed web-data platform for retrieving JavaScript-rendered pages, extracting structured data with natural-language instructions, crawling websites, querying Google search results, and running reusable scraping workflows.
Haystack applications frequently need current information that is not available in a model's training data or an existing document store. This integration gives Haystack Pipelines and Agents a supported way to acquire that information without requiring users to build and operate browser automation, proxy rotation, anti-bot handling, extraction prompts, and result-storage infrastructure themselves.
The integration supports use cases including:
The proposed integration is implemented in PR #3898.
Adoption signals
GitHub stars of the main repository: MrScraper is primarily a commercial hosted service, so its core platform is not represented by one open-source main repository. The official Python SDK repository currently has 0 stars. The MrScraper GitHub organization maintains nine public repositories covering its Python and Node.js SDKs, LangChain integration, n8n node, CLI, MCP server, and agent plugins. These repositories are new, so their current star counts are not yet representative of adoption of the hosted platform.
PyPI downloads in the last 30 days: The official
mrscraper-sdkreceived 4,784 downloads in the last 30 days as of September 8, 2026 (live monthly-download badge, PyPI Stats).Release activity: The latest
mrscraper-sdkrelease is 0.3.0, published September 8, 2026. Six versions have been published since its initial release on March 9, 2026: 0.1.0, 0.1.1, 0.1.2, 0.2.0, 0.2.1, and 0.3.0. Releases have accompanied API and SDK development, with updates approximately every one to three months after the initial release.Maintenance: MrScraper is actively maintained by a dedicated commercial company. The Python SDK received commits on September 8, 2026, and the official MCP server received commits on September 7, 2026. The company also maintains product documentation, SDKs, integrations, and support for the hosted API. MrScraper's LinkedIn page lists approximately 5,900 followers and a company size of 51–200 employees.
Haystack community demand: We have not found an earlier Haystack Discord thread or issue requesting the integration. This proposal and PR #3898 were initiated by the MrScraper team to make the existing demand for real-time web access and structured web extraction available natively to Haystack users.
Anything else: MrScraper already provides integrations for LangChain, n8n, Zapier, and MCP-compatible agents, as well as official Python and Node.js SDKs and a CLI. The n8n package received approximately 694 npm downloads during the latest reported 30-day period. These integrations demonstrate demand for using MrScraper in agent and automation workflows comparable to the proposed Haystack integration.
The public SDK and integration repositories are relatively new. We expect adoption to grow because AI Agents increasingly need reliable access to current, structured web data, while operating browser infrastructure, proxies, retries, and site-specific extraction logic remains expensive for individual application teams.
If helpful for evaluating adoption, the MrScraper team can also provide non-public aggregate metrics such as active accounts, monthly API requests, extracted records, customer growth, or enterprise usage directly to the Haystack maintainers.
Detailed design
The integration is distributed as the independent
mrscraper-haystackpackage under thehaystack_integrationsnamespace. It depends onhaystack-aiandhttpxand communicates with the documented MrScraper API.It exposes 15 Haystack components organized into six capability groups:
Account
MrScraperGetAccountInfoDiscovery
MrScraperCrawlWebsiteUrlsMrScraperSearchGoogleSerpExtraction
MrScraperExtractPageByPromptMrScraperExtractListingsMrScraperExtractStructuredDataMrScraperFetchRenderedHtmlResults
MrScraperGetResultsMrScraperGetLatestResultsMrScraperGetResultDetailScraper creation
MrScraperCreatePromptScraperMrScraperCreateListingScraperMrScraperCreateWebsiteCrawlScraperSaved scraper execution
MrScraperRunExistingScraperMrScraperRunExistingScraperBatchEvery component provides both
run()and true asynchronousrun_async()methods. The asynchronous methods use asynchronous HTTP operations rather than wrapping blocking calls.All components expose one stable Haystack output:
{"result": value}