io.github.ScrapeGraphAI/scrapegraph-mcp
PYPI · SCRAPEGRAPH-MCP · SCANNED SEP 20
AI-powered web scraping and data extraction capabilities through ScrapeGraph API
Available components
How this component scores in each security and reliability category. Every signal is checked automatically from public evidence about the published package, including repeated runs of it in an isolated sandbox, and we only credit what we can confirm. How we score → Why this is hard to score →
Supply Chain Security99
- No malware found by supply-chain analysis.Pass
- No known CVEs affecting this package version or its production dependencies.Pass
- Runs hatchling.build at install time, a recognised native-build step with no shell scripting around it. View diagnostics → Pass
- 4 of 34 dependencies flagged as unhealthy. View diagnostics → Partial
Provenance & Transparency44
- Source repository is publicly reachable at the declared URL. View diagnostics → Pass
- Provenance check failed: no build-provenance attestation is published. See how to fix → View diagnostics → Fail
- Clear OSI-approved license (MIT).Pass
- Actively maintained (last published 304 days ago).Pass
- Disclosure check failed: no security disclosure policy was found in the source repository. See how to fix → Fail
Schema Quality & AI Usability73
- 100% of prompts and resources have a non-trivial description (not blank, and not just the item's name).Pass
- AI-judged instruction clarity (excellent).Pass
- Context-footprint check failed: tool/resource definitions use about 5489 tokens (~457/item across 12 items; 8 tools + 4 resources), over budget; trim descriptions and params. See how to fix → Fail
- Usage-examples check failed: none of the tools include examples. See how to fix → Fail
Stability & Change Management87
- Stability observed for 26 of 30 days with no destabilising changes; credit accrues until the full window elapses.Partial
Tool Coverage91
- 100% of tools have a non-trivial description (not blank, and not just the tool's name).Pass
- 70% of tool parameters carry a description.Partial
- Structured output schemas are declared (100% of tools); any adoption earns full credit.Pass
Tool Safety100
- No prompt-injection markers were found in the server instructions, tool names or descriptions we captured.Pass
- All 1 tool(s) whose name or description implies an irreversible operation declare an MCP destructiveHint annotation.Pass
- An AI judge read all 9 captured unit(s) of tool text and found none that tries to manipulate the model reading it.Pass
Capabilities100
- Implements a current MCP spec version (2026-07-28).Pass
How do I install the io.github.ScrapeGraphAI/scrapegraph-mcp server?
io.github.ScrapeGraphAI/scrapegraph-mcp runs locally as a PyPI package, launched with uvx scrapegraph-mcp. Ready-made configuration for Claude, Cursor, VS Code, Codex and 5 more is on this page, copied from each client's own documentation.
pypi · scrapegraph-mcp
claude mcp add scrapegraphai-scrapegraph-mcp -- uvx scrapegraph-mcp
{
"mcpServers": {
"scrapegraphai-scrapegraph-mcp": {
"command": "uvx",
"args": [
"scrapegraph-mcp"
]
}
}
} {
"servers": {
"scrapegraphai-scrapegraph-mcp": {
"command": "uvx",
"args": [
"scrapegraph-mcp"
]
}
}
} codex mcp add scrapegraphai-scrapegraph-mcp -- uvx scrapegraph-mcp
{
"$schema": "https://opencode.ai/config.json",
"mcp": {
"scrapegraphai-scrapegraph-mcp": {
"type": "local",
"command": [
"uvx",
"scrapegraph-mcp"
],
"enabled": true
}
}
} openclaw mcp add scrapegraphai-scrapegraph-mcp --command uvx --arg scrapegraph-mcp
mcp_servers:
scrapegraphai-scrapegraph-mcp:
command: "uvx"
args: ["scrapegraph-mcp"] {
"McpServers": {
"scrapegraphai-scrapegraph-mcp": {
"Transport": "stdio",
"Command": "uvx",
"Arguments": [
"scrapegraph-mcp"
]
}
}
} assistant mcp add scrapegraphai-scrapegraph-mcp -t stdio -c uvx -a scrapegraph-mcp
{
"mcpServers": {
"scrapegraphai-scrapegraph-mcp": {
"command": "uvx",
"args": [
"scrapegraph-mcp"
]
}
}
} Every change we have recorded for this component, newest first. Security-relevant changes are always shown. ▲ marks a change for the better, ▼ a change for the worse; unmarked changes are neutral.
- 19 Sept 26 +1
No change was recorded against any check on this day. Stability & Change Management went from 80 to 83. That category is still filling its 30-day observation window: 24 days of observed history at the previous scan, 25 at this one. The score rises as the window fills, whether or not the server changes.
- 18 Sept 26 +12
- Malware scan: unverified → pass ▲ security
- Stability: pass → 0.80 functional
- 17 Sept 26 −15
- Malware scan: pass → unverified ▼ security
- Stability: 0.97 → pass security
- 16 Sept 26 +1
No change was recorded against any check on this day. Stability & Change Management went from 93 to 97. That category is still filling its 30-day observation window: 28 days of observed history at the previous scan, 29 at this one. The score rises as the window fills, whether or not the server changes.
- 15 Sept 26 +15
- Malware scan: unverified → pass ▲ security
- 14 Sept 26 −14
- Malware scan: pass → unverified ▼ security
- 12 Sept 26 +16
- Malware scan: unverified → pass ▲ security
- 11 Sept 26 −18
- Malware scan: pass → unverified ▼ security
- Stability: pass → 0.80 functional
Diagnostic detail from the automated scan of this channel: what the scanner observed at each step, so you can see exactly where a check passed or failed. It is informational only and never changes the trust score.
Captured 20 Sept 2026 · Analysed pypi/scrapegraph-mcp@1.0.1
Provenance No attestation
The registry publishes no build provenance for this version, so there is nothing to verify.
| Result | No attestation |
|---|---|
| Ecosystem | pypi |
Background: How many MCP packages publish verified provenance →
Install scripts 1 script
| Hook | Tier | Command |
|---|---|---|
| build_backend | allowlisted | hatchling.build |
Background: Why install scripts are a supply-chain risk →
Dependencies 34 packages
| Packages resolved | 34 |
|---|---|
| Stale | 4 |
| Tree resolution | Complete |
Background: SBOMs and build attestations, explained →
The tools this component advertises to a client, with an estimated token cost for each. Expand a tool to see its parameters and schema. The per-tool counts are indicative and are not scored directly; the schema's total context footprint is one signal in Schema Quality & AI Usability. A tool's description is untrusted text the model reads on every call, which is what makes this list a security surface and not just an inventory: how tool poisoning works →
agentic_scrapper Agentic Scrapper ~1,379
Execute complex multi-step web scraping workflows with AI-powered automation. This tool runs an intelligent agent that can navigate websites, interact with forms and buttons, follow multi-step workflows, and extract structured data. Ideal for complex scraping scenarios requiring user interaction simulation, form submissions, or multi-page navigation flows. Supports custom output schemas and step-by-step instructions. Variable credit cost based on complexity. Can perform actions on the website (non-read-only, non-idempotent). The agent accepts flexible input formats for steps (list or JSON string) and output_schema (dict or JSON string) to accommodate different client implementations.
| Name | Type | Req | Description |
|---|---|---|---|
| ai_extraction | – | – | Enable AI-powered extraction mode for intelligent data parsing. - Default: true (recommended for most use cases) - Options: * true: Uses advanced AI to intelligently extract and structure data… |
| output_schema | – | – | Desired output structure for extracted data. - Can be provided as a dictionary or JSON string - Defines the format and structure of the final extracted data - Helps ensure consistent, predictable out… |
| persistent_session | – | – | Maintain session state between steps. - Default: false (each step starts fresh) - Options: * true: Keeps cookies, login state, and session data between steps - Essential for authenticated workf… |
| steps | – | – | Step-by-step instructions for the agent. - Can be provided as a list of strings or JSON array string - Provides detailed, sequential instructions for the automation workflow - Each step should be a c… |
| timeout_seconds | – | – | Maximum time to wait for the entire workflow. - Default: 120 seconds (2 minutes) - Recommended ranges: * 60-120: Simple workflows (2-5 steps) * 180-300: Medium complexity (5-10 steps) * 300-600… |
| url | string | yes | The target website URL where the agentic scraping workflow should start. - Must include protocol (http:// or https://) - Should be the starting page for your automation workflow - The agent will begi… |
| user_prompt | – | – | High-level instructions for what the agent should accomplish. - Describes the overall goal and desired outcome of the automation - Should be clear and specific about what you want to achieve - Works… |
Structured output declared, but exposes no named fields.
No examples provided.
markdownify Markdownify ~184
Convert a webpage into clean, formatted markdown. This tool fetches any webpage and converts its content into clean, readable markdown format. Useful for extracting content from documentation, articles, and web pages for further processing. Costs 2 credits per page. Read-only operation with no side effects.
| Name | Type | Req | Description |
|---|---|---|---|
| website_url | string | yes | The complete URL of the webpage to convert to markdown format. - Must include protocol (http:// or https://) - Supports most web content types (HTML, articles, documentation) - Works with both static… |
Structured output declared, but exposes no named fields.
No examples provided.
scrape Scrape ~397
Fetch raw page content from any URL with optional JavaScript rendering. This tool performs basic web scraping to retrieve the raw HTML content of a webpage. Optionally enable JavaScript rendering for Single Page Applications (SPAs) and sites with heavy client-side rendering. Lower cost than AI extraction (1 credit/page). Read-only operation with no side effects.
| Name | Type | Req | Description |
|---|---|---|---|
| render_heavy_js | – | – | Enable full JavaScript rendering for dynamic content. - Default: false (faster, lower cost, works for most static sites) - Set to true for sites that require JavaScript execution to display content -… |
| website_url | string | yes | The complete URL of the webpage to scrape. - Must include protocol (http:// or https://) - Returns raw HTML content of the page - Works with both static and dynamic websites - Examples: * https://e… |
Structured output declared, but exposes no named fields.
No examples provided.
searchscraper Searchscraper ~546
Perform AI-powered web searches with structured data extraction. This tool searches the web based on your query and uses AI to extract structured information from the search results. Ideal for research, competitive analysis, and gathering information from multiple sources. Each website searched costs 10 credits (default 3 websites = 30 credits). Read-only operation but results may vary over time (non-idempotent).
| Name | Type | Req | Description |
|---|---|---|---|
| num_results | – | – | Number of websites to search and extract data from. - Default: 3 websites (costs 30 credits total) - Range: 1-20 websites (recommended to stay under 10 for cost efficiency) - Each website costs 10 cr… |
| number_of_scrolls | – | – | Number of infinite scrolls per searched webpage. - Default: 0 (no scrolling on search result pages) - Range: 0-10 scrolls per page - Useful when search results point to pages with dynamic content loa… |
| user_prompt | string | yes | Search query or natural language instructions for information to find. - Can be a simple search query or detailed extraction instructions - The AI will search the web and extract relevant data from f… |
Structured output declared, but exposes no named fields.
No examples provided.
sitemap Sitemap ~277
Extract and discover the complete sitemap structure of any website. This tool automatically discovers all accessible URLs and pages within a website, providing a comprehensive map of the site's structure. Useful for understanding site architecture before crawling or for discovering all available content. Very cost-effective at 1 credit per request. Read-only operation with no side effects.
| Name | Type | Req | Description |
|---|---|---|---|
| website_url | string | yes | The base URL of the website to extract sitemap from. - Must include protocol (http:// or https://) - Should be the root domain or main section you want to map - The tool will discover all accessible… |
Structured output declared, but exposes no named fields.
No examples provided.
smartcrawler_fetch_results Smartcrawler Fetch Results ~122
Retrieve the results of an asynchronous SmartCrawler operation. This tool fetches the results from a previously initiated crawling operation using the request_id. The crawl request processes asynchronously in the background. Keep polling this endpoint until the status field indicates 'completed'. While processing, you'll receive status updates. Read-only operation that safely retrieves results without side effects.
| Name | Type | Req | Description |
|---|---|---|---|
| request_id | string | yes | The unique request ID returned by smartcrawler_initiate. Use this to retrieve the crawling results. Keep polling until status is 'completed'. Example: 'req_abc123xyz' |
Structured output declared, but exposes no named fields.
No examples provided.
smartcrawler_initiate Smartcrawler Initiate ~1,129
Initiate an asynchronous multi-page web crawling operation with AI extraction or markdown conversion. This tool starts an intelligent crawler that discovers and processes multiple pages from a starting URL. Choose between AI Extraction Mode (10 credits/page) for structured data or Markdown Mode (2 credits/page) for content conversion. The operation is asynchronous - use smartcrawler_fetch_results to retrieve results. Creates a new crawl request (non-idempotent, non-read-only). SmartCrawler supports two modes: - AI Extraction Mode: Extracts structured data based on your prompt from every crawled page - Markdown Conversion Mode: Converts each page to clean markdown format
| Name | Type | Req | Description |
|---|---|---|---|
| depth | – | – | Maximum depth of link traversal from the starting URL. - Default: unlimited (will follow links until max_pages or no more links) - Depth levels: * 0: Only the starting URL (no link following) * 1… |
| extraction_mode | string | – | Extraction mode for processing crawled pages. - Default: "ai" - Options: * "ai": AI-powered structured data extraction (10 credits per page) - Uses the prompt to extract specific data from each… |
| max_pages | – | – | Maximum number of pages to crawl in total. - Default: unlimited (will crawl until no more links or depth limit) - Recommended ranges: * 10-20: Testing and small sites * 50-100: Medium sites and f… |
| prompt | – | – | AI prompt for data extraction. - REQUIRED when extraction_mode is 'ai' - Ignored when extraction_mode is 'markdown' - Describes what data to extract from each crawled page - Applied consistently acro… |
| same_domain_only | – | – | Whether to crawl only within the same domain. - Default: true (recommended for most use cases) - Options: * true: Only crawl pages within the same domain as starting URL - Prevents following ex… |
| url | string | yes | The starting URL to begin crawling from. - Must include protocol (http:// or https://) - The crawler will discover and process linked pages from this starting point - Should be a page with links to o… |
Structured output declared, but exposes no named fields.
No examples provided.
smartscraper Smartscraper ~1,328
Extract structured data from a webpage, HTML, or markdown using AI-powered extraction. This tool uses advanced AI to understand your natural language prompt and extract specific structured data from web content. Supports three input modes: URL scraping, local HTML processing, or local markdown processing. Ideal for extracting product information, contact details, article metadata, or any structured content. Costs 10 credits per page. Read-only operation. Args: user_prompt (str): Natural language instructions describing what data to extract. - Be specific about the fields you want for better results - Use clear, descriptive language about the target data - Examples: * "Extract product name, price, description, and availability status" * "Find all contact methods: email addresses, phone numbers, and social media links" * "Get article title, author, publication date, and summary" * "Extract all job listings with title, company, location, and salary" - Tips for better results: * Specify exact field names you want * Mention data types (numbers, dates, URLs, etc.) * Include context about where data might be located website_url (Optional[str]): The complete URL of the webpage to scrape. - Mutually exclusive with website_html and website_markdown - Must include protocol (http:// or https://) - Supports dynamic and static content - Examples: * https://example.com/products/item * https://news.site.com/article/123 * https://company.com/contact - Default: None (must provide one of the three input sources) website_html (Optional[str]): Raw HTML content to process locally. - Mutually exclusive with website_url and website_markdown - Maximum size: 2MB…
| Name | Type | Req | Description |
|---|---|---|---|
| number_of_scrolls | – | – | – |
| output_schema | – | – | – |
| render_heavy_js | – | – | – |
| stealth | – | – | – |
| total_pages | – | – | – |
| user_prompt | string | yes | – |
| website_html | – | – | – |
| website_markdown | – | – | – |
| website_url | – | – | – |
Structured output declared, but exposes no named fields.
No examples provided.
What is the io.github.ScrapeGraphAI/scrapegraph-mcp server?
io.github.ScrapeGraphAI/scrapegraph-mcp is listed in the public MCP registry as io.github.ScrapeGraphAI/scrapegraph-mcp. AI-powered web scraping and data extraction capabilities through ScrapeGraph API. This page covers its PyPI package (scrapegraph-mcp).
Is the io.github.ScrapeGraphAI/scrapegraph-mcp server safe to use?
io.github.ScrapeGraphAI/scrapegraph-mcp scores 81 out of 100 on VerifyMCP. We found no known CVEs affecting it as of 20 September 2026. That is a record of what we were able to check automatically, not an endorsement. The category breakdown on this page shows every signal behind the number, including the ones we could not confirm.
What tools does the io.github.ScrapeGraphAI/scrapegraph-mcp server expose?
io.github.ScrapeGraphAI/scrapegraph-mcp exposes 8 tools: markdownify, smartscraper, searchscraper, smartcrawler_initiate, smartcrawler_fetch_results, and 3 more. Their descriptions and schemas cost roughly 5,362 tokens of context every time the server is loaded.
Is the io.github.ScrapeGraphAI/scrapegraph-mcp server still maintained?
io.github.ScrapeGraphAI/scrapegraph-mcp is still listed as active in the MCP registry. We last reached this channel on 20 September 2026. Those dates come from our own scans of the registry and the channel itself, not from anything the publisher announced.
What licence is the io.github.ScrapeGraphAI/scrapegraph-mcp server under?
io.github.ScrapeGraphAI/scrapegraph-mcp declares the MIT licence, which is OSI-approved. That covers the source only, and says nothing about the cost of any service it calls.