Data Profiler
PYPI · MCP-DATA-PROFILER · SCANNED SEP 21
Profile local CSV, Parquet, JSON and Excel files into a compact data-quality summary.
Available components
How this component scores in each security and reliability category. Every signal is checked automatically from public evidence about the published package, including repeated runs of it in an isolated sandbox, and we only credit what we can confirm. How we score → Why this is hard to score →
Supply Chain Security99
- No malware found by supply-chain analysis.Pass
- No known CVEs affecting this package version or its production dependencies.Pass
- Runs hatchling.build at install time, a recognised native-build step with no shell scripting around it. View diagnostics → Pass
- 3 of 37 dependencies flagged as unhealthy. View diagnostics → Partial
Provenance & Transparency32
- Source repository is publicly reachable at the declared URL. View diagnostics → Pass
- Provenance check failed: no build-provenance attestation is published. See how to fix → View diagnostics → Fail
- License check failed: the license (MIT License) isn't a recognized OSI-approved license. See how to fix → Fail
- Actively maintained (last published 48 days ago).Pass
- Disclosure check failed: no security disclosure policy was found in the source repository. See how to fix → Fail
Schema Quality & AI Usability61
- AI-judged instruction clarity (excellent).Pass
- Context-footprint check failed: tool/resource definitions use about 457 tokens (~457/item across 1 items; 1 tools + 0 resources), over budget; trim descriptions and params. See how to fix → Fail
- Usage-examples check failed: none of the tools include examples. See how to fix → Fail
Stability & Change Management90
- Stability observed for 27 of 30 days with no destabilising changes; credit accrues until the full window elapses.Partial
Tool Coverage71
- 100% of tools have a non-trivial description (not blank, and not just the tool's name).Pass
- 0% of tool parameters carry a description.Fail
- Structured output schemas are declared (100% of tools); any adoption earns full credit.Pass
Tool Safety100
- No prompt-injection markers were found in the server instructions, tool names or descriptions we captured.Pass
- We read all 1 captured tool definition(s), and no name or description among them implies an irreversible operation.Pass
- An AI judge read all 2 captured unit(s) of tool text and found none that tries to manipulate the model reading it.Pass
Capabilities100
- Implements a current MCP spec version (2026-07-28).Pass
How do I install the Data Profiler MCP server?
Data Profiler runs locally as a PyPI package, launched with uvx mcp-data-profiler. Ready-made configuration for Claude, Cursor, VS Code, Codex and 5 more is on this page, copied from each client's own documentation.
pypi · mcp-data-profiler
claude mcp add ridadata-mcp-data-profiler -- uvx mcp-data-profiler
{
"mcpServers": {
"ridadata-mcp-data-profiler": {
"command": "uvx",
"args": [
"mcp-data-profiler"
]
}
}
} {
"servers": {
"ridadata-mcp-data-profiler": {
"command": "uvx",
"args": [
"mcp-data-profiler"
]
}
}
} codex mcp add ridadata-mcp-data-profiler -- uvx mcp-data-profiler
{
"$schema": "https://opencode.ai/config.json",
"mcp": {
"ridadata-mcp-data-profiler": {
"type": "local",
"command": [
"uvx",
"mcp-data-profiler"
],
"enabled": true
}
}
} openclaw mcp add ridadata-mcp-data-profiler --command uvx --arg mcp-data-profiler
mcp_servers:
ridadata-mcp-data-profiler:
command: "uvx"
args: ["mcp-data-profiler"] {
"McpServers": {
"ridadata-mcp-data-profiler": {
"Transport": "stdio",
"Command": "uvx",
"Arguments": [
"mcp-data-profiler"
]
}
}
} assistant mcp add ridadata-mcp-data-profiler -t stdio -c uvx -a mcp-data-profiler
{
"mcpServers": {
"ridadata-mcp-data-profiler": {
"command": "uvx",
"args": [
"mcp-data-profiler"
]
}
}
} Every change we have recorded for this component, newest first. Security-relevant changes are always shown. ▲ marks a change for the better, ▼ a change for the worse; unmarked changes are neutral.
- 21 Sept 26 +15
- Malware scan: unverified → pass ▲ security
- 20 Sept 26 +1
No change was recorded against any check on this day. Stability & Change Management went from 83 to 87. That category is still filling its 30-day observation window: 25 days of observed history at the previous scan, 26 at this one. The score rises as the window fills, whether or not the server changes.
- 18 Sept 26 −18
- Malware scan: pass → unverified ▼ security
- Stability: pass → 0.80 functional
- 17 Sept 26 +16
- Malware scan: unverified → pass ▲ security
- Stability: 0.97 → pass security
- 15 Sept 26 +1
No change was recorded against any check on this day. Stability & Change Management went from 90 to 93. That category is still filling its 30-day observation window: 27 days of observed history at the previous scan, 28 at this one. The score rises as the window fills, whether or not the server changes.
- 14 Sept 26 −15
- Malware scan: pass → unverified ▼ security
- 13 Sept 26 +1
No change was recorded against any check on this day. Stability & Change Management went from 83 to 87. That category is still filling its 30-day observation window: 25 days of observed history at the previous scan, 26 at this one. The score rises as the window fills, whether or not the server changes.
- 11 Sept 26 −3
- Stability: pass → 0.80 functional
Diagnostic detail from the automated scan of this channel: what the scanner observed at each step, so you can see exactly where a check passed or failed. It is informational only and never changes the trust score.
Captured 21 Sept 2026 · Analysed pypi/mcp-data-profiler@0.1.1
Provenance No attestation
The registry publishes no build provenance for this version, so there is nothing to verify.
| Result | No attestation |
|---|---|
| Ecosystem | pypi |
Background: How many MCP packages publish verified provenance →
Install scripts 1 script
| Hook | Tier | Command |
|---|---|---|
| build_backend | allowlisted | hatchling.build |
Background: Why install scripts are a supply-chain risk →
Dependencies 37 packages
| Packages resolved | 37 |
|---|---|
| Stale | 2 |
| No linked repository | 1 |
| Tree resolution | Complete |
Background: SBOMs and build attestations, explained →
The tools this component advertises to a client, with an estimated token cost for each. Expand a tool to see its parameters and schema. The per-tool counts are indicative and are not scored directly; the schema's total context footprint is one signal in Schema Quality & AI Usability. A tool's description is untrusted text the model reads on every call, which is what makes this list a security surface and not just an inventory: how tool poisoning works →
profile_dataset ~414
Summarise the structure and quality of a local data file. Call this whenever you need to understand a dataset — its columns, types, ranges, missing values, or quality problems — before analysing it, writing code against it, or answering questions about it. Prefer this over reading the file directly: it returns a compact summary instead of raw rows, so it works on files far too large to read, at a small fraction of the tokens. Reports per column: dtype, null count and percentage, distinct count, sample values, quartiles for numbers, date ranges, and the most frequent values for categories. Flags likely problems: all-null and constant columns, probable ID columns, mixed types, and numbers or dates that were stored as text. Args: path: Path to the file. Supports .csv, .tsv, .parquet, .json, .jsonl, .xlsx, and .xls. sample_rows: Profile at most this many rows. Pass null to read every row, which is slower on large files but makes all statistics exact. The result always states whether it was sampled. max_columns: Describe at most this many columns, so the response stays small on very wide tables. The true column count is always reported. top_k: How many of the most frequent values to list per categorical column. sheet: For Excel workbooks, the name of the sheet to profile. Defaults to the first sheet, which is often a title or notes page rather than the data. The result lists every available sheet, so if the one profiled looks empty or wrong, call again naming another. Returns: A profile with file info, shape, per-column detail, and duplicate row count.
| Name | Type | Req | Description |
|---|---|---|---|
| max_columns | integer | – | – |
| path | string | yes | – |
| sample_rows | – | – | – |
| sheet | – | – | – |
| top_k | integer | – | – |
Structured output declared, but exposes no named fields.
No examples provided.
What is the Data Profiler MCP server?
Data Profiler is an MCP server listed in the public MCP registry as io.github.Ridadata/mcp-data-profiler. Profile local CSV, Parquet, JSON and Excel files into a compact data-quality summary. This page covers its PyPI package (mcp-data-profiler).
Is the Data Profiler MCP server safe to use?
Data Profiler scores 75 out of 100 on VerifyMCP. We found no known CVEs affecting it as of 21 September 2026. That is a record of what we were able to check automatically, not an endorsement. The category breakdown on this page shows every signal behind the number, including the ones we could not confirm.
What tools does the Data Profiler MCP server expose?
Data Profiler exposes 1 tool: profile_dataset. Their descriptions and schemas cost roughly 414 tokens of context every time the server is loaded.
Is the Data Profiler MCP server still maintained?
Data Profiler is still listed as active in the MCP registry. We last reached this channel on 21 September 2026. Those dates come from our own scans of the registry and the channel itself, not from anything the publisher announced.