io.github.vola-trebla/flakiness-knowledge-graph-mcp
NPM · FLAKINESS-KNOWLEDGE-GRAPH-MCP · SCANNED AUG 3
MCP server + Playwright reporter that builds a flakiness knowledge graph from test run history
Available components
How this component scores in each security and reliability category. Every signal is checked automatically from public evidence about the published package, including repeated runs of it in an isolated sandbox, and we only credit what we can confirm. How we score →
Supply Chain Security87
- No malware found by supply-chain analysis.Pass
- Only part of the dependency tree could be resolved (95 of 99), so this covers what we could see, not the whole tree.Partial
- No install/post-install scripts declared.Pass
- Only part of the dependency tree could be resolved (95 of 99), so this covers what we could see, not the whole tree. View diagnostics → Partial
Provenance & Transparency45
- Source repository is publicly reachable at the declared URL. View diagnostics → Pass
- Provenance check failed: no build-provenance attestation is published. See how to fix → View diagnostics → Fail
- Clear OSI-approved license (MIT).Pass
- Actively maintained (last published 76 days ago).Pass
- Disclosure check failed: no security disclosure policy was found in the source repository. See how to fix → Fail
Schema Quality & AI Usability77
- AI-judged instruction clarity (excellent).Pass
- Tool/resource definitions use about 833 tokens (~104/item across 8 items; 8 tools + 0 resources), lean.Pass
- Usage-examples check failed: none of the tools include examples. See how to fix → Fail
Stability & Change Management23
- Stability observed for 7 of 30 days with no destabilising changes; credit accrues until the full window elapses.Partial
Tool Coverage100
- 100% of tools have a non-trivial description (not blank, and not just the tool's name).Pass
- 100% of tool parameters carry a description.Pass
Capabilities100
- Implements a supported MCP spec version (2025-11-25); the latest is 2026-07-28.Pass
Add this component to your MCP client. Where a client-specific snippet is available, pick your client below and copy it straight into your config; otherwise use the connection detail shown.
npm · flakiness-knowledge-graph-mcp
claude mcp add vola-trebla-flakiness-knowledge-graph-mcp -- npx -y flakiness-knowledge-graph-mcp
codex mcp add vola-trebla-flakiness-knowledge-graph-mcp -- npx -y flakiness-knowledge-graph-mcp
{
"$schema": "https://opencode.ai/config.json",
"mcp": {
"vola-trebla-flakiness-knowledge-graph-mcp": {
"type": "local",
"command": [
"npx",
"-y",
"flakiness-knowledge-graph-mcp"
],
"enabled": true
}
}
} openclaw mcp add vola-trebla-flakiness-knowledge-graph-mcp --command npx --arg -y --arg flakiness-knowledge-graph-mcp
mcp_servers:
vola-trebla-flakiness-knowledge-graph-mcp:
command: "npx"
args: ["-y", "flakiness-knowledge-graph-mcp"] {
"mcpServers": {
"vola-trebla-flakiness-knowledge-graph-mcp": {
"command": "npx",
"args": [
"-y",
"flakiness-knowledge-graph-mcp"
]
}
}
} Every change we have recorded for this component, newest first. Security-relevant changes are always shown. ▲ marks a change for the better, ▼ a change for the worse; unmarked changes are neutral.
- 2 Aug 26 +48
- Provenance: unverified → fail ▼ security
- Known CVEs: unverified → partial ▲ security
- Install scripts: unverified → pass ▲ security
- Malware scan: unverified → pass ▲ security
- MCP protocol: unverified → pass ▲ functional
- Schema quality: unverified → excellent ▲ functional
- License: unverified → pass ▲ functional
- Dependency health: unverified → partial ▲ functional
- Stability: unverified → 0.20 ▲ functional
- Maintenance: unverified → pass ▲ functional
- Licence: MIT functional
- 1 Aug 26 −7
- We updated how we score, so this day's move reflects our rubric, not a change to the server See what changed → functional
- 31 Jul 26 −18
- Malware scan: pass → unverified ▼ security
- 27 Jul 26 46
First indexed and scored.
Diagnostic detail from the automated scan of this channel: what the scanner observed at each step, so you can see exactly where a check passed or failed. It is informational only and never changes the trust score.
Captured 3 Aug 2026 · Analysed npm/[email protected]
Provenance none
Ecosystem: npm · Outcome: none
Dependencies 95 packages
95 packages in the resolved dependency tree · 95 deprecated · 29 stale.
The dependency tree was only partially resolved, so these counts may be incomplete.
The tools this component advertises to a client, with an estimated token cost for each. Expand a tool to see its parameters and schema. The per-tool counts are indicative and are not scored directly; the schema's total context footprint is one signal in Schema Quality & AI Usability.
cluster_semantic_error_trees ~157
Groups test failures by semantic error similarity rather than raw string prefix. Strips dynamic values (UUIDs, numeric IDs, hashes, timestamps, URLs) via regex, then applies Levenshtein fuzzy matching to merge errors that differ only in minor dynamic fragments. Classifies each cluster by taxonomy (TimeoutError, AssertionError, NetworkError, ReferenceError). Use instead of get_error_groups when failure messages contain dynamic IDs or selector attributes that make identical root causes look different.
| Name | Type | Req | Description |
|---|---|---|---|
| db_path | string | yes | Absolute path to the flakiness.db SQLite file |
| min_instances | integer | — | Minimum number of failure instances to include a cluster |
| since_days | integer | — | Only include failures from the last N days |
No output schema declared.
No examples provided.
correlate_git_commit_flakiness ~153
Finds the exact point where a test transitioned from stable to flaky (or back), and returns the git commit SHA, branch, and author at that transition. Requires the Playwright reporter to be running in a CI environment where GITHUB_SHA / CI_COMMIT_SHA / CIRCLE_SHA1 env vars are set. Use to answer: which commit broke this test, and who authored it?
| Name | Type | Req | Description |
|---|---|---|---|
| db_path | string | yes | Absolute path to the flakiness.db SQLite file |
| min_stable_runs | integer | — | Consecutive passes required to consider a test 'stable' before a transition (default 3) |
| since_days | integer | — | Only look at runs from the last N days |
No output schema declared.
No examples provided.
get_error_groups ~106
Groups failing tests by similar error messages to surface systemic failures. Use to answer: are 10 tests failing because of the same broken endpoint or shared root cause?
| Name | Type | Req | Description |
|---|---|---|---|
| db_path | string | yes | Absolute path to the flakiness.db SQLite file |
| limit | integer | — | Max error groups to return |
| min_failures | integer | — | Minimum number of failures sharing the same error to include |
| since_days | integer | — | Only include failures from the last N days |
No output schema declared.
No examples provided.
get_failure_patterns ~70
Breaks down failure rates by browser and OS combination. Use to answer: does this test only fail on Firefox? Only on Windows?
| Name | Type | Req | Description |
|---|---|---|---|
| db_path | string | yes | Absolute path to the flakiness.db SQLite file |
| since_days | integer | — | Only include runs from the last N days |
No output schema declared.
No examples provided.
get_flakiness_trend ~92
Returns the daily flakiness rate for a specific test over the last N days. Use to answer: is this test getting worse, better, or staying the same?
| Name | Type | Req | Description |
|---|---|---|---|
| days | integer | — | Number of days to look back |
| db_path | string | yes | Absolute path to the flakiness.db SQLite file |
| test_id | string | yes | Test ID from get_flaky_tests |
No output schema declared.
No examples provided.
get_flaky_tests ~105
Returns tests ranked by flakiness rate (failed+flaky / total runs). Use to answer: which tests are the most unreliable?
| Name | Type | Req | Description |
|---|---|---|---|
| db_path | string | yes | Absolute path to the flakiness.db SQLite file |
| limit | integer | — | Max tests to return |
| min_runs | integer | — | Minimum number of runs to consider a test (filters out one-off failures) |
| since_days | integer | — | Only include runs from the last N days |
No output schema declared.
No examples provided.
get_slow_tests ~59
Returns tests ranked by average duration. Use to answer: which tests are slowing down the CI pipeline?
| Name | Type | Req | Description |
|---|---|---|---|
| db_path | string | yes | Absolute path to the flakiness.db SQLite file |
| limit | integer | — | Max tests to return |
No output schema declared.
No examples provided.
get_test_history ~91
Returns the run history for a specific test — status, duration, error, retry count, browser, and OS for each run. Use to answer: is this test getting worse over time?
| Name | Type | Req | Description |
|---|---|---|---|
| db_path | string | yes | Absolute path to the flakiness.db SQLite file |
| limit | integer | — | Max runs to return |
| test_id | string | yes | Test ID from get_flaky_tests |
No output schema declared.
No examples provided.