Skip to content
verify mcp Beta VerifyMCP is currently in beta. If you notice any issues, get in touch and we’ll put it right.

io.github.charliemtnez/ghost-inspector-mcp

NPM · GHOST-INSPECTOR-MCP · SCANNED OCT 1

Analyze, validate and safely update Ghost Inspector end-to-end browser tests

Available components

+15 this week 81 Trust /100
Trust breakdown (7 categories)

How this component scores in each security and reliability category. Every signal is checked automatically from public evidence about the published package, including repeated runs of it in an isolated sandbox, and we only credit what we can confirm. How we score → Why this is hard to score →

Supply Chain Security98
  • No malware found by supply-chain analysis.Pass
  • No known CVEs affecting this package version or its production dependencies.Pass
  • No install/post-install scripts declared.Pass
  • 31 of 92 dependencies flagged as unhealthy. View diagnostics → Partial
Provenance & Transparency48
  • Source repository is publicly reachable at the declared URL. View diagnostics → Pass
  • Provenance check failed: no build-provenance attestation is published. See how to fix → View diagnostics → Fail
  • Clear OSI-approved license (MIT).Pass
  • Actively maintained (last published 2 days ago).Pass
  • Publishes a security disclosure policy (SECURITY.md).Pass
Schema Quality & AI Usability64
  • AI-judged instruction clarity (excellent).Pass
  • Context-footprint check failed: tool/resource definitions use about 8013 tokens (~364/item across 22 items; 22 tools + 0 resources), over budget; trim descriptions and params. See how to fix → Fail
  • Usage-examples check failed: none of the tools include examples. See how to fix → Fail
Stability & Change Management90
  • Stability observed for 27 of 30 days with no destabilising changes; credit accrues until the full window elapses.Partial
Tool Coverage100
  • 100% of tools have a non-trivial description (not blank, and not just the tool's name).Pass
  • 100% of tool parameters carry a description.Pass
Tool Safety100
  • No prompt-injection markers were found in the server instructions, tool names or descriptions we captured.Pass
  • All 1 tool(s) whose name or description implies an irreversible operation declare an MCP destructiveHint annotation.Pass
  • An AI judge read all 23 captured unit(s) of tool text and found none that tries to manipulate the model reading it.Pass
Capabilities100
  • Implements a supported MCP spec version (2025-11-25); the latest is 2026-07-28.Pass
Install

How do I install the io.github.charliemtnez/ghost-inspector-mcp server?

io.github.charliemtnez/ghost-inspector-mcp runs locally as an npm package, launched with npx -y ghost-inspector-mcp. Ready-made configuration for Claude, Cursor, VS Code, Codex and 5 more is on this page, copied from each client's own documentation.

npm · ghost-inspector-mcp

# add to Claude Code
claude mcp add charliemtnez-ghost-inspector-mcp -- npx -y ghost-inspector-mcp
// .cursor/mcp.json
{
  "mcpServers": {
    "charliemtnez-ghost-inspector-mcp": {
      "command": "npx",
      "args": [
        "-y",
        "ghost-inspector-mcp"
      ]
    }
  }
}
// .vscode/mcp.json
{
  "servers": {
    "charliemtnez-ghost-inspector-mcp": {
      "command": "npx",
      "args": [
        "-y",
        "ghost-inspector-mcp"
      ]
    }
  }
}
# add to Codex CLI
codex mcp add charliemtnez-ghost-inspector-mcp -- npx -y ghost-inspector-mcp
// opencode.json
{
  "$schema": "https://opencode.ai/config.json",
  "mcp": {
    "charliemtnez-ghost-inspector-mcp": {
      "type": "local",
      "command": [
        "npx",
        "-y",
        "ghost-inspector-mcp"
      ],
      "enabled": true
    }
  }
}
# add to OpenClaw
openclaw mcp add charliemtnez-ghost-inspector-mcp --command npx --arg -y --arg ghost-inspector-mcp
# ~/.hermes/config.yaml
mcp_servers:
  charliemtnez-ghost-inspector-mcp:
    command: "npx"
    args: ["-y", "ghost-inspector-mcp"]
// ~/.netclaw/config/netclaw.json
{
  "McpServers": {
    "charliemtnez-ghost-inspector-mcp": {
      "Transport": "stdio",
      "Command": "npx",
      "Arguments": [
        "-y",
        "ghost-inspector-mcp"
      ]
    }
  }
}
# add to Vellum
assistant mcp add charliemtnez-ghost-inspector-mcp -t stdio -c npx -a -y ghost-inspector-mcp
// mcp.json
{
  "mcpServers": {
    "charliemtnez-ghost-inspector-mcp": {
      "command": "npx",
      "args": [
        "-y",
        "ghost-inspector-mcp"
      ]
    }
  }
}
Changelog

Every change we have recorded for this component, newest first. Security-relevant changes are always shown. ▲ marks a change for the better, ▼ a change for the worse; unmarked changes are neutral.

  • 30 Sept 26 +1

    No change was recorded against any check on this day. Stability & Change Management went from 83 to 87. That category is still filling its 30-day observation window: 25 days of observed history at the previous scan, 26 at this one. The score rises as the window fills, whether or not the server changes.

  • 29 Sept 26 +15
    • Malware scan: unverified → pass ▲ security
    • Package version: 0.2.0 → 0.4.1 functional
  • 28 Sept 26 −18
    • Malware scan: pass → unverified ▼ security
    • Schema quality: 7190 → 8013 ▼ functional
    • Stability: pass → 0.80 functional
    • Package version: 0.3.0 → 0.4.1 functional
    • Package version: 0.3.0 → 0.4.0 functional
    • We updated how we score, so this day's move reflects our rubric, not a change to the server See what changed → functional
  • 27 Sept 26 +1
    • Stability: 0.97 → pass security
  • 25 Sept 26 +16
    • We updated how we score, so this day's move reflects our rubric, not a change to the server See what changed → functional
  • 24 Sept 26 −15
    • Malware scan: pass → unverified ▼ security
    • Schema quality: 325 → 359 ▼ functional
    • Package version: 0.2.0 → 0.3.0 functional
  • 22 Sept 26 +1

    No change was recorded against any check on this day. Stability & Change Management went from 80 to 83. That category is still filling its 30-day observation window: 24 days of observed history at the previous scan, 25 at this one. The score rises as the window fills, whether or not the server changes.

  • 21 Sept 26 −3
    • Stability: pass → 0.80 functional
Diagnostics

Diagnostic detail from the automated scan of this channel: what the scanner observed at each step, so you can see exactly where a check passed or failed. It is informational only and never changes the trust score.

Captured 1 Oct 2026 · Analysed npm/ghost-inspector-mcp@0.4.1

Provenance No attestation

The registry publishes no build provenance for this version, so there is nothing to verify.

Result No attestation
Ecosystem npm

Background: How many MCP packages publish verified provenance →

Dependencies 92 packages
Packages resolved 92
Stale 31
Tree resolution Complete

Background: SBOMs and build attestations, explained →

MCP tools · 22 exposed · ~7,676 tokens

The tools this component advertises to a client, with an estimated token cost for each. Expand a tool to see its parameters and schema. The per-tool counts are indicative and are not scored directly; the schema's total context footprint is one signal in Schema Quality & AI Usability. A tool's description is untrusted text the model reads on every call, which is what makes this list a security surface and not just an inventory: how tool poisoning works →

Tool Tokens
gi_accept_screenshot ~279

Makes the latest result's screenshot the baseline that every later run of this test is compared against. 🔴 The API offers no way to restore an earlier baseline; the response returns `previousBaselineResult` so it can at least be found again. `expectedResultId` is required: the result whose screenshot you looked at, from gi_screenshot_status. The accept is refused if a newer run has landed since, if the latest run is still going, or if its comparison passed or did not run (there is nothing to accept then). The check is read-then-accept with no compare-and-swap, so a run landing in the moment between them could still slip through: it catches a stale id, not a genuine race. After the accept the accepted result is re-read: `verification` confirms it reads comparison passing and is now `currentBaselineResult`, the image the next run will be compared against. Its own `screenshotCompareBaselineResult` keeps naming the old image, which is what it was measured against, not the baseline. Accepting does not move `dateUpdated`, so it does not invalidate a token you already hold.

NameTypeReqDescription
expectedResultIdstringyeslatestResult.id from gi_screenshot_status: the result whose screenshot you reviewed.
testIdstringyesThe 24-character test id.

No output schema declared.

No examples provided.

gi_create_suite ~236

Creates an empty suite, optionally inside a folder. The folder is honoured at creation, so no follow-up move is needed. Refuses when a suite of the same name already exists in the same place, because Ghost Inspector allows the duplicate and nothing distinguishes the two afterwards. Pass allowDuplicateName only when the repetition is genuinely intended. 🔴 Getting the name right matters more than usual: this server never exposes suite deletion, because DELETE /suites/{id}/ cascades to every test inside with no version history and no recycle bin. Folders have no delete route in the API at all. Anything created here is tidied up by hand, in the web UI. Creating adds and overwrites nothing, so no concurrency token applies.

NameTypeReqDescription
allowDuplicateNameboolean–Proceed even though a suite of this name already exists here.
folderstring–Folder id to create it in. Omit to leave it unfiled.
namestringyesSuite name. Must be unique where it is being created.
organizationstring–Organization id. Defaults to GHOST_INSPECTOR_ORG_ID.

No output schema declared.

No examples provided.

gi_duplicate_test ~404

Copies an existing test, then places it in a suite and renames it in one call. Returns the new test with its steps and its dateUpdated. 🔴 This is how a test comes into existence here, and it is not a create. Ghost Inspector has no endpoint that builds a test from nothing — POST /tests/ is the listing wearing a POST — so a source test is mandatory and there is no way around that. Pick the closest existing test and adapt the copy with gi_update_test. 🔴 The copy's schedule is cleared unless keepSchedule is set. Whether a copy inherits its source's testFrequency is not verified, and in an account whose tests submit live forms against production, an inherited schedule means an unattended run posting real data. The uncertain case is pinned to the safe direction. The copy carries the source's steps verbatim, including any `execute` steps: it imports the same modules, so editing those modules still affects it. If the copy is placed in a suite with a different viewport or browser, selectors that resolved for the source may not resolve for it — validate with gi_validate_test before trusting it. If the copy is made but placing or renaming it fails, the response says so and returns the id, because the copy is already real and needs cleaning up. A copy keeps the source's `dateCreated` to the millisecond, so `dateCreated` cannot date a copy or tell it apart from its source.

NameTypeReqDescription
keepScheduleboolean–Keep any inherited schedule. Off by default; leaving it off is the safe choice.
namestring–New name. Defaults to "<source> (Copy)".
sourceTestIdstringyesThe test to copy. Required — there is no create.
startUrlstring–Where the copy starts. Defaults to the source's startUrl.
suiteIdstring–Suite to place it in. Defaults to the source's suite.

No output schema declared.

No examples provided.

gi_failure_groups ~234

Read-only. For every red test (modules excluded), finds its onset, the first failure after its last green run, and groups onsets that follow each other within `windowHours`, across suites and folders. Many tests going red within hours usually share one cause: a deploy, a shared module, a page change. Each group lists its tests with ids and the most common errors and targets, with numbers and quoted text normalised. A red test with no green run within `maxRunsPerTest` goes to `onsetUnknown` rather than being guessed. Costs one request per 50 runs per red test; narrow with folder or suite on a large account. Staleness is not considered: use gi_stale_tests for that.

NameTypeReqDescription
folderstring–Folder id, or part of its name.
maxRunsPerTestinteger–How far back to look for a green run per test. Default 200.
suitestring–Suite id, or part of its name.
windowHoursnumber–Largest gap between consecutive onsets in one group. Default 12.

No output schema declared.

No examples provided.

gi_find_tests ~333

Read-only. Finds tests and returns their ids, which every other tool takes, with suite, folder, importOnly, passing and screenshotCompare. 🔴 `passing` is the functional result only. A test can be green there while its screenshot comparison fails: read `screenshotCompare.passing` for that, or pass `screenshotFailing: true` to list only those. Comparison settings a test leaves at null are inherited from its suite, and `screenshotCompare.enabled` resolves that. `name`, `folder` and `suite` are matched against the listing: case-insensitive substrings, or an exact id for folder and suite. They are cheap, three requests. `step` searches each remaining test's own steps (command exact, target across every fallback selector, value as a substring) and costs one request per test left after the other filters, so narrow first on a large account. A match inside a module is reported on the module, not on the tests that import it; gi_module_usage lists those.

NameTypeReqDescription
folderstring–Folder id, or part of its name.
limitinteger–Most results to return. Default 50; `total` is always the full count.
namestring–Part of the test's name.
screenshotFailingboolean–Only tests whose screenshot comparison is enabled and failing. Cheap: read from the listing.
stepobject–Match tests by what their own steps do. All given fields must match one step.
suitestring–Suite id, or part of its name.

No output schema declared.

No examples provided.

gi_get_test ~314

Returns a single test's stored definition, identity and current state: steps, startUrl, suite, whether it is a module, its last run, and its `dateUpdated`. 🔴 `dateUpdated` is the concurrency token. gi_update_test requires it as `expectedDateUpdated` and refuses the write if the record moved since you read it. Call this first and pass the value through. Do not discover the token by sending a deliberately wrong one and reading the correct value off the refusal — that defeats the guard, which exists to prove the edit was composed against the definition that is actually stored. The steps returned are the test's OWN steps. An `execute` step names an imported module in `value` and is not expanded here, so the definition you edit may be smaller than the run you observed: a result expands every module inline. If the step you need to fix came from a module, edit that module's test, not this one.

NameTypeReqDescription
expandModulesboolean–Also return `expanded`: every step a run would execute, modules inlined, each with ownerId, ownerName, indexInOwner (its position in its owner's own list), rootIndex and the combined condition. It li…
testIdstring–The 24-character test id.
testIdsarray–Up to 20 test ids instead of testId. Each comes back with its own result or error; one failure does not sink the rest.

No output schema declared.

No examples provided.

gi_inventory ~243

Read-only tour of the whole account: every folder, the suites inside it, and per-suite counts of passing / failing / module / not-yet-run tests, plus the names of the failing ones. Start here — no other question about this account can be answered without knowing what is in it. Import-only tests (modules: shared steps that other tests import, the equivalent of a function) are counted in their own bucket and never as failures. Marking a test import-only deletes its stored results, so every module looks permanently unrun; folding that into a failure count invents breakage that does not exist and aims cleanup at the steps the live tests all share. Read the `notes` field before drawing conclusions. Fetches roughly 440 KB from the API and returns a summary of it, so ask for the whole account rather than probing folder by folder. Totals always describe the entire account even when a filter narrows the listing.

NameTypeReqDescription
failingOnlyboolean–List only suites with at least one failing test. Totals stay account-wide.
folderstring–Case-insensitive substring of a folder name. Omit to see every folder.

No output schema declared.

No examples provided.

gi_module_usage ~393

Answers the one question the API cannot: if I edit this module, which tests break? Builds the reverse index of `execute` steps — for every imported test, its direct importers and its full transitive blast radius through nested chains. Run this BEFORE editing any module. Also surfaces three things that only appear once the index exists: import-only tests nobody imports (dead, or a test that lost its caller and is silently not running); imported tests NOT flagged import-only, which run standalone *and* inside their importers, so an edit changes both paths; and execute steps pointing at ids that no longer exist. 🔴 It also finds `vacuousTests`: tests that execute no steps at all, because their definition is only `execute` calls and the chain bottoms out in empty modules. Those pass — nothing can fail — so the dashboard shows them green while they assert nothing, which is worse than a red test and invisible any other way. Emptying one shared module does this to every test that imports it. This is the expensive tool. `steps` is absent from the test listing, so it costs one request per test in the account — a few seconds for a few hundred tests, at deliberately low concurrency because the rate limit is undisclosed. Call it once and work from the result rather than per module. A definition that cannot be read is counted in `scanned.unreadable`, never skipped silently, because a missing definition understates a blast radius.

NameTypeReqDescription
folderstring–Folder id, or part of its name. Lists only the modules its tests import; counts stay account-wide.
modulestring–Case-insensitive substring of a module name. Narrows the listing and names every importer instead of capping the list.
suitestring–Suite id, or part of its name. Lists only the modules its tests import; with folder, both apply.

No output schema declared.

No examples provided.

gi_move_suite ~213

Moves a suite, with all of its tests, into another folder. Reversible — unlike everything else on the write path — and the response carries `previousFolder` so the undo is one call. `expectedCurrentFolder` is required: state the folder you believe the suite is in, and the move is refused if it is somewhere else, which is the case where you are about to move the wrong suite. The result is verified by re-reading: the folder must have changed and `testCount` must be identical, since a move should never detach a test. Note that `DELETE /folders/{id}/` does not exist, so a folder left empty by a move can only be removed from the web UI. Folder names have to be right the first time.

NameTypeReqDescription
expectedCurrentFolderstringyesThe folder id you believe this suite is in right now. Refused if it is not.
folderIdstringyesDestination folder id.
suiteIdstringyesSuite to move.

No output schema declared.

No examples provided.

gi_move_test ~234

Moves one test into another suite, steps untouched. Reversible: the response carries `previousSuite`, so the undo is one call. This is also how to retire a test: move it to a suite that has no schedule. Deleting a test is not offered, because Ghost Inspector keeps no version history and no recycle bin. `expectedCurrentSuite` is required: the suite id you believe the test is in, from gi_get_test or gi_find_tests. The move is refused if it is elsewhere. 🔴 A test runs with its suite. If the destination is scheduled, the test starts running on that schedule, unattended; if it submits a form, each run is a real submission. The response says whether the destination is scheduled. Settings the test leaves at null (browser, viewport, screenshot comparison and threshold) and the suite's variables change with the move too.

NameTypeReqDescription
expectedCurrentSuitestringyesThe suite id you believe this test is in right now. Refused if it is not.
suiteIdstringyesDestination suite id.
testIdstringyesTest to move.

No output schema declared.

No examples provided.

gi_plan_test ~336

Read-only. Returns exactly what gi_validate_test would send: modules inlined, {{variables}} resolved from `variables`, the suite and the organization, and all three submit-guard layers applied. Nothing is executed, only definitions are read: no browser starts, and no organization id is needed. Use it before any validation of a test that touches production. `plan` lists the steps in order, numbered as every other report numbers them. `guard.stoppedAt` is the static cut, where a submit-shaped step became an assertion on its target. `guard.probedClicks` counts the clicks that will be probed in the browser before they run. `variables` shows what was resolved, what is left for a step that sets it at run time, and what has no value. `wouldRefuse` is non-null when gi_validate_test would refuse the run before sending anything, and says why.

NameTypeReqDescription
browserstring–Override, e.g. "chrome".
definitionobject–Ad-hoc definition to plan instead of an existing test.
stopBeforeinteger–Stop before this plan step. Can only cut earlier than the guard's own cut.
suiteIdstring–With `definition`: the suite whose configuration and variables it runs with.
testIdstring–Existing test to plan. Its suite's configuration and variables are replicated.
variablesobject–Variable values, e.g. {"subdomain": "www"}. Win over the suite's and the organization's.
viewportstring–Override, e.g. "1280x800".

No output schema declared.

No examples provided.

gi_propose_repair ~297

Turns a diagnosis into a concrete argument: what to change, in which test, and the token needed to write it. Read-only — it applies nothing and returns steps for you to validate first. 🔴 This server cannot see the page. It has the definition, the error and the Ghost Inspector contract, so proposals come in two kinds and are never blurred. `applicable` proposals carry a rewritten step and come from rules that hold whatever the page contains — an assertTextPresent with no target always fails, an eval without an explicit return is always undefined. `advisory` ones name a real problem that cannot be fixed without looking at the DOM, and deliberately stop there rather than inventing a selector. Refuses outright on a stale diagnosis. A failure that predates a change describes a definition that is no longer stored, so a repair built on it would overwrite whatever replaced it, with no version history to recover. 🔴 `editTarget` is frequently NOT the test you asked about. Results expand imported modules inline, so the failing step often belongs to a module — and `proposedSteps` is that module's full step list, not this test's. Editing a module affects every test importing it. Validate `proposedSteps` with gi_validate_test before writing. A proposal that has not been run is a hypothesis.

NameTypeReqDescription
testIdstringyesThe failing test. Its diagnosis drives the proposal.

No output schema declared.

No examples provided.

gi_run_test ~343

Executes a test exactly as saved and waits for the verdict. Use it to confirm a repair actually worked, or to get a fresh result when the stored one is stale. 🔴 Nothing is truncated here. Unlike gi_validate_test, which runs a throwaway copy and stops before anything can submit, this runs the real test against the real startUrl — in most accounts, production. If the test fills and submits a form, this posts a real record into whatever that form feeds, and nothing here can withdraw it. Because of that, a test that submits is refused unless `confirmSubmit` is set. The check inlines imported modules first, since a test whose steps are only `execute` calls hides its submit inside one. A test that submits nothing runs without the flag. If a chain cannot be fully expanded it counts as submitting — an unnecessary confirmation is cheaper than an unintended record. If you only need to know whether the selectors still resolve, this is the wrong tool: gi_validate_test answers that without submitting. 🔴 A wait that expires is NOT a failure. Browser runs take 20-70 seconds and slow ones take longer; the response returns the result id and says the run is still going. Read the outcome with gi_test_result rather than concluding the test failed.

NameTypeReqDescription
confirmSubmitboolean–Required only when the test contains a step that could submit a form. Setting it means you accept a real submission against a real environment.
testIdstringyesThe 24-character test id. Import-only tests cannot run.
timeoutMsinteger–How long to wait before handing back the result id. Default 240000.

No output schema declared.

No examples provided.

gi_screenshot_diff ~248

Read-only. Downloads a result's full-size screenshot and a baseline's, compares them pixel by pixel, and returns the changed regions as horizontal bands (rows and columns in the full-size image, largest first), followed by PNG crops of the largest ones: baseline above a red rule, this result below. Look at the crops before gi_accept_screenshot: accepting cannot be undone. `changedShare` is this tool's own pixel count, not Ghost Inspector's `difference`, which uses an unpublished method, so the two will not match. A page that grew or content that moved down shows up as one band covering everything below the move; the notes say so. Each screenshot is a few MB, so this costs two downloads.

NameTypeReqDescription
againststring–comparedAgainst (default): the image the result was measured against, which explains its difference. currentBaseline: what the next run will be compared with.
cropsinteger–How many regions to return as images. Default 3; 0 for the regions alone.
resultIdstring–The result to inspect. Default: the latest.
testIdstringyesThe 24-character test id.

No output schema declared.

No examples provided.

gi_screenshot_status ~295

Read-only. The test's screenshot-comparison settings and its latest result's comparison: enabled, passing, the measured difference against the threshold, the current screenshot (`screenshotUrl`) and the difference image (`diffUrl`). Two images are kept apart because they stop being the same one the moment a screenshot is accepted. `comparedAgainst` is the image the latest result was measured against. `currentBaseline` is the one the next run will be compared against: the newest result whose comparison passed. Accepting flips a result to passing, so right after an accept the latest result reads passing with a difference above its threshold (`acceptedManually: true`), it is itself the baseline, and `comparedAgainst` still names the old image. `test.enabled` and `test.threshold` are the settings in force: a test that stores null inherits both from its suite, and the threshold stored on such a test is not the one applied. To find which tests to look at, gi_find_tests with `screenshotFailing: true` and a suite. To see what changed, gi_screenshot_diff. `latestResult.id` is what gi_accept_screenshot takes as `expectedResultId`.

NameTypeReqDescription
testIdstring–The 24-character test id.
testIdsarray–Up to 20 test ids instead of testId. Each comes back with its own result or error.

No output schema declared.

No examples provided.

gi_stale_tests ~340

Call this BEFORE diagnosing or editing any red test. Splits failures into two piles by comparing the whole `execute` chain's `dateUpdated` against each test's last run. `staleFailures` are red tests whose definition or module chain changed AFTER the failing run. The failure describes a version that no longer exists — a colleague may already have fixed it and the test simply has not run again. Editing on top of one destroys their work, and Ghost Inspector keeps no version history of steps. One level deep is not enough here, because modules nest; the whole chain is walked. `genuineFailures` have had no change since the failing run, so the failure still describes the current definition. Start there, oldest first. 🔴 Re-running is not free advice: many Ghost Inspector suites submit real forms against production. Confirm what a test does before triggering it. Also reports the case nobody looks for: passing tests whose chain changed after their last run, whose green result describes the old definition and proves nothing about the current one. Import-only modules are excluded rather than evaluated, since they have no results to compare against. Costs one request per test, a few seconds for a few hundred tests.

NameTypeReqDescription
folderstring–Folder id, or part of its name. Narrows the report to that folder's tests.
includePassesboolean–List the passing-but-unverified tests too. Off by default because it is the long bucket; the count is always reported.
suitestring–Suite id, or part of its name. Narrows the report to that suite's tests; with folder, both apply.

No output schema declared.

No examples provided.

gi_test_history ~226

Read-only. Walks a test's results newest first, 50 per page, and returns each run's verdict and failing step (command, error, the selector that resolved), plus `lastPass` and `firstFail`, the oldest failure of the current red streak, which dates a regression. 🔴 Read `horizon` before concluding anything. Ghost Inspector purges old results. `exhausted: true` means retention ended there and nothing older exists. `exhausted: false` means more history exists than was asked for, and `streakMayContinue` says the red streak may have begun earlier: raise `runs`. A run with `passing: null` is in flight, never a failure. Modules are refused: import-only deletes their results. For why the latest run failed, gi_test_result maps the failing step back to its definition.

NameTypeReqDescription
runsinteger–How many runs to walk back. Default 50, at most 500; one request per 50.
testIdstringyesThe 24-character test id.

No output schema declared.

No examples provided.

gi_test_result ~587

The last run of one test: the step that failed, its error, and whether the result can be trusted at all. Read-only. This is the starting point for repairing a failure — gi_stale_tests tells you which tests to look at, this tells you what happened in one of them. 🔴 Read `verdict` and `staleness` BEFORE the error. A verdict of `stale` means the test or one of its imported modules changed after this run, so the failure describes a definition that is no longer stored. Diagnosing from it is diagnosing from nothing, and a colleague may already have fixed it. Re-run the test and read the fresh result instead. 🔴 `failingStep.resolvedTarget` is the selector that resolved, NOT what the test looks for. Ghost Inspector collapses an authored fallback array to the one it used, and sometimes normalises it so it matches nothing in the definition textually. `authoredTargets` is what was actually written. Reporting the resolved one as the intent is a real and easy misreading. 🔴 `failingStep.ownedBy` names the test that contributed the step. Results expand imported modules inline, so the failing step frequently belongs to a module rather than to the test you asked about — that module is what needs editing, and editing it affects every test that imports it. The step's position in the result is meaningless against the definition; use `ownedBy.sequenceInOwner`. `failingStep.mapping` says how that position was found. `position`: the current definition was expanded locally and lines up with the result step for step. `stored sequence`: it did not, and the result's own stored position was used because the run recorded a distinct one for every step of that owner. `unmapped`: neither held, so sequenceInOwner and authoredTargets are unknown — never guessed. A result's own `extra.source.sequence` is copied from the stored `sequence` field, which a client that omits it leaves at 0 on every step, so it is not trusted on its own. Other cases it distinguishes rather than blurring: a module…

NameTypeReqDescription
runsBackinteger–0 (default) is the latest run. Higher walks backwards, within the retained window.
testIdstring–The 24-character test id.
testIdsarray–Up to 20 test ids instead of testId. Each comes back with its own result or error; one failure does not sink the rest.

No output schema declared.

No examples provided.

gi_update_test ~632

Replaces a test's steps, renames it, or changes its startUrl. 🔴 Ghost Inspector keeps NO version history of steps and no recycle bin, so this is permanent. Four guards run on every call and none can be turned off. (1) The whole `execute` chain's `dateUpdated` is compared against the test's last run; if anything changed after it, the test is stale and the call is refused, because a fix diagnosed from a failure that describes a deleted version is diagnosed from nothing — and a colleague may already have fixed it. (2) The complete prior definition, credentials removed, is saved to `backupFile` (under GHOST_INSPECTOR_BACKUP_DIR, owner-only) on refusals too, with a `backupSummary` beside it. The file is the rollback; clients without filesystem access should pass verbose:true to get it inline as `backup`. If the file cannot be written it comes back inline anyway. (3) The change is applied. (4) The record is re-read and diffed, both that what was sent landed exactly and that every field you did not send is untouched — `HTTP 200` proves neither. `expectedDateUpdated` is required: state the `dateUpdated` you believe is current, and the write is refused if the record has moved since. Read it with gi_get_test first. Every response carries the record's current `dateUpdated`; after an applied write that is the token for the next edit, so a series of edits needs no re-read in between. After a refusal, re-read and recompose rather than resending. Each step's `sequence` is overwritten with its position. Results map a failure back to its step through that field, and a list stored without it maps every failure to step 0. Before overwriting a red test, prefer gi_validate_test, which runs the current definition without saving and without submitting a form. If the staleness guard trips and you have genuinely verified the current state, `confirmStaleDiagnosis` proceeds — read what the refusal says first.

NameTypeReqDescription
confirmStaleDiagnosisboolean–Proceed even though the chain changed after the last run. Only after verifying the current definition yourself; otherwise you may be overwriting someone else's fix.
expectedDateUpdatedstringyesThe dateUpdated you read from this test. Proof you have seen its current state; the write is refused if it no longer matches.
namestring–New name. Renaming does not move the test or break importers, which reference it by id.
startUrlstring–New start URL. Read back after the write like everything else sent. A module's startUrl is never visited: its importer decides where it starts.
stepsarray–Replacement step list, in order. Replaces the whole array — send every step you want kept.
testIdstringyesTest to change.
verboseboolean–Also return the prior definition inline as `backup`. For clients that cannot read `backupFile`.

No output schema declared.

No examples provided.

gi_vacuous_tests ~332

Finds tests that pass while verifying nothing. Green is the dangerous colour: a red test gets investigated, a hollow green one sits there while every report says coverage is fine. Three classes, kept apart because conflating them hides two of them. `executesNothing` runs zero steps — its definition is only `execute` calls into empty modules. `assertsNothing` runs its steps and contains no assertion anywhere in the chain, so it can only fail if a step errors; act on this one first, it is usually the largest. `worthChecking` is a SHORTLIST, not a verdict: their single assertion is the final step, so if its target also exists on the page the test starts from, it passes with the feature completely broken. 🔴 That third class cannot be settled from the definition — the selector is simply present on the starting page and no earlier step mentions it. Scanning for a repeated target finds none of them. To decide, run the test with the decisive action removed and see whether the assertion still passes; if it does, the test proves nothing. Modules are excluded before counting: import-only deletes results, so including them would condemn the shared layer every live test depends on. Costs one request per test; with folder or suite, one per test in scope plus the modules they import.

NameTypeReqDescription
folderstring–Folder id, or part of its name. Narrows the report to that folder's tests.
suitestring–Suite id, or part of its name. Narrows the report to that suite's tests; with folder, both apply.

No output schema declared.

No examples provided.

gi_validate_test ~993

Runs a test definition through on-demand execution, which executes it and discards it — nothing in the account is created or changed. Use it to check that a selector chain still resolves before editing a test, and to check a definition you are authoring before saving it. 🔴 It drives a real browser against a real URL, so it is an action with real-world effects even though nothing is saved. Modules are inlined first, because a test whose steps are just `execute` calls hides its submit inside a module. Then three guard layers apply, and none can be turned off or extended past its cut: (A) Static: the run is truncated at the first click on a submit-shaped target, Enter keypress, or eval/assertEval/extractEval script or step condition that could submit or send data (.submit(, requestSubmit(, .click(, dispatchEvent(, fetch(, XMLHttpRequest, sendBeacon(, $.ajax, $.post, axios). That step becomes an assertElementVisible on its target, so the chain is verified, including that the control is reachable. `stopBefore` can move this cut earlier, never later. (B) In the browser: before every remaining click, a wait on its target and a probe. The probe stops the run if the element is a form's submit button or input, a control inside a form that is not a field, or cannot be resolved from the top document; every step is gated on that stop. Accepted false positive: a type=submit "Continue" inside a form stops the run. (C) Tripwire: armed before every step on every page, click or not, it blocks submit events, form.submit(), non-GET fetch and XHR, and sendBeacon, and `guard.blockedRequests` lists them. It does not stop the run. Residual gaps: a script that saved window.fetch or form.submit before the page's first step ran, a WebSocket, anything inside a child frame, and data sent by a GET (a pixel or a navigation). Step numbers in `plan`, `steps` and `guard` count plan steps; the injected ones never shift them. There is no way to make this tool submit; that stays a deliberate curl.…

NameTypeReqDescription
browserstring–Override, e.g. "chrome". Omit to replicate the suite's.
definitionobject–Ad-hoc definition to validate instead of an existing test.
dryRunboolean–Deprecated: use gi_plan_test, which does exactly this and is read-only. Reports what would run and sends nothing.
stopBeforeinteger–Stop before this plan step (as numbered by gi_plan_test). Can only cut earlier than the guard's own cut; a later value is ignored.
suiteIdstring–With `definition`: the suite whose configuration and variables it runs with.
testIdstring–Existing test to validate. Its suite's configuration and variables are replicated.
variablesobject–Variable values, e.g. {"subdomain": "www"}. Win over the suite's and the organization's.
verboseboolean–Keep `plan` after a run and return every console entry. Off by default to keep the response small.
viewportstring–Override, e.g. "1280x800". Omit to use the test's, then the suite's.

No output schema declared.

No examples provided.

gi_whoami ~164

Confirms the configured API key works, lists the organizations it can reach, and reports whether writing is enabled. Read-only and safe to call first when diagnosing setup. 🔴 Call this before concluding that this server cannot modify anything. Every tool is registered whether or not its gate is open, so a tool being listed says nothing about whether it will run. `writesEnabled` and `runsEnabled` are the authority — GHOST_INSPECTOR_ALLOW_WRITES and GHOST_INSPECTOR_ALLOW_RUNS. When one is false the operator must set the matching variable and restart this server; it cannot be turned on from a tool call. Returns each organization's id — export the one you want as GHOST_INSPECTOR_ORG_ID to enable on-demand validation runs.

Input schema present but exposes no named parameters.

No output schema declared.

No examples provided.

Common questions

What is the io.github.charliemtnez/ghost-inspector-mcp server?

io.github.charliemtnez/ghost-inspector-mcp is listed in the public MCP registry as io.github.charliemtnez/ghost-inspector-mcp. Analyze, validate and safely update Ghost Inspector end-to-end browser tests. This page covers its npm package (ghost-inspector-mcp).

Is the io.github.charliemtnez/ghost-inspector-mcp server safe to use?

io.github.charliemtnez/ghost-inspector-mcp scores 81 out of 100 on VerifyMCP. We found no known CVEs affecting it as of 1 October 2026. It declares no install or post-install scripts. That is a record of what we were able to check automatically, not an endorsement. The category breakdown on this page shows every signal behind the number, including the ones we could not confirm.

What tools does the io.github.charliemtnez/ghost-inspector-mcp server expose?

io.github.charliemtnez/ghost-inspector-mcp exposes 22 tools: gi_whoami, gi_inventory, gi_module_usage, gi_stale_tests, gi_find_tests, and 17 more. Their descriptions and schemas cost roughly 7,676 tokens of context every time the server is loaded.

Is the io.github.charliemtnez/ghost-inspector-mcp server still maintained?

io.github.charliemtnez/ghost-inspector-mcp is still listed as active in the MCP registry. We last reached this channel on 1 October 2026. Those dates come from our own scans of the registry and the channel itself, not from anything the publisher announced.

What licence is the io.github.charliemtnez/ghost-inspector-mcp server under?

io.github.charliemtnez/ghost-inspector-mcp declares the MIT licence, which is OSI-approved. That covers the source only, and says nothing about the cost of any service it calls.