Skip to content
verify mcp Beta VerifyMCP is currently in beta. If you notice any issues, email [email protected] and we’ll put it right.

GoldenMatch

REMOTE · GOLDENMATCH-MCP-PRODUCTION.UP.RAILWAY.APP · 2 COMPONENTS · SCANNED AUG 3

Find duplicate records in 30 seconds. Zero-config entity resolution, 97.2% F1 out of the box.

+4 this week 66 Trust /100
Trust breakdown (6 categories)

How this component scores in each security and reliability category. Every signal is checked automatically against the live server, and we only credit what we can confirm. How we score →

Endpoint Security57
Transport & Reachability100
Schema Quality & AI Usability76
  • 100% of prompts and resources have a non-trivial description (not blank, and not just the item's name).Pass
  • AI-judged instruction clarity (good).Pass
  • Tool/resource definitions use about 7291 tokens (~94/item across 77 items; 77 tools + 0 resources), lean.Pass
  • Usage-examples check failed: none of the tools include examples. See how to fix → Fail
Stability & Change Management27
  • Stability observed for 8 of 30 days with no destabilising changes; credit accrues until the full window elapses.Partial
Tool Coverage86
  • 100% of tools have a non-trivial description (not blank, and not just the tool's name).Pass
  • 59% of tool parameters carry a description.Partial
Capabilities100
  • Implements a supported MCP spec version (2025-11-25); the latest is 2026-07-28.Pass
Install

Add this component to your MCP client. Where a client-specific snippet is available, pick your client below and copy it straight into your config; otherwise use the connection detail shown.

remote · goldenmatch-mcp-production.up.railway.app

# add to Claude Code
claude mcp add --transport http benseverndev-oss-goldenmatch https://goldenmatch-mcp-production.up.railway.app/mcp/
# ~/.codex/config.toml
[mcp_servers.benseverndev-oss-goldenmatch]
url = "https://goldenmatch-mcp-production.up.railway.app/mcp/"
// opencode.json
{
  "$schema": "https://opencode.ai/config.json",
  "mcp": {
    "benseverndev-oss-goldenmatch": {
      "type": "remote",
      "url": "https://goldenmatch-mcp-production.up.railway.app/mcp/",
      "enabled": true
    }
  }
}
# add to OpenClaw
openclaw mcp add benseverndev-oss-goldenmatch --url https://goldenmatch-mcp-production.up.railway.app/mcp/ --transport streamable-http
# ~/.hermes/config.yaml
mcp_servers:
  benseverndev-oss-goldenmatch:
    url: "https://goldenmatch-mcp-production.up.railway.app/mcp/"
// mcp.json
{
  "mcpServers": {
    "benseverndev-oss-goldenmatch": {
      "type": "http",
      "url": "https://goldenmatch-mcp-production.up.railway.app/mcp/"
    }
  }
}

The mcpServers block is a cross-client convention. Remote transports vary, so check your client's docs.

Changelog

Every change we have recorded for this component, newest first. Security-relevant changes are always shown. ▲ marks a change for the better, ▼ a change for the worse; unmarked changes are neutral.

  • 3 Aug 26 +1

    No change was recorded against any check on this day. Stability & Change Management went from 23 to 27. That category is still filling its 30-day observation window: 7 days of observed history at the previous scan, 8 at this one. The score rises as the window fills, whether or not the server changes.

  • 1 Aug 26 +1

    No change was recorded against any check on this day. Stability & Change Management went from 17 to 20. That category is still filling its 30-day observation window: 5 days of observed history at the previous scan, 6 at this one. The score rises as the window fills, whether or not the server changes.

  • 31 Jul 26 0
    • We updated how we score, so this day's move reflects our rubric, not a change to the server See what changed → functional
  • 30 Jul 26 +1
    • We updated how we score, so this day's move reflects our rubric, not a change to the server See what changed → functional
  • 28 Jul 26 +1

    No change was recorded against any check on this day. Stability & Change Management went from 3 to 7. That category is still filling its 30-day observation window: 1 days of observed history at the previous scan, 2 at this one. The score rises as the window fills, whether or not the server changes.

  • 27 Jul 26 +1
    • We updated how we score, so this day's move reflects our rubric, not a change to the server See what changed → functional
  • 26 Jul 26 61

    First indexed and scored.

Diagnostics

Diagnostic detail from the automated scan of this channel: what the scanner observed at each step, so you can see exactly where a check passed or failed. It is informational only and never changes the trust score.

Captured 3 Aug 2026 · Probed https://goldenmatch-mcp-production.up.railway.app/mcp/

TLS valid

Negotiated TLS 1.3 with TLS_AES_128_GCM_SHA256 .

Subject Issuer Valid from Valid until Key Signature Serial
CN=*.up.railway.app CN=YE1,O=Let's Encrypt,C=US 29 Jul 2026 27 Oct 2026 ECDSA 256 ECDSA-SHA384 6da79bb561da3efeb0e751ca21abd3999fe
SANs: *.up.railway.app, up.railway.app
CN=YE1,O=Let's Encrypt,C=US (CA) CN=Root YE,O=ISRG,C=US 3 Sept 2025 2 Sept 2028 ECDSA 384 ECDSA-SHA384 5ddd70dd31f801c85c186a7a04b80afe
CN=Root YE,O=ISRG,C=US (CA) CN=ISRG Root X2,O=Internet Security Research Group,C=US 13 May 2026 2 Sept 2032 ECDSA 384 ECDSA-SHA384 872165fc34b6e5fba8add5b3705fb53a
CN=ISRG Root X2,O=Internet Security Research Group,C=US (CA) CN=ISRG Root X1,O=Internet Security Research Group,C=US 13 May 2026 2 Sept 2032 ECDSA 384 SHA256-RSA 6c8f1dc727c7117f7baf853ac980f9cd
DNSSEC insecure

Validation of goldenmatch-mcp-production.up.railway.app. Not signed

Zone DS Keys Algorithms Outcome
. trust_anchor 20326, 38696 8, 8 Verified
app. present 23684 8 Verified
railway.app. absent Unsigned (proven) parent-signed NSEC/NSEC3 proves an unsigned delegation
Authentication No authorisation required

The endpoint answered without asking for a token. Anyone who knows the URL can reach it.

Result No authorisation required
HTTP status 200
Transports 2 probes
Transport URL Outcome Status Location
streamable-http https://goldenmatch-mcp-production.up.railway.app/mcp/ Verified 200
http (plaintext) http://goldenmatch-mcp-production.up.railway.app/mcp/ HTTPS enforced 301 https://goldenmatch-mcp-production.up.railway.app/mcp/
MCP tools — 77 exposed · ~7,291 tokens

The tools this component advertises to a client, with an estimated token cost for each. Expand a tool to see its parameters and schema. The per-tool counts are indicative and are not scored directly; the schema's total context footprint is one signal in Schema Quality & AI Usability.

Tool Tokens
add_correction ~276

Add a Learning Memory correction. Two shapes: - pair-level: decision='approve' or 'reject', requires id_a + id_b - field-level (v1.18.2+): decision='field_correct', requires cluster_id + field_name + corrected_value Source is 'agent' with trust=0.5 (lower than human steward 1.0). Pair (id_a, id_b) is canonicalized to (min, max) before storage.

NameTypeReqDescription
cluster_idintegerField-level: cluster_id the correction targets.
corrected_valuestringField-level: the value the reviewer changed it to.
datasetstringyesDataset identifier (e.g. file path). Required, non-empty.
decisionstringyes
field_namestringField-level: the column being corrected.
id_aintegerPair-level: first row id. Field-level: ignored.
id_bintegerPair-level: second row id. Field-level: ignored.
matchkey_namestring
original_valuestringField-level: the value build_golden_record chose.
pathstringSQLite memory DB path. Default: .goldenmatch/memory.db
reasonstring

No output schema declared.

No examples provided.

agent_approve_reject ~66

Approve or reject a review queue pair

NameTypeReqDescription
decided_bystringyes
decisionstringyes
id_aintegeryes
id_bintegeryes
job_namestringyes
reasonstring

No output schema declared.

No examples provided.

agent_compare_strategies ~115

Compare ER strategies on your data

NameTypeReqDescription
encodingstringEncoding of *_content (default base64)
file_contentstringAlternative to file_path: file bytes (base64 default, or raw with encoding='text')
file_pathstring
filenamestringOriginal filename when using file_content
ground_truthstring
ground_truth_contentstringAlternative to ground_truth: base64/text bytes
ground_truth_namestring

No output schema declared.

No examples provided.

agent_deduplicate ~137

Run full ER pipeline with confidence gating and reasoning

NameTypeReqDescription
configobject
encodingstringEncoding of *_content (default base64)
exclude_columnsarrayColumn names to skip across GoldenMatch + GoldenFlow + auto-config. Optional. Layered with config.exclude_columns when both are set. force_include (env var) rescues from any opt-out path.
file_contentstringAlternative to file_path: file bytes (base64 default, or raw with encoding='text')
file_pathstring
filenamestringOriginal filename when using file_content

No output schema declared.

No examples provided.

agent_explain_cluster ~27

Explain why records are in the same cluster

NameTypeReqDescription
cluster_idintegeryes

No output schema declared.

No examples provided.

agent_explain_pair ~48

Natural language explanation for a record pair

NameTypeReqDescription
exactarray
fuzzyobject
record_aobjectyes
record_bobjectyes

No output schema declared.

No examples provided.

agent_match_sources ~157

Match two files with intelligent strategy selection

NameTypeReqDescription
configobject
encodingstringEncoding of *_content (default base64)
exclude_columnsarrayColumn names to skip across GoldenMatch + GoldenFlow + auto-config. Optional. Layered with config.exclude_columns when both are set. force_include (env var) rescues from any opt-out path.
file_astring
file_a_contentstringAlternative to file_a: base64/text bytes
file_a_namestring
file_bstring
file_b_contentstringAlternative to file_b: base64/text bytes
file_b_namestring

No output schema declared.

No examples provided.

agent_review_queue ~23

Get borderline pairs awaiting approval

NameTypeReqDescription
job_namestringyes

No output schema declared.

No examples provided.

analyze_blocking ~79

Diagnose blocking on the loaded dataset: returns ranked blocking key candidates with block counts, max block size, total candidate comparisons, and estimated recall. Use it to explain why matching is slow or produces too many candidate pairs.

NameTypeReqDescription
limitintegerTop N suggestions
sample_sizeinteger
target_block_sizeinteger

No output schema declared.

No examples provided.

analyze_data ~80

Profile data, detect domain, recommend ER strategy

NameTypeReqDescription
encodingstringEncoding of *_content (default base64)
file_contentstringAlternative to file_path: file bytes (base64 default, or raw with encoding='text')
file_pathstring
filenamestringOriginal filename when using file_content

No output schema declared.

No examples provided.

auto_configure ~180

Run AutoConfigController on a CSV; return the committed GoldenMatchConfig (incl. negative_evidence / Path Y when chosen) plus telemetry — stop_reason, health, decision trace, indicator column priors. Programmatic equivalent of `goldenmatch autoconfig`.

NameTypeReqDescription
constraintsobject
encodingstringEncoding of *_content (default base64)
exclude_columnsarrayColumn names to skip across GoldenMatch + GoldenFlow + auto-config. Optional. Layered with config.exclude_columns when both are set. force_include (env var) rescues from any opt-out path.
file_contentstringAlternative to file_path: file bytes (base64 default, or raw with encoding='text')
file_pathstring
filenamestringOriginal filename when using file_content

No output schema declared.

No examples provided.

certify_recall ~156

Estimate match RECALL without ground truth (unsupervised). Treats each auto-configured matchkey/pass as a decorrelated system and uses capture-recapture over their overlaps to estimate how many true matches were missed. Returns a point estimate (a safe lower bound additionally needs a small labelled audit; see `goldenmatch evaluate --certify --audit-out`). Needs >=3 decorrelated systems.

NameTypeReqDescription
encodingstringEncoding of *_content (default base64)
file_contentstringAlternative to file_path: file bytes (base64 default, or raw with encoding='text')
file_pathstringDataset to dedupe + certify
filenamestringOriginal filename when using file_content

No output schema declared.

No examples provided.

compare_clusters ~171

Compare two ER clustering outcomes on the same dataset without ground truth (CCMS): classifies each cluster as unchanged / merged / partitioned / overlapping and returns the Talburt-Wang Index. Both inputs are JSON cluster files (as written by export-style output).

NameTypeReqDescription
clusters_a_contentstringAlternative to clusters_a_path: base64/text bytes (JSON, use encoding='text')
clusters_a_namestring
clusters_a_pathstringBaseline clusters JSON
clusters_b_contentstringAlternative to clusters_b_path: base64/text bytes (JSON, use encoding='text')
clusters_b_namestring
clusters_b_pathstringComparison clusters JSON
encodingstringEncoding of *_content (default base64)

No output schema declared.

No examples provided.

config_weaknesses ~117

Diagnose weaknesses in the loaded run's auto-config: columns admitted that shouldn't be (source/provenance labels, per-row IDs), oversized or shared-value blocks, null sinks, low-signal matchkeys, and over-merging. Returns ranked findings, each with a plain-English explanation + a concrete fix, plus a one-paragraph summary.

NameTypeReqDescription
max_findingsintegerMax findings to return, ranked by severity (default 6).
phrasingstringWording style for the findings (default plain).

No output schema declared.

No examples provided.

controller_telemetry ~54

Return the AutoConfigController telemetry from the most recent `auto_configure` or `agent_deduplicate` call in this MCP session. Same JSON shape as the web /api/v1/controller/telemetry endpoint.

Input schema present but exposes no named parameters.

No output schema declared.

No examples provided.

create_domain ~227

Create a custom domain extraction rulebook. Define patterns for a specific data domain (medical devices, automotive parts, real estate, etc.).

NameTypeReqDescription
attribute_patternsobjectNamed regex patterns for domain attributes (e.g. {'size': '\\b(\\d+mm)\\b'})
brand_patternsarrayBrand/manufacturer names to extract (e.g. ['Medtronic', 'Abbott'])
identifier_patternsobjectNamed regex patterns for domain identifiers (e.g. {'ndc': '\\b(\\d{5}-\\d{4}-\\d{2})\\b'})
namestringyesDomain name (e.g. 'medical_devices', 'automotive_parts')
scopestringSave locally (.goldenmatch/domains/) or globally (~/.goldenmatch/domains/). Default: local.
signalsarrayyesColumn name keywords that trigger this domain (e.g. ['ndc', 'fda', 'implant'])
stop_wordsarrayWords to strip during name normalization

No output schema declared.

No examples provided.

dedupe ~75

Alias for `find_duplicates`. Find duplicate matches for a record. Provide field values to search against the loaded dataset.

NameTypeReqDescription
recordobjectyesRecord fields to match (e.g. {"name": "John Smith", "zip": "10001"})
top_kintegerMax results to return (default 5)

No output schema declared.

No examples provided.

documents_ingest ~94

Extract records from documents (PDF/image) against a target schema into rows ready for dedupe_df. Returns records + an ingest report.

NameTypeReqDescription
backendstring
drop_emptyboolean
modelstring
out_pathstringoptional CSV/parquet to also write
pathsarrayyes
schemaobjectyesschema JSON: {'fields':[...]}

No output schema declared.

No examples provided.

documents_suggest_schema ~49

Propose a target extraction schema (JSON) from a sample document image/PDF.

NameTypeReqDescription
backendstring
modelstring
sample_pathstringyes

No output schema declared.

No examples provided.

evaluate ~85

Score the loaded run against ground-truth pairs. Loads a ground-truth CSV (id_a,id_b columns) and returns precision, recall, and F1 for the current clustering.

NameTypeReqDescription
col_astring
col_bstring
ground_truth_pathstringyesCSV of true match pairs (columns id_a,id_b or idA,idB).

No output schema declared.

No examples provided.

explain_cluster ~33

Alias for `agent_explain_cluster`. Explain why records are in the same cluster

NameTypeReqDescription
cluster_idintegeryes

No output schema declared.

No examples provided.

explain_match ~45

Explain why two records match or don't match. Shows per-field score breakdown.

NameTypeReqDescription
record_aobjectyesFirst record fields
record_bobjectyesSecond record fields

No output schema declared.

No examples provided.

explain_pair ~52

Alias for `explain_match`. Explain why two records match or don't match. Shows per-field score breakdown.

NameTypeReqDescription
record_aobjectyesFirst record fields
record_bobjectyesSecond record fields

No output schema declared.

No examples provided.

explain_routing ~67

Human-readable explanation of why each stage is routed the way it is, with the driver-RAM projection that drove it.

NameTypeReqDescription
clusterobject
driver_mem_gbnumber
estimated_pair_countintegeryes
n_rowsintegeryes

No output schema declared.

No examples provided.

export_results ~44

Export matching results to a file (CSV or JSON).

NameTypeReqDescription
formatstringOutput format (default csv)
output_pathstringyesFile path to save results

No output schema declared.

No examples provided.

find_duplicates ~69

Find duplicate matches for a record. Provide field values to search against the loaded dataset.

NameTypeReqDescription
recordobjectyesRecord fields to match (e.g. {"name": "John Smith", "zip": "10001"})
top_kintegerMax results to return (default 5)

No output schema declared.

No examples provided.

fix_quality ~177

Run GoldenCheck scan and apply fixes to a CSV file. Returns the fixed data summary and a manifest of all fixes applied. Requires goldencheck: pip install goldenmatch[quality]

NameTypeReqDescription
domainstringOptional domain hint (healthcare, finance, ecommerce)
encodingstringEncoding of *_content (default base64)
file_contentstringAlternative to file_path: file bytes (base64 default, or raw with encoding='text')
file_pathstringPath to the CSV file to fix
filenamestringOriginal filename when using file_content
fix_modestringFix aggressiveness: safe (conservative) or moderate (balanced). Default: safe
output_pathstringOptional path to save the fixed CSV. If omitted, returns summary only.

No output schema declared.

No examples provided.

get_cluster ~36

Get details of a specific cluster: all member records and their field values.

NameTypeReqDescription
cluster_idintegeryesCluster ID to look up

No output schema declared.

No examples provided.

get_golden_record ~32

Get the merged golden (canonical) record for a cluster.

NameTypeReqDescription
cluster_idintegeryesCluster ID

No output schema declared.

No examples provided.

get_stats ~24

Get dataset statistics: record count, cluster count, match rate, cluster sizes.

Input schema present but exposes no named parameters.

No output schema declared.

No examples provided.

identity_audit ~83

Export the append-only identity audit log in commit order: every event with actor / trust / timestamp / reason, so a reviewer can reconstruct exactly which actor changed what, when, and why. Optionally filtered by dataset / actor.

NameTypeReqDescription
actorstring
datasetstring
limitinteger
pathstring

No output schema declared.

No examples provided.

identity_audit_seal ~133

Anchor the append-only audit log with a tamper-evidence seal: a chained sha256 root over every event since the last seal. Cheap and idempotent (a no-op when nothing new has been logged). Run it periodically (or after a batch of stewardship actions) so the history becomes provably untampered. Optionally scoped to a dataset. Publish/mirror the returned root_hash to make tampering detectable by an external party.

NameTypeReqDescription
actorstringPrincipal sealing the log. Defaults to 'agent'.
datasetstring
pathstringIdentity DB path

No output schema declared.

No examples provided.

identity_audit_verify ~99

Verify the append-only audit log against its seal chain. Replays the per-event content hashes and the seal roots to detect content edits, deletion, reordering, and insertion of any sealed event. Returns {ok, events_checked, seals_checked} plus the ids of any content mismatches / broken seals / missing sealed events. Optionally scoped to a dataset.

NameTypeReqDescription
datasetstring
pathstringIdentity DB path

No output schema declared.

No examples provided.

identity_claim ~133

Claim a record into an identity, moving it out of any prior entity ('this record belongs to that identity'). Emits a provenance-stamped `claimed` event on both the gaining and losing entities.

NameTypeReqDescription
actorstringPrincipal, e.g. 'agent:claude'. Defaults to 'agent'.
entity_idstringyesEntity to claim the record into
pathstring
reasonstring
record_idstringyesrecord id in `{source}:{source_pk}` form
trustnumberTrust in [0,1]. Default by actor prefix.

No output schema declared.

No examples provided.

identity_conflicts ~32

List evidence edges marked `conflicts_with`.

NameTypeReqDescription
datasetstring
pathstring

No output schema declared.

No examples provided.

identity_history ~39

Return the temporal event log for an identity.

NameTypeReqDescription
entity_idstringyes
limitinteger
pathstring

No output schema declared.

No examples provided.

identity_list ~52

List identities, optionally filtered by dataset/status.

NameTypeReqDescription
datasetstring
limitinteger
offsetinteger
pathstring
statusstring

No output schema declared.

No examples provided.

identity_merge ~159

Manually merge two identities. All records from `absorb_entity_id` are reassigned to `keep_entity_id`. The merge events are stamped with `actor`/`trust` provenance so the audit log records who merged these and on what authority.

NameTypeReqDescription
absorb_entity_idstringyes
actorstringPrincipal making the change, e.g. 'agent:claude' or 'steward:alice'. Defaults to 'agent'.
keep_entity_idstringyes
pathstring
reasonstring
trustnumberTrust of the actor in [0,1]. Defaults by actor prefix (steward 1.0, agent 0.5).

No output schema declared.

No examples provided.

identity_profile ~74

MDM profile of one entity: record count + per-source breakdown, golden record, confidence, conflict count, canonical version (structural-event count), and first/last activity. Returns {found: false} when no such entity exists.

NameTypeReqDescription
entity_idstringyes
pathstringIdentity DB path

No output schema declared.

No examples provided.

identity_resolve ~70

Resolve a record_id to its durable identity. Returns the full identity view (members, evidence edges, recent events) or null when no identity exists for that record.

NameTypeReqDescription
pathstringIdentity DB path
record_idstringyesrecord id in `{source}:{source_pk}` form

No output schema declared.

No examples provided.

identity_resolve_conflict ~186

Adjudicate a `conflicts_with` pair: 'same' keeps the entity intact, 'distinct' splits the second record out into a new identity, 'defer' only logs. Records a durable mediation verdict + event with actor/trust provenance, and stops the conflict re-surfacing in the open-conflicts queue.

NameTypeReqDescription
actorstringPrincipal, e.g. 'steward:alice'. Defaults to 'agent'.
applybooleanAct on the verdict (split on 'distinct'); false = log only.
datasetstring
pathstring
reasonstring
record_a_idstringyes
record_b_idstringyes
resolutionstringyes
trustnumberTrust in [0,1]. Default by actor prefix.

No output schema declared.

No examples provided.

identity_show ~69

Fetch the full detail of one identity by entity_id: its member records, evidence edges, and recent event log. Returns {found: false} when no such entity exists.

NameTypeReqDescription
entity_idstringyes
event_limitinteger
pathstringIdentity DB path

No output schema declared.

No examples provided.

identity_split ~118

Split a subset of records off an identity into a brand-new identity. The original keeps the remaining records. The split events carry `actor`/`trust` provenance.

NameTypeReqDescription
actorstringPrincipal making the change, e.g. 'agent:claude'. Defaults to 'agent'.
entity_idstringyes
pathstring
reasonstring
record_idsarrayyes
trustnumberTrust of the actor in [0,1]. Default by actor prefix.

No output schema declared.

No examples provided.

identity_stats ~63

Graph-level summary / health stats: entities by status, total records, records-per-entity distribution, conflict total, source mix, and the largest entities. Optionally scoped to a dataset.

NameTypeReqDescription
datasetstring
pathstringIdentity DB path

No output schema declared.

No examples provided.

identity_worklist ~68

Prioritized steward worklist: active entities needing attention (open conflicts and/or confidence below weak_confidence), highest conflict count first.

NameTypeReqDescription
datasetstring
limitinteger
pathstringIdentity DB path
weak_confidencenumber

No output schema declared.

No examples provided.

incremental ~172

Match a batch of new records against an existing base dataset (without re-running the whole base). Returns matched (new_row_id, base_row_id, score) pairs plus counts. Auto-configures from the base file if no config is given.

NameTypeReqDescription
base_filestringExisting base dataset path
base_file_contentstringAlternative to base_file: base64/text bytes
base_file_namestring
configstringOptional config YAML path
encodingstringEncoding of *_content (default base64)
new_recordsstringNew records file to match in
new_records_contentstringAlternative to new_records: base64/text bytes
new_records_namestring
thresholdnumberOptional threshold override

No output schema declared.

No examples provided.

learn_thresholds ~97

Force a MemoryLearner pass over accumulated corrections. Returns the list of LearnedAdjustments produced (matchkey_name, threshold, sample_size, learned_at). Requires >= 10 corrections per matchkey before threshold tuning fires; otherwise returns an empty list.

NameTypeReqDescription
matchkey_namestringOptional: learn only for this matchkey.
pathstringSQLite memory DB path. Default: .goldenmatch/memory.db

No output schema declared.

No examples provided.

lineage ~82

Field-level provenance for the loaded run: for each scored pair, the per-field scores that produced the match, plus cluster id. Optionally write a lineage JSON to a directory.

NameTypeReqDescription
max_pairsinteger
natural_languageboolean
output_dirstringIf set, write lineage JSON here and return the path instead of inline records.

No output schema declared.

No examples provided.

lint_routing ~89

Flag config/env overrides that force a slow path (e.g. CLUSTERING_THRESHOLD=0 when the edge set fits driver RAM). ERROR at scale; would_refuse mirrors the runtime guard.

NameTypeReqDescription
clusterobject
driver_mem_gbnumber
envobject
estimated_pair_countintegeryes
n_rowsintegeryes

No output schema declared.

No examples provided.

list_clusters ~58

List duplicate clusters found in the dataset. Returns cluster IDs, sizes, and member counts.

NameTypeReqDescription
limitintegerMax clusters to return (default 20)
min_sizeintegerMinimum cluster size to include (default 2)

No output schema declared.

No examples provided.