# NARA Catalog (pypi · nara-catalog-mcp)

Genealogical research in the US National Archives Catalog: records, transcriptions, page images.

- Trust score: 66/100 (medium)
- Registry status: active
- Liveness: live
- Owner verified: no
- Last scored: 2026-10-01

## Components

- pypi · `nara-catalog-mcp`: 66/100 (this document), [markdown](https://verifymcp.io/servers/ianderso-nara-catalog-mcp/nara-catalog-mcp.md), [page](https://verifymcp.io/servers/ianderso-nara-catalog-mcp/nara-catalog-mcp)

## Channel facts

- Registry: `pypi`
- Package: `nara-catalog-mcp`
- Version: `1.0.2`
- Transport: `stdio`

## Trust breakdown

How this component scores in each security and reliability category. Every signal is checked automatically from public evidence about the published package, including repeated runs of it in an isolated sandbox, and we only credit what we can confirm. Scores are 0–100 per category. Scoring method: https://verifymcp.io/docs/scoring (what has changed: https://verifymcp.io/docs/scoring/changelog)

Scored 2026-10-01.

- **Supply Chain Security**: 50/100
  - Malware scan not yet available for this package.
  - No known CVEs affecting this package version or its production dependencies.
  - Runs hatchling.build at install time, a recognised build step with no custom scripting around it.
  - 1 of 32 dependencies flagged as unhealthy.
- **Provenance & Transparency**: 100/100
  - Source repository is publicly reachable at the declared URL.
  - Cryptographically verified build provenance (signed, bound to ianderso/nara-catalog-mcp).
  - Clear OSI-approved license (MIT).
  - Actively maintained (last published 2 days ago).
  - Publishes a security disclosure policy (SECURITY.md).
- **Schema Quality & AI Usability**: 70/100
  - AI-judged instruction clarity (excellent).
  - Context-footprint check failed: tool/resource definitions use about 4250 tokens (~236/item across 18 items; 18 tools + 0 resources), over budget; trim descriptions and params.
  - Usage-examples check failed: none of the tools include examples.
- **Stability & Change Management**: 0/100
  - Stability not yet verified: not enough scan history yet (needs a 30-day window).
- **Tool Coverage**: 100/100
  - 100% of tools have a non-trivial description (not blank, and not just the tool's name).
  - 100% of tool parameters carry a description.
- **Tool Safety**: 100/100
  - No prompt-injection markers were found in the server instructions, tool names or descriptions we captured.
  - We read all 18 captured tool definition(s), and no name or description among them implies an irreversible operation.
  - An AI judge read all 19 captured unit(s) of tool text and found none that tries to manipulate the model reading it.
- **Capabilities**: 100/100
  - Implements a current MCP spec version (2026-07-28).

**Unverified: 1 category.** A category scored 0 because we could not verify it: a data source with nothing on this package, evidence we could not reach, or a check we could not run. We only credit what we can confirm.

## Install

### How do I install the NARA Catalog MCP server?

NARA Catalog runs locally as a PyPI package, launched with uvx nara-catalog-mcp. Ready-made configuration for Claude, Cursor, VS Code, Codex and 5 more is on this page, copied from each client's own documentation.

### Claude

```bash
claude mcp add ianderso-nara-catalog-mcp -- uvx nara-catalog-mcp
```

### Cursor

```json
{
  "mcpServers": {
    "ianderso-nara-catalog-mcp": {
      "command": "uvx",
      "args": [
        "nara-catalog-mcp"
      ]
    }
  }
}
```

### VS Code

```json
{
  "servers": {
    "ianderso-nara-catalog-mcp": {
      "command": "uvx",
      "args": [
        "nara-catalog-mcp"
      ]
    }
  }
}
```

### Codex

```bash
codex mcp add ianderso-nara-catalog-mcp -- uvx nara-catalog-mcp
```

### opencode

```json
{
  "$schema": "https://opencode.ai/config.json",
  "mcp": {
    "ianderso-nara-catalog-mcp": {
      "type": "local",
      "command": [
        "uvx",
        "nara-catalog-mcp"
      ],
      "enabled": true
    }
  }
}
```

### OpenClaw

```bash
openclaw mcp add ianderso-nara-catalog-mcp --command uvx --arg nara-catalog-mcp
```

### Hermes

```yaml
mcp_servers:
  ianderso-nara-catalog-mcp:
    command: "uvx"
    args: ["nara-catalog-mcp"]
```

### Netclaw

```json
{
  "McpServers": {
    "ianderso-nara-catalog-mcp": {
      "Transport": "stdio",
      "Command": "uvx",
      "Arguments": [
        "nara-catalog-mcp"
      ]
    }
  }
}
```

### Vellum

```bash
assistant mcp add ianderso-nara-catalog-mcp -t stdio -c uvx -a nara-catalog-mcp
```

### Other

```json
{
  "mcpServers": {
    "ianderso-nara-catalog-mcp": {
      "command": "uvx",
      "args": [
        "nara-catalog-mcp"
      ]
    }
  }
}
```

## Changelog

Every change recorded for this component, newest first. Days that predate change tracking, or that we cannot explain, say so: "we were watching and nothing happened" and "we were not watching" are different claims.

### 2026-09-28 (score 66)

First indexed and scored.

## MCP tools (18)

### `search_records` (~235 tokens)

Search the National Archives Catalog for records.

Start with `title` and a specific phrase; the Catalog holds tens of
millions of descriptions and a broad `query` will bury the useful hit. A
result is a lead: check the hierarchy and dates against what you already
know before reading the images.

Returns the total number of matches and a page of summaries, each with its
NAID, hierarchy, holding unit and image count. Use `search_records_advanced`
when you need dates, a record group, an M-number or digitised-only.

Input parameters:

- `limit` (integer): Maximum records to return (1-100).
- `page` (integer): Page of results, 1-based. The API pages rather than offsetting; beyond 10,000 results it needs cursor pagination.
- `query` (string): Full-text search across the description. Broader and noisier than title; use it when a title search finds nothing.
- `title` (string): Words to match in the record title, e.g. 'Hall pension' or a person's name. Titles of case files usually carry the name.

### `search_records_advanced` (~1139 tokens)

Search the Catalog with the filters that narrow a common name.

Every parameter is optional but at least one is required. The filters that
earn their keep for research are the date range, `microform_publication`
for an M-number you already cite, `record_group_number` or `ancestor_naid`
to stay inside one body of records, and `available_online` when you intend
to read pages rather than order copies.

Results are the same summaries `search_records` returns. Past 10,000 hits
use the `next_search_after` cursor rather than `page`; a common surname
passes that boundary easily.

Input parameters:

- `ancestor_naid` (string): NAID of an ancestor node: returns only records below it in the hierarchy. Use it to search inside one series.
- `available_online`: True for digitised records only -- those whose pages you can read now rather than order from a reading room.
- `collection_identifier` (string): Collection identifier, the Presidential-library equivalent of a record group.
- `comments_exist`: True for records carrying researcher comments.
- `congress_number`: Records of one numbered Congress, e.g. 55 for 1897-99. Private relief bills, petitions and claims naming individuals sit in the records of Congress.
- `contributions_exist`: True for records carrying any contribution at all: a transcription, tag or comment. Someone has already worked on them.
- `control_numbers` (string): Any identifier NARA attaches to a record: accession number, local identifier, microfilm publication, NAID, transfer number or variant control number. For a citation whose kind you cannot name.
- `creators` (string): The agency or person who created the records, matched against the creator headings.
- `data_source` (string): 'description' for archival descriptions, 'authority' for authority records (people, organisations, topics). This narrows a search and cannot be one on its own.
- `end_date` (string): Latest date to include, in the same format as start_date.
- `exact` (boolean): Match title, local_identifier and microform_publication in full and exactly, instead of by words within them. Use it when a words match returns too much, or you hold the complete title.
- `exact_date` (string): A single date, YYYY-MM-DD. Cannot be combined with start_date or end_date.
- `geographic_reference` (string): Place the records are about, matched against geographic subject headings, e.g. 'Franklin County (Pa.)'.
- `include_extracted_text` (boolean): Fold each hit's OCR text into the response, saving a call per hit: each record gains an extracted_text list with one entry per page that carries text, NARA's own OCR or a partner's, capped at 2000 ch…
- `level_of_description` (string): One of recordGroup, collection, series, fileUnit, item. A case file is usually a fileUnit; a single page is an item.
- `limit` (integer): Maximum records to return (1-100).
- `local_identifier` (string): The archives' own identifier for the record, as printed in finding aids.
- `microform_publication` (string): Microfilm publication number, e.g. 'M804' for Revolutionary War pension applications or 'T624' for the 1910 census. This is the citation genealogists actually carry.
- `page` (integer): Page of results, 1-based. Ignored when search_after is given, and unusable past 10,000 results.
- `person_or_org` (string): A person or organisation named in the description, either as its subject or in a role such as creator. Distinct from `creators`, which is the record's creating body only.
- `query` (string): Full-text search across the whole description.
- `record_group_number` (string): Record group number, e.g. '15' for Veterans Affairs. Scopes the search to one agency's records.
- `recurring_day` (string): Day as DD. Normally used with recurring_month.
- `recurring_month` (string): Month as MM. With recurring_day, finds records dated to that day in any year -- a birthday across every census.
- `reference_units` (string): Name of the archive holding the paper, e.g. 'National Archives at St. Louis'. Comma-separate several.
- `search_after` (string): Cursor for paging past 10,000 results. You MUST pass '*' for the first page, then the 'next_search_after' value from each response. Starting from an ordinary search does not work: without '*' the res…
- `start_date` (string): Earliest date to include, as YYYY, YYYY-MM or YYYY-MM-DD. Use the same precision as end_date. A surname search is usually only workable once it is bounded to a lifetime.
- `tags_exist`: True for records carrying citizen tags.
- `title` (string): Words to match in the record title.
- `transcriptions_exist`: True for records someone has transcribed; False for ones nobody has. Transcribed records are searchable by their text.
- `type_of_materials` (string): Material type, e.g. 'Textual Records', 'Photographs and other Graphic Materials', 'Maps and Charts', 'Moving Images'.

### `get_record` (~127 tokens)

Read one Catalog record in full, by NAID.

Returns the scope and content note, the full hierarchy, the holding
reference units, and every page-image URL. Use this once a search has given
you a NAID worth pursuing.

Input parameters:

- `naid` (string, required): The record's NAID, e.g. '54765873'.
- `refresh` (boolean): True re-reads from the Catalog instead of the cache, spending one call and replacing the cached copy. The cache never expires on its own, so use this when the answer may have changed since you last a…

### `get_record_images` (~157 tokens)

List a record's page images in page order, each with its object id.

These are the evidence. A catalog description summarises a file; it does
not tell you what any individual page says, so read the images before
citing anything to this record. The object id is what get_extracted_text
and get_transcriptions entries point at; pass it, or the page number, to
download_page_image.

Input parameters:

- `naid` (string, required): The record's NAID.
- `refresh` (boolean): True re-reads from the Catalog instead of the cache, spending one call and replacing the cached copy. The cache never expires on its own, so use this when the answer may have changed since you last a…

### `get_extracted_text` (~277 tokens)

Read the OCR text NARA machine-extracted from a record's page images.

This is a lead, not evidence. OCR was run over scans of handwriting,
carbon copies and microfilm: it drops handwritten pages entirely, and it
turns one surname into another silently. Use it to find which page matters,
then open that page image with `get_record_images` and cite what you saw
there -- never cite the OCR text itself.

Returns one entry per digital object, in page order, with the text
truncated to `max_chars`.

Input parameters:

- `limit` (integer): Maximum digital objects to return (1-100).
- `max_chars` (integer): Characters of text to return per page. 0 returns all of it, which for a long file unit is tens of thousands.
- `naid` (string, required): The record's NAID.
- `object_id` (string): Limit to one digital object (one scanned page). Omit it to get the text of every page of the record.
- `page` (integer): Page of results, 1-based.
- `refresh` (boolean): True re-reads from the Catalog instead of the cache, spending one call and replacing the cached copy. The cache never expires on its own, so use this when the answer may have changed since you last a…

### `get_transcriptions` (~209 tokens)

Read the citizen transcriptions of a record's pages.

Volunteers have transcribed handwritten pension files, service records and
letters that OCR cannot touch, which makes this the fastest way into a
document in copperplate. It is still a lead, not evidence: a transcription
is one stranger's reading, unreviewed, and names are exactly where such a
reading goes wrong. Check the page image before citing a name or date you
found here, and cite the image.

Returns one entry per transcription with its text, its author and the page
it belongs to.

Input parameters:

- `max_chars` (integer): Characters of each transcription to return. 0 returns the whole thing.
- `naid` (string, required): The record's NAID.
- `refresh` (boolean): True re-reads from the Catalog instead of the cache, spending one call and replacing the cached copy. The cache never expires on its own, so use this when the answer may have changed since you last a…

### `get_tags` (~144 tokens)

Read the citizen tags on a record.

On genealogical records tags are very often the names of the people who
appear inside the file -- the widow, the children, the witnesses -- which
the title does not carry. A tag is a stranger's reading of the document and
is not evidence: follow it to the page image and cite what the image shows.

Input parameters:

- `naid` (string, required): The record's NAID.
- `refresh` (boolean): True re-reads from the Catalog instead of the cache, spending one call and replacing the cached copy. The cache never expires on its own, so use this when the answer may have changed since you last a…

### `get_comments` (~146 tokens)

Read other researchers' comments on a record.

Comments often say what a file actually contains, where a related file
sits, or that the description is wrong. None of it is verified by NARA:
it is correspondence between researchers, useful for finding the next
record and never a source in itself.

Input parameters:

- `max_chars` (integer): Characters of each comment to return.
- `naid` (string, required): The record's NAID.
- `refresh` (boolean): True re-reads from the Catalog instead of the cache, spending one call and replacing the cached copy. The cache never expires on its own, so use this when the answer may have changed since you last a…

### `search_transcriptions` (~209 tokens)

Find records whose citizen transcriptions mention something.

This searches the documents rather than the catalogue, which is the
difference between finding a pension file titled with the veteran's name
and finding the file that names his widow, his children and the neighbours
who swore to the marriage. Only transcribed records are reachable this way,
so silence here means nobody has transcribed it, not that it does not exist.

Returns record summaries. A transcription is a stranger's reading and is
not evidence: open the matching page images to see what was written.

Input parameters:

- `contributor` (string): Restrict to one contributor's screen name.
- `limit` (integer): Maximum records to return (1-100).
- `naid` (string): Restrict to transcriptions of one record.
- `page` (integer): Page of results, 1-based.
- `query` (string): Words to find in transcribed text. Accepts AND, OR, NOT, wildcards and "exact phrases".

### `search_tags` (~166 tokens)

Find records that carry a given citizen tag.

Tags are short, so an exact `tag_is` on a surname is often sharper than a
title search: a volunteer who read the file tagged the people in it. What
comes back is what a stranger thought the document said -- a lead to the
page image, and not evidence.

Input parameters:

- `contributor` (string): Restrict to one contributor's screen name.
- `limit` (integer): Maximum records to return (1-100).
- `naid` (string): Restrict to tags on one record.
- `page` (integer): Page of results, 1-based.
- `query` (string): Words to find in citizen tags.
- `tag_is` (string): Match one exact tag rather than words within tags.

### `search_extracted_text` (~146 tokens)

Find records whose extracted text mentions something.

This reaches text contributed by NARA's digitisation partners over the
digital objects -- the searchable layer under the scans. It is machine
output and is not evidence: it misses handwriting, it mangles names, and a
hit means a page probably says this, not that it does. Read the page image
before citing anything you find here.

Input parameters:

- `limit` (integer): Maximum records to return (1-100).
- `page` (integer): Page of results, 1-based.
- `query` (string, required): Words to find in extracted text. Accepts AND, OR, NOT, wildcards and "exact phrases".

### `search_comments` (~130 tokens)

Find records other researchers have commented on.

Useful for picking up where someone else stopped: a comment naming a
surname often marks a file that a researcher has already read. Unverified
by NARA, so it points at records rather than settling anything.

Input parameters:

- `contributor` (string): Restrict to one contributor's screen name.
- `limit` (integer): Maximum records to return (1-100).
- `naid` (string): Restrict to comments on one record.
- `page` (integer): Page of results, 1-based.
- `query` (string): Words to find in comments.

### `browse_children` (~150 tokens)

List a record's immediate children, one level down the hierarchy.

The Catalog nests record group, then series, then file unit, then item.
Searching finds a node; this walks down from it, which is how you get from
a series you trust to the file unit for one person, and how you find out
what else sits alongside a file you already have. A record's own ancestors
come back from `get_record`.

Input parameters:

- `limit` (integer): Maximum children to return (1-100).
- `page` (integer): Page of results, 1-based.
- `parent_naid` (string, required): NAID of the parent: a record group, collection, series or file unit.

### `get_online_availability` (~99 tokens)

List where else this record is available online.

NARA records what has been digitised and published elsewhere, including by
commercial partners. Use it to resolve a hint from a subscription site back
to the archival original: the NAID and the reference unit are what make a
citation refindable, and a partner's index entry is not a substitute for
the page.

Input parameters:

- `naid` (string, required): The record's NAID.

### `get_partner_digital_objects` (~152 tokens)

List digital object ids a partner has matched to this record.

NARA indexes metadata supplied by commercial partners — Ancestry among
them — against its own records. A hit tells you the partner holds imagery
for this NAID, which is worth knowing when NARA's own pages are not
online.

An empty list is the common answer and is not an error. It means no
partner metadata has been matched, not that no partner holds the record.

The ids are pointers into the partner's index, not a citation. Cite the
archival record by NAID and reference unit.

Input parameters:

- `naid` (string, required): The record's NAID, e.g. '54765873'.

### `search_by_contribution_text` (~298 tokens)

Search what people wrote on records, and get the records back.

This is the difference between searching a catalogue and searching the
documents. A pension file titled only with the veteran's name will name
his widow, his children and his witnesses in its transcribed text — none
of which a title search reaches.

Unlike `search_transcriptions` and its siblings, which return the
contributions themselves, this filters the main index and returns full
record summaries. Use this when you want the record; use those when you
want to read what a particular volunteer wrote.

A transcription is one volunteer's reading and OCR is a machine's. Both
are leads, not evidence — open the page image before citing anything.

Input parameters:

- `available_online`: Restrict to digitised records only.
- `comment_text` (string): Words to find in researchers' comments on records.
- `extracted_text` (string): Words to find in OCR text, including text NARA's partners contributed.
- `limit` (integer): Maximum records (1-100).
- `page` (integer): Page of results, 1-based.
- `tag_text` (string): Words to find in citizen tags, which on genealogical records are usually the names of people appearing in them.
- `transcription_text` (string): Words to find in volunteers' transcriptions of the handwriting. Accepts AND, OR, NOT, wildcards and "exact phrases".

### `download_page_image` (~283 tokens)

Download one page of a record so it can actually be read.

This is the step that turns a catalogue hit into evidence. The image is
written to disk rather than returned inline — a page scan runs to several
megabytes.

NARA's media URLs are open and need no key, so this costs nothing
against your API allowance. One catalogue call resolves the page list;
the download itself is not an API call.

A pension file can run to sixty pages and the page you need is rarely the
first. Use get_extracted_text or get_transcriptions to find which page
carries the fact, then fetch it by the object_id they name.

Input parameters:

- `destination` (string, required): Absolute path of a new file to write, e.g. '/tmp/hall-w17050-p51.jpg'. The directory must already exist, and the file must not: nothing is ever overwritten.
- `naid` (string, required): The record's NAID, e.g. '54765873'.
- `object_id` (string): The digital object to download, as get_extracted_text, get_transcriptions and get_record_images name it. An alternative to page: pass one or the other.
- `page`: Which page to download, 1-based, in the order get_record_images lists them. Defaults to 1 when object_id is not given.

### `api_budget` (~66 tokens)

Report the key's Catalog API spend, this session and this month.

A key is capped per month and a long sweep can exhaust it. Live calls
are kept in a ledger beside the cache, so the month's figure survives
restarts. Cached repeats cost nothing and are not counted.

## Diagnostics

Captured diagnostic sections: Provenance, Install scripts, Dependencies. The full working is on the page: https://verifymcp.io/servers/ianderso-nara-catalog-mcp/nara-catalog-mcp#diagnostics

## Score history

- 2026-10-01: 66
- 2026-09-30: 66
- 2026-09-29: 66
- 2026-09-28: 66

## Common questions

### What is the NARA Catalog MCP server?

NARA Catalog is an MCP server listed in the public MCP registry as io.github.ianderso/nara-catalog-mcp. Genealogical research in the US National Archives Catalog: records, transcriptions, page images. This page covers its PyPI package (nara-catalog-mcp).

### Is the NARA Catalog MCP server safe to use?

NARA Catalog scores 66 out of 100 on VerifyMCP. We found no known CVEs affecting it as of 1 October 2026. Its build provenance is signed and verified. That is a record of what we were able to check automatically, not an endorsement. The category breakdown on this page shows every signal behind the number, including the ones we could not confirm.

### What tools does the NARA Catalog MCP server expose?

NARA Catalog exposes 18 tools: search_records, search_records_advanced, get_record, get_record_images, get_extracted_text, and 13 more. Their descriptions and schemas cost roughly 4,133 tokens of context every time the server is loaded.

### Is the NARA Catalog MCP server still maintained?

NARA Catalog is still listed as active in the MCP registry. We last reached this channel on 1 October 2026. Those dates come from our own scans of the registry and the channel itself, not from anything the publisher announced.

### What licence is the NARA Catalog MCP server under?

NARA Catalog declares the MIT licence, which is OSI-approved. That covers the source only, and says nothing about the cost of any service it calls.

## Links

- PyPI project: https://pypi.org/project/nara-catalog-mcp/
- Socket report: https://socket.dev/pypi/package/nara-catalog-mcp
- Repository: https://github.com/ianderso/nara-catalog-mcp
- Changelog RSS feed: https://verifymcp.io/servers/ianderso-nara-catalog-mcp/nara-catalog-mcp.xml
- Changelog JSON feed: https://verifymcp.io/servers/ianderso-nara-catalog-mcp/nara-catalog-mcp.json
- HTML version of this page: https://verifymcp.io/servers/ianderso-nara-catalog-mcp/nara-catalog-mcp
