# com.yaktool/yaktool (npm · yaktool-mcp)

73 local deterministic tools: GTIN, feeds, CSV-XLSX, Digital Link QR, e-invoices. Rules cited.

- Trust score: 73/100 (medium)
- Change this week: +4
- Registry status: active
- Liveness: live
- Owner verified: no
- Last scored: 2026-09-20

## Components

- npm · `yaktool-mcp`: 73/100 (this document), [markdown](https://verifymcp.io/servers/com-yaktool-yaktool/yaktool-mcp.md), [page](https://verifymcp.io/servers/com-yaktool-yaktool/yaktool-mcp)

## Channel facts

- Registry: `npm`
- Package: `yaktool-mcp`
- Version: `0.4.2`
- Transport: `stdio`

## Trust breakdown

How this component scores in each security and reliability category. Every signal is checked automatically from public evidence about the published package, including repeated runs of it in an isolated sandbox, and we only credit what we can confirm. Scores are 0–100 per category. Scoring method: https://verifymcp.io/docs/scoring (what has changed: https://verifymcp.io/docs/scoring/changelog)

Scored 2026-09-20.

- **Supply Chain Security**: 100/100
  - No malware found by supply-chain analysis.
  - No known CVEs affecting this package version or its production dependencies.
  - No install/post-install scripts declared.
  - No production dependencies, so there is no dependency health to assess.
- **Provenance & Transparency**: 19/100
  - Repository check failed: no source repository is declared.
  - Provenance check failed: no build-provenance attestation is published.
  - Clear OSI-approved license (MIT).
  - Actively maintained (last published 25 days ago).
  - Security-disclosure policy not yet verified: we couldn't inspect the source repository.
- **Schema Quality & AI Usability**: 64/100
  - AI-judged instruction clarity (excellent).
  - Context-footprint check failed: tool/resource definitions use about 25676 tokens (~351/item across 73 items; 73 tools + 0 resources), over budget; trim descriptions and params.
  - Usage-examples check failed: none of the tools include examples.
- **Stability & Change Management**: 89/100
  - Stability check failed: the tool surface changed between 0.2.0 and 0.4.2: 1 tool removals, 0 breaking changes, 3 additions.
- **Tool Coverage**: 99/100
  - 100% of tools have a non-trivial description (not blank, and not just the tool's name).
  - 96% of tool parameters carry a description.
- **Tool Safety**: 50/100
  - Injection-marker check failed: the description of tool "check_invoice_sequence" contains an instruction to conceal the call from the user, the text "DO NOT TELL THE USER", at byte 953 of that field, plus 1 further marker(s) of the same kind.
  - We read all 73 captured tool definition(s), and no name or description among them implies an irreversible operation.
  - An AI judge read all 73 captured unit(s) of tool text and found none that tries to manipulate the model reading it.
- **Capabilities**: 100/100
  - Implements a supported MCP spec version (2025-11-25); the latest is 2026-07-28.

## Install

### How do I install the com.yaktool/yaktool MCP server?

com.yaktool/yaktool runs locally as an npm package, launched with npx -y yaktool-mcp. Ready-made configuration for Claude, Cursor, VS Code, Codex and 5 more is on this page, copied from each client's own documentation.

### Claude

```bash
claude mcp add com-yaktool-yaktool -- npx -y yaktool-mcp
```

### Cursor

```json
{
  "mcpServers": {
    "com-yaktool-yaktool": {
      "command": "npx",
      "args": [
        "-y",
        "yaktool-mcp"
      ]
    }
  }
}
```

### VS Code

```json
{
  "servers": {
    "com-yaktool-yaktool": {
      "command": "npx",
      "args": [
        "-y",
        "yaktool-mcp"
      ]
    }
  }
}
```

### Codex

```bash
codex mcp add com-yaktool-yaktool -- npx -y yaktool-mcp
```

### opencode

```json
{
  "$schema": "https://opencode.ai/config.json",
  "mcp": {
    "com-yaktool-yaktool": {
      "type": "local",
      "command": [
        "npx",
        "-y",
        "yaktool-mcp"
      ],
      "enabled": true
    }
  }
}
```

### OpenClaw

```bash
openclaw mcp add com-yaktool-yaktool --command npx --arg -y --arg yaktool-mcp
```

### Hermes

```yaml
mcp_servers:
  com-yaktool-yaktool:
    command: "npx"
    args: ["-y", "yaktool-mcp"]
```

### Netclaw

```json
{
  "McpServers": {
    "com-yaktool-yaktool": {
      "Transport": "stdio",
      "Command": "npx",
      "Arguments": [
        "-y",
        "yaktool-mcp"
      ]
    }
  }
}
```

### Vellum

```bash
assistant mcp add com-yaktool-yaktool -t stdio -c npx -a -y yaktool-mcp
```

### Other

```json
{
  "mcpServers": {
    "com-yaktool-yaktool": {
      "command": "npx",
      "args": [
        "-y",
        "yaktool-mcp"
      ]
    }
  }
}
```

## Changelog

Every change recorded for this component, newest first. Days that predate change tracking, or that we cannot explain, say so: "we were watching and nothing happened" and "we were not watching" are different claims.

### 2026-09-20 (score 73, +1)

No change was recorded against any check on this day. Stability & Change Management went from 85 to 89.

### 2026-09-18 (score 72, +1)

No change was recorded against any check on this day. Stability & Change Management went from 79 to 82.

### 2026-09-16 (score 71, +1)

No change was recorded against any check on this day. Stability & Change Management went from 72 to 75.

### 2026-09-14 (score 70, +1)

No change was recorded against any check on this day. Stability & Change Management went from 65 to 69.

### 2026-09-11 (score 69, +1)

No change was recorded against any check on this day. Stability & Change Management went from 55 to 59.

### 2026-09-09 (score 68, +1)

No change was recorded against any check on this day. Stability & Change Management went from 49 to 52.

### 2026-09-07 (score 67, +1)

No change was recorded against any check on this day. Stability & Change Management went from 42 to 45.

### 2026-09-05 (score 66, +1)

No change was recorded against any check on this day. Stability & Change Management went from 35 to 39.

## MCP tools (73)

### `validate_gtin` (~277 tokens)

Validate a GTIN / UPC / EAN barcode

Validate a single GTIN (GTIN-8 / GTIN-12 UPC-A / GTIN-13 EAN-13 / GTIN-14) against 26 documented rules: check digit (Mod-10), input cleaning (Excel scientific-notation damage, lost leading zeros, separators, full-width digits), format equivalences, GS1 prefix attribution, restricted/coupon/ISBN/ISSN ranges, and placeholder detection. Returns a structured ValidationResult: { input, normalized, valid, format, findings[{ruleId, severity, message, fix?, sourceUrl?}], meta }. `valid` means no error-level findings; warnings (in-store codes, placeholders) do not make it false. Pass the raw value as-is — do NOT pre-clean it; cleaning is part of what gets validated. BULK: pass `values` (an array, max 1000) instead of the single value to check a whole column in one call. The result is then { rows, summary } rather than a single result. Do not loop this tool over a column one value at a time.

Input parameters:

- `value` (string): The barcode value exactly as found in the source data (uncleaned).
- `values` (array): Bulk form: one value per array element. Use this instead of looping the single-value form.

### `calculate_check_digit` (~289 tokens)

Calculate / complete a GTIN check digit

Compute the GS1 Mod-10 check digit for a GTIN payload and return the completed code with the full step-by-step working (digits × 3/1 weights → sum → next multiple of ten). Length semantics (ambiguity is made explicit, never guessed): 7/11 digits → completed as GTIN-8/GTIN-12; 12/13 digits → BOTH interpretations returned (payload to complete AND already-complete code to verify), payload first; 8/14 digits → verify-only (the value already contains a check digit — use validate_gtin for full validation). Input cleaning (Excel damage, separators, full-width digits) matches validate_gtin. Returns { input, normalized, findings, interpretations[{kind: "payload"|"complete", ...}] }. BULK: pass `values` (an array, max 1000) instead of the single value to check a whole column in one call. The result is then { rows, summary } rather than a single result. Do not loop this tool over a column one value at a time.

Input parameters:

- `value` (string): Digits to complete (7/11/12/13, without check digit) or a complete GTIN to verify — raw, uncleaned.
- `values` (array): Bulk form: one value per array element. Use this instead of looping the single-value form.

### `analyze_amazon_search_terms` (~368 tokens)

Analyze Amazon backend search terms (byte budget + rules)

Check Amazon backend search terms against the UTF-8 BYTE limit (under 250 bytes, i.e. 249 max) and 13 documented rules. CRITICAL for listing-generation agents: the limit is bytes, NOT characters — accented Latin costs 2 bytes per character, CJK 3, emoji 4 — and exceeding it can stop Amazon indexing the ENTIRE field, not just the overflow. Also flags wasted budget (duplicate words, commas as separators, singular+plural pairs, words already in the title), disallowed content (ASINs, subjective claims like "best"), and invisible/full-width characters. Optional `title` enables title-overlap detection; optional `competitorBrands` enables branded-term detection (this tool cannot know every trademark — an empty result is not clearance). Returns { bytes, byteLimit, bytesRemaining, characters, tokens[{text,bytes}], normalized, findings[], ok }. Use `normalized` as the cleaned value to submit. BULK: pass `values` (an array, max 1000) instead of the single value to check a whole column in one call. The result is then { rows, summary } rather than a single result. Do not loop this tool over a column one value at a time.

Input parameters:

- `byteLimit` (integer): Override the byte limit for a marketplace with different rules (default 249).
- `competitorBrands` (array): Brand names that must not appear (e.g. competitor brands).
- `searchTerms` (string): The backend search terms string, exactly as it would be submitted.
- `title` (string): Product title, to detect keywords that are already indexed.
- `values` (array): Bulk form: one value per array element. Use this instead of looping the single-value form.

### `convert_gtin` (~346 tokens)

Convert a GTIN between UPC-12, EAN-13 and GTIN-14

Convert a barcode between formats: GTIN-12 (UPC-A), GTIN-13 (EAN-13) and GTIN-14 (case code). The code is VALIDATED FIRST — an invalid GTIN is refused rather than converted, so a broken number is not propagated across channels. Check digits are recalculated whenever the indicator digit changes (GTIN-14 ⇄ GTIN-13); zero-padding (UPC ⇄ EAN-13) preserves the check digit. Two conversions are impossible by GS1 rules and are refused with an explanation: GTIN-8 → anything (it is an independent number range, NOT a short EAN-13), and GTIN-13 → UPC-12 when the code does not start with 0. Returns { input, normalized, sourceFormat, target, converted?, ok, note?, findings[] }. BULK: pass `values` (an array, max 1000) instead of the single value to check a whole column in one call. The result is then { rows, summary } rather than a single result. Do not loop this tool over a column one value at a time.

Input parameters:

- `indicator` (integer): Indicator digit for GTIN-14 targets: 1-8 = packaging level, 9 = variable measure. Default 1.
- `target` (string, required): Target format. GTIN-8 is not a valid target — those codes are allocated by GS1, not constructed.
- `value` (string): The barcode to convert, raw and uncleaned.
- `values` (array): Bulk form: one value per array element. Use this instead of looping the single-value form.

### `validate_google_feed` (~289 tokens)

Validate a Google Shopping product feed before upload

Pre-flight a Google Merchant Center product feed (CSV or TSV text) against 28 documented rules and get every fixable problem at once, instead of discovering them as disapprovals after upload. Covers: file structure (delimiter, header, duplicate/unknown columns, ragged rows), required attributes (id/title/description/link/image_link/availability/price/condition), product identifiers (FULL GS1 validation of the gtin column including spreadsheet damage, plus the identifier_exists x gtin x brand x mpn logic and duplicate ids), formats and enums (price "9.99 USD" form, availability/condition values, absolute URLs, title/description length), and content quality (all-caps titles, promotional text, placeholder values). SCOPE LIMIT — anything requiring a website visit (landing-page price match, image reachability, policy judgement) is NOT checked, so a clean result is not a guarantee of approval. Returns { rows[{line, id, findings[], ok}], fileFindings[], summary{totalRows, okRows, errorRows, warningRows, byRule, detectedDelimiter, columns, unknownColumns} }. Read summary.byRule first: it shows which single rule affects the most rows.

Input parameters:

- `feed` (string, required): The feed file contents as text (CSV or TSV, header row first). Pass it raw — delimiter and encoding quirks are part of what gets checked.

### `validate_shopify_csv` (~302 tokens)

Validate a Shopify product CSV before importing

Pre-flight a Shopify product CSV against 24 documented rules. Shopify reports problems only AFTER the import, and the most damaging one it never reports at all: if a Handle in the file already exists in the store, the import OVERWRITES that live product (title, description, pricing, variants) and still reports success. Pass `existingHandles` (the Handle column of a store export) to detect exactly which live products this import would replace — without it, that check stays silent rather than guessing. Also checks: header/required columns, handle format, non-adjacent handle blocks (which silently drop variants), variant option structure and duplicate combinations, price/boolean/inventory-policy/URL formats, duplicate SKUs, FULL GS1 validation of the barcode column, and dirty data (smart quotes, invisible characters, spreadsheet damage). Returns { rows[{line, handle, isProductRow, findings[], ok}], products[{handle, lines, variantCount}], fileFindings[], summary{productCount, errorRows, warningRows, overwriteCount, byRule, ...} }. Check summary.overwriteCount first — a non-zero value means live products will be replaced.

Input parameters:

- `csv` (string, required): The product CSV contents as text, header row first. Pass it raw — encoding and quoting quirks are part of what gets checked.
- `existingHandles` (array): Handles already in the store (from Products → Export). Enables precise silent-overwrite detection.

### `validate_isbn` (~305 tokens)

Validate and convert an ISBN (10 ⇄ 13)

Validate an ISBN and return BOTH forms at once. ISBN-10 and ISBN-13 use DIFFERENT checksum algorithms — ISBN-10 is modulo-11 with weights 10..2 where a check value of 10 is written as the letter X; ISBN-13 is the GS1 modulo-10 checksum. Applying the wrong one is the most common ISBN bug. CRITICAL: ISBNs beginning 979 have NO 10-character form (the prefix was introduced after ISBN-10 was retired), so `isbn10` is absent for them — do not synthesise one. The 979-0 sub-range is ISMN (printed music), not a book. Also identifies the registration group (the issuing agency — not the country of printing or the language of the text) and flags structurally implausible values. Returns { input, normalized, format, valid, isbn13, isbn10?, group?, findings[], meta? }. BULK: pass `values` (an array, max 1000) instead of the single value to check a whole column in one call. The result is then { rows, summary } rather than a single result. Do not loop this tool over a column one value at a time.

Input parameters:

- `value` (string): The ISBN in either form, with or without hyphens — raw, uncleaned.
- `values` (array): Bulk form: one value per array element. Use this instead of looping the single-value form.

### `lint_listing_title` (~312 tokens)

Check a product title against marketplace rules

Check a product listing title against 16 marketplace rules and per-platform length limits (Amazon, eBay, Etsy, Walmart, Google Shopping, Shopify SEO). ESSENTIAL for listing-generation agents: titles are rejected or suppressed for length, ALL-CAPS words, emoji, trademark symbols, promotional wording ("free shipping", "best price"), subjective claims, contact details, repeated words, and invisible/full-width/non-Latin characters left over from spreadsheets and translation. Returns `spans` — character ranges (CODE POINT offsets, not UTF-16) marking exactly where each problem sits, so you can fix the specific words rather than rewriting the title. Also returns per-platform verdicts with characters remaining, plus the byte count. Pass `platforms` to narrow the length check to the marketplaces you actually sell on. Returns { input, characters, bytes, limits[], findings[], spans[], ok }. BULK: pass `values` (an array, max 1000) instead of the single value to check a whole column in one call. The result is then { rows, summary } rather than a single result. Do not loop this tool over a column one value at a time.

Input parameters:

- `platforms` (array): Marketplaces to check length against (e.g. ["Amazon","eBay"]). Defaults to all.
- `title` (string): The product title exactly as it would be submitted.
- `values` (array): Bulk form: one value per array element. Use this instead of looping the single-value form.

### `check_main_image` (~341 tokens)

Check product-image measurements against Amazon main-image rules

Judge a product image against 12 Amazon main-image rules. This tool does NOT decode images — you supply MEASUREMENTS (dimensions, edge pixel samples, transparency, the bounding box of non-white content) and it applies the rules, so it works anywhere without an image pipeline. Key rule: the background must be RGB 255,255,255 exactly. A deviation of 3 or more levels (e.g. 250,250,250 — invisible to the eye) is an ERROR, because the automated scan measures it; 1-2 levels is a warning. Also checks minimum/maximum dimensions, aspect ratio, transparency, frame fill and whether content touches the border. SCOPE LIMIT: text, logos, watermarks, borders and whether the photo shows the right product are NOT checked — that needs vision, not measurement. The frame-fill figure is an estimate from the bounding box, not Amazon's algorithm. A clean result is not a guarantee of approval. Returns { width, height, fillRatio, worstSample?, findings[], ok }.

Input parameters:

- `contentBounds` (required): Bounding box of non-white content in pixel coordinates, or null if the image is blank.
- `edgeSamples` (array, required): Background pixel samples from the image edges — corners and edge midpoints give the most reliable verdict.
- `hasTransparency` (boolean, required): Whether the image contains any pixel with alpha < 255.
- `height` (integer, required): Image height in pixels.
- `name` (string): File name, for display only.
- `nonWhiteRatio` (number, required): Fraction of pixels that are not background white.
- `width` (integer, required): Image width in pixels.

### `convert_upc_e` (~244 tokens)

Convert UPC-E ⇄ UPC-A (zero suppression)

Expand a UPC-E (8 digits) to its UPC-A (12 digits) form, or compress a UPC-A to UPC-E when the standard allows it. CRITICAL details most implementations get wrong: (1) a UPC-E check digit is computed from the EXPANDED UPC-A, never from the 8 digits themselves; (2) which zeros are suppressed is determined by the LAST DATA DIGIT, and the four compression rules are mutually exclusive — applying them in the wrong order compresses one UPC-A into two different UPC-E values; (3) a UPC-E must start with number system 0 or 1 — an 8-digit code starting with anything else is an EAN-8/GTIN-8, an INDEPENDENT GS1 range that expands to nothing; (4) most UPC-A codes have NO UPC-E form, and this tool says so instead of inventing one. Returns { input, normalized, detected, valid, upcE?, upcA?, rule?, findings[] }.

Input parameters:

- `value` (string, required): A UPC-E (8 digits) or UPC-A (12 digits), with or without hyphens.

### `interpret_eight_digit_barcode` (~118 tokens)

Disambiguate an 8-digit barcode (UPC-E vs GTIN-8)

An 8-digit barcode is ambiguous: it may be a UPC-E (a compressed 12-digit UPC-A) or a GTIN-8/EAN-8 (an independent GS1 range for small packages that expands to nothing). They use different validation and mean different things. This tool returns BOTH readings with validity for each, rather than guessing — use it before assuming which one you have. Returns an array of { format, valid, upcA?, explanation }.

Input parameters:

- `digits` (string, required): Exactly 8 digits.

### `detect_invisible_characters` (~411 tokens)

Find and explain invisible or look-alike characters in text

Scan text for characters that are invisible or that impersonate ordinary ones, and explain each: Unicode name, where it typically comes from, and what it breaks. Catches zero-width spaces, the BOM (the usual cause of "missing header" import errors), non-breaking and narrow spaces, smart quotes, en/em dashes and minus signs, soft hyphens, exotic line separators, bidirectional overrides, full-width Latin forms, variation selectors and control characters. Use this when two strings that look identical fail to match, when an import complains about a column that is clearly present, or before storing text that came from a spreadsheet, a PDF or a web page. Positions are CODE POINT offsets. `cleaned` applies each character's safe replacement — note that characters whose meaning is contextual (zero-width joiners inside emoji sequences) are deliberately KEPT, so `cleaned` is not a blanket strip. The result also carries `bidi`, the BALANCE of bidirectional controls, and that is the part worth acting on: listing "there is a U+202E here" does not explain anything, but an UNTERMINATED opener does — everything from that point to the end of the string displays differently from how it is stored, which is the mechanism behind Trojan-Source-style text. `bidi.unterminated` and `bidi.strayClosers` give the positions; `bidi.affectedFrom` is where the divergence starts. Note U+202C closes only embeddings and overrides while U+2069 closes only isolates — the wrong closer does not balance. These characters are legitimate in Arabic and Hebrew text, so report the imbalance as a fact, not as intent. Returns { input, characters, findings[{info{code,name,category,severity,origin,consequence,replacement?}, positions[], count}], cleaned, changed, hasErrors }.

Input parameters:

- `text` (string, required): The text to inspect, exactly as stored — do not pre-clean it.

### `identify_tracking_carrier` (~247 tokens)

Identify the carrier from a tracking number

Work out which carrier a tracking number belongs to from its format, and VERIFY its check digit where the carrier publishes an algorithm (UPS 1Z, USPS Intelligent Mail package barcode, and UPU S10 for international post). A verified check digit is evidence; a matching shape is only a guess, and the result labels which one you have (`checkDigitValid` is true/false when verifiable, and absent when the carrier publishes no algorithm). AMBIGUITY IS EXPLICIT: many formats are just "N digits" and are shared by several carriers, so ALL matching candidates are returned, sorted with verified ones first — do not treat the first entry as certain when `hasVerified` is false. S10 numbers also decode the country of origin from their final two letters. SCOPE: this identifies formats only — it does NOT track packages or return delivery status, which would require each carrier's API. Returns { input, normalized, candidates[{carrier, format, confidence, checkDigitValid?, checkDigitNote?, origin?, note?}], hasVerified, message? }.

Input parameters:

- `trackingNumber` (string, required): The tracking number, with or without spaces and hyphens.

### `generate_barcode_modules` (~281 tokens)

Generate the module pattern for a barcode symbol

Encode a GTIN as a barcode symbol and return the MODULE PATTERN (a string of 1s and 0s, where 1 is a bar and 0 is a space) plus the layout data needed to render it. Supports EAN-13, UPC-A, EAN-8 and ITF-14; the symbology is detected from the digit count unless you specify it. The check digit is VALIDATED FIRST and generation is REFUSED when it is wrong (with the correct number in `suggestion`) — a barcode encoding a wrong number scans as a different product or not at all, which is worse than having no barcode. Rendering to an image is left to the caller: the response gives `modules`, `moduleCount`, `textGroups` (how the human-readable digits are grouped under the symbol) and `guardRanges` (which modules must be drawn as longer guard bars). The encoding tables are verified in our test suite by an independent decoder over thousands of random round-trips. Returns { symbology, value, modules, moduleCount, textGroups[], guardRanges[] } or { error, suggestion? }.

Input parameters:

- `symbology` (string): Force a symbology instead of detecting it from the digit count.
- `value` (string, required): The complete GTIN including its check digit.

### `validate_vin` (~283 tokens)

Validate a VIN check digit and decode its structure

Validate a 17-character Vehicle Identification Number and decode its sections (WMI/region, descriptor, year code, plant, serial). Three things implementations get wrong: (1) the letter-to-number table is NOT alphabetical — J restarts at 1 and S restarts at 2, and I, O and Q never appear in a VIN at all, so finding one means it was mis-transcribed (we suggest the corrected reading); (2) the MODEL YEAR IS AMBIGUOUS BY DESIGN — position 10 uses 30 codes that repeat, so "A" means 1980, 2010 and 2040. We return ALL candidates with a `likely` flag based on the North American position-7 convention; never report a single year as certain; (3) the check digit is mandatory in North America (FMVSS 115) but frequently unenforced elsewhere, so a failure is strong evidence of a transcription error, NOT proof the VIN is fake. A valid check digit confirms the transcription only — it says nothing about whether the vehicle exists or its title status. Returns { input, normalized, valid, checkDigitValid, expectedCheckDigit, structure{...}, yearCandidates[{year, likely}], findings[] }.

Input parameters:

- `vin` (string, required): The 17-character VIN, with or without spaces and hyphens.

### `validate_imei` (~249 tokens)

Validate an IMEI or IMEISV

Validate a device IMEI (15 digits, Luhn check digit) or recognise an IMEISV (16 digits). CRITICAL distinction: an IMEISV has NO check digit — its final two digits are a Software Version Number — so running Luhn over a 16-digit value and reporting "invalid" is a category error, not a finding. This tool identifies which form you have and only checks the checksum where one exists. A 14-digit value is an IMEI missing its check digit, and the completed number is returned. The first 8 digits are the TAC (model identifier); resolving it to an actual device needs the licensed GSMA TAC database, so the code is reported without a guessed model name. A valid check digit proves the digits are consistent — it says nothing about blacklist or stolen status, which live in carrier and GSMA databases. Returns { input, normalized, valid, type: "IMEI"|"IMEISV", tac, serial, checkDigit?, expectedCheckDigit?, softwareVersion?, findings[] }.

Input parameters:

- `imei` (string, required): The IMEI or IMEISV, with or without spaces, hyphens or slashes.

### `validate_container_number` (~275 tokens)

Validate a shipping container number (ISO 6346)

Validate an 11-character freight container number against ISO 6346 and decode its parts (owner code, equipment category, serial, check digit). Two traps that make hand-written implementations reject valid numbers: (1) the letter values start at A=10 and SKIP EVERY MULTIPLE OF 11, so B=12 (not 11), L=23 (not 22), V=34 (not 33) — an alphabetical mapping is wrong; (2) the weights are POWERS OF TWO (1,2,4,…,512), not position numbers. The fourth letter is the equipment category and may only be U (freight container), J (detachable equipment) or Z (trailer/chassis). The standard has a known weakness: a remainder of 10 is written as 0, so some numbers ending in 0 are genuinely ambiguous — `usedTenAsZero` flags when yours is one of them. Owner prefixes are registered with the BIC; we report the code without a guessed company name. Returns { input, normalized, valid, ownerCode, category{code,description}, serial, checkDigit, expectedCheckDigit, usedTenAsZero, findings[] }.

Input parameters:

- `containerNumber` (string, required): The 11-character container number, with or without the printed spacing.

### `parse_gs1_element_string` (~397 tokens)

Parse a GS1-128 element string into its Application Identifiers

Decode a GS1 element string (the data under a GS1-128 / logistics barcode) into its Application Identifiers, using 30 known AI definitions. Accepts the human-readable bracketed form "(01)09501101020917(17)250101(10)LOT1" and raw scanner output where FNC1 appears as the GS control character (U+001D). THE RULE MOST PARSERS GET WRONG: an AI has a predefined length when its FIRST TWO DIGITS appear in the GS1 predefined-length table (00,01,02,03,04,11-19,20,31-36,41) — those are followed directly by the next AI with NO separator. Every other AI is variable-length and MUST be terminated by FNC1 unless it is last. Ignore this and a batch number silently swallows the serial number after it, producing data that looks fine and is wrong. This tool reports a missing separator explicitly. Also validates GTIN/SSCC/GLN check digits inside the string, YYMMDD dates (where day 00 legitimately means END OF MONTH), the decimal indicator on measurement AIs (3103 = 3 implied decimals), the AI character set (lowercase is not permitted), and repeated AIs. Two-digit years resolve on a sliding window; pass `referenceYear` for a deterministic reading. Unknown AIs stop the parse rather than being guessed at, because an unknown AI has no known data length. Returns { input, normalized, elements[{ai, title, value, interpreted?, start, end, findings[]}], findings[], ok }.

Input parameters:

- `elementString` (string, required): The GS1 element string, bracketed or raw. Pass it exactly as scanned — separators are part of what gets checked.
- `referenceYear` (integer): Year used to resolve two-digit dates (GS1 sliding window). Defaults to the current year.

### `diff_product_feeds` (~366 tokens)

Compare two product feed versions aligned by SKU

Compare two versions of a product feed (CSV or TSV) ALIGNED BY KEY rather than by line. This matters because product exports do not preserve row order — a re-sorted or partially deleted export makes a line-based diff report that everything changed when nothing did. Rows here are matched on id/sku/handle/gtin (auto-detected, or pass `keyColumn`), so reordering yields zero differences. Returns per-row classification (added / removed / changed / unchanged) and, for changed rows, FIELD-LEVEL before/after pairs with numeric deltas and percentages for price and stock, plus semantic notes (stock hit zero, availability flipped, image replaced, value cleared). WARNINGS matter: duplicate keys make any key-based comparison unreliable and are reported explicitly, as are columns present in only one file (which are excluded from comparison rather than reported as changes). Use `ignoreColumns` for volatile fields such as export timestamps. Returns { keyColumn, keySource, columns{both,onlyBefore,onlyAfter}, rows[], summary{added,removed,changed,unchanged,byField,duplicateKeys*}, warnings[{ruleId,severity,message,fix?}] }. Structural warnings carry stable rule ids (DIFF-K01 no key column, DIFF-K02 key missing on one side, DIFF-K03 duplicate keys, DIFF-C01 different columns) — branch on the id rather than matching the English text.

Input parameters:

- `after` (string, required): The newer feed export, as text.
- `before` (string, required): The older feed export, as text.
- `ignoreColumns` (array): Columns whose changes should be ignored (export timestamps and similar).
- `keyColumn` (string): Column to align rows by. Auto-detected when omitted.

### `fix_mojibake` (~244 tokens)

Diagnose and repair encoding damage (mojibake)

Repair text mangled by an encoding mix-up — "cafÃ©" back to "café", "donâ€™t" back to "don't". This does NOT use a substitution table: it infers and executes the actual encoding chain (UTF-8 bytes read as CP1252 or ISO-8859-1), so any character in any script repairs the same way, including ones no lookup table would carry. Every candidate is SELF-VERIFIED — the repaired text is pushed back through the same broken chain and only returned if it reproduces your input exactly. Handles double encoding and reports how many layers were undone. IMPORTANT LIMIT: U+FFFD replacement characters (�) mean a decoder already discarded those bytes. That damage is NOT reversible by any tool — the result will say so rather than guessing. Also worth telling the user: repairing the text fixes today's file, but the import that produced it will corrupt the next one too. Returns { input, detected, candidates[{repaired, chain, misreadAs, layers, score}], findings[], signatures[] }.

Input parameters:

- `text` (string, required): The mangled text, exactly as stored.

### `detect_csv_dialect` (~229 tokens)

Detect the delimiter, quoting and encoding traits of a CSV

Work out what dialect a delimited file actually is: delimiter, quoting, line endings, byte-order mark, whether the first row is a header, and which rows are ragged. Detection is by COLUMN-COUNT CONSISTENCY across the whole file with RFC 4180 quoting honoured — not by counting separators in the first line, which picks the wrong delimiter on the most common real file there is (a semicolon-separated European export whose descriptions are full of commas). Each candidate separator is returned with its consistency score, so the verdict can be checked rather than trusted. Use before parsing an unfamiliar feed, or when an import complains about a header column that is visibly present (that symptom is almost always a BOM). Returns { delimiter, delimiterName, candidates[{delimiter,name,consistency,columns}], lineEnding, hasBom, usesQuotes, looksLikeHeader, columns[], rowCount, raggedRows[], findings[] }.

Input parameters:

- `text` (string, required): The file contents, or its first few dozen lines. Pass it raw — the BOM and line endings are part of what gets detected.

### `check_number_format` (~325 tokens)

Check a price or number for ambiguous formatting

Decide whether a numeric value has one reading or several, and validate any currency attached to it. THE TRAP THIS EXISTS FOR: a single separator followed by exactly three digits is genuinely ambiguous — "1,234" is 1234 in English and 1.234 in German, a thousandfold apart, and no parser errors on it. When that happens `ambiguous` is true and EVERY reading is returned in `interpretations`; do NOT pick one silently, surface the ambiguity to the user. Unambiguous values (two decimal places, both separators present, or none at all) return a single `value`. Currency: ISO 4217 codes are validated with their correct minor-unit count — note that JPY and KRW have ZERO decimal places and BHD/KWD have THREE, so "1500.00 JPY" is malformed. A bare symbol is flagged because $ alone is used by more than twenty currencies. Returns { input, currency?, numericPart, interpretations[{convention,value,canonical}], value?, ambiguous, findings[] }. BULK: pass `values` (an array, max 1000) instead of the single value to check a whole column in one call. The result is then { rows, summary } rather than a single result. Do not loop this tool over a column one value at a time.

Input parameters:

- `value` (string): The price or number as written, including any currency code or symbol.
- `values` (array): Bulk form: one value per array element. Use this instead of looping the single-value form.

### `check_date_format` (~401 tokens)

Check a date for ambiguity, or decode a spreadsheet serial

Decide what a date value actually says, and whether it says more than one thing. THE TRAP THIS EXISTS FOR: when both components are 12 or lower, "01/02/2026" is 1 February in most of the world and 2 January in the United States — a month apart, and no parser errors on it. When that happens `ambiguous` is true and EVERY reading is returned in `readings` with its weekday; do NOT pick one silently, surface the ambiguity to the user. Also decodes values that are dates in disguise: a bare number in the spreadsheet serial range (45658 is 1 January 2025), where serial 60 is the non-existent 29 February 1900 that spreadsheets deliberately reproduce, and the legacy 1904 Mac epoch puts the same serial 1462 days away; and 10- or 13-digit Unix timestamps. Validates impossible dates against the full Gregorian leap rule, flags two-digit years and separator-less compact dates, checks Merchant Center date ranges (start/end) including ranges that end before they start, and warns when a time carries no zone — Merchant Center does NOT simply assume UTC there, it uses the target country default for text/XML feeds and UTC only for the API. Returns { input, kind, readings[{convention,iso,spelled,weekday}], ambiguous, iso?, time?, findings[] }. BULK: pass `values` (an array, max 1000) instead of the single value to check a whole column in one call. The result is then { rows, summary } rather than a single result. Do not loop this tool over a column one value at a time.

Input parameters:

- `value` (string): The date as written — any format, including a bare serial number or a start/end range.
- `values` (array): Bulk form: one value per array element. Use this instead of looping the single-value form.

### `diagnose_excel_damage` (~420 tokens)

Diagnose what a spreadsheet broke in a column of data

Take a column of values that came out of a spreadsheet and report, per value, what damage it carries and — the part that matters — WHETHER IT CAN STILL BE REPAIRED. Each row gets `recoverability`: "intact", "recoverable" (a determinate fix exists and is given in `repaired`), or "unrecoverable" (the original data is not in this file and no tool can restore it). Use that field to decide whether to fix in place or tell the user to re-export from the source. THE TRAP THIS EXISTS FOR: expanding scientific notation is NOT the same as repairing it. "8.71E+12" describes a 13-digit number but carries only 3 of its digits; expanding it to 8710000000000 invents the other 10. This tool counts carried vs lost digits and refuses to offer a repair when digits were destroyed — do not expand such values yourself. Also detects: leading zeros stripped (only when the column width makes it detectable — at least three same-length numeric cells), rounding past the 15 significant digits a spreadsheet keeps, codes coerced to dates (MAR1 becomes 1-Mar and CANNOT be reversed — `originalCandidates` lists the possibilities), floating-point tails, spreadsheet error values, display formatting baked into values (accounting brackets mean NEGATIVE), booleans, and invisible whitespace. SECURITY: cells beginning with =, +, -, @, tab, CR, LF or the full-width variants ＝ ＋ － ＠ are executed as formulas when the file is opened (OWASP CSV Injection). These are reported as errors, not style notes. Returns { rows[{line,input,recoverability,repaired?,originalCandidates?,findings[]}], modalNumericLength?, summary{total,intact,recoverable,unrecoverable,byRule[]} }.

Input parameters:

- `column` (string, required): The column of values, one per line, exactly as exported — do NOT pre-clean it; the damage is what gets diagnosed.

### `check_gs1_digital_link` (~336 tokens)

Check a GS1 Digital Link URI against the current standard

Parse a GS1 Digital Link URI, report every conformance problem, and return the corrected URI in `corrected`. TWO RECENT CHANGES INVALIDATED MOST EXISTING URIS, and neither makes the URI look wrong — do NOT assume an older URI is fine: (1) convenience alphabetic names such as /gtin/ were deprecated in v1.2 and REMOVED COMPLETELY in v1.3.0, so they are non-conformant, not merely discouraged; (2) as of v1.4.0 the GTIN must be expressed with 14 digits — a GTIN-8/12/13 needs leading-zero padding, even though the code and its check digit are unchanged. Also enforces: the path grammar of all sixteen primary keys (01 GTIN, 00 SSCC, 414 GLN, 8006 ITIP and the rest), key qualifier ORDER (for a GTIN: 22 then 10 then 21 — the standard gives the reversed form as an explicit counter-example), qualifiers that are required rather than optional (415 needs 8020), the GS1 Mod-10 check digit, and the rule that data attributes SHALL live in the query string rather than the path. Compressed URIs are detected but NOT decoded — compression is a separate GS1 standard; the result says so instead of guessing. Returns { input, stem?, primary?, qualifiers[], attributes[], segments[], corrected?, conformant, findings[] }.

Input parameters:

- `uri` (string, required): The Digital Link URI exactly as it appears, including the scheme if present.

### `build_gs1_digital_link` (~306 tokens)

Build a conformant GS1 Digital Link URI from a GTIN

Construct a Digital Link URI that is conformant by construction: the GTIN is padded to the 14 digits v1.4.0 requires, and any key qualifiers are emitted in the order the standard mandates (consumer product variant, then batch/lot, then serial number) regardless of the order you pass them in. Data attributes are placed in the query string, which is where the standard requires them. NOTE: this is the quick, unvalidated builder — it does not verify the check digit, does not validate qualifier character sets, and produces no QR. For a validated build with a refuse-on-error gate, print-size advice and the QR image, use generate_digital_link_qr. Returns { uri } — feed it to check_gs1_digital_link if you want the conformance report as well.

Input parameters:

- `attributes` (object): Data attributes as AI → value, e.g. { "17": "261001" }. They go in the query string.
- `cpv` (string): Consumer product variant, AI 22.
- `gtin` (string, required): The GTIN — 8, 12, 13 or 14 digits; it will be padded to 14.
- `lot` (string): Batch or lot number, AI 10.
- `serial` (string): Serial number, AI 21.
- `stem` (string): Resolver origin and optional path prefix (default https://id.gs1.org).

### `check_image_metadata` (~276 tokens)

Check an image for privacy-leaking EXIF metadata

Read the EXIF metadata out of a JPEG or TIFF supplied as base64 and report what it leaks. THE FINDING THAT MATTERS: GPS coordinates. A product photo taken at home carries the seller's home address to a few metres, and many marketplaces pass it through unchanged — if `gps` is present, tell the user before the image is published anywhere. Also reports camera and lens SERIAL NUMBERS (a stable identifier linking every photo from that device), personal names in the artist/copyright/owner fields, capture timestamps, editing software, an embedded thumbnail (which may still show the image BEFORE a crop), and an orientation flag other than 1 (feed processors that ignore it display the photo sideways). Do NOT paste the coordinates into a map service on the user's behalf — that discloses the location you were asked to check for. Only JPEG and TIFF are parsed; other containers return format with a "not parsed" finding rather than a false all-clear. Returns { format, hasExif, byteOrder?, tags[{tag,name,value,ifd}], gps?{latitude,longitude,latitudeDms,longitudeDms,altitude?}, thumbnailBytes?, findings[] }.

Input parameters:

- `imageBase64` (string, required): The image file as base64. Only the first metadata segments are read; the pixels are ignored.

### `check_country_code` (~406 tokens)

Check or convert an ISO 3166-1 country code

Resolve a country identifier and classify it. THE TRAP: **UK is not a country code.** It is not in the registry at all — ISO 3166-1 assigns GB — and it is the most common reason a country column is rejected. Never emit UK; if a user writes it, correct it to GB. Returns `status` with five distinct values, which need different handling: "assigned" (an ordinary country code, safe for a feed), "special" (registered but not a country — EU, UN, EZ, AC; feeds reject these), "private-use" (AA/ZZ, an escaped internal placeholder), "deprecated" (withdrawn; `suggestion` names the successor when the registry does, and is ABSENT when the territory split into several — do not invent one), and "unknown" (not in the registry). Also converts other forms: alpha-3 (DEU), numeric (276, and 8 for a code whose leading zeros a spreadsheet ate), and country names in either the IANA or UN wording. Data is generated from the IANA Language Subtag Registry joined to the UN M49 list and cross-checked against ICU. Note TW is a valid ISO code but absent from the UN list, so its alpha-3 and numeric fields are empty rather than guessed. Returns { input, normalized, status, entry?, suggestion?{code,reason}, findings[] }. BULK: pass `values` (an array, max 1000) instead of the single value to check a whole column in one call. The result is then { rows, summary } rather than a single result. Do not loop this tool over a column one value at a time.

Input parameters:

- `value` (string): The country code, alpha-3, numeric code or country name, as written.
- `values` (array): Bulk form: one value per array element. Use this instead of looping the single-value form.

### `check_hreflang` (~339 tokens)

Check hreflang annotations and produce the block every page must carry

Validate a set of hreflang annotations. THE TRAP: invalid annotations are IGNORED SILENTLY by search engines — no error, no warning, nothing in any report. "en-UK" is the classic case: UK is not an ISO 3166-1 region code (the United Kingdom is GB), so the annotation is thrown away and the targeting simply does not happen. Never emit en-UK. Rules enforced: language must be ISO 639-1 and region ISO 3166-1 alpha-2 (validated against the IANA registry); a country code alone is NOT valid because the language is never inferred from the country; underscores are not hyphens; each set must reference its own page; no value may point at two URLs. THE USEFUL OUTPUT is `canonicalBlock`: because every version must list itself and all the others, the correct set of tags is IDENTICAL on every page. Emit that block on all of them and the reciprocity requirement is satisfied by construction — rather than trying to patch pages one at a time. Honest boundary: whether the other pages actually carry the block cannot be checked here, because nothing is fetched. Do not claim return links are verified. Accepts link tags pasted from page source, or "value URL" pairs one per line. Returns { entries[{value,normalized,href,valid,...}], canonicalBlock?, findings[] }.

Input parameters:

- `annotations` (string, required): The hreflang link tags, or one "value URL" pair per line.
- `selfUrl` (string): This page's own URL, so the self-reference can be checked.

### `validate_awb` (~304 tokens)

Validate an air waybill number

Check an 11-digit air waybill number: three-digit IATA airline prefix, seven-digit serial, and an UNWEIGHTED MODULO-7 check digit — the serial divided by seven, remainder is the check digit. Nothing is weighted and nothing alternates. USEFUL SHORTCUT: because the check digit is a remainder modulo seven it can only be 0-6, so **a number ending in 7, 8 or 9 is invalid regardless of airline or serial** and needs no arithmetic to reject. Use this to screen a batch before doing anything else. This tool does NOT map the airline prefix to a carrier name — that table is maintained by IATA and not published openly, and an unverified table would give confidently wrong answers about real shipments. Do not fill it in from memory either. Returns { input, normalized, formatted?, airlinePrefix?, serial?, checkDigit?, expectedCheckDigit?, valid, findings[] }. BULK: pass `values` (an array, max 1000) instead of the single value to check a whole column in one call. The result is then { rows, summary } rather than a single result. Do not loop this tool over a column one value at a time.

Input parameters:

- `value` (string): The air waybill number, with or without hyphens or spaces.
- `values` (array): Bulk form: one value per array element. Use this instead of looping the single-value form.

### `detect_mixed_script` (~262 tokens)

Find characters that are not what they look like

Detect characters in listing text that render like ordinary Latin text but are not. Two distinct problems: (1) FULL-WIDTH PUNCTUATION — ，。（）！？ produced by CJK input methods. They look like ordinary punctuation, break search matching, and are a mechanical substitution, so `suggested` carries the corrected text. (2) LOOKALIKE LETTERS — Cyrillic а е о р с and Greek ο are visually IDENTICAL to Latin letters. A word containing one is a different string: search misses it, deduplication misses it, spellcheck stays quiet. This is invisible to any amount of proofreading, so it can only be found mechanically. Leftover words in another script are reported but NOT changed — deciding to translate, transliterate or delete them requires knowing what they say. Do not substitute anything there. Spans use CODE-POINT offsets, not UTF-16 indices, so they stay correct with emoji present. Returns { input, scripts[], spans[{start,end,text,script,ruleId}], suggested?, findings[] }.

Input parameters:

- `expectedScript` (string): The script the text should be in (default Latin). Anything outside it is reported as leftover.
- `text` (string, required): The listing text to examine.

### `convert_units` (~342 tokens)

Convert dimensions or weights for a feed, and report what rounding costs

Convert length or mass values, one per line, and report BOTH the exact result and the value rounded to the precision a feed field accepts — plus what that rounding cost as a percentage. THE POINT: the conversion factors are EXACT by definition (an inch is exactly 0.0254 m, a pound exactly 0.453 592 37 kg, since 1959), so any error in a converted value came from rounding, not from the conversion. One pound rounded to two decimals is 0.45 kg — about 0.8 % light, always in the same direction, so it accumulates across a consignment rather than cancelling out. Report the cost to the user rather than only the rounded number. Batch mode also flags a column that MIXES units, which is worse than a wrong one: a feed field carries a single unit for the whole column, so every row that does not match is silently mislabelled while each number remains individually valid. Length and mass only. Volume, area and temperature are absent on purpose rather than approximated. Returns { rows[{line,input,exactValue?,rounded?,roundingErrorPercent?,findings[]}], summary{total,converted,failed,worstErrorPercent,unitsSeen}, findings[] }.

Input parameters:

- `decimals` (integer): Decimal places the target field accepts (default 2).
- `to` (string, required): Target unit code: mm, cm, m, in, ft, yd, g, kg, lb, oz.
- `values` (string, required): One value per line, e.g. "1 lb" or "10 in". The unit may be attached or spelled out.

### `check_utm` (~251 tokens)

Check a set of campaign URLs for tagging problems and report fragmentation

Check UTM tagging across a SET of campaign URLs, one per line. Pass them all together — the most damaging problem is invisible in any single URL. THE TRAP: parameter values are CASE SENSITIVE. utm_source=Meta and utm_source=meta are two different sources and become two separate rows in every report, while each URL looks perfectly correct on its own. The same happens with separator style (spring_sale vs spring-sale). This tool reports those collisions and lists every value each parameter took. Also per URL: missing required parameters (utm_source, utm_medium, utm_campaign — a missing one reports as "(not set)", it does not error); misspelled utm_ parameters, which are silently IGNORED so the URL looks tagged and reports nothing; parameters placed after a # fragment, which are never sent anywhere; duplicate parameters; upper case and spaces in values. When emitting UTM values yourself, lower-case them — it is the only habit that prevents the split. Returns { urls[], valuesByParameter[{parameter,values[]}], findings[] }.

Input parameters:

- `urls` (string, required): One campaign URL per line. Pass the whole set, not one at a time.

### `build_utm_url` (~175 tokens)

Build a correctly tagged campaign URL

Construct a campaign URL with the required parameters. Values are lower-cased and spaces become hyphens by default, because parameter values are case sensitive and that is the only habit that keeps one channel from splitting into several rows. Returns { uri } — feed it to check_utm if you also want the report.

Input parameters:

- `campaign` (string, required): utm_campaign — the campaign name.
- `content` (string): utm_content — creative variant.
- `id` (string): utm_id — campaign id.
- `medium` (string, required): utm_medium — the channel type, e.g. cpc, email.
- `source` (string, required): utm_source — where the traffic comes from.
- `term` (string): utm_term — paid keyword.
- `url` (string, required): The landing page URL.

### `check_colour_space` (~307 tokens)

Check whether an image is CMYK, RGB, greyscale or indexed

Read the colour space out of a JPEG or PNG header supplied as base64. Only the header is needed, so send the FIRST 64 KB rather than the whole file. THE FINDING THAT MATTERS: a CMYK JPEG. It looks completely normal in every viewer — nothing is visibly wrong — and marketplaces then reject it or display the colours inverted. Files that came from print artwork are the usual source, and converting the file format does not convert the colour model. If `colourSpace` is "CMYK" or "YCCK", tell the user to re-export as RGB before uploading anywhere. Detection is from the JPEG frame header (component count: 1 greyscale, 3 colour, 4 CMYK) and the Adobe APP14 transform byte, or from the PNG IHDR colour-type byte. No pixels are decoded. PNG has no CMYK at all, so never report one for a PNG; PNG problems are different — an indexed palette bands on product photography, and an alpha channel is composited against different backgrounds on different platforms. This says what colour model the file DECLARES, not whether the colours are right, and not whether the background is white — that needs pixel measurement. Returns { format, colourSpace, components?, bitDepth?, hasAlpha, iccProfileBytes?, adobeTransform?, findings[] }.

Input parameters:

- `imageBase64` (string, required): The first part of the image file as base64. 64 KB is more than enough.

### `generate_slug` (~296 tokens)

Turn product titles into URL slugs and platform handles

Convert one title per line into a URL slug, and report the two things generic slug functions miss. AMBIGUITY: some letters have two accepted romanisations — German expands ü to ue, most software just drops the accent, so "Müller" is legitimately either "muller" or "mueller". Both are returned in `alternatives`; do NOT pick one silently, ask which convention the store uses. COLLISIONS: two titles that reduce to the same slug are not rejected by Shopify — it appends a number, so the second product quietly becomes "name-1" and every link already published for it points at the other product. Nothing in the import report mentions this. `actualHandle` shows what the platform would really create, and `collided` marks the affected rows. Characters from Han, Cyrillic, Arabic and other scripts are DROPPED rather than romanised, because romanisation is language-specific and a guess would produce a confident wrong URL. Removed characters are listed, since dropping punctuation is how two distinct products end up fighting over one URL. Returns { rows[{line,input,slug,alternatives[],actualHandle,collided,findings[]}], summary{total,collisions,ambiguous,empty} }.

Input parameters:

- `maxLength` (integer): Length past which a slug is flagged as long (default 70).
- `titles` (string, required): One product title per line.

### `check_product_jsonld` (~285 tokens)

Validate Product structured data, values included

Check a Product JSON-LD block. Accepts the bare JSON, a whole page of HTML (the ld+json block is extracted), or an @graph. Beyond the required properties (name; and at least one of offers, review, aggregateRating; and price inside an offer), this VALIDATES THE VALUES using the same engines as the rest of this server: price must be a NUMBER — "19,99" is not 19.99 and is the single most common silent failure, because it reads correctly to a European author; priceCurrency is checked against ISO 4217 including its minor-unit count, so "1500.00 JPY" is flagged (the yen has no decimal places); gtin goes through the full barcode engine, so a wrong check digit, a restricted in-store range or a placeholder code is caught — a Product asserting an identifier that does not exist is worse than one asserting none. GTIN problems keep the barcode engine's own severity rather than being re-graded here. Also: aggregateRating with no review or rating count behind it, which is a documented cause of manual action; availability that is free text rather than a schema.org URL; relative image URLs. Returns { input, data?, type?, valid, findings[] }.

Input parameters:

- `markup` (string, required): The JSON-LD block, or the page HTML containing it.

### `build_product_jsonld` (~229 tokens)

Generate Product JSON-LD that passes validation

Build a Product JSON-LD block. The price is emitted as a bare number string with the currency in priceCurrency (the shape the specification requires), the currency is upper-cased, and availability defaults to the schema.org URL form rather than free text. If no price is given, `offers` is omitted entirely rather than emitted empty — an empty offer is worse than none. Feed the result to check_product_jsonld to see the report as well. Returns { jsonLd }.

Input parameters:

- `availability` (string): InStock, OutOfStock, PreOrder or BackOrder (default InStock).
- `brand` (string)
- `gtin` (string): The barcode; it is validated by check_product_jsonld.
- `image` (string): Absolute URL.
- `name` (string, required): Product name — required.
- `price` (string): Bare number, e.g. "19.99". No currency symbol, no comma.
- `priceCurrency` (string): Three-letter ISO 4217 code.
- `sku` (string)
- `url` (string)

### `check_description_html` (~308 tokens)

Check product description HTML for what a marketplace will strip

Check the HTML of a product description against the things that go wrong on every marketplace. THE TRAP: active content is removed SILENTLY. eBay's wording is that JavaScript, Flash, plug-ins and form actions are "no longer displayed in listings" — not rejected, not reported, just absent, so the seller sees a description with a hole in it and no explanation anywhere. Scripts, iframes, forms, embeds and INLINE EVENT HANDLERS (onclick, onerror) all go. The event handlers matter most because they hide inside otherwise ordinary elements and survive a copy-paste from a site template. SECOND TRAP: an unclosed tag rarely damages the description — it damages the page AROUND it, because the marketplace template inherits the open element, so the symptom appears far from the cause. `sanitised` removes only the active content. Styles, links and tables are REPORTED AND LEFT ALONE: whether to keep them depends on the layout, and deleting a seller's formatting automatically would be its own silent failure. Do not strip those on the user's behalf either. This does NOT know each marketplace's exact allowed-tag list — Amazon's lives behind a Seller Central login. It checks what goes wrong everywhere. Returns { input, tags[], spans[{start,end,tag,ruleId}], textLength, sanitised?, findings[] }. Spans use code-point offsets.

Input parameters:

- `html` (string, required): The description HTML, exactly as it would be submitted.

### `check_postcode` (~332 tokens)

Check postcodes against the format each country uses

Check one postcode per line against the pattern its country normally uses. A line may name its own country ("SW1A 1AA, GB"); otherwise `defaultCountry` applies. READ THIS BEFORE ACTING ON THE RESULT: this checks FORMAT, not existence. The patterns come from the address metadata Chrome and Android use — widely deployed, but NOT the postal authorities, whose own tables sit behind a paid service. A mismatch is therefore a WARNING, never an error, and where the country publishes an official lookup the finding names it. Do not tell a user their address is invalid on this basis; tell them it does not match the usual pattern and point at the lookup. The summary separates two different amounts of work: `formatting` rows are batch-fixable (SW1A1AA vs SW1A 1AA — accepted either way, but they do not match each other when systems compare addresses exactly), while `mismatched` rows need a human with the real lookup. It also knows the countries that have NO postcode system at all — there, a value in the field is itself the bug, usually a checkout form that made the field mandatory everywhere. Returns { rows[{line,input,country,normalized,matches,corrected?,format?,findings[]}], summary{total,matching,formatting,mismatched} }.

Input parameters:

- `defaultCountry` (string, required): ISO 3166-1 alpha-2 code for rows that do not name a country.
- `postcodes` (string, required): One postcode per line, optionally followed by ", CC" to set that row's country.

### `check_lithium_battery` (~469 tokens)

Calculate lithium battery watt hours and the passenger-carriage band

Work out watt hours and place a battery in the published passenger-carriage bands (up to 100 Wh: no approval; 101-160 Wh: airline approval and at most two spares; above 160 Wh: not permitted in passenger baggage; lithium metal uses grams of lithium at 2 g and 8 g). THE COMMON ERROR IS THE VOLTAGE. A capacity in mAh is rated at the CELL voltage (3.6-3.7 V for lithium-ion), not the 5 V a USB port outputs. A 20 000 mAh pack is 74 Wh at 3.7 V and 100 Wh at 5 V — the difference between "carry it freely" and "ask the airline". If a Wh figure is printed on the battery, use `ratedWattHours` and skip the calculation entirely. If no voltage is given, EVERY common nominal reading is returned in `readings` and the verdict assumes the highest — do not present a single number as the answer in that case, present the range and ask which voltage the battery is marked with. Always mention that spare batteries and power banks travel in the cabin only, never checked. SCOPE: these are the published FAA rules for what a PASSENGER may carry. CARGO shipping is governed by the IATA Dangerous Goods Regulations and its packing instructions (PI 965-970), a paid standard this server does not carry — do NOT reconstruct that classification from memory, it is a safety matter. Refer the user to their carrier or freight forwarder. This is a calculator plus a summary of published rules, not a compliance determination. Returns { chemistry, milliampHours?, lithiumGrams?, readings[{volts,basis,wattHours}], carriage, findings[] }.

Input parameters:

- `chemistry` (string): Default lithium-ion.
- `lithiumGrams` (number): Lithium content in grams, for lithium metal batteries.
- `milliampHours` (number): Capacity in mAh.
- `ratedWattHours` (number): The Wh figure printed on the battery — preferred over calculating.
- `spare` (boolean): Carried as a spare rather than installed in a device (default true).
- `volts` (number): CELL nominal voltage, not the output voltage.

### `validate_gs1_key` (~369 tokens)

Validate SSCC-18 and GLN numbers

Validate the two GS1 keys that are not products: the 18-digit SSCC on a logistic unit and the 13-digit GLN identifying a party or place. Checks the Mod-10 check digit and reports what the number can be. THE MOST IMPORTANT THING THIS RETURNS IS AN AMBIGUITY. A GLN and a GTIN-13 are indistinguishable as numbers — same length, same check digit algorithm, same prefix pool. For 13 digits the `readings` array therefore holds BOTH interpretations with the condition under which each applies. Do not tell a user "this is a GLN"; tell them it is a well-formed 13-digit GS1 key and that the field it sits in decides which. Two damage patterns worth knowing: a 17-digit value is almost always an SSCC whose leading zero a spreadsheet ate (if restoring it validates, the finding gives the recovered number), and a 20-digit value starting 00 is an SSCC still carrying the AI (00) prefix from a GS1-128 scan, which is stripped automatically. Also note the SSCC extension digit is NOT the packaging-level indicator that opens a GTIN-14 — it has no meaning at all. This tool will NOT split the number into company prefix and reference: a GS1 Company Prefix is 7-12 digits and the boundary lives only in GS1 GEPIR. Do not infer it. Returns { rows[{line,input,normalized,valid,keyType,checkDigit,prefixOrigin?,readings[],findings[]}], summary{total,valid,invalid,duplicates} }.

Input parameters:

- `expect` (string): Force a key type; omit to infer from length.
- `numbers` (string, required): One SSCC or GLN per line.

### `validate_media_codes` (~409 tokens)

Validate ISSN journal numbers and ISRC recording codes

Check a mixed column of ISSNs and ISRCs; each line is routed by shape, so the caller does not have to sort them first. THE TWO CODES DIFFER IN WHAT CAN BE KNOWN. An ISSN carries a modulus 11 check digit (the eighth character, which may be X), so it can genuinely be checked. An ISRC carries NO check digit of any kind — the only question answerable about one is whether it is the right SHAPE. Never report an ISRC as "valid" meaning verified; say well formed. ISSN ↔ barcode: a periodical EAN-13 is 977 + the first SEVEN ISSN digits + a two-digit variant + a GS1 check digit. The ISSN check character is NOT in the barcode, so converting back requires recomputing it — the classic error in serials data. Pass a 977 barcode in and the recovered ISSN comes back with the character recomputed. ISRC prefixes: do NOT reject QM, QZ, QT, QN, ZZ or CP for "not being a country". IFPI allocates the whole five-character prefix and these are deliberate non-country allocations; rejecting them is a common and wrong behaviour. ISRC year: two digits, defined by IFPI as the year of ASSIGNMENT, and the scheme began in 1989 — so a value between the current year and 89 has no reading as an assignment year. The finding then names both possibilities (typo, or back catalogue coded with the recording year) without choosing. Returns { rows[{line,kind,issn?,isrc?,valid}], summary{total,issn,isrc,valid,invalid} }.

Input parameters:

- `codes` (string, required): One ISSN, ISRC or 977 periodical barcode per line.
- `referenceYear` (integer): The current year, used to read the ISRC two-digit year. Pass it — this server does not read the clock.

### `validate_regulated_codes` (~363 tokens)

Verify CAS Registry numbers and convert NDC drug codes

Check a column of CAS Registry Numbers (chemicals) or NDC codes (US drug products). Set `kind` — the two shapes overlap. CAS numbers carry a check digit (each digit multiplied by its position from the right, modulo 10), so a typo is genuinely detectable; the finding names the expected digit. A pass means well formed, NOT that the number names the substance on the label. THE CRITICAL NDC BEHAVIOUR: an NDC has NO check digit, and a ten-digit NDC WITHOUT hyphens is genuinely ambiguous. The FDA assigns 4-4-2, 5-3-2 or 5-4-1, and the hyphens are the only thing that distinguishes them — so the same ten digits pad into THREE DIFFERENT eleven-digit billing codes (0002735001 → 00002-7350-01 or 00027-0350-01 or 00027-3500-01). `readings` then holds all three. DO NOT pick one and present it as the answer: that is picking a different drug. Ask the user for the hyphenated form off the package or the FDA NDC Directory. Also: the eleven-digit form is a HIPAA BILLING format, not the NDC. The FDA-listed NDC is ten digits, and the padding is not reversible because a leading zero might be padding or might be part of the code. Returns { rows[{line,kind,cas?,ndc?,valid}], summary{total,valid,invalid,ambiguous} } — `ambiguous` counts rows needing a human.

Input parameters:

- `codes` (string, required): One code per line.
- `kind` (string, required): Which identifier these are.

### `validate_iban` (~573 tokens)

Validate IBANs and SWIFT/BIC codes

Check one IBAN per line against ISO 7064 mod-97 AND the fixed length and BBAN character layout of the country. STATE THIS WHENEVER YOU REPORT A RESULT: a passing IBAN is structurally consistent and NOTHING MORE. It does not mean the account exists, is open, belongs to anyone in particular, or can receive a payment. The single most common reason a person checks an IBAN is that an invoice arrived with changed bank details — in that situation a "valid" result is not reassurance, and you should say so and suggest confirming by phone on a previously known number. Never let a pass here read as clearance to pay. The length check is the useful part: every country has exactly one IBAN length, so a failure is reported as "Germany is 22 characters, this is 21" rather than "checksum failed", and a character-class failure names the position ("character 9 must be a digit"). Prefer relaying those specifics over a bare verdict. Do NOT substitute `expectedCheckDigits` into the IBAN to make it pass — that hides an error in the account number rather than fixing it. IBAN is not universal: for US, CA, AU, NZ, JP, CN, IN, ZA, MX, HK and SG the result is IBAN-K01 explaining that the country uses a different scheme and naming it. If a user asks for "the IBAN" of a US account, the answer is that no such thing exists — give them the ABA routing number, account number and SWIFT/BIC instead. BIC: pass `bic` instead to check a SWIFT/BIC. Two findings there are worth leading with. BIC-T01 — a "0" in the SECOND position of the location code marks a TEST BIC, not one for the live network; it looks exactly like a normal BIC and nothing reveals it until a payment fails. BIC-C02 — SWIFT assigned XK to Kosovo, which has no ISO 3166-1 code, so a Kosovan BIC is valid despite failing a naive "country must be ISO" check; do not tell a user their XK code is wrong. ISO 9362 is a paid standard and both iso.org and swift.com refuse automated requests, so the BIC structure cited is a public transcription rath…

Input parameters:

- `bic` (string): A SWIFT/BIC to check instead.
- `ibans` (string): One IBAN per line; spaces and lower case are fine.

### `validate_tax_ids` (~673 tokens)

Check EU VAT numbers and EORI numbers against their published formats and check digits

Check a column of EU VAT numbers or EORI numbers. Set `kind`. FORMAT AND CHECK DIGIT, NEVER REGISTRATION. There is no EU-wide VAT database — VIES asks each member state in turn — and this server makes no network requests. A pass means the value fits its country's published shape and, for most member states, its check digit is internally consistent; it does NOT mean the number is registered, belongs to that business, or is still active. Before a zero-rated invoice, the user must check VIES. Say this when you report a pass — for every result, not just failures. THE CHECK-DIGIT ALGORITHMS ARE NOT FROM OFFICIAL GOVERNMENT SOURCES, AND YOU MUST SAY SO WHEN RELAYING A RESULT. Most EU tax authorities do not publish their check-digit algorithm at all (the European Commission's own TIN compilation notes Greece's as "not publicly available" — typical, not exceptional). The algorithms here were cross-checked against python-stdnum (a mature, source-visible open-source library), then independently re-derived and verified against real worked examples by this project. That establishes the arithmetic is correct; it does not make it an official specification. Never tell a user a VAT-C02 pass means the number is "verified" or "confirmed" by the tax authority — say it is arithmetically consistent, sourced as above. CHECK-DIGIT COVERAGE IS PARTIAL BY DESIGN. Only the company/legal-entity branch is checked for countries whose scheme also covers individuals (personal-taxpayer numbers reuse a separate national ID checksum and are out of scope). `checksumValid` is `true`/`false` when checked, and `undefined` (VAT-C03) when this country or this branch has no implemented algorithm — undefined is NOT a pass, it means "not further checked", and you must not present it as validated. THE TWO HIGH-VALUE FORMAT FINDINGS ARE PREFIX ERRORS. Greece uses EL, not the ISO code GR — every other member state uses its ISO code, so GR is the natural thing to type and is always wrong; the finding hands back…

Input parameters:

- `kind` (string, required): Which identifier these are.
- `numbers` (string, required): One number per line.

### `lint_bullet_points` (~457 tokens)

Check listing bullet points and screen them for risky claim language

Lint listing bullets (one per line) for length, all-caps, HTML, off-platform contact details, promotional wording and repeated openings — and screen the copy for the claim language that appears in FDA and FTC enforcement. TWO BULLET FINDINGS MATTER MORE THAN THE REST. BUL-P01 (an email address, phone number or URL) is enforced against the SELLING ACCOUNT, not the listing — treat it as urgent. BUL-C01 means the surplus bullets are not displayed AND not indexed, so copy that lives only there is invisible to search too. THE CLAIM SCREEN IS NOT A COMPLIANCE CHECK, and you must say so whenever you report its output. A hit is not a violation; a clean result is not clearance. Regulators read whole claims in context and no word list is exhaustive. Never tell a user their copy is compliant on the basis of this tool, and never give legal advice about it. Each claim hit carries a `why`. RELAY THE WHY, not just the word. A tool that only flags "cure" teaches people to write "heal", which changes nothing about the risk and makes the copy worse. The disease rule exists because the FDA states only a drug may claim to "diagnose, treat, cure or prevent any disease" — those four verbs are the regulator's own list. "FDA approved" is flagged because the FDA does not approve supplements or cosmetics at all, so it is a false statement about a federal agency rather than an exaggeration. maxBullets and maxCharacters are CATEGORY-SPECIFIC and live behind marketplace logins. The defaults (5, 500) are common general values, not a specification — if the user knows their category limits, pass them. Returns { bullets[], lengths[], claims[{text,bullet,start,end,category,why}], findings[] }; offsets are code points.

Input parameters:

- `bullets` (string, required): One bullet per line; leading -, * or • are stripped.
- `maxBullets` (integer): Bullets the category displays (default 5).
- `maxCharacters` (integer): Character limit per bullet (default 500).
- `skipClaims` (boolean): Turn off claim screening for categories where it does not apply.

### `audit_column` (~345 tokens)

Find duplicate and inconsistently written values in a column

Audit one column (one value per line, or a CSV first column) for values that are the same thing. Answers in THREE degrees, and the distinction is the point: (1) `identical` — the same string on several rows; in a SKU or id column this is a hard error because the second row either overwrites the first or is rejected. (2) `presentation` — differ only in case, whitespace, full-width characters or invisible characters. The merge is unambiguous, and `normalise_column` applies it. Canonical form is the most-used spelling, or the one written first if they tie. (3) `synonym` — L and Large, Grey and Gray. THESE ARE NEVER MERGED AUTOMATICALLY and you must not merge them on the user's behalf either: in some categories the short and long forms genuinely differ, and applying the merge is a merchandising decision. Present them and ask. The synonym table is deliberately conservative — it holds only pairs that are the same word, and specifically does NOT contain Small/XS, which genuinely differ. Google's size and colour values are RECOMMENDATIONS, not requirements, so word synonym findings as observations, never as errors. Returns { total, distinct, groups[{canonical,variants[{value,lines[]}],basis,collapses}], exactDuplicateLines[], findings[] }, groups sorted by how many distinct values merging would remove.

Input parameters:

- `column` (string, required): One value per line.
- `normalise` (boolean): Also return the column with the unambiguous presentation merges applied.
- `skipSynonyms` (boolean): Turn off synonym suggestions — do this for SKU and id columns.

### `calculate_dim_weight` (~447 tokens)

Dimensional weight and the break-even box size

Compute dimensional (volumetric) weight across divisors and, more usefully, the BREAK-EVEN box size at which volume stops setting the price. LEAD WITH THE BREAK-EVEN, not the dimensional weight. A carrier calculator already answers "what does this parcel weigh dimensionally"; what a person packing goods needs is "how small does the box have to be" — `breakEvenScale` is the factor on the current linear dimensions, so multiply the longest side by it to get an actionable target. Carriers round EACH SIDE up to a whole unit before multiplying, then round the result up. `dimensionalRounded` does that; `dimensionalPlain` is the multiply-then-round-once figure most online calculators use, and it is smaller. Quote the rounded one. THE DIVISOR IS NOT VERIFIED DATA. It varies by carrier, service, destination and contract, and negotiated accounts often have their own. This project could not verify the published values by automated request on 2026-08-23 (UPS refused the connection, FedEx serves the figure only via JavaScript, DHL returned 403). Present presets as a starting point and ask the user for the divisor on their rate card — pass it as `customDivisor` and only that is used. Never present the output as a shipping quote. Divisors are unit-specific: 139 and 166 are inches/pounds, 5000 and 6000 are centimetres/kilograms. Mismatched presets are SKIPPED rather than computed, because feeding centimetres into 139 gives an answer wrong by an order of magnitude that looks entirely plausible. Returns { readings[{preset,dimensionalRounded,dimensionalPlain,billable,volumetric,breakEvenScale}], findings[] }.

Input parameters:

- `actualWeight` (number, required)
- `customDivisor` (number): The divisor from the user’s own rate card; overrides all presets.
- `dimUnit` (string, required): Unit the dimensions are in.
- `height` (number, required)
- `length` (number, required)
- `massUnit` (string, required): Unit the weight is in.
- `width` (number, required)

### `check_compliance_statement` (~433 tokens)

Check Prop 65, allergen and CE/UKCA/FCC statement wording

Check the WORDING of a statement that law prescribes, element by element. Set `kind`. NOT LEGAL ADVICE, and it cannot tell you whether a statement is REQUIRED — that depends on formulation, exposure and jurisdiction, none of which text reveals. Say this when you report results. THE THREE KINDS DIFFER IN HOW CERTAIN THEY ARE, so report them differently: prop65 — the regulation (27 CCR §25603) fixes near-verbatim text, so all four elements (signal word, named chemical, the harm it is listed for, www.P65Warnings.ca.gov) are checked separately. Report WHICH element is missing: dropping the URL and dropping the chemical name are different mistakes. The yellow triangle symbol is a GRAPHIC requirement that no text check can see — mention it, but never fail text for it. allergen — the nine major US allergens are statutory (milk, eggs, fish, Crustacean shellfish, tree nuts, peanuts, wheat, soybeans, sesame), but whether a product must declare one is not checkable here. Watch for ALG-S01: sesame became the ninth allergen after most label templates were written, and those templates still circulate — this is the most common way a careful allergen statement is now out of date. Also: "may contain" is a VOLUNTARY advisory and does not substitute for the Contains declaration. marking — CE, UKCA and FCC are NOT text problems. They are declarations that conformity assessment was completed and a technical file exists. No wording change can create one, and claiming a marking the product does not hold is an offence. Never help a user reword their way into a marking claim; ask for the directive, standard or FCC ID instead. CE does not imply UKCA or the reverse. Returns { kind, elements[{name,present,found?,expected}], complete, findings[] }. `expected` is copy-ready text for a missing element.

Input parameters:

- `kind` (string, required): Which kind of statement this is.
- `statement` (string, required): The statement text as it appears on the listing or label.

### `check_open_graph` (~273 tokens)

Check Open Graph meta tags

Check Open Graph markup pasted as HTML. Reports the four required properties (og:title, og:type, og:image, og:url), the attribute each tag used, and the problems that a browser never reveals. THE FINDING THAT MATTERS MOST IS OG-A01: the protocol requires `property=`, and almost every other meta tag uses `name=`, so `name="og:title"` is the natural thing to write. The page looks identical, naive validators searching for the string "og:title" find it, and the share card comes out blank. This tool therefore counts a name= tag as ABSENT, exactly as a scraper does — report it that way rather than saying the tag is present but wrong. OG-U01 is the other invisible one: a relative og:image resolves fine in a browser (which has the page URL) and not at all for a scraper (which does not). THIS TOOL FETCHES NOTHING. It does not request og:image or og:url, so it cannot say whether the image exists, what size it is, or whether the URL 404s. Do not imply otherwise. Returns { tags[{property,content,attribute}], required[{property,present,content?}], valid, findings[] }.

Input parameters:

- `html` (string, required): The meta tags, or a whole <head>.

### `check_email_template` (~307 tokens)

Check email merge fields before a send

Find merge fields in an email template and report the three failures that get noticed after the send. TPL-U01 (unclosed field) is the expensive one: the fragment goes out literally to every recipient, and in an editor it looks nearly right. TPL-X01 means two merge syntaxes are present — Mailchimp *|FIELD|* alongside {{field}}, say — which is almost always a block pasted from a template built on another platform, and one of the two will not be replaced. TPL-F01 is NOT a syntax error: "Hi {{first_name}}," is perfectly valid and renders as "Hi ," for every recipient with a blank record, and there are always some. When you report it, show the user what the empty case reads like rather than quoting the rule. Recognises *|FIELD|*, {{ field }}, %%FIELD%%, ${field} and [FIELD], reads fallbacks (*|FNAME:there|*, {{ x | default: "y" }}), and counts delimiter balance PER SYNTAX so one platform’s markers are not mistaken for another’s unclosed field. These are this project’s own engineering rules — no email platform publishes a formal grammar for merge fields. Say so rather than presenting them as a standard. Returns { variables[{raw,name,syntax,start,end,fallback?}], syntaxes[], unclosed[], findings[] }; offsets are code points.

Input parameters:

- `template` (string, required): The email template text or HTML.

### `check_textile_label` (~407 tokens)

Check garment fibre content and care wording

Check a fibre content declaration against 16 CFR 303 and a care instruction against 16 CFR 423. THE RULE MOST TOOLS MISS IS TEX-P02: a fibre under 5 percent of total weight must be disclosed as "other fiber", NOT by name. Adding to 100 is not sufficient — "97% cotton, 3% spandex" totals correctly and is still a violation. There IS an exception: a small fibre may be named if it has a clearly established and definite functional significance at that amount (the regulation's own example is 96% acetate / 4% spandex). That is a fact about the product, not the text, so ASK THE USER and pass `functionalFibres` — never assume it. TEX-N01: fibres must carry the generic name. Lycra is spandex, Tencel is lyocell, viscose is rayon. A trade name beside the generic one is fine; instead of it is a violation. CARE SYMBOLS ARE DELIBERATELY NOT COVERED and you should say so if asked: ISO 3758 is a paid standard and the symbols are GINETEX trade marks in many countries, so this project neither reproduces nor translates them. What is checked is the WORDS 16 CFR 423 requires — cleaning method, water temperature where washing is instructed, drying, bleaching, ironing. Do not fill the gap by reciting a symbol chart from memory. Returns { fibres[{raw,percent,name,generic,other}], total, careCovered[{aspect,present,found?}], findings[] }.

Input parameters:

- `careInstruction` (string, required): The care instruction line; pass an empty string to skip.
- `fibreContent` (string, required): The fibre content line, e.g. "60% cotton, 40% polyester".
- `functionalFibres` (array): Fibres under 5% that genuinely have a function at that amount — ask the user, do not assume.

### `audit_media_specs` (~368 tokens)

Audit a batch of images and videos against marketplace limits

Given file metadata (name, pixel width and height, byte size, format, kind, optional duration and alpha flag), report which files fail on which platform. THE OUTPUT IS A MATRIX and should be relayed as one: the useful answer before a bulk upload is "6 of these 40 fail on eBay only", not a per-file verdict. `summary` gives the failing count per platform. The same file legitimately passes one marketplace and fails another — a 12.5 MB photo is fine for Google Shopping and over eBay’s limit; a WebP is accepted by Google and eBay and not by Amazon. Do not collapse this into a single pass/fail. METADATA ONLY: this does not look at pixels, so it says nothing about background, crop or how much of the frame the product fills — that is a different tool (check_main_image). VIDEO: a browser can read duration and frame size but NOT the codec, and most platform video specs require a particular codec. Every video result carries MED-V02 saying so; repeat that caveat rather than letting a pass read as "ready to upload". Duration limits vary by platform and programme, so `maxVideoSeconds` is an input — ask the user. Amazon’s values are the general marketplace ones (Seller Central is login-gated and requirements vary by category); Google and eBay values come from their published pages. Returns { files[{file,verdicts[{platform,passes,reasons[],notes[]}],passesAll,findings[]}], summary[{platform,failing,total}], findings[] }.

Input parameters:

- `files` (array, required): Metadata for each file.
- `maxVideoSeconds` (number): Duration limit for the target placement.
- `platforms` (array): Limit to these platforms; omit for all.

### `convert_shoe_size` (~334 tokens)

Convert shoe sizes from the unit definitions

Convert between EU, UK, US men’s, US women’s and foot length in centimetres. Each reading carries the arithmetic that produced it. LEAD WITH `footLengthCm`, not the converted size. Foot length is the only quantity that does not change between brands, and it is what the user should measure and match against the brand’s own chart. A size number is an inference from it. The unit arithmetic is exact — a Paris point is 2/3 cm, a barleycorn is 1/3 inch — but the last allowance and the US/UK offsets are trade convention, and brands deviate by more than the conversion moves. Always pass on SHO-B01 rather than presenting a converted size as definitive. US men’s and US women’s are DIFFERENT SCALES about a size and a half apart; if the user says only "US 8", ask which. APPAREL SIZES ARE NOT SUPPORTED, deliberately. There is no derivable relationship between S/M/L and any national numeric apparel scale — only each brand’s block. If asked to convert clothing sizes, call `explain_apparel_sizing` or relay the same point: DO NOT produce a conversion chart from memory. Those charts are one or two brands’ tables presented as universal, and the returns land on the seller. Tell them to publish the garment’s actual measurements instead. Returns { input, footLengthCm?, readings[{system,label,value,working}], findings[] }.

Input parameters:

- `system` (string, required): Which scale the value is in.
- `value` (number, required): The size to convert from.

### `explain_apparel_sizing` (~123 tokens)

Explain why apparel sizes cannot be converted

Returns the reasoned position on cross-brand apparel size conversion, for when a user asks for a clothing size chart. Use this INSTEAD OF generating a conversion table. S/M/L and national numeric apparel scales have no derivable relationship; the universal charts online are individual brands’ tables presented as general truth. Producing one from memory states something uncertain in a confident voice and the returns are paid for by the seller. The constructive answer is in the finding’s `fix`: publish the garment’s actual measurements beside the size label. Returns { findings[] }.

### `lookup_un_number` (~392 tokens)

Look up dangerous goods entries by name or UN number

Search the Hazardous Materials Table by product name (perfumery, lithium ion, aerosol) or by UN/NA number. Name search is the primary path — people have a product, not a number. THIS IS A LOOKUP, NOT A CLASSIFICATION, and you must say so. Which entry a product falls under depends on its actual composition, concentration, packaging and mode of transport — a dangerous goods classification done by someone qualified from the product’s safety data sheet. Never tell a user their product "is UN####" on the basis of a word match. ONE NUMBER OFTEN COVERS SEVERAL ENTRIES with different packing groups and different per-package limits — UN1950 (aerosols) has separate flammable, non-flammable and corrosive entries. All of them are returned; DO NOT pick one. Getting the number right is not the same as getting the entry right, and filling the number in while taking the wrong limit is how a shipment is refused with the paperwork apparently complete. A "no match" result is NOT evidence the goods ship as ordinary cargo — many hazardous items are listed under a chemical name rather than the one on the box. Suggest the safety data sheet. UNN-F01 marks entries forbidden on passenger or cargo aircraft: that is not a paperwork question, the goods cannot travel that way. Rows come from 49 CFR 172.101, the US table that adopts the international list — the UN’s own publication and the US regulator’s site both refuse automated access. For international movements say the mode’s own regulations must be confirmed. Returns { query, matches[{id,number,name,hazardClass,packingGroup,labels,passengerLimit,cargoLimit}], matchedBy, findings[] }.

Input parameters:

- `limit` (integer): Maximum entries to return (default 25).
- `query` (string, required): A product name or a UN/NA number.

### `check_incoterm` (~429 tokens)

Check an Incoterms trade term against the transport mode

Check how a trade term is written and whether it is valid for the mode of transport. Pass the term as it appears in the contract and, if known, the mode. THE FINDING THAT MATTERS IS INC-M01. FAS, FOB, CFR and CIF are SEA AND INLAND WATERWAY rules whose risk transfer is "goods on board the vessel". For a container that moment does not happen as anyone imagines — the box is handed over at a terminal days earlier and the seller is still carrying risk while the buyer believes they took it. For air freight the moment does not exist. Both sides think they agreed a transfer point and did not agree the same one. Report this as a contract problem, not a style note. The answer for containers is FCA, or CPT/CIP if the seller pays carriage. If the mode is NOT stated, this check does not fire — do not assume a mode on the user’s behalf; ask them. Also flags: a term with no named place (incomplete — "FOB" alone does not say where risk passes); no edition stated (DAT became DPU, CIP insurance changed in 2020); retired terms (DAT, DDU, DAF, DES, DEQ); EXW, which leaves export clearance in the seller’s country to a buyer who often has no entity there; and DDP, which puts import clearance and duty on a seller who often has no registration in the destination. COPYRIGHT: the text of the Incoterms® rules and the buyer/seller obligation matrix are ICC copyright. This tool does not reproduce them and NEITHER SHOULD YOU — do not generate a who-pays-what table from memory in response to its output. Point the user to ICC for the rules themselves. Returns { input, code?, entry?, namedPlace?, version?, findings[] }.

Input parameters:

- `mode` (string): How the goods actually travel. Omit if unknown — do not guess.
- `term` (string, required): The term as written, e.g. "FOB Shanghai Incoterms 2020".

### `validate_marketplace_feed` (~343 tokens)

Validate a product feed for Meta, TikTok Shop, Walmart or Amazon

Check a tab- or comma-separated feed (first line is the header) against a destination platform’s profile. THE FINDING THAT JUSTIFIES THIS TOOL IS PFD-X01: a value that is CORRECT ON ANOTHER PLATFORM. Google Shopping writes in_stock with an underscore; Meta writes "in stock" with a space. A working Google feed copied to Meta has the availability field on every row and the value rejected on every row, and what the platform reports is usually just that the product is not live. When you see PFD-X01, lead with "this is the <other platform> spelling" — never send the user to check a spelling that was never wrong. PFD-D01 deserves emphasis too: Meta states that where a content ID repeats, ALL instances are ignored. A duplicate id does not overwrite — BOTH products disappear from the catalogue. ONLY THE META PROFILE IS FIELD-VERIFIED, from Meta’s publicly readable catalog fields reference. TikTok Shop, Walmart and Amazon carry ONLY constraints checkable without their documentation (unique id, core fields present, numeric price, absolute URLs) because their references are login-gated or change per release. Those runs emit PFD-N01 — relay it. A clean result on an unverified profile is NOT a statement that the feed will upload, and you must not fill the gap with a field list from memory. Returns { platform, profile, rows[{line,id?,findings[]}], summary{total,withErrors,duplicateIds}, findings[] }.

Input parameters:

- `feed` (string, required): The feed with a header row; tab or comma separated.
- `platform` (string, required): Destination platform.

### `check_invoice_sequence` (~365 tokens)

Find gaps and duplicates in a column of invoice numbers

Check a column of invoice references for missing numbers, duplicates, width drift and out-of-order rows. THE HARD PART IS SERIES DETECTION, not gap finding. Real references carry prefixes, year segments and branch codes (INV-2026-0001, 2026/001, FV/1/2026), and one column often holds SEVERAL independent sequences. This tool infers the pattern and splits them; counting several series as one is how a healthy set of books comes out looking like hundreds of gaps. When you report results, show the split so the user can confirm it. REPORT A GAP AND A DUPLICATE DIFFERENTLY. A gap is NOT automatically a violation — cancelling an invoice legitimately leaves a hole, and most jurisdictions want the void recorded rather than the number reused. Present gaps as a list to reconcile against cancelled documents. A DUPLICATE has no innocent reading: two documents claim the same identity and any reference to that number is ambiguous. That one is an error. DO NOT TELL THE USER WHAT THEIR JURISDICTION REQUIRES. Numbering rules differ by country — some mandate an unbroken sequence, some only that gaps be explainable, and even the definition of a series varies. This tool reports what is in the column; that is all it can support. Returns { total, unparsed[{raw,line}], series[{key,example,count,min,max,gaps[{from,to}],missing,duplicates[{number,lines[]}],widths[],outOfOrder[]}], findings[] }. Gaps are collapsed into ranges.

Input parameters:

- `maxGapRanges` (integer): Cap on gap ranges listed per series (default 20).
- `references` (string, required): One invoice reference per line, or a CSV first column.

### `check_peppol_invoice` (~349 tokens)

Check a UBL invoice against the published Peppol BIS Billing 3.0 rules

Validate a Peppol BIS Billing 3.0 UBL invoice supplied as XML. FINDINGS CARRY THE OFFICIAL PEPPOL RULE IDENTIFIERS (BR-01, UBL-CR-002). Relay them verbatim — the whole point is that the user can open the Peppol specification and read the rule themselves. Do NOT paraphrase a rule into your own numbering or wording. REPORT THE SCOPE EVERY TIME. Only a subset of the published rules is evaluated (the result gives `evaluated` and `total`); the rest are conditional rules, monetary total calculations and code list checks that need the full semantic model. A clean result is NOT a full Peppol validation and you must say so — money depends on this answer. CHECK THE IDENTITY FIRST. The most common real-world rejection is not a business rule but a wrong cbc:CustomizationID: the document is well formed and declares a specification the receiver is not running. `identity.isPeppolBis3` answers that; lead with it. CII documents are identified and declined rather than checked badly — the rules here are the UBL ones. EN 16931 itself is a paid CEN standard and is not quoted; what is cited is Peppol's public restatement. Do not fill the gap with EN 16931 text from memory. Returns { identity{rootElement,kind,syntax,customizationId?,profileId?,isPeppolBis3}, evaluated, total, violations[{rule,detail}], findings[] }.

Input parameters:

- `limit` (integer): Cap on violations returned (default 200).
- `xml` (string, required): The invoice XML.

### `convert_feed` (~397 tokens)

Convert a product feed between Google Shopping and Meta

Convert a tab- or comma-separated feed (first line is the header) between Google Shopping and Meta catalog dialects. THE POINT OF THIS TOOL IS WHAT IT REFUSES TO DO. Google supports four availability values (in_stock, out_of_stock, preorder, backorder); Meta supports two. `preorder` and `backorder` have NO equivalent, and neither substitute is free: "in stock" advertises goods the seller does not have, "out of stock" refuses an order they would have taken. Those cells are left EMPTY and listed in `rows[].undecided`. DO NOT FILL THEM IN. Not in the output, not in your reply, not "reasonably". It is a merchandising decision that belongs to the seller, and a converter that picks one looks more finished while having decided something on their behalf. Present the trade-off and ask. CNV-T01 covers the values that DO map but are spelled differently (in_stock with an underscore on Google, "in stock" with a space on Meta) — that difference is what silently breaks a reused feed. Only Google and Meta are supported, because both publish their value sets openly. Other marketplaces keep theirs behind a seller login or change them per release; converting to those would mean guessing, and a wrong conversion is worse than none because it looks finished. Do not extend this to another platform from memory. This CONVERTS, it does not VALIDATE — run the result through validate_marketplace_feed afterwards. Returns { from, to, header[], rows[{line,values,undecided[{field,from,guidance}]}], needingDecision, findings[] } plus `output` when `render` is set.

Input parameters:

- `feed` (string, required): The feed with a header row.
- `from` (string, required): Source dialect.
- `render` (boolean): Also return the converted feed as text.
- `to` (string, required): Target dialect.

### `repair_feed` (~441 tokens)

Repair a damaged product feed

Repair a tab- or comma-separated feed (first line is the header) in one pass: mojibake, spreadsheet damage, invisible characters, stray whitespace and inconsistent capitalisation. IT ONLY APPLIES REPAIRS THAT HAVE EXACTLY ONE POSSIBLE ANSWER. Mojibake is reversed only when re-encoding the repaired text reproduces the original bytes — several encoding chains produce plausible output and only one round-trips. TWO LISTS COME BACK AND THE SECOND MATTERS AS MUCH AS THE FIRST. `repairs` is what changed; `leftAlone` is every cell it refused to guess at, with the reason and, where they exist, the competing candidates. Relay `leftAlone` to the user — do NOT resolve those cells yourself. A cell a spreadsheet turned into a date could have come from several originals, and choosing needs to know what the product is. UNRECOVERABLE IS NOT THE SAME AS FINE. Scientific notation (8.71E+12) destroys digits permanently; those cells appear in `leftAlone` as needing the source system, not in `repairs`. Never reconstruct the missing digits — a barcode you completed from memory is worse than a visibly broken one. IT WILL NOT MERGE SYNONYMS. L and Large stay separate: whether they mean the same thing is a merchandising decision. Use audit_column to show the seller those sets. Capitalisation IS unified within a column (presentation, never which value a cell holds); pass skipCaseUnification for columns where case carries meaning, such as SKUs. REPAIRING IS NOT VALIDATING — run the result through validate_marketplace_feed afterwards. Returns { header[], rows[], repairs[{row,column,before,after,kind,because}], leftAlone[{row,column,value,because,candidates?}], findings[] } plus `output` when `render` is set.

Input parameters:

- `feed` (string, required): The feed with a header row.
- `render` (boolean): Also return the repaired feed as text.
- `skipCaseUnification` (boolean): Leave capitalisation alone (use for columns where case is meaningful, like SKUs).

### `check_hs_code` (~509 tokens)

Check an HS / HTS / CN commodity code

Check whether a tariff code exists, what its heading officially covers, and whether it is the right length for a destination. IT DOES NOT CLASSIFY GOODS AND NEITHER SHOULD YOU. It cannot tell the user which code their product needs — that depends on material and use and is decided by the General Rules of Interpretation. Do NOT supply a code from your own knowledge when this tool says one is wrong; send the user to their customs broker or the destination tariff search. A confidently wrong tariff code is a penalty, not a typo. EXISTENCE IS CHECKED AGAINST REAL DATA — every published heading and subheading, extracted from the US tariff (the six-digit level is identical in every WCO country by treaty). `HS-H01`/`HS-S01` mean the code is wrong in every country, not just one. RELAY `headingDescription` TO THE USER. Confirming that the official heading text matches the goods is the cheapest check available, and a well-formed code for the wrong goods passes every format validator. LEADING ZEROS: chapters 01–09 (animals, meat, fish, dairy, vegetables, fruit, coffee, tea, spices) lose their zero in spreadsheets. An odd digit count is that, and `HS-X01` gives the restored code. DIGIT COUNTS ARE ONLY GIVEN FOR DESTINATIONS WITH A CITED SOURCE (international 6, United States 10, EU 8). Other countries use other lengths; do NOT fill them in from memory — say the requirement is not covered here. Chapters 98/99 are one country’s own special provisions, so they are reported as scoped, not as invalid. Chapter 77 is empty by design. Returns { input, normalized?, formatted?, valid, parts?{chapter,heading,subheading?,national?}, headingDescription?, findings[] }. BULK: pass `values` (an array, max 1000) instead of the single value to check a whole column in one call. The result is then { rows, summary } rather than a single result. Do not loop this tool over a column one value at a time.

Input parameters:

- `code` (string): The commodity code. Dots and spaces are ignored.
- `destination` (string): Check the digit count for this destination. Omit if unknown — do not assume.
- `values` (array): Bulk form: one value per array element. Use this instead of looping the single-value form.

### `check_origin_declaration` (~438 tokens)

Check a preferential origin declaration against the prescribed wording

Compare the origin declaration on a commercial invoice against the text EU Regulation 2015/2447 prescribes, in every prescribed language version. THE WORDING IS FIXED BY LAW AND THIS CHECK IS LITERAL. A sentence that means the same thing is not the same thing: customs in the importing country can refuse the preference on the wording alone, and the duty then falls on the buyer. Do NOT reassure the user that a reworded declaration is fine because the meaning is clear. THREE DECLARATIONS HAVE NEARLY IDENTICAL OPENINGS and it is the ending that says which arrangement is being claimed: Annex 22-13 (general) stops at "preferential origin"; Annex 22-09 (GSP) continues with the Generalised System of Preferences clause; Annex 22-07 (REX statement on origin) must also state the origin criterion. `ORG-X01` means two of them have been spliced together. DELETING THE BRACKETED AUTHORISATION NUMBER IS CORRECT for an exporter who is not an approved exporter — the regulation requires it. Do not tell the user to put it back. What is wrong is leaving a number copied from someone else’s template (`ORG-A01` reports the number so they can confirm it is theirs). IT DOES NOT DECIDE WHETHER THE GOODS QUALIFY. Preferential origin depends on the rules of origin, the processing done and the materials used. Never answer that question from the wording, and never draft an origin claim for goods you have not been told qualify. SCOPE: only the Union Customs Code declarations. Individual free trade agreements — the EU–UK agreement among them — prescribe their own wording in their own annexes; say so rather than checking those against this text. Returns { input, matched?{annex,label,lang,code,prescribed}, score, diff[{op,text,slot?}], slots[{index,role,content,placeholder}], findings[] }. `diff` is word level: op is same | missing | extra | slot.

Input parameters:

- `text` (string, required): The declaration text as it appears on the invoice.

### `check_supplier_declaration` (~494 tokens)

Pick and check an EU supplier's declaration

Choose the right one of the four supplier declarations EU Regulation 2015/2447 prescribes, and check a long-term declaration's validity period. THE FOUR FORMS ARE A 2x2: preferential origin status or not, one-off or long-term. Annex 22-15 / 22-16 / 22-17 / 22-18. Using a preferential form for goods containing non-originating materials is the common and costly error, because the manufacturer downstream builds their own origin claims on it. THE TWO LONG-TERM FORMS ARE WORDED DIFFERENTLY AND YOU MUST NOT FLATTEN THEM. Annex 22-16 says the period "shall not exceed 24 months or 12 months if the declaration was issued retrospectively". Annex 22-18 says only that it "should not exceed 24 months" — softer wording, and NO retrospective rule at all. Never apply the 12-month retrospective limit to 22-18, and never state it as a general rule for long-term supplier declarations; `SUP-R01` reports that silence deliberately. The verbatim footnote comes back in `form.periodNote` — quote that rather than paraphrasing. IT DOES NOT DECIDE WHETHER THE GOODS QUALIFY for preferential origin. That is an input, from the rules of origin for the arrangement. Do not infer origin status from the form choice, and do not tell a user their goods qualify. Dates are YYYY-MM-DD and are validated as real calendar dates, never rounded. Exactly 24 months is allowed — the annex says not to EXCEED it. Returns { form{annex,key,title,body,notes[],periodNote,binding}, period?{from,to,months,limitMonths,fromRetrospective,latestAllowed,withinLimit}, shipmentCovered?, findings[] }.

Input parameters:

- `issuedOn` (string): YYYY-MM-DD. Earlier than validFrom means the declaration is retrospective.
- `longTerm` (boolean, required): Is this a long-term declaration covering regular supplies?
- `preferential` (boolean, required): Do the goods already have preferential origin status? This is an input, not something the tool determines.
- `shipmentDate` (string): YYYY-MM-DD. Checks whether that shipment falls inside the validity period.
- `validFrom` (string): Long-term only. YYYY-MM-DD.
- `validTo` (string): Long-term only. YYYY-MM-DD.

### `check_lei` (~446 tokens)

Check the structure and check digit of an LEI

Check a Legal Entity Identifier (ISO 17442) for length, character set, the reserved pair at positions five and six, and the modulo-97 check digit. WELL FORMED IS NOT REGISTERED, AND YOU MUST SAY SO. A passing check digit means the twenty characters are internally consistent. It does NOT mean the code was ever issued, does not say which entity holds it, and does not say whether the registration lapsed — a dead LEI passes exactly as a live one does. Never tell a user their LEI is "valid" without that qualification, and never name the entity behind an LEI from your own knowledge. Send them to the GLEIF search. THE CHECK DIGIT IS WEAKER THAN IT LOOKS: over seven thousand single-character mutations of real codes, about three in a thousand still passed. CHECK DIGITS RUN 02 TO 98. A code ending 00, 01 or 99 was not produced by the rule at all (`LEI-D01`), which usually means the check digits were never calculated. SOURCING, STATED HONESTLY: ISO 17442 is a paid standard and is not quoted. GLEIF's own code-structure page could not be reached anonymously (two URLs returned an identical generic page). The structure rules here are MEASURED against 800 real codes from GLEIF's public API across 16 countries and 16 issuing organisations. If you relay these rules, relay that basis too rather than implying a specification citation. Returns { input, normalized?, valid, parts?[{label,value,meaning}], expectedCheckDigits?, findings[] }. BULK: pass `values` (an array, max 1000) instead of the single value to check a whole column in one call. The result is then { rows, summary } rather than a single result. Do not loop this tool over a column one value at a time.

Input parameters:

- `code` (string): The LEI. Spaces, hyphens and lower case are normalised away.
- `values` (array): Bulk form: one value per array element. Use this instead of looping the single-value form.

### `identify_einvoice` (~410 tokens)

Identify an e-invoice XML: syntax, kind, profile, and which validator applies

Answer "what is this file" for an EN 16931 e-invoice before validating it: is it UBL or CII syntax, an invoice or credit note, and which profile does its CustomizationID declare — Peppol BIS 3.0, XRechnung (with version), the EN 16931 core, or something unrecognised. THIS IS THE FRONT DOOR, NOT A VALIDATOR. It does not check business rules; it identifies the document and names which validator to send it to. `profile.routesTo` gives the slug of the matching validator on this site when one exists. THE CUSTOMIZATION ID IS THE KEY FIELD. The most common real-world rejection is not a broken rule but a document sent to a receiver running a different profile from the one it declares. `IDD-C02` names the recognised profile; `IDD-C03` means the declared profile is not one recognised here (a national profile not yet covered, a newer version, or a typo). THE CUSTOMIZATION IDs ARE VERIFIED VALUES, NOT GUESSES. Peppol from the Peppol spec; XRechnung extracted from KoSIT's public Schematron (github.com/itplr-kosit/xrechnung-schematron). XRechnung is matched as a family so both the current xeinkauf.de and the legacy xoev-de namespaces are recognised, and the version is extracted. Do not add or infer other profile URNs from memory. CII is IDENTIFIED but the invoice validators here read UBL (IDD-M01). A document that is not well formed is reported with the position (IDD-P01) and still identified as far as it parsed. Returns { input, wellFormed, syntax, kind, rootElement?, namespaceConfirms?, customizationId?, profileId?, profile?{id,label,routesTo?,version?}, findings[] }.

Input parameters:

- `xml` (string, required): The e-invoice XML (UBL or CII).

### `check_creditor_reference` (~356 tokens)

Check an ISO 11649 RF creditor reference

Validate an RF creditor reference — the structured payment reference on European invoices — by its ISO 7064 MOD 97-10 check digit, the same checksum IBAN uses. A VALID REFERENCE IS NOT A MATCHED PAYMENT. The check digit proves the reference was copied without a slip; it does not mean the reference exists in a creditor's system or that a payment will reconcile. Never tell a user their payment will be matched — this is a structural check, not a lookup. IT IS NOT AN IBAN. RF references and IBANs share the checksum and both open with two letters plus two check digits, which is exactly why they get confused. An RF reference identifies a payment, starts with the literal letters RF and is at most 25 characters; do not route it to IBAN validation. SOURCING: ISO 11649 is a paid standard and is not reproduced; the checksum is public ISO 7064 MOD 97-10, self-proved against the standard's published example RF18 5390 0754 7034. When the check digits are wrong, `expectedCheckDigits` gives the correct pair — but relay the advice to re-read the source, since the error may be in the reference body, not the check digits. BULK: pass `values` (an array, max 1000) instead of `value` to check a column in one call; the result is then { rows, summary }. Returns { input, normalized?, valid, reference?, expectedCheckDigits?, findings[] }.

Input parameters:

- `value` (string): The RF reference. Spaces and lower case are normalised away.
- `values` (array): Bulk form: one reference per array element.

### `check_xrechnung` (~447 tokens)

Check a German XRechnung invoice against KoSIT's official BR-DE rules

Check an XRechnung invoice (UBL or CII) against the German CIUS additions to EN 16931. THESE ARE NOT PEPPOL'S RULES. XRechnung (BR-DE-*) and Peppol BIS Billing (BR-*, UBL-CR-*, PEPPOL-*) are siblings, each adding different rules on top of the same EN 16931 core. A document can pass one and fail the other. If unsure which applies, use identify_einvoice first — its `profile.id` tells you. ONLY UNCONDITIONAL RULES ARE EVALUATED: 13 of 56 for UBL, 12 for CII. Most BR-DE rules are conditional on other field values (VAT code branching, payment method, computed amounts) and are NOT evaluated — evaluating them wrong would be worse than not evaluating them, so they are catalogued (id + official German text) but not checked. `evaluated`/`total` says exactly how far this went; a clean result is not a full XRechnung validation. THE OFFICIAL RULE TEXT IS GERMAN ONLY — KoSIT has never published an English translation. Finding messages include an English description written and checked by this project (NOT an official translation), the verbatim German original, and the official EN 16931 English business term embedded in it (e.g. "Seller city" (BT-37) — that quoted term is KoSIT's own citation, not our translation). Do not present the English description as KoSIT's own wording. RULES ARE CONDITIONAL ON OPTIONAL GROUPS WHERE THE GERMAN TEXT SAYS SO (e.g. delivery-address rules only apply if a delivery address is given at all — most domestic invoices have none and that is correct, not a violation). Never tell a user a rule was violated when the underlying business group is simply absent; the checker already accounts for this. Returns { identity{syntax,kind,rootElement,customizationId?}, evaluated, total, violations[{rule,detail}], findings[] }.

Input parameters:

- `xml` (string, required): The invoice XML, UBL or CII.

### `convert_csv_to_xlsx` (~379 tokens)

Convert a CSV file to .xlsx without destroying identifiers

Convert a local CSV/TSV file to .xlsx the safe way: columns that hold identifiers are auto-detected and written as TEXT cells so nothing gets reinterpreted — leading zeros survive, 12+ digit codes never become 8.71E+12, GTIN columns (recognised by check digit) stay verbatim, phone numbers keep their +, date-like text is not silently committed to one reading, and formula-looking values (=, @) become inert text. The protections are published as 7 rules (PROT-*); the result reports exactly which columns were protected and why, plus per-value fallbacks in unprotected columns. Values that are plain safe numbers (prices, quantities) stay real numbers, so the sheet still calculates. The source is also pre-checked against known Excel damage signatures: cells already mangled by a previous spreadsheet round-trip (scientific notation corpses, suspected lost zeros) are WARNED about, never "repaired" — this tool refuses to invent digits (XLC-W01). Dialect (comma/semicolon/tab/pipe, quoting, embedded newlines) is detected automatically. Encoding: UTF-8, with Windows-1252 fallback (reported when used). GUARANTEES: the input file is never modified; an existing output file is never overwritten. Returns { outcome, path, out?, encoding, report } where report lists protections [{header, ruleId, reason, examples}], cellCounts, damage precheck and findings.

Input parameters:

- `out` (string): Absolute output path (.xlsx). Default: alongside the input as <name>.xlsx. Never overwrites.
- `path` (string, required): Absolute path to the CSV/TSV file.
- `sheetName` (string): Worksheet name (default "Sheet1"; illegal characters are replaced, max 31 chars).

### `convert_xlsx_to_csv` (~504 tokens)

Convert an .xlsx workbook to CSV without silent data changes

Convert a local .xlsx file to CSV, one file per worksheet, with every liberty disclosed instead of taken silently: numbers are NEVER emitted in scientific notation (exponent-stored values are expanded exactly by string arithmetic, not floating point); date cells become unambiguous ISO 8601 text using the epoch the workbook itself declares (including the legacy Mac 1904 system — the same serial is 1,462 days apart between epochs); formulas contribute their last calculated value and are counted (XLC-R03); merged ranges flatten with the value in the top-left cell (XLC-R04); embedded images/charts are counted as not-carried-over (XLC-R06); leading/trailing spaces in cells survive verbatim. Every disclosure is a numbered rule (9 XLC-* rules published). If the output still contains values a spreadsheet would damage on double-click (leading zeros, 12+ digits), that risk is flagged (XLC-W08) — the CSV is correct, the danger is in how it gets opened next. Legacy .xls and password-protected workbooks are refused with a clear explanation (XLC-E07), not mis-parsed. GUARANTEES: the input file is never modified; existing output files are never overwritten (the call fails before writing anything if any target exists). Defaults follow Excel-survival practice: comma, minimal quoting, CRLF, WITH a UTF-8 BOM (double-clicked no-BOM UTF-8 is mojibake in Excel); all four are overridable. Returns { outcome, path, files[{sheet, out, rows, columns}], report } with findings, epoch, counts and any structural problems.

Input parameters:

- `bom` (boolean): Write a UTF-8 BOM (default true — Excel double-click reads BOM-less UTF-8 as ANSI and mangles every non-ASCII character; some Unix tools dislike the BOM, hence the switch).
- `delimiter` (string): Output delimiter, default comma.
- `lineEnding` (string): Line ending, default CRLF (the Excel-ecosystem default).
- `outDir` (string): Absolute directory for the CSV output (default: the input's directory). One sheet → <name>.csv; several → <name>-<sheet>.csv each.
- `path` (string, required): Absolute path to the .xlsx file.
- `quoting` (string): minimal (default) quotes only fields that need it; all quotes every field.

### `generate_digital_link_qr` (~554 tokens)

Generate a Sunrise-2027-ready GS1 Digital Link QR code (validated)

Build a conformant GS1 Digital Link URI for a GTIN (optionally with CPV (22), batch/lot (10) and serial (21) qualifiers) and render it as a QR code — with a hard gate: if the GTIN fails the full validation engine, or a qualifier value falls outside the standard's 82-character set, NOTHING is generated (8 published QRG-* rules explain every refusal). A printed QR cannot be patched, so this tool refuses rather than warns. Conformance details handled for you: GTIN padded to 14 digits (required since URI syntax v1.4.0), qualifiers emitted in the mandatory 22→10→21 order, reserved characters percent-encoded per the standard's own table, verified round-trip against this project's Digital Link parser. The result includes print-size advice for retail POS scanning derived from GS1's published X-dimension range (0.396–0.990 mm) and the (modules + 8) × X formula — the quiet zone is included in the numbers and in the SVG. Error correction level M is this project's default (GS1 mandates none). Sunrise 2027 context, honestly stated: it is an industry capability milestone (retail POS able to scan 2D by end of 2027), not a legal mandate; during the transition GS1 US describes on-pack marking as mandatory 1D plus optional 2D. For a quick unvalidated URI (no QR, no gate) use build_gs1_digital_link instead. Returns { uri?, conformant, gtin14?, checks{gtin,qualifiers,domain}, findings[], qr?{moduleCount, modules[] (rows of 1/0)}, sizeAdvice?{minMm,exampleMm,maxMm}, svg? }.

Input parameters:

- `cpv` (string): AI 22 Consumer product variant (max 20 chars from the GS1 82-character set).
- `domain` (string): Resolver origin, default https://id.gs1.org (canonical). A brand domain is conformant too — the result explains the operational difference.
- `gtin` (string, required): The GTIN (8/12/13/14 digits; spaces and hyphens are cleaned). Refused if the check digit fails.
- `includeSvg` (boolean): Include the rendered SVG string (default true). The SVG already carries the 4-module quiet zone.
- `lot` (string): AI 10 batch/lot (max 20 chars from the GS1 82-character set).
- `serial` (string): AI 21 serial number (max 20 chars from the GS1 82-character set).

## Diagnostics

Captured diagnostic sections: Provenance, Dependencies. The full working is on the page: https://verifymcp.io/servers/com-yaktool-yaktool/yaktool-mcp#diagnostics

## Score history

- 2026-09-20: 73
- 2026-09-19: 72
- 2026-09-18: 72
- 2026-09-17: 71
- 2026-09-16: 71
- 2026-09-15: 70
- 2026-09-14: 70
- 2026-09-13: 69
- 2026-09-12: 69
- 2026-09-11: 69
- 2026-09-10: 68
- 2026-09-09: 68
- 2026-09-08: 67
- 2026-09-07: 67
- 2026-09-06: 66
- 2026-09-05: 66
- 2026-09-04: 65
- 2026-09-03: 65
- 2026-09-02: 64
- 2026-09-01: 64
- 2026-08-31: 63
- 2026-08-30: 63
- 2026-08-29: 62
- 2026-08-28: 62
- 2026-08-27: 62
- 2026-08-26: 61
- 2026-08-25: 62
- 2026-08-24: 50

## Common questions

### What is the com.yaktool/yaktool MCP server?

com.yaktool/yaktool is an MCP server listed in the public MCP registry as com.yaktool/yaktool. 73 local deterministic tools: GTIN, feeds, CSV-XLSX, Digital Link QR, e-invoices. Rules cited. This page covers its npm package (yaktool-mcp).

### Is the com.yaktool/yaktool MCP server safe to use?

com.yaktool/yaktool scores 73 out of 100 on VerifyMCP. We found no known CVEs affecting it as of 20 September 2026. It declares no install or post-install scripts. That is a record of what we were able to check automatically, not an endorsement. The category breakdown on this page shows every signal behind the number, including the ones we could not confirm.

### What tools does the com.yaktool/yaktool MCP server expose?

com.yaktool/yaktool exposes 73 tools: validate_gtin, calculate_check_digit, analyze_amazon_search_terms, convert_gtin, validate_google_feed, and 68 more. Their descriptions and schemas cost roughly 25,676 tokens of context every time the server is loaded.

### Is the com.yaktool/yaktool MCP server still maintained?

com.yaktool/yaktool is still listed as active in the MCP registry. We last reached this channel on 20 September 2026. Those dates come from our own scans of the registry and the channel itself, not from anything the publisher announced.

### What licence is the com.yaktool/yaktool MCP server under?

com.yaktool/yaktool declares the MIT licence, which is OSI-approved. That covers the source only, and says nothing about the cost of any service it calls.

## Links

- npm package: https://www.npmjs.com/package/yaktool-mcp
- Socket report: https://socket.dev/npm/package/yaktool-mcp
- Changelog RSS feed: https://verifymcp.io/servers/com-yaktool-yaktool/yaktool-mcp.xml
- Changelog JSON feed: https://verifymcp.io/servers/com-yaktool-yaktool/yaktool-mcp.json
- HTML version of this page: https://verifymcp.io/servers/com-yaktool-yaktool/yaktool-mcp
