Skip to content
verify mcp Beta VerifyMCP is currently in beta. If you notice any issues, email [email protected] and we’ll put it right.

IA-QA — 130+ QA & Dev Tools for AI Agents

REMOTE · WWW.IA-QA.COM · SCANNED AUG 3

130+ QA & dev tools for AI agents: prompt injection, RAG testing, VLM eval, guardrails. Free.

Available components

+6 this week 76 Trust /100
Trust breakdown (6 categories)

How this component scores in each security and reliability category. Every signal is checked automatically against the live server, and we only credit what we can confirm. How we score →

Endpoint Security83
Transport & Reachability100
Schema Quality & AI Usability70
  • AI-judged instruction clarity (excellent).Pass
  • Context-footprint check failed: tool/resource definitions use about 20450 tokens (~136/item across 150 items; 150 tools + 0 resources), over budget; trim descriptions and params. See how to fix → Fail
  • Usage-examples check failed: none of the tools include examples. See how to fix → Fail
Stability & Change Management27
  • Stability observed for 8 of 30 days with no destabilising changes; credit accrues until the full window elapses.Partial
Tool Coverage100
  • 100% of tools have a non-trivial description (not blank, and not just the tool's name).Pass
  • 100% of tool parameters carry a description.Pass
  • Structured output schemas are declared (100% of tools); any adoption earns full credit.Pass
Capabilities100
  • Implements a supported MCP spec version (2025-11-25); the latest is 2026-07-28.Pass
Install

Add this component to your MCP client. Where a client-specific snippet is available, pick your client below and copy it straight into your config; otherwise use the connection detail shown.

remote · www.ia-qa.com

# add to Claude Code
claude mcp add --transport http jcjamet-ia-qa-toolbox https://www.ia-qa.com/mcp
# ~/.codex/config.toml
[mcp_servers.jcjamet-ia-qa-toolbox]
url = "https://www.ia-qa.com/mcp"
// opencode.json
{
  "$schema": "https://opencode.ai/config.json",
  "mcp": {
    "jcjamet-ia-qa-toolbox": {
      "type": "remote",
      "url": "https://www.ia-qa.com/mcp",
      "enabled": true
    }
  }
}
# add to OpenClaw
openclaw mcp add jcjamet-ia-qa-toolbox --url https://www.ia-qa.com/mcp --transport streamable-http
# ~/.hermes/config.yaml
mcp_servers:
  jcjamet-ia-qa-toolbox:
    url: "https://www.ia-qa.com/mcp"
// mcp.json
{
  "mcpServers": {
    "jcjamet-ia-qa-toolbox": {
      "type": "http",
      "url": "https://www.ia-qa.com/mcp"
    }
  }
}

The mcpServers block is a cross-client convention. Remote transports vary, so check your client's docs.

Changelog

Every change we have recorded for this component, newest first. Security-relevant changes are always shown. ▲ marks a change for the better, ▼ a change for the worse; unmarked changes are neutral.

  • 2 Aug 26 +1

    No change was recorded against any check on this day. Stability & Change Management went from 20 to 23. That category is still filling its 30-day observation window: 6 days of observed history at the previous scan, 7 at this one. The score rises as the window fills, whether or not the server changes.

  • 31 Jul 26 +3
    • We updated how we score, so this day's move reflects our rubric, not a change to the server See what changed → functional
  • 30 Jul 26 +1
    • We updated how we score, so this day's move reflects our rubric, not a change to the server See what changed → functional
  • 29 Jul 26 0
    • Tool “analyze_diff_bugs” rewrote its description, which is the text the model reads security
    • Tool “run_pr_gate_pipeline” rewrote its description, which is the text the model reads security
    • “find_tool” added an optional parameter “max_results” cosmetic
  • 28 Jul 26 +1

    No change was recorded against any check on this day. Stability & Change Management went from 3 to 7. That category is still filling its 30-day observation window: 1 days of observed history at the previous scan, 2 at this one. The score rises as the window fills, whether or not the server changes.

  • 27 Jul 26 +1
    • We updated how we score, so this day's move reflects our rubric, not a change to the server See what changed → functional
  • 26 Jul 26 69

    First indexed and scored.

Diagnostics

Diagnostic detail from the automated scan of this channel: what the scanner observed at each step, so you can see exactly where a check passed or failed. It is informational only and never changes the trust score.

Captured 3 Aug 2026 · Probed https://www.ia-qa.com/mcp

TLS valid

Negotiated TLS 1.3 with TLS_AES_128_GCM_SHA256 .

Subject Issuer Valid from Valid until Key Signature Serial
CN=www.ia-qa.com CN=YR1,O=Let's Encrypt,C=US 11 Jun 2026 9 Sept 2026 RSA 2048 SHA256-RSA 6b2b8b5f2547d58014f06ba311ff41342ce
SANs: www.ia-qa.com
CN=YR1,O=Let's Encrypt,C=US (CA) CN=Root YR,O=ISRG,C=US 3 Sept 2025 2 Sept 2028 RSA 2048 SHA256-RSA a20253f15f2691c05dc1ce13b9bcca4e
CN=Root YR,O=ISRG,C=US (CA) CN=ISRG Root X1,O=Internet Security Research Group,C=US 13 May 2026 2 Sept 2032 RSA 4096 SHA256-RSA f24b6d17f9d9ad7cb1c9fea78782699f
DNSSEC secure

Validation of www.ia-qa.com. Secure

Zone DS Keys Algorithms Outcome
. trust_anchor 20326, 38696 8, 8 Verified
com. present 19718 13 Verified
ia-qa.com. present 52852 8 Verified
www.ia-qa.com. Verified address RRset verified with the apex keys
Authentication No authorisation required

The endpoint answered without asking for a token. Anyone who knows the URL can reach it.

Result No authorisation required
HTTP status 200
Header Value
strict-transport-security max-age=63072000; includeSubDomains; preload
content-security-policy default-src 'self';script-src 'self' 'unsafe-inline' 'unsafe-eval' https://www.googletagmanager.com;style-src 'self' 'unsafe-inline' https://fonts.googleapis.com;img-src 'self' data: blob: https:;font-src 'self' https://fonts.gstatic.com data:;connect-src 'self' https: wss: http://localhost:11434 data:;media-src 'self' blob:;object-src 'none';frame-ancestors 'self';base-uri 'self';form-action 'self';script-src-attr 'none';upgrade-insecure-requests
x-content-type-options nosniff
x-frame-options SAMEORIGIN
referrer-policy no-referrer
permissions-policy camera=(), microphone=(), geolocation=(), payment=(), usb=(), display-capture=()
Transports 2 probes
Transport URL Outcome Status Location
streamable-http https://www.ia-qa.com/mcp Verified 200
http (plaintext) http://www.ia-qa.com/mcp HTTPS enforced 301 https://www.ia-qa.com/mcp
MCP tools — 150 exposed · ~20,209 tokens

The tools this component advertises to a client, with an estimated token cost for each. Expand a tool to see its parameters and schema. The per-tool counts are indicative and are not scored directly; the schema's total context footprint is one signal in Schema Quality & AI Usability.

Tool Tokens
ab_test_report ~77

Generate an A/B test report comparing two prompts or model configurations. Accepts arrays of scores and returns statistical comparison: mean, median, std deviation, winner, and improvement percentage.

NameTypeReqDescription
variant_aobjectyesFirst variant configuration with name and score array
variant_bobjectyesSecond variant configuration with name and score array
NameTypeReqDescription
countnumber
improvement_percent
max
meanstring
medianstring
min
recommendationnumber
std_devstring
variant_aobject
variant_bobject
winner

No examples provided.

analyze_diff_bugs ~188

Pattern-based diff linter: flags a fixed set of risky shapes in changed code — query-string interpolation (SQL/Cypher/Mongo injection shape), shell interpolation, eval/new Function, empty catch blocks, regex built from a variable, fewer catch blocks than before, and named authorization guards that disappeared. Every finding cites the line that produced it. It does NOT do data-flow analysis: it cannot follow a value to a sink, across functions or files, and an empty result is not a safety verdict (the response lists what it did not analyse). Advisory triage — use a static analyser for a real security gate.

NameTypeReqDescription
contextstringOptional PR title or feature context for better analysis
version1stringOriginal code (before changes). If omitted, only the new version is analysed.
version2stringyesNew/modified code (after changes)
NameTypeReqDescription
bugsarray
disclaimerstring
notAnalysedarray
overallRiskstring
rulesAppliednumber
scannedLinesnumber
totalSuggestionsnumber

No examples provided.

analyze_responses ~209

Semantically analyze N already-produced model outputs for the SAME task (the MCP counterpart to the LLM Sandbox). Without a reference: computes consensus — pairwise cosine agreement, the most-representative output, and the outlier. With a `reference` (ground truth): also ranks every output by closeness (token cosine + ROUGE-L composite) and names the closest. Deterministic, no LLM, no key — gate-able in CI. You bring the outputs (2+). For a 2-way head-to-head with structural JSON diff use compare_responses instead.

NameTypeReqDescription
referencestringOptional ground-truth answer. If set, each output is also ranked by closeness to it and the closest one is named.
responsesarrayyesThe outputs to analyze (same task, N models/prompts/versions). Each item is a plain string or { "label": "GPT-4o", "text": "..." }. At least 2 required.
NameTypeReqDescription
consensus
count
reference_ranking
summarystring

No examples provided.

base64_decode ~55

Decode a Base64 string back to UTF-8 text. Use for inspecting Base64-encoded API responses, JWT payload claims, config file values, or attachment data.

NameTypeReqDescription
inputstringyesBase64 string to decode
NameTypeReqDescription
decoded

No examples provided.

base64_encode ~58

Encode a UTF-8 string to Base64. Use when you need to embed binary data, multi-line text, or special characters safely inside JSON fields, HTTP headers, or data URIs.

NameTypeReqDescription
inputstringyesText to encode
NameTypeReqDescription
encodedstring

No examples provided.

bias_detect ~86

Analyse a set of LLM responses generated from the same prompt template but with different demographic variants (gender, origin, age, tone). Returns a bias score (0-100), sentiment analysis per variant, pairwise Jaccard similarity, and a human-readable verdict. No API key needed — runs entirely locally.

NameTypeReqDescription
responsesarrayyesArray of variant responses to compare for bias
NameTypeReqDescription
avgSimilaritystring
biasScore
lengthCVstring
minSimilaritystring
negative
pairwiseSimilarities
positive
ratio
sentimentVariancestring
sentiments
verdict

No examples provided.

bm25_score ~127

Compute BM25 relevance score between a query and one or more documents. BM25 is the industry-standard keyword-based ranking algorithm used in Elasticsearch, OpenSearch, and Weaviate hybrid search. Returns ranked results with normalized scores.

NameTypeReqDescription
bnumberLength normalization factor (default: 0.75)
documentsarrayyesArray of documents to rank
k1numberTerm frequency saturation (default: 1.5)
querystringyesThe search query
top_knumberReturn top K results (default: all)
NameTypeReqDescription
avg_doc_lengthnumber
b
bm25_scorenumber
doc_lengthnumber
doc_preview
documents_countnumber
index
k1
query
results

No examples provided.

build_rag_prompt ~163

Assemble a complete RAG (Retrieval-Augmented Generation) prompt from retrieved context chunks and a user query. Handles token budgeting, citation numbering, system instruction injection, and source attribution.

NameTypeReqDescription
chunksarrayyesRetrieved context chunks with .text (required), .source (optional), .score (optional)
cite_sourcesbooleanAdd [1], [2] citation numbers (default: true)
languagestringResponse language instruction (e.g. "French", "Spanish")
max_context_tokensnumberMax tokens for context section (default: 2000)
querystringyesThe user question to answer
system_instructionstringCustom system instruction (default: standard RAG grounding instruction)
NameTypeReqDescription
chunks_includednumber
chunks_truncatednumber
context_tokens_estimatenumber
included_chunks
prompt
system_prompt
total_tokens_estimatenumber

No examples provided.

calculate_readability ~59

Calculate readability scores: Flesch Reading Ease, Flesch-Kincaid Grade Level, Coleman-Liau Index, and Automated Readability Index. Useful for evaluating LLM output quality.

NameTypeReqDescription
inputstringyesText to analyze for readability
NameTypeReqDescription
automated_readability_indexnumber
coleman_liau_indexnumber
flesch_kincaid_gradenumber
flesch_reading_ease
level
statsobject

No examples provided.

case_convert ~106

Convert a string between naming conventions: camelCase, PascalCase, snake_case, kebab-case, UPPER_SNAKE_CASE, dot.case, Title Case. Essential for code generation and refactoring.

NameTypeReqDescription
inputstringyesString to convert (e.g., "myVariableName", "my-css-class")
tostringyesTarget case: "camel", "pascal", "snake", "kebab", "upper_snake", "dot", "title"
NameTypeReqDescription
from_words
result
target_case

No examples provided.

check_contrast_ratio ~72

Calculate WCAG 2.1 contrast ratio between two colors. Returns ratio and compliance for AA/AAA normal and large text.

NameTypeReqDescription
backgroundstringyesBackground color in hex (e.g., "#ffffff")
foregroundstringyesForeground color in hex (e.g., "#333333")
NameTypeReqDescription
AAA_largeboolean
AAA_normalboolean
AA_largeboolean
AA_normalboolean
backgroundobject
foregroundobject
ratio
ratio_textstring

No examples provided.

color_convert ~109

Convert a color between HEX, RGB, and HSL formats. Use when translating design tokens between CSS notations, verifying color accessibility, or normalizing color values from user input. Accepts #rrggbb, #rgb, rgb(r,g,b), or hsl(h,s%,l%).

NameTypeReqDescription
inputstringyesColor value to convert, e.g. "#ff6b6b", "rgb(255,107,107)", "hsl(0,100%,71%)"
NameTypeReqDescription
b
g
hex
hslstring
input
r
rgbstring

No examples provided.

compare_models ~102

Compare 2-5 AI models side by side: context window, pricing, multimodal, reasoning capabilities, and provider. Returns a comparison table with a recommendation based on your use case.

NameTypeReqDescription
modelsarrayyesArray of 2-5 model names (e.g. ["gpt-4o","claude-3.5-sonnet","gemini-2.0-flash"])
use_casestringOptimize recommendation for this criterion
NameTypeReqDescription
cost_per_1k_totalstring
model
models_comparednumber
recommendation
rows
use_case

No examples provided.

compare_responses ~350

Compare two ALREADY-PRODUCED outputs (e.g. model A vs model B on the same task) side by side. Returns deterministic metrics (token cosine, ROUGE-L, Jaccard, length/structure deltas, JSON diff) and a verdict. If a `reference` (ground truth) is given, scores each output against it and picks the closer one. If `model` + `api_key` are given, an LLM judge also picks a qualitative winner for the task. No re-execution — you bring the outputs.

NameTypeReqDescription
api_keystringOptional API key for the judge model (BYOK). Used only for the judge call; never stored.
check_jsonbooleanTry to parse as JSON and compare structurally (keys, types, values)
label_astringLabel for output A (e.g. "GPT-4o", "v1.0")
label_bstringLabel for output B (e.g. "GPT-5-nano", "v1.1")
modelstringOptional judge model id (BYOK). When set with api_key, an LLM judge picks a qualitative winner.
referencestringOptional ground-truth / expected answer. If set, each output is scored against it and the closer one wins (deterministic).
response_astringyesFirst output (e.g. model A's answer)
response_bstringyesSecond output (e.g. model B's answer)
taskstringThe task/prompt both outputs were answering — used by the LLM judge for context
NameTypeReqDescription
judge
labelA
labelB
metrics
summarystring
verdict

No examples provided.

consistency_check ~143

Compare multiple LLM responses to the same prompt and detect inconsistencies using Jaccard word-overlap similarity and fact drift (number comparison). Fast, deterministic, no API key needed. Limitations: relies on surface-level word matching — "Paris is the capital of France" vs "Paris is the French capital" may score low despite semantic equivalence. For true semantic consistency, use run_semantic_tests with embedding mode. Essential for determinism testing.

NameTypeReqDescription
check_factsbooleanCheck for contradictory numbers/facts across responses (default: true)
responsesarrayyesArray of 2+ LLM responses to compare (same prompt, different runs)
NameTypeReqDescription
avg_similarity
fact_contradiction
fact_drift
length_variance_percent
pairwise_scores
response_countnumber
verdict

No examples provided.

context_window_check ~106

Given an array of message objects [{role, content}], estimate total token usage and check if it fits in the target model's context window. Warns about truncation risk.

NameTypeReqDescription
max_output_tokensnumberReserved tokens for output (default: 4096)
messagesarrayyesArray of messages (system/user/assistant)
modelstringyesTarget model name (e.g. gpt-4o, claude-3.5-sonnet)
NameTypeReqDescription
breakdownobject
charsnumber
context_window
fits
index
message_countnumber
model
per_message
reserved_output_tokens
role
tokens
total_input_tokens
total_tokens
utilization_percentstring
warnings

No examples provided.

conversation_analyze ~53

Analyze a multi-turn conversation for context retention, topic drift, instruction following, and repetition. Accepts messages array [{role, content}]. Essential for chatbot QA.

NameTypeReqDescription
messagesarrayyesConversation messages in order
NameTypeReqDescription
assistant_messagesnumber
avg_response_lengthnumber
context_retention
has_system_promptboolean
repetition_detectedboolean
repetitions
topic_drift
turn_countnumber
user_messagesnumber

No examples provided.

cookie_security_audit ~126

Audit the security attributes of cookies set by any URL. Fetches the URL and inspects all Set-Cookie headers for: HttpOnly, Secure, SameSite, Domain scope, Path scope, Max-Age/Expires, __Host-/__Secure- prefixes. Flags insecure patterns: missing HttpOnly on session cookies, missing Secure flag, SameSite=None without Secure, overly broad Domain, and excessive TTL. Returns per-cookie grades and an overall security score (0–100).

NameTypeReqDescription
urlstringyesFull URL to audit (e.g. https://example.com/login)
NameTypeReqDescription
cookiesarray
cookies_foundnumber
domain
host_prefix
httpOnly
issues
max_age
messagestring
name
path
sameSite
scorenumber
secure
secure_prefix
url

No examples provided.

cors_checker ~114

Check the CORS configuration of a URL the same way a browser would. Returns the main response status, all Access-Control-* headers, the tested origin, and the preflight OPTIONS response. Use this for direct CORS debugging, not just security auditing.

NameTypeReqDescription
methodstringHTTP method to simulate (default: GET)
originstringOrigin header to simulate (default: https://yourdomain.com)
urlstringyesFull URL to test, e.g. https://api.example.com/resource
NameTypeReqDescription
allHeaders
corsHeaders
method
preflight
status
testedOrigin
url

No examples provided.

cors_test ~131

Test a URL for CORS misconfigurations. Sends preflight (OPTIONS) and cross-origin requests with various Origin headers to detect: wildcard origins with credentials, origin reflection (echoing any origin), null origin acceptance, subdomain wildcard bypass, and missing Vary headers. Returns risk level (safe/low/medium/high/critical), per-test results, and fix recommendations. Essential for API security audits.

NameTypeReqDescription
originstringCustom Origin header to test (default: tests multiple origins automatically)
urlstringyesFull URL to test (e.g. https://api.example.com/endpoint)
NameTypeReqDescription
origins_testednumber
risk_level
tests
total_findingsnumber
url

No examples provided.

cot_analyzer ~104

Analyze a Chain-of-Thought (CoT) or reasoning trace from an LLM. Detects step count, logical flow, conclusion presence, backtracking, and estimates reasoning depth. Useful for o1/o3/DeepSeek-R1 evaluation.

NameTypeReqDescription
expected_conclusionstringExpected final answer to check against (optional)
reasoningstringyesThe CoT / reasoning trace text (e.g. from <think> tags or step-by-step output)
NameTypeReqDescription
backtracking_signals
conclusion_matches_expected
has_conclusionboolean
markers
reasoning_depth
reasoning_depth_label
step_countnumber
total_charsnumber
total_linesnumber

No examples provided.

count_code_lines ~109

Count lines of code: total, code lines, comment lines, blank lines, and comment density. Supports JS/TS, Python, Java/C/C++, Ruby, Go, Shell, HTML/XML, and CSS.

NameTypeReqDescription
inputstringyesSource code to analyze
languagestringLanguage hint: "js", "ts", "py", "java", "c", "rb", "go", "sh", "html", "css" (auto-detect if omitted)
NameTypeReqDescription
blank_lines
code_lines
code_to_comment_ratiostring
comment_density
comment_lines
language
total_lines

No examples provided.

count_tokens ~77

Estimate the token count of a text string using the cl100k_base approximation (~4 chars/token). Call this BEFORE sending any text to an LLM API to check if it fits within the model context window and to estimate cost. Returns token estimate, character count, and word count.

NameTypeReqDescription
inputstringyesText to count tokens for
NameTypeReqDescription
chars
tokens_estimate
words

No examples provided.

create_confluence_page ~237

Create a new Confluence page from the output of jira_to_test_suite. Formats Gherkin, E2E steps, API tests, and test data as a properly structured Confluence page with code blocks and tables. STATEFUL — creates a new page in the specified space.

NameTypeReqDescription
confluence_base_urlstringyesAtlassian base URL
confluence_emailstringyesAtlassian account email
confluence_tokenstringyesAtlassian API token
issue_keystringSource Jira issue key (for the page title and source link)
issue_urlstringSource Jira issue URL (added as a link in the page)
parent_page_idstringOptional parent page ID — page will be created as a child of this page
space_keystringyesConfluence space key where the page will be created, e.g. "QA", "ENG"
test_suiteobjectyesThe test_suite object from jira_to_test_suite result
titlestringPage title. Defaults to "Test Plan: {issue_key}"
NameTypeReqDescription
page_idstring
page_urlstring
successboolean
titlestring

No examples provided.

cron_parse ~63

Parse a cron expression into a human-readable schedule description. Supports standard 5-field cron (minute hour day month weekday).

NameTypeReqDescription
expressionstringyesCron expression (e.g., "0 9 * * 1-5", "*/15 * * * *")
NameTypeReqDescription
expression
fieldsobject
human_readablestring

No examples provided.

cron_validator ~106

Validate a 5-field cron expression, explain the schedule, and preview the next execution times. Use this to debug cron jobs before they reach production. Returns parsed fields, a human-readable description, and upcoming ISO timestamps.

NameTypeReqDescription
expressionstringyesCron expression with 5 fields, e.g. "*/15 9-18 * * 1-5"
next_runs_countnumberHow many upcoming runs to return (1-50, default: 10)
NameTypeReqDescription
expression
fieldsobject
human_readable
next_runs
validboolean

No examples provided.

decode_jwt ~80

Decode a JWT (JSON Web Token) and return its header and payload without verifying the signature. Also reports whether the token is expired and the exact expiry date. Use to inspect claims (sub, iss, exp, roles) during debugging or when integrating with an auth provider.

NameTypeReqDescription
tokenstringyesThe JWT string to decode (header.payload.signature)
NameTypeReqDescription
expired
expiresAt
header
notestring
payload

No examples provided.

detect_language ~80

Detect the natural language of a text using n-gram frequency analysis and common word markers. Supports 15 languages: English, French, Spanish, German, Italian, Portuguese, Dutch, Russian, Chinese, Japanese, Korean, Arabic, Polish, Turkish, Swedish.

NameTypeReqDescription
inputstringyesText to detect language from (min 20 chars for accuracy)
NameTypeReqDescription
confidencenumber
lang
languagestring
matched
methodstring
namestring
score
top_candidatesarray

No examples provided.

detect_secrets ~97

Scan code or config files for hardcoded secrets: AWS keys, GitHub tokens, OpenAI/Anthropic API keys, Stripe secrets, JWTs, database connection strings, and generic passwords. Returns findings with severity. Run before every commit.

NameTypeReqDescription
filenamestringOptional filename for context (e.g. ".env", "config.js")
inputstringyesCode or config content to scan (max 500KB)
NameTypeReqDescription
filename
findings
recommendation
risk_level
total_findingsnumber

No examples provided.

diff_mappings ~199

Diff a baseline page mapping against a current one and return a CI-style verdict: PASS / FIX / BLOCK, plus per-element drift (ok, renamed, healable, ambiguous, lost, added, rebound). Pure and deterministic — provide two mappings as JSON with "elements" arrays of {role, name, selector, context?}. Use the companion @ia-qa/self-healing package (npm install -g @ia-qa/self-healing) to capture mappings from your app via its local MCP server ia-qa-heal-mcp, or paste the snippet from ia-qa.com/devtools/selector-drift into your browser console.

NameTypeReqDescription
afterobjectyesCurrent page mapping: same shape as before, captured after the UI change.
beforeobjectyesBaseline page mapping: { page, url, capturedAt, elements: [{role, name, selector, context?}] }. Captured before a UI change.
NameTypeReqDescription
addedarray
countsobject
rowsarray
verdict

No examples provided.

diff_text ~116

Compute a unified line-by-line diff between two text strings (LCS algorithm). Returns added/removed/unchanged line counts and formatted diff hunks with configurable context lines (0–20). Use to compare versions of prompts, configs, code snippets, or any text where you need to see exactly what changed.

NameTypeReqDescription
astringyesOriginal (before) text
bstringyesModified (after) text
contextnumberContext lines around each change (0–20, default: 3)
NameTypeReqDescription
addednumber
diffstring
removednumber
unchangednumber

No examples provided.

embedding_similarity ~186

Compute text similarity using local algorithms (Bag of Words, TF-IDF, Character N-grams). No API key needed — runs entirely in-process. NOT real embeddings: for true semantic similarity with vector embeddings, use run_semantic_tests with mode="embeddings" and your OpenAI API key. Supports single pair or batch mode with pipe-separated pairs. Useful for RAG retrieval testing, semantic search evaluation, and text deduplication.

NameTypeReqDescription
batcharrayBatch mode: array of { text_a, text_b } pairs. Overrides text_a/text_b if provided.
methodsarrayAlgorithms to use (default: all three). Options: "bow", "tfidf", "ngram"
text_astringFirst text to compare (single-pair mode)
text_bstringSecond text to compare (single-pair mode)
NameTypeReqDescription
countnumber
modestring
results
scores
text_a
text_b

No examples provided.

env_parse ~93

Parse a .env file content into a JSON object. Handles quoted values (single and double), inline comments, export prefix, and escaped sequences (\n, \t inside double quotes). Returns all key-value pairs. Use in CI/CD pipelines, agent config loaders, or when processing dotenv files programmatically.

NameTypeReqDescription
inputstringyes.env file content to parse (e.g. the output of `cat .env`)
NameTypeReqDescription
countnumber
vars

No examples provided.

escape_html ~67

Escape HTML special characters (&, <, >, ", ') to their safe HTML entities. ALWAYS call this before inserting any user-provided or LLM-generated content into an HTML template to prevent cross-site scripting (XSS) attacks.

NameTypeReqDescription
inputstringyesString to HTML-escape
NameTypeReqDescription
escaped
original_lengthstring

No examples provided.

estimate_llm_cost ~162

Estimate the API cost in USD for a given model and token counts. Supports all major 2024–2026 models: GPT-4o, GPT-4.1, o3, o4-mini, Claude Opus 4, Claude Sonnet 4/4.5, Gemini 2.5 Pro/Flash, DeepSeek V3/R1, Grok 3, and legacy models.

NameTypeReqDescription
input_tokensnumberyesNumber of input/prompt tokens
modelstringyesModel name, e.g. "gpt-4o", "claude-3.5-sonnet", "deepseek-v3"
output_tokensnumberNumber of output/completion tokens (default: 0)
NameTypeReqDescription
input_cost_usdstring
input_tokens
model
output_cost_usdstring
output_tokens
ratesobject
total_cost_usdstring

No examples provided.

extract_json_from_text ~88

Extract the first valid JSON object or array embedded in chaotic LLM output (surrounded by markdown fences, prose, or explanatory text). Handles ```json blocks and inline JSON. Call this whenever an LLM returns structured data mixed with explanation text instead of raw JSON.

NameTypeReqDescription
inputstringyesRaw text (e.g., LLM output) that may contain a JSON object or array
NameTypeReqDescription
json
sourcestring

No examples provided.

extract_json_path ~88

Extract a value from a JSON string using dot-notation path (e.g., "user.address.city", "items.0.name", "meta.tags"). Supports array index access via numeric path segments.

NameTypeReqDescription
inputstringyesA valid JSON string to traverse
pathstringyesDot-notation path, e.g. "user.address.city" or "items.0.name"
NameTypeReqDescription
path
type
value

No examples provided.

extract_links ~69

Extract all URLs, email addresses, and domain names from text. Returns categorized and deduplicated results. Useful for content auditing, link checking, and web scraping validation.

NameTypeReqDescription
inputstringyesText to extract links from
typesarrayTypes to extract (default: all three)
NameTypeReqDescription
total

No examples provided.

extract_todos ~111

Extract TODO, FIXME, HACK, BUG, NOTE, OPTIMIZE, and custom tags from any source code or text. Returns line numbers, tag types, and message text. Essential for technical debt auditing.

NameTypeReqDescription
include_contextbooleanInclude full line text (default: true)
inputstringyesCode or text to scan
tagsarrayCustom tags to add (default set: TODO, FIXME, HACK, NOTE, BUG, OPTIMIZE, XXX)
NameTypeReqDescription
counts
has_criticalboolean
items
totalnumber

No examples provided.

fetch_confluence_page ~200

Fetch a Confluence page and return its content as clean Markdown. Accepts a numeric page_id or a full page URL. Optionally lists direct child pages. BYOK — credentials transit in-memory only, never stored.

NameTypeReqDescription
confluence_base_urlstringyesAtlassian base URL, e.g. "https://mycompany.atlassian.net"
confluence_emailstringyesAtlassian account email (same credentials as Jira)
confluence_tokenstringyesAtlassian API token
include_childrenbooleanList direct child pages (id + title) (default: false)
page_idstringConfluence page ID (numeric string), e.g. "123456789"
page_urlstringFull Confluence page URL (alternative to page_id), e.g. "https://mycompany.atlassian.net/wiki/spaces/ENG/pages/123456789"
NameTypeReqDescription
childrenarray
markdownstring
page_idstring
titlestring
urlstring

No examples provided.

fetch_jira_issue ~201

Fetch a complete Jira issue: summary, description converted to Markdown, status, assignee, priority, labels, custom fields, and optionally comments and attachment metadata. BYOK — credentials transit in-memory only, never stored on ia-qa.com.

NameTypeReqDescription
fieldsarraySpecific Jira field names to return. Omit for all standard fields.
include_attachmentsbooleanInclude attachment metadata list (default: false)
include_commentsbooleanInclude issue comments, up to 20 (default: true)
issue_keystringyesJira issue key, e.g. "PROJ-123"
jira_base_urlstringyesAtlassian base URL, e.g. "https://mycompany.atlassian.net"
jira_emailstringyesAtlassian account email
jira_tokenstringyesAtlassian API token (from id.atlassian.com > Security > API tokens)
NameTypeReqDescription
assigneestring
descriptionstring
keystring
labelsarray
prioritystring
reporterstring
statusstring
summarystring
typestring
urlstring

No examples provided.

fetch_veille_feed ~118

Fetch the latest QA & AI/LLM articles aggregated from curated RSS sources (Google Testing Blog, DEV.to Testing/QA/AI/LLM/Agents, Hugging Face Blog, Simon Willison). Perfect for agents monitoring the QA & AI landscape.

NameTypeReqDescription
categorystringFilter: "qa" (testing/quality), "ai" (AI/LLM/agents), "all" (default — both)
limitnumberMax articles to return (default: 20, max: 50)
NameTypeReqDescription
articlesarray
category
sources_queriednumber
total_foundnumber

No examples provided.

few_shot_formatter ~107

Format few-shot examples for LLM prompts. Converts example pairs into formatted blocks. Supports chat format (User/Assistant), XML tags, Markdown, or plain text.

NameTypeReqDescription
examplesarrayyesArray of {input, output} pairs
formatstringOutput format (default: chat)
input_labelstringLabel for input (default: User / <input>)
output_labelstringLabel for output (default: Assistant / <output>)
NameTypeReqDescription
example_countnumber
format
formatted
token_estimatenumber

No examples provided.

find_tool ~197

Search available MCP tools by keyword or category before calling them. Returns matching tool names, descriptions, and optionally their inputSchemas. Call this when you are unsure which tool to use or want to explore the catalogue. Categories: data, encoding, text, llm, qa, rag, dev, security, web.

NameTypeReqDescription
categorystringOptional: filter by category — data | encoding | text | llm | qa | rag | dev | security | web
max_resultsnumberMaximum tools to return (default 10, max 50). Results are ranked by IDF-weighted relevance, so common words like "test" do not inflate the list.
querystringyesKeyword(s) to search in tool name and description (e.g. "cors", "token", "vector", "json")
with_schemabooleanSet true to include inputSchema in results (default: false)
NameTypeReqDescription
category
countnumber
hint
query
score
tool
toolsarray
total_matchesnumber
truncatedboolean

No examples provided.

fix_gherkin ~290

Fix Gherkin syntax warnings from a jira_to_test_suite result. Takes the current gherkin text and the _gherkin_warnings array, calls your LLM to fix ONLY the flagged issues (adds missing Given/When/Then steps, etc.), and returns the corrected Gherkin. Lightweight — uses ~300-500 tokens vs ~5k for a full regeneration. Requires BYOK LLM key.

NameTypeReqDescription
api_keystringyesYour own LLM provider API key (BYOK) — OpenAI "sk-…", Anthropic "sk-ant-…", Google "AIzaSy…", or Groq "gsk_…". There is no server-side key for this tool: if you do not have one, do not call it and do…
gherkinstringyesThe current Gherkin text from the jira_to_test_suite result (test_suite.gherkin).
modelstringyesLLM model to use for the fix, e.g. "gpt-4o-mini". Must belong to the provider whose key you passed in api_key.
warningsarrayyesThe _gherkin_warnings array from the jira_to_test_suite result.
NameTypeReqDescription
fixed_gherkinstring
latency_msnumber
model_usedstring
remaining_warningsarray
warnings_afternumber
warnings_beforenumber

No examples provided.