IA-QA — 130+ QA & Dev Tools for AI Agents
REMOTE · WWW.IA-QA.COM · SCANNED AUG 3
130+ QA & dev tools for AI agents: prompt injection, RAG testing, VLM eval, guardrails. Free.
Available components
How this component scores in each security and reliability category. Every signal is checked automatically against the live server, and we only credit what we can confirm. How we score →
Endpoint Security83
- The endpoint's TLS certificate is valid, in date, and uses a strong key. View diagnostics → Pass
- No authorisation is required to call this server. Every tool declares its destructiveHint and none is destructive, so open access doesn't expose one. See how to fix → View diagnostics → Partial
- HTTPS is enforced; there's no plaintext access path. View diagnostics → Pass
- The HSTS (Strict-Transport-Security) header is present. View diagnostics → Pass
- DNSSEC is configured correctly; the domain's records validate against the full chain to the root. View diagnostics → Pass
Transport & Reachability100
- Verified streamable-http transport via a live MCP handshake. View diagnostics → Pass
Schema Quality & AI Usability70
- AI-judged instruction clarity (excellent).Pass
- Context-footprint check failed: tool/resource definitions use about 20450 tokens (~136/item across 150 items; 150 tools + 0 resources), over budget; trim descriptions and params. See how to fix → Fail
- Usage-examples check failed: none of the tools include examples. See how to fix → Fail
Stability & Change Management27
- Stability observed for 8 of 30 days with no destabilising changes; credit accrues until the full window elapses.Partial
Tool Coverage100
- 100% of tools have a non-trivial description (not blank, and not just the tool's name).Pass
- 100% of tool parameters carry a description.Pass
- Structured output schemas are declared (100% of tools); any adoption earns full credit.Pass
Capabilities100
- Implements a supported MCP spec version (2025-11-25); the latest is 2026-07-28.Pass
Add this component to your MCP client. Where a client-specific snippet is available, pick your client below and copy it straight into your config; otherwise use the connection detail shown.
remote · www.ia-qa.com
claude mcp add --transport http jcjamet-ia-qa-toolbox https://www.ia-qa.com/mcp
[mcp_servers.jcjamet-ia-qa-toolbox] url = "https://www.ia-qa.com/mcp"
{
"$schema": "https://opencode.ai/config.json",
"mcp": {
"jcjamet-ia-qa-toolbox": {
"type": "remote",
"url": "https://www.ia-qa.com/mcp",
"enabled": true
}
}
} openclaw mcp add jcjamet-ia-qa-toolbox --url https://www.ia-qa.com/mcp --transport streamable-http
mcp_servers:
jcjamet-ia-qa-toolbox:
url: "https://www.ia-qa.com/mcp" {
"mcpServers": {
"jcjamet-ia-qa-toolbox": {
"type": "http",
"url": "https://www.ia-qa.com/mcp"
}
}
} The mcpServers block is a cross-client convention. Remote transports vary, so check your client's docs.
Every change we have recorded for this component, newest first. Security-relevant changes are always shown. ▲ marks a change for the better, ▼ a change for the worse; unmarked changes are neutral.
- 2 Aug 26 +1
No change was recorded against any check on this day. Stability & Change Management went from 20 to 23. That category is still filling its 30-day observation window: 6 days of observed history at the previous scan, 7 at this one. The score rises as the window fills, whether or not the server changes.
- 31 Jul 26 +3
- We updated how we score, so this day's move reflects our rubric, not a change to the server See what changed → functional
- 30 Jul 26 +1
- We updated how we score, so this day's move reflects our rubric, not a change to the server See what changed → functional
- 29 Jul 26 0
- Tool “analyze_diff_bugs” rewrote its description, which is the text the model reads security
- Tool “run_pr_gate_pipeline” rewrote its description, which is the text the model reads security
- “find_tool” added an optional parameter “max_results” cosmetic
- 28 Jul 26 +1
No change was recorded against any check on this day. Stability & Change Management went from 3 to 7. That category is still filling its 30-day observation window: 1 days of observed history at the previous scan, 2 at this one. The score rises as the window fills, whether or not the server changes.
- 27 Jul 26 +1
- We updated how we score, so this day's move reflects our rubric, not a change to the server See what changed → functional
- 26 Jul 26 69
First indexed and scored.
Diagnostic detail from the automated scan of this channel: what the scanner observed at each step, so you can see exactly where a check passed or failed. It is informational only and never changes the trust score.
Captured 3 Aug 2026 · Probed https://www.ia-qa.com/mcp
TLS valid
Negotiated TLS 1.3 with TLS_AES_128_GCM_SHA256 .
| Subject | Issuer | Valid from | Valid until | Key | Signature | Serial |
|---|---|---|---|---|---|---|
| CN=www.ia-qa.com | CN=YR1,O=Let's Encrypt,C=US | 11 Jun 2026 | 9 Sept 2026 | RSA 2048 | SHA256-RSA | 6b2b8b5f2547d58014f06ba311ff41342ce |
| SANs: www.ia-qa.com | ||||||
| CN=YR1,O=Let's Encrypt,C=US (CA) | CN=Root YR,O=ISRG,C=US | 3 Sept 2025 | 2 Sept 2028 | RSA 2048 | SHA256-RSA | a20253f15f2691c05dc1ce13b9bcca4e |
| CN=Root YR,O=ISRG,C=US (CA) | CN=ISRG Root X1,O=Internet Security Research Group,C=US | 13 May 2026 | 2 Sept 2032 | RSA 4096 | SHA256-RSA | f24b6d17f9d9ad7cb1c9fea78782699f |
DNSSEC secure
Validation of www.ia-qa.com. — Secure
| Zone | DS | Keys | Algorithms | Outcome |
|---|---|---|---|---|
| . | trust_anchor | 20326, 38696 | 8, 8 | Verified |
| com. | present | 19718 | 13 | Verified |
| ia-qa.com. | present | 52852 | 8 | Verified |
| www.ia-qa.com. | Verified address RRset verified with the apex keys |
Authentication No authorisation required
The endpoint answered without asking for a token. Anyone who knows the URL can reach it.
| Result | No authorisation required |
|---|---|
| HTTP status | 200 |
| Header | Value |
|---|---|
| strict-transport-security | max-age=63072000; includeSubDomains; preload |
| content-security-policy | default-src 'self';script-src 'self' 'unsafe-inline' 'unsafe-eval' https://www.googletagmanager.com;style-src 'self' 'unsafe-inline' https://fonts.googleapis.com;img-src 'self' data: blob: https:;font-src 'self' https://fonts.gstatic.com data:;connect-src 'self' https: wss: http://localhost:11434 data:;media-src 'self' blob:;object-src 'none';frame-ancestors 'self';base-uri 'self';form-action 'self';script-src-attr 'none';upgrade-insecure-requests |
| x-content-type-options | nosniff |
| x-frame-options | SAMEORIGIN |
| referrer-policy | no-referrer |
| permissions-policy | camera=(), microphone=(), geolocation=(), payment=(), usb=(), display-capture=() |
Transports 2 probes
| Transport | URL | Outcome | Status | Location |
|---|---|---|---|---|
| streamable-http | https://www.ia-qa.com/mcp | Verified | 200 | |
| http (plaintext) | http://www.ia-qa.com/mcp | HTTPS enforced | 301 | https://www.ia-qa.com/mcp |
The tools this component advertises to a client, with an estimated token cost for each. Expand a tool to see its parameters and schema. The per-tool counts are indicative and are not scored directly; the schema's total context footprint is one signal in Schema Quality & AI Usability.
ab_test_report ~77
Generate an A/B test report comparing two prompts or model configurations. Accepts arrays of scores and returns statistical comparison: mean, median, std deviation, winner, and improvement percentage.
| Name | Type | Req | Description |
|---|---|---|---|
| variant_a | object | yes | First variant configuration with name and score array |
| variant_b | object | yes | Second variant configuration with name and score array |
| Name | Type | Req | Description |
|---|---|---|---|
| count | number | — | — |
| improvement_percent | — | — | — |
| max | — | — | — |
| mean | string | — | — |
| median | string | — | — |
| min | — | — | — |
| recommendation | number | — | — |
| std_dev | string | — | — |
| variant_a | object | — | — |
| variant_b | object | — | — |
| winner | — | — | — |
No examples provided.
analyze_diff_bugs ~188
Pattern-based diff linter: flags a fixed set of risky shapes in changed code — query-string interpolation (SQL/Cypher/Mongo injection shape), shell interpolation, eval/new Function, empty catch blocks, regex built from a variable, fewer catch blocks than before, and named authorization guards that disappeared. Every finding cites the line that produced it. It does NOT do data-flow analysis: it cannot follow a value to a sink, across functions or files, and an empty result is not a safety verdict (the response lists what it did not analyse). Advisory triage — use a static analyser for a real security gate.
| Name | Type | Req | Description |
|---|---|---|---|
| context | string | — | Optional PR title or feature context for better analysis |
| version1 | string | — | Original code (before changes). If omitted, only the new version is analysed. |
| version2 | string | yes | New/modified code (after changes) |
| Name | Type | Req | Description |
|---|---|---|---|
| bugs | array | — | — |
| disclaimer | string | — | — |
| notAnalysed | array | — | — |
| overallRisk | string | — | — |
| rulesApplied | number | — | — |
| scannedLines | number | — | — |
| totalSuggestions | number | — | — |
No examples provided.
analyze_responses ~209
Semantically analyze N already-produced model outputs for the SAME task (the MCP counterpart to the LLM Sandbox). Without a reference: computes consensus — pairwise cosine agreement, the most-representative output, and the outlier. With a `reference` (ground truth): also ranks every output by closeness (token cosine + ROUGE-L composite) and names the closest. Deterministic, no LLM, no key — gate-able in CI. You bring the outputs (2+). For a 2-way head-to-head with structural JSON diff use compare_responses instead.
| Name | Type | Req | Description |
|---|---|---|---|
| reference | string | — | Optional ground-truth answer. If set, each output is also ranked by closeness to it and the closest one is named. |
| responses | array | yes | The outputs to analyze (same task, N models/prompts/versions). Each item is a plain string or { "label": "GPT-4o", "text": "..." }. At least 2 required. |
| Name | Type | Req | Description |
|---|---|---|---|
| consensus | — | — | — |
| count | — | — | — |
| reference_ranking | — | — | — |
| summary | string | — | — |
No examples provided.
base64_decode ~55
Decode a Base64 string back to UTF-8 text. Use for inspecting Base64-encoded API responses, JWT payload claims, config file values, or attachment data.
| Name | Type | Req | Description |
|---|---|---|---|
| input | string | yes | Base64 string to decode |
| Name | Type | Req | Description |
|---|---|---|---|
| decoded | — | — | — |
No examples provided.
base64_encode ~58
Encode a UTF-8 string to Base64. Use when you need to embed binary data, multi-line text, or special characters safely inside JSON fields, HTTP headers, or data URIs.
| Name | Type | Req | Description |
|---|---|---|---|
| input | string | yes | Text to encode |
| Name | Type | Req | Description |
|---|---|---|---|
| encoded | string | — | — |
No examples provided.
bias_detect ~86
Analyse a set of LLM responses generated from the same prompt template but with different demographic variants (gender, origin, age, tone). Returns a bias score (0-100), sentiment analysis per variant, pairwise Jaccard similarity, and a human-readable verdict. No API key needed — runs entirely locally.
| Name | Type | Req | Description |
|---|---|---|---|
| responses | array | yes | Array of variant responses to compare for bias |
| Name | Type | Req | Description |
|---|---|---|---|
| avgSimilarity | string | — | — |
| biasScore | — | — | — |
| lengthCV | string | — | — |
| minSimilarity | string | — | — |
| negative | — | — | — |
| pairwiseSimilarities | — | — | — |
| positive | — | — | — |
| ratio | — | — | — |
| sentimentVariance | string | — | — |
| sentiments | — | — | — |
| verdict | — | — | — |
No examples provided.
bm25_score ~127
Compute BM25 relevance score between a query and one or more documents. BM25 is the industry-standard keyword-based ranking algorithm used in Elasticsearch, OpenSearch, and Weaviate hybrid search. Returns ranked results with normalized scores.
| Name | Type | Req | Description |
|---|---|---|---|
| b | number | — | Length normalization factor (default: 0.75) |
| documents | array | yes | Array of documents to rank |
| k1 | number | — | Term frequency saturation (default: 1.5) |
| query | string | yes | The search query |
| top_k | number | — | Return top K results (default: all) |
| Name | Type | Req | Description |
|---|---|---|---|
| avg_doc_length | number | — | — |
| b | — | — | — |
| bm25_score | number | — | — |
| doc_length | number | — | — |
| doc_preview | — | — | — |
| documents_count | number | — | — |
| index | — | — | — |
| k1 | — | — | — |
| query | — | — | — |
| results | — | — | — |
No examples provided.
build_rag_prompt ~163
Assemble a complete RAG (Retrieval-Augmented Generation) prompt from retrieved context chunks and a user query. Handles token budgeting, citation numbering, system instruction injection, and source attribution.
| Name | Type | Req | Description |
|---|---|---|---|
| chunks | array | yes | Retrieved context chunks with .text (required), .source (optional), .score (optional) |
| cite_sources | boolean | — | Add [1], [2] citation numbers (default: true) |
| language | string | — | Response language instruction (e.g. "French", "Spanish") |
| max_context_tokens | number | — | Max tokens for context section (default: 2000) |
| query | string | yes | The user question to answer |
| system_instruction | string | — | Custom system instruction (default: standard RAG grounding instruction) |
| Name | Type | Req | Description |
|---|---|---|---|
| chunks_included | number | — | — |
| chunks_truncated | number | — | — |
| context_tokens_estimate | number | — | — |
| included_chunks | — | — | — |
| prompt | — | — | — |
| system_prompt | — | — | — |
| total_tokens_estimate | number | — | — |
No examples provided.
calculate_readability ~59
Calculate readability scores: Flesch Reading Ease, Flesch-Kincaid Grade Level, Coleman-Liau Index, and Automated Readability Index. Useful for evaluating LLM output quality.
| Name | Type | Req | Description |
|---|---|---|---|
| input | string | yes | Text to analyze for readability |
| Name | Type | Req | Description |
|---|---|---|---|
| automated_readability_index | number | — | — |
| coleman_liau_index | number | — | — |
| flesch_kincaid_grade | number | — | — |
| flesch_reading_ease | — | — | — |
| level | — | — | — |
| stats | object | — | — |
No examples provided.
case_convert ~106
Convert a string between naming conventions: camelCase, PascalCase, snake_case, kebab-case, UPPER_SNAKE_CASE, dot.case, Title Case. Essential for code generation and refactoring.
| Name | Type | Req | Description |
|---|---|---|---|
| input | string | yes | String to convert (e.g., "myVariableName", "my-css-class") |
| to | string | yes | Target case: "camel", "pascal", "snake", "kebab", "upper_snake", "dot", "title" |
| Name | Type | Req | Description |
|---|---|---|---|
| from_words | — | — | — |
| result | — | — | — |
| target_case | — | — | — |
No examples provided.
check_contrast_ratio ~72
Calculate WCAG 2.1 contrast ratio between two colors. Returns ratio and compliance for AA/AAA normal and large text.
| Name | Type | Req | Description |
|---|---|---|---|
| background | string | yes | Background color in hex (e.g., "#ffffff") |
| foreground | string | yes | Foreground color in hex (e.g., "#333333") |
| Name | Type | Req | Description |
|---|---|---|---|
| AAA_large | boolean | — | — |
| AAA_normal | boolean | — | — |
| AA_large | boolean | — | — |
| AA_normal | boolean | — | — |
| background | object | — | — |
| foreground | object | — | — |
| ratio | — | — | — |
| ratio_text | string | — | — |
No examples provided.
color_convert ~109
Convert a color between HEX, RGB, and HSL formats. Use when translating design tokens between CSS notations, verifying color accessibility, or normalizing color values from user input. Accepts #rrggbb, #rgb, rgb(r,g,b), or hsl(h,s%,l%).
| Name | Type | Req | Description |
|---|---|---|---|
| input | string | yes | Color value to convert, e.g. "#ff6b6b", "rgb(255,107,107)", "hsl(0,100%,71%)" |
| Name | Type | Req | Description |
|---|---|---|---|
| b | — | — | — |
| g | — | — | — |
| hex | — | — | — |
| hsl | string | — | — |
| input | — | — | — |
| r | — | — | — |
| rgb | string | — | — |
No examples provided.
compare_models ~102
Compare 2-5 AI models side by side: context window, pricing, multimodal, reasoning capabilities, and provider. Returns a comparison table with a recommendation based on your use case.
| Name | Type | Req | Description |
|---|---|---|---|
| models | array | yes | Array of 2-5 model names (e.g. ["gpt-4o","claude-3.5-sonnet","gemini-2.0-flash"]) |
| use_case | string | — | Optimize recommendation for this criterion |
| Name | Type | Req | Description |
|---|---|---|---|
| cost_per_1k_total | string | — | — |
| model | — | — | — |
| models_compared | number | — | — |
| recommendation | — | — | — |
| rows | — | — | — |
| use_case | — | — | — |
No examples provided.
compare_responses ~350
Compare two ALREADY-PRODUCED outputs (e.g. model A vs model B on the same task) side by side. Returns deterministic metrics (token cosine, ROUGE-L, Jaccard, length/structure deltas, JSON diff) and a verdict. If a `reference` (ground truth) is given, scores each output against it and picks the closer one. If `model` + `api_key` are given, an LLM judge also picks a qualitative winner for the task. No re-execution — you bring the outputs.
| Name | Type | Req | Description |
|---|---|---|---|
| api_key | string | — | Optional API key for the judge model (BYOK). Used only for the judge call; never stored. |
| check_json | boolean | — | Try to parse as JSON and compare structurally (keys, types, values) |
| label_a | string | — | Label for output A (e.g. "GPT-4o", "v1.0") |
| label_b | string | — | Label for output B (e.g. "GPT-5-nano", "v1.1") |
| model | string | — | Optional judge model id (BYOK). When set with api_key, an LLM judge picks a qualitative winner. |
| reference | string | — | Optional ground-truth / expected answer. If set, each output is scored against it and the closer one wins (deterministic). |
| response_a | string | yes | First output (e.g. model A's answer) |
| response_b | string | yes | Second output (e.g. model B's answer) |
| task | string | — | The task/prompt both outputs were answering — used by the LLM judge for context |
| Name | Type | Req | Description |
|---|---|---|---|
| judge | — | — | — |
| labelA | — | — | — |
| labelB | — | — | — |
| metrics | — | — | — |
| summary | string | — | — |
| verdict | — | — | — |
No examples provided.
consistency_check ~143
Compare multiple LLM responses to the same prompt and detect inconsistencies using Jaccard word-overlap similarity and fact drift (number comparison). Fast, deterministic, no API key needed. Limitations: relies on surface-level word matching — "Paris is the capital of France" vs "Paris is the French capital" may score low despite semantic equivalence. For true semantic consistency, use run_semantic_tests with embedding mode. Essential for determinism testing.
| Name | Type | Req | Description |
|---|---|---|---|
| check_facts | boolean | — | Check for contradictory numbers/facts across responses (default: true) |
| responses | array | yes | Array of 2+ LLM responses to compare (same prompt, different runs) |
| Name | Type | Req | Description |
|---|---|---|---|
| avg_similarity | — | — | — |
| fact_contradiction | — | — | — |
| fact_drift | — | — | — |
| length_variance_percent | — | — | — |
| pairwise_scores | — | — | — |
| response_count | number | — | — |
| verdict | — | — | — |
No examples provided.
context_window_check ~106
Given an array of message objects [{role, content}], estimate total token usage and check if it fits in the target model's context window. Warns about truncation risk.
| Name | Type | Req | Description |
|---|---|---|---|
| max_output_tokens | number | — | Reserved tokens for output (default: 4096) |
| messages | array | yes | Array of messages (system/user/assistant) |
| model | string | yes | Target model name (e.g. gpt-4o, claude-3.5-sonnet) |
| Name | Type | Req | Description |
|---|---|---|---|
| breakdown | object | — | — |
| chars | number | — | — |
| context_window | — | — | — |
| fits | — | — | — |
| index | — | — | — |
| message_count | number | — | — |
| model | — | — | — |
| per_message | — | — | — |
| reserved_output_tokens | — | — | — |
| role | — | — | — |
| tokens | — | — | — |
| total_input_tokens | — | — | — |
| total_tokens | — | — | — |
| utilization_percent | string | — | — |
| warnings | — | — | — |
No examples provided.
conversation_analyze ~53
Analyze a multi-turn conversation for context retention, topic drift, instruction following, and repetition. Accepts messages array [{role, content}]. Essential for chatbot QA.
| Name | Type | Req | Description |
|---|---|---|---|
| messages | array | yes | Conversation messages in order |
| Name | Type | Req | Description |
|---|---|---|---|
| assistant_messages | number | — | — |
| avg_response_length | number | — | — |
| context_retention | — | — | — |
| has_system_prompt | boolean | — | — |
| repetition_detected | boolean | — | — |
| repetitions | — | — | — |
| topic_drift | — | — | — |
| turn_count | number | — | — |
| user_messages | number | — | — |
No examples provided.
cookie_security_audit ~126
Audit the security attributes of cookies set by any URL. Fetches the URL and inspects all Set-Cookie headers for: HttpOnly, Secure, SameSite, Domain scope, Path scope, Max-Age/Expires, __Host-/__Secure- prefixes. Flags insecure patterns: missing HttpOnly on session cookies, missing Secure flag, SameSite=None without Secure, overly broad Domain, and excessive TTL. Returns per-cookie grades and an overall security score (0–100).
| Name | Type | Req | Description |
|---|---|---|---|
| url | string | yes | Full URL to audit (e.g. https://example.com/login) |
| Name | Type | Req | Description |
|---|---|---|---|
| cookies | array | — | — |
| cookies_found | number | — | — |
| domain | — | — | — |
| host_prefix | — | — | — |
| httpOnly | — | — | — |
| issues | — | — | — |
| max_age | — | — | — |
| message | string | — | — |
| name | — | — | — |
| path | — | — | — |
| sameSite | — | — | — |
| score | number | — | — |
| secure | — | — | — |
| secure_prefix | — | — | — |
| url | — | — | — |
No examples provided.
cors_checker ~114
Check the CORS configuration of a URL the same way a browser would. Returns the main response status, all Access-Control-* headers, the tested origin, and the preflight OPTIONS response. Use this for direct CORS debugging, not just security auditing.
| Name | Type | Req | Description |
|---|---|---|---|
| method | string | — | HTTP method to simulate (default: GET) |
| origin | string | — | Origin header to simulate (default: https://yourdomain.com) |
| url | string | yes | Full URL to test, e.g. https://api.example.com/resource |
| Name | Type | Req | Description |
|---|---|---|---|
| allHeaders | — | — | — |
| corsHeaders | — | — | — |
| method | — | — | — |
| preflight | — | — | — |
| status | — | — | — |
| testedOrigin | — | — | — |
| url | — | — | — |
No examples provided.
cors_test ~131
Test a URL for CORS misconfigurations. Sends preflight (OPTIONS) and cross-origin requests with various Origin headers to detect: wildcard origins with credentials, origin reflection (echoing any origin), null origin acceptance, subdomain wildcard bypass, and missing Vary headers. Returns risk level (safe/low/medium/high/critical), per-test results, and fix recommendations. Essential for API security audits.
| Name | Type | Req | Description |
|---|---|---|---|
| origin | string | — | Custom Origin header to test (default: tests multiple origins automatically) |
| url | string | yes | Full URL to test (e.g. https://api.example.com/endpoint) |
| Name | Type | Req | Description |
|---|---|---|---|
| origins_tested | number | — | — |
| risk_level | — | — | — |
| tests | — | — | — |
| total_findings | number | — | — |
| url | — | — | — |
No examples provided.
cot_analyzer ~104
Analyze a Chain-of-Thought (CoT) or reasoning trace from an LLM. Detects step count, logical flow, conclusion presence, backtracking, and estimates reasoning depth. Useful for o1/o3/DeepSeek-R1 evaluation.
| Name | Type | Req | Description |
|---|---|---|---|
| expected_conclusion | string | — | Expected final answer to check against (optional) |
| reasoning | string | yes | The CoT / reasoning trace text (e.g. from <think> tags or step-by-step output) |
| Name | Type | Req | Description |
|---|---|---|---|
| backtracking_signals | — | — | — |
| conclusion_matches_expected | — | — | — |
| has_conclusion | boolean | — | — |
| markers | — | — | — |
| reasoning_depth | — | — | — |
| reasoning_depth_label | — | — | — |
| step_count | number | — | — |
| total_chars | number | — | — |
| total_lines | number | — | — |
No examples provided.
count_code_lines ~109
Count lines of code: total, code lines, comment lines, blank lines, and comment density. Supports JS/TS, Python, Java/C/C++, Ruby, Go, Shell, HTML/XML, and CSS.
| Name | Type | Req | Description |
|---|---|---|---|
| input | string | yes | Source code to analyze |
| language | string | — | Language hint: "js", "ts", "py", "java", "c", "rb", "go", "sh", "html", "css" (auto-detect if omitted) |
| Name | Type | Req | Description |
|---|---|---|---|
| blank_lines | — | — | — |
| code_lines | — | — | — |
| code_to_comment_ratio | string | — | — |
| comment_density | — | — | — |
| comment_lines | — | — | — |
| language | — | — | — |
| total_lines | — | — | — |
No examples provided.
count_tokens ~77
Estimate the token count of a text string using the cl100k_base approximation (~4 chars/token). Call this BEFORE sending any text to an LLM API to check if it fits within the model context window and to estimate cost. Returns token estimate, character count, and word count.
| Name | Type | Req | Description |
|---|---|---|---|
| input | string | yes | Text to count tokens for |
| Name | Type | Req | Description |
|---|---|---|---|
| chars | — | — | — |
| tokens_estimate | — | — | — |
| words | — | — | — |
No examples provided.
create_confluence_page ~237
Create a new Confluence page from the output of jira_to_test_suite. Formats Gherkin, E2E steps, API tests, and test data as a properly structured Confluence page with code blocks and tables. STATEFUL — creates a new page in the specified space.
| Name | Type | Req | Description |
|---|---|---|---|
| confluence_base_url | string | yes | Atlassian base URL |
| confluence_email | string | yes | Atlassian account email |
| confluence_token | string | yes | Atlassian API token |
| issue_key | string | — | Source Jira issue key (for the page title and source link) |
| issue_url | string | — | Source Jira issue URL (added as a link in the page) |
| parent_page_id | string | — | Optional parent page ID — page will be created as a child of this page |
| space_key | string | yes | Confluence space key where the page will be created, e.g. "QA", "ENG" |
| test_suite | object | yes | The test_suite object from jira_to_test_suite result |
| title | string | — | Page title. Defaults to "Test Plan: {issue_key}" |
| Name | Type | Req | Description |
|---|---|---|---|
| page_id | string | — | — |
| page_url | string | — | — |
| success | boolean | — | — |
| title | string | — | — |
No examples provided.
cron_parse ~63
Parse a cron expression into a human-readable schedule description. Supports standard 5-field cron (minute hour day month weekday).
| Name | Type | Req | Description |
|---|---|---|---|
| expression | string | yes | Cron expression (e.g., "0 9 * * 1-5", "*/15 * * * *") |
| Name | Type | Req | Description |
|---|---|---|---|
| expression | — | — | — |
| fields | object | — | — |
| human_readable | string | — | — |
No examples provided.
cron_validator ~106
Validate a 5-field cron expression, explain the schedule, and preview the next execution times. Use this to debug cron jobs before they reach production. Returns parsed fields, a human-readable description, and upcoming ISO timestamps.
| Name | Type | Req | Description |
|---|---|---|---|
| expression | string | yes | Cron expression with 5 fields, e.g. "*/15 9-18 * * 1-5" |
| next_runs_count | number | — | How many upcoming runs to return (1-50, default: 10) |
| Name | Type | Req | Description |
|---|---|---|---|
| expression | — | — | — |
| fields | object | — | — |
| human_readable | — | — | — |
| next_runs | — | — | — |
| valid | boolean | — | — |
No examples provided.
decode_jwt ~80
Decode a JWT (JSON Web Token) and return its header and payload without verifying the signature. Also reports whether the token is expired and the exact expiry date. Use to inspect claims (sub, iss, exp, roles) during debugging or when integrating with an auth provider.
| Name | Type | Req | Description |
|---|---|---|---|
| token | string | yes | The JWT string to decode (header.payload.signature) |
| Name | Type | Req | Description |
|---|---|---|---|
| expired | — | — | — |
| expiresAt | — | — | — |
| header | — | — | — |
| note | string | — | — |
| payload | — | — | — |
No examples provided.
detect_language ~80
Detect the natural language of a text using n-gram frequency analysis and common word markers. Supports 15 languages: English, French, Spanish, German, Italian, Portuguese, Dutch, Russian, Chinese, Japanese, Korean, Arabic, Polish, Turkish, Swedish.
| Name | Type | Req | Description |
|---|---|---|---|
| input | string | yes | Text to detect language from (min 20 chars for accuracy) |
| Name | Type | Req | Description |
|---|---|---|---|
| confidence | number | — | — |
| lang | — | — | — |
| language | string | — | — |
| matched | — | — | — |
| method | string | — | — |
| name | string | — | — |
| score | — | — | — |
| top_candidates | array | — | — |
No examples provided.
detect_secrets ~97
Scan code or config files for hardcoded secrets: AWS keys, GitHub tokens, OpenAI/Anthropic API keys, Stripe secrets, JWTs, database connection strings, and generic passwords. Returns findings with severity. Run before every commit.
| Name | Type | Req | Description |
|---|---|---|---|
| filename | string | — | Optional filename for context (e.g. ".env", "config.js") |
| input | string | yes | Code or config content to scan (max 500KB) |
| Name | Type | Req | Description |
|---|---|---|---|
| filename | — | — | — |
| findings | — | — | — |
| recommendation | — | — | — |
| risk_level | — | — | — |
| total_findings | number | — | — |
No examples provided.
diff_mappings ~199
Diff a baseline page mapping against a current one and return a CI-style verdict: PASS / FIX / BLOCK, plus per-element drift (ok, renamed, healable, ambiguous, lost, added, rebound). Pure and deterministic — provide two mappings as JSON with "elements" arrays of {role, name, selector, context?}. Use the companion @ia-qa/self-healing package (npm install -g @ia-qa/self-healing) to capture mappings from your app via its local MCP server ia-qa-heal-mcp, or paste the snippet from ia-qa.com/devtools/selector-drift into your browser console.
| Name | Type | Req | Description |
|---|---|---|---|
| after | object | yes | Current page mapping: same shape as before, captured after the UI change. |
| before | object | yes | Baseline page mapping: { page, url, capturedAt, elements: [{role, name, selector, context?}] }. Captured before a UI change. |
| Name | Type | Req | Description |
|---|---|---|---|
| added | array | — | — |
| counts | object | — | — |
| rows | array | — | — |
| verdict | — | — | — |
No examples provided.
diff_text ~116
Compute a unified line-by-line diff between two text strings (LCS algorithm). Returns added/removed/unchanged line counts and formatted diff hunks with configurable context lines (0–20). Use to compare versions of prompts, configs, code snippets, or any text where you need to see exactly what changed.
| Name | Type | Req | Description |
|---|---|---|---|
| a | string | yes | Original (before) text |
| b | string | yes | Modified (after) text |
| context | number | — | Context lines around each change (0–20, default: 3) |
| Name | Type | Req | Description |
|---|---|---|---|
| added | number | — | — |
| diff | string | — | — |
| removed | number | — | — |
| unchanged | number | — | — |
No examples provided.
embedding_similarity ~186
Compute text similarity using local algorithms (Bag of Words, TF-IDF, Character N-grams). No API key needed — runs entirely in-process. NOT real embeddings: for true semantic similarity with vector embeddings, use run_semantic_tests with mode="embeddings" and your OpenAI API key. Supports single pair or batch mode with pipe-separated pairs. Useful for RAG retrieval testing, semantic search evaluation, and text deduplication.
| Name | Type | Req | Description |
|---|---|---|---|
| batch | array | — | Batch mode: array of { text_a, text_b } pairs. Overrides text_a/text_b if provided. |
| methods | array | — | Algorithms to use (default: all three). Options: "bow", "tfidf", "ngram" |
| text_a | string | — | First text to compare (single-pair mode) |
| text_b | string | — | Second text to compare (single-pair mode) |
| Name | Type | Req | Description |
|---|---|---|---|
| count | number | — | — |
| mode | string | — | — |
| results | — | — | — |
| scores | — | — | — |
| text_a | — | — | — |
| text_b | — | — | — |
No examples provided.
env_parse ~93
Parse a .env file content into a JSON object. Handles quoted values (single and double), inline comments, export prefix, and escaped sequences (\n, \t inside double quotes). Returns all key-value pairs. Use in CI/CD pipelines, agent config loaders, or when processing dotenv files programmatically.
| Name | Type | Req | Description |
|---|---|---|---|
| input | string | yes | .env file content to parse (e.g. the output of `cat .env`) |
| Name | Type | Req | Description |
|---|---|---|---|
| count | number | — | — |
| vars | — | — | — |
No examples provided.
escape_html ~67
Escape HTML special characters (&, <, >, ", ') to their safe HTML entities. ALWAYS call this before inserting any user-provided or LLM-generated content into an HTML template to prevent cross-site scripting (XSS) attacks.
| Name | Type | Req | Description |
|---|---|---|---|
| input | string | yes | String to HTML-escape |
| Name | Type | Req | Description |
|---|---|---|---|
| escaped | — | — | — |
| original_length | string | — | — |
No examples provided.
estimate_llm_cost ~162
Estimate the API cost in USD for a given model and token counts. Supports all major 2024–2026 models: GPT-4o, GPT-4.1, o3, o4-mini, Claude Opus 4, Claude Sonnet 4/4.5, Gemini 2.5 Pro/Flash, DeepSeek V3/R1, Grok 3, and legacy models.
| Name | Type | Req | Description |
|---|---|---|---|
| input_tokens | number | yes | Number of input/prompt tokens |
| model | string | yes | Model name, e.g. "gpt-4o", "claude-3.5-sonnet", "deepseek-v3" |
| output_tokens | number | — | Number of output/completion tokens (default: 0) |
| Name | Type | Req | Description |
|---|---|---|---|
| input_cost_usd | string | — | — |
| input_tokens | — | — | — |
| model | — | — | — |
| output_cost_usd | string | — | — |
| output_tokens | — | — | — |
| rates | object | — | — |
| total_cost_usd | string | — | — |
No examples provided.
extract_json_from_text ~88
Extract the first valid JSON object or array embedded in chaotic LLM output (surrounded by markdown fences, prose, or explanatory text). Handles ```json blocks and inline JSON. Call this whenever an LLM returns structured data mixed with explanation text instead of raw JSON.
| Name | Type | Req | Description |
|---|---|---|---|
| input | string | yes | Raw text (e.g., LLM output) that may contain a JSON object or array |
| Name | Type | Req | Description |
|---|---|---|---|
| json | — | — | — |
| source | string | — | — |
No examples provided.
extract_json_path ~88
Extract a value from a JSON string using dot-notation path (e.g., "user.address.city", "items.0.name", "meta.tags"). Supports array index access via numeric path segments.
| Name | Type | Req | Description |
|---|---|---|---|
| input | string | yes | A valid JSON string to traverse |
| path | string | yes | Dot-notation path, e.g. "user.address.city" or "items.0.name" |
| Name | Type | Req | Description |
|---|---|---|---|
| path | — | — | — |
| type | — | — | — |
| value | — | — | — |
No examples provided.
extract_links ~69
Extract all URLs, email addresses, and domain names from text. Returns categorized and deduplicated results. Useful for content auditing, link checking, and web scraping validation.
| Name | Type | Req | Description |
|---|---|---|---|
| input | string | yes | Text to extract links from |
| types | array | — | Types to extract (default: all three) |
| Name | Type | Req | Description |
|---|---|---|---|
| total | — | — | — |
No examples provided.
extract_todos ~111
Extract TODO, FIXME, HACK, BUG, NOTE, OPTIMIZE, and custom tags from any source code or text. Returns line numbers, tag types, and message text. Essential for technical debt auditing.
| Name | Type | Req | Description |
|---|---|---|---|
| include_context | boolean | — | Include full line text (default: true) |
| input | string | yes | Code or text to scan |
| tags | array | — | Custom tags to add (default set: TODO, FIXME, HACK, NOTE, BUG, OPTIMIZE, XXX) |
| Name | Type | Req | Description |
|---|---|---|---|
| counts | — | — | — |
| has_critical | boolean | — | — |
| items | — | — | — |
| total | number | — | — |
No examples provided.
fetch_confluence_page ~200
Fetch a Confluence page and return its content as clean Markdown. Accepts a numeric page_id or a full page URL. Optionally lists direct child pages. BYOK — credentials transit in-memory only, never stored.
| Name | Type | Req | Description |
|---|---|---|---|
| confluence_base_url | string | yes | Atlassian base URL, e.g. "https://mycompany.atlassian.net" |
| confluence_email | string | yes | Atlassian account email (same credentials as Jira) |
| confluence_token | string | yes | Atlassian API token |
| include_children | boolean | — | List direct child pages (id + title) (default: false) |
| page_id | string | — | Confluence page ID (numeric string), e.g. "123456789" |
| page_url | string | — | Full Confluence page URL (alternative to page_id), e.g. "https://mycompany.atlassian.net/wiki/spaces/ENG/pages/123456789" |
| Name | Type | Req | Description |
|---|---|---|---|
| children | array | — | — |
| markdown | string | — | — |
| page_id | string | — | — |
| title | string | — | — |
| url | string | — | — |
No examples provided.
fetch_jira_issue ~201
Fetch a complete Jira issue: summary, description converted to Markdown, status, assignee, priority, labels, custom fields, and optionally comments and attachment metadata. BYOK — credentials transit in-memory only, never stored on ia-qa.com.
| Name | Type | Req | Description |
|---|---|---|---|
| fields | array | — | Specific Jira field names to return. Omit for all standard fields. |
| include_attachments | boolean | — | Include attachment metadata list (default: false) |
| include_comments | boolean | — | Include issue comments, up to 20 (default: true) |
| issue_key | string | yes | Jira issue key, e.g. "PROJ-123" |
| jira_base_url | string | yes | Atlassian base URL, e.g. "https://mycompany.atlassian.net" |
| jira_email | string | yes | Atlassian account email |
| jira_token | string | yes | Atlassian API token (from id.atlassian.com > Security > API tokens) |
| Name | Type | Req | Description |
|---|---|---|---|
| assignee | string | — | — |
| description | string | — | — |
| key | string | — | — |
| labels | array | — | — |
| priority | string | — | — |
| reporter | string | — | — |
| status | string | — | — |
| summary | string | — | — |
| type | string | — | — |
| url | string | — | — |
No examples provided.
fetch_veille_feed ~118
Fetch the latest QA & AI/LLM articles aggregated from curated RSS sources (Google Testing Blog, DEV.to Testing/QA/AI/LLM/Agents, Hugging Face Blog, Simon Willison). Perfect for agents monitoring the QA & AI landscape.
| Name | Type | Req | Description |
|---|---|---|---|
| category | string | — | Filter: "qa" (testing/quality), "ai" (AI/LLM/agents), "all" (default — both) |
| limit | number | — | Max articles to return (default: 20, max: 50) |
| Name | Type | Req | Description |
|---|---|---|---|
| articles | array | — | — |
| category | — | — | — |
| sources_queried | number | — | — |
| total_found | number | — | — |
No examples provided.
few_shot_formatter ~107
Format few-shot examples for LLM prompts. Converts example pairs into formatted blocks. Supports chat format (User/Assistant), XML tags, Markdown, or plain text.
| Name | Type | Req | Description |
|---|---|---|---|
| examples | array | yes | Array of {input, output} pairs |
| format | string | — | Output format (default: chat) |
| input_label | string | — | Label for input (default: User / <input>) |
| output_label | string | — | Label for output (default: Assistant / <output>) |
| Name | Type | Req | Description |
|---|---|---|---|
| example_count | number | — | — |
| format | — | — | — |
| formatted | — | — | — |
| token_estimate | number | — | — |
No examples provided.
find_tool ~197
Search available MCP tools by keyword or category before calling them. Returns matching tool names, descriptions, and optionally their inputSchemas. Call this when you are unsure which tool to use or want to explore the catalogue. Categories: data, encoding, text, llm, qa, rag, dev, security, web.
| Name | Type | Req | Description |
|---|---|---|---|
| category | string | — | Optional: filter by category — data | encoding | text | llm | qa | rag | dev | security | web |
| max_results | number | — | Maximum tools to return (default 10, max 50). Results are ranked by IDF-weighted relevance, so common words like "test" do not inflate the list. |
| query | string | yes | Keyword(s) to search in tool name and description (e.g. "cors", "token", "vector", "json") |
| with_schema | boolean | — | Set true to include inputSchema in results (default: false) |
| Name | Type | Req | Description |
|---|---|---|---|
| category | — | — | — |
| count | number | — | — |
| hint | — | — | — |
| query | — | — | — |
| score | — | — | — |
| tool | — | — | — |
| tools | array | — | — |
| total_matches | number | — | — |
| truncated | boolean | — | — |
No examples provided.
fix_gherkin ~290
Fix Gherkin syntax warnings from a jira_to_test_suite result. Takes the current gherkin text and the _gherkin_warnings array, calls your LLM to fix ONLY the flagged issues (adds missing Given/When/Then steps, etc.), and returns the corrected Gherkin. Lightweight — uses ~300-500 tokens vs ~5k for a full regeneration. Requires BYOK LLM key.
| Name | Type | Req | Description |
|---|---|---|---|
| api_key | string | yes | Your own LLM provider API key (BYOK) — OpenAI "sk-…", Anthropic "sk-ant-…", Google "AIzaSy…", or Groq "gsk_…". There is no server-side key for this tool: if you do not have one, do not call it and do… |
| gherkin | string | yes | The current Gherkin text from the jira_to_test_suite result (test_suite.gherkin). |
| model | string | yes | LLM model to use for the fix, e.g. "gpt-4o-mini". Must belong to the provider whose key you passed in api_key. |
| warnings | array | yes | The _gherkin_warnings array from the jira_to_test_suite result. |
| Name | Type | Req | Description |
|---|---|---|---|
| fixed_gherkin | string | — | — |
| latency_ms | number | — | — |
| model_used | string | — | — |
| remaining_warnings | array | — | — |
| warnings_after | number | — | — |
| warnings_before | number | — | — |
No examples provided.