io.github.apasztetnik/el-buen-agente-mcp
REMOTE · EL-BUEN-AGENTE-MCP-PRODUCTION.UP.RAILWAY.APP · SCANNED AUG 3
La guía 'El Buen Agente' como 18 tools para evaluar, mejorar y construir agentes LLM. En español.
Available components
How this component scores in each security and reliability category. Every signal is checked automatically against the live server, and we only credit what we can confirm. How we score →
Endpoint Security57
- The endpoint's TLS certificate is valid, in date, and uses a strong key. View diagnostics → Pass
- Authorisation not fully verified: no authorisation is required to call this server, and 18 tool(s) never declared a destructiveHint. The MCP spec treats an absent hint as destructive by default, so we cannot call this surface safe. See how to fix → View diagnostics → Unverified
- HTTPS is enforced; there's no plaintext access path. View diagnostics → Pass
- HSTS check failed: the Strict-Transport-Security header is absent. See how to fix → View diagnostics → Fail
- DNSSEC check failed: this domain isn't protected by DNSSEC. See how to fix → View diagnostics → Fail
Transport & Reachability100
- Verified streamable-http transport via a live MCP handshake. View diagnostics → Pass
Schema Quality & AI Usability81
- 100% of prompts and resources have a non-trivial description (not blank, and not just the item's name).Pass
- AI-judged instruction clarity (good).Pass
- Context-footprint check failed: tool/resource definitions use about 3380 tokens (~105/item across 32 items; 18 tools + 14 resources), over budget; trim descriptions and params. See how to fix → Fail
- Usage-examples check failed: none of the tools include examples. See how to fix → Fail
Stability & Change Management27
- Stability observed for 8 of 30 days with no destabilising changes; credit accrues until the full window elapses.Partial
Tool Coverage100
- 100% of tools have a non-trivial description (not blank, and not just the tool's name).Pass
- 100% of tool parameters carry a description.Pass
- Structured output schemas are declared (6% of tools); any adoption earns full credit.Pass
Capabilities100
- Implements a supported MCP spec version (2025-11-25); the latest is 2026-07-28.Pass
Add this component to your MCP client. Where a client-specific snippet is available, pick your client below and copy it straight into your config; otherwise use the connection detail shown.
remote · el-buen-agente-mcp-production.up.railway.app
claude mcp add --transport http apasztetnik-el-buen-agente-mcp https://el-buen-agente-mcp-production.up.railway.app/mcp
[mcp_servers.apasztetnik-el-buen-agente-mcp] url = "https://el-buen-agente-mcp-production.up.railway.app/mcp"
{
"$schema": "https://opencode.ai/config.json",
"mcp": {
"apasztetnik-el-buen-agente-mcp": {
"type": "remote",
"url": "https://el-buen-agente-mcp-production.up.railway.app/mcp",
"enabled": true
}
}
} openclaw mcp add apasztetnik-el-buen-agente-mcp --url https://el-buen-agente-mcp-production.up.railway.app/mcp --transport streamable-http
mcp_servers:
apasztetnik-el-buen-agente-mcp:
url: "https://el-buen-agente-mcp-production.up.railway.app/mcp" {
"mcpServers": {
"apasztetnik-el-buen-agente-mcp": {
"type": "http",
"url": "https://el-buen-agente-mcp-production.up.railway.app/mcp"
}
}
} The mcpServers block is a cross-client convention. Remote transports vary, so check your client's docs.
Every change we have recorded for this component, newest first. Security-relevant changes are always shown. ▲ marks a change for the better, ▼ a change for the worse; unmarked changes are neutral.
- 3 Aug 26 +1
No change was recorded against any check on this day. Stability & Change Management went from 23 to 27. That category is still filling its 30-day observation window: 7 days of observed history at the previous scan, 8 at this one. The score rises as the window fills, whether or not the server changes.
- 31 Jul 26 0
- We updated how we score, so this day's move reflects our rubric, not a change to the server See what changed → functional
- 30 Jul 26 0
- We updated how we score, so this day's move reflects our rubric, not a change to the server See what changed → functional
- 29 Jul 26 +1
No change was recorded against any check on this day. Stability & Change Management went from 7 to 10. That category is still filling its 30-day observation window: 2 days of observed history at the previous scan, 3 at this one. The score rises as the window fills, whether or not the server changes.
- 28 Jul 26 +1
No change was recorded against any check on this day. Stability & Change Management went from 3 to 7. That category is still filling its 30-day observation window: 1 days of observed history at the previous scan, 2 at this one. The score rises as the window fills, whether or not the server changes.
- 27 Jul 26 0
- We updated how we score, so this day's move reflects our rubric, not a change to the server See what changed → functional
- 26 Jul 26 65
First indexed and scored.
Diagnostic detail from the automated scan of this channel: what the scanner observed at each step, so you can see exactly where a check passed or failed. It is informational only and never changes the trust score.
Captured 3 Aug 2026 · Probed https://el-buen-agente-mcp-production.up.railway.app/mcp
TLS valid
Negotiated TLS 1.3 with TLS_AES_128_GCM_SHA256 .
| Subject | Issuer | Valid from | Valid until | Key | Signature | Serial |
|---|---|---|---|---|---|---|
| CN=*.up.railway.app | CN=YE1,O=Let's Encrypt,C=US | 29 Jul 2026 | 27 Oct 2026 | ECDSA 256 | ECDSA-SHA384 | 6da79bb561da3efeb0e751ca21abd3999fe |
| SANs: *.up.railway.app, up.railway.app | ||||||
| CN=YE1,O=Let's Encrypt,C=US (CA) | CN=Root YE,O=ISRG,C=US | 3 Sept 2025 | 2 Sept 2028 | ECDSA 384 | ECDSA-SHA384 | 5ddd70dd31f801c85c186a7a04b80afe |
| CN=Root YE,O=ISRG,C=US (CA) | CN=ISRG Root X2,O=Internet Security Research Group,C=US | 13 May 2026 | 2 Sept 2032 | ECDSA 384 | ECDSA-SHA384 | 872165fc34b6e5fba8add5b3705fb53a |
| CN=ISRG Root X2,O=Internet Security Research Group,C=US (CA) | CN=ISRG Root X1,O=Internet Security Research Group,C=US | 13 May 2026 | 2 Sept 2032 | ECDSA 384 | SHA256-RSA | 6c8f1dc727c7117f7baf853ac980f9cd |
DNSSEC insecure
Validation of el-buen-agente-mcp-production.up.railway.app. — Not signed
| Zone | DS | Keys | Algorithms | Outcome |
|---|---|---|---|---|
| . | trust_anchor | 20326, 38696 | 8, 8 | Verified |
| app. | present | 23684 | 8 | Verified |
| railway.app. | absent | Unsigned (proven) parent-signed NSEC/NSEC3 proves an unsigned delegation |
Authentication No authorisation required
The endpoint answered without asking for a token. Anyone who knows the URL can reach it.
| Result | No authorisation required |
|---|---|
| HTTP status | 200 |
Transports 2 probes
| Transport | URL | Outcome | Status | Location |
|---|---|---|---|---|
| streamable-http | https://el-buen-agente-mcp-production.up.railway.app/mcp | Verified | 200 | |
| http (plaintext) | http://el-buen-agente-mcp-production.up.railway.app/mcp | HTTPS enforced | 301 | https://el-buen-agente-mcp-production.up.railway.app/mcp |
The tools this component advertises to a client, with an estimated token cost for each. Expand a tool to see its parameters and schema. The per-tool counts are indicative and are not scored directly; the schema's total context footprint is one signal in Schema Quality & AI Usability.
aplicar_challenger §5: Relación con el humano: copiloto, reviewer, challenger ~136
Aplica el patrón challenger/red-team a la definición o a una decisión del agente: contraargumentos basados en datos, autocrítica y gate de calidad. Útil como segunda pasada adversarial. Recibe la definición del agente y devuelve un brief de evaluación estructurado.
| Name | Type | Req | Description |
|---|---|---|---|
| agent_definition | string | yes | Definición completa del agente a mejorar: system prompt, frontmatter, configuración, descripción de tools y cualquier doc de diseño. Cuanto más completa, mejor la evaluación. |
| language | string | — | Idioma de la respuesta (default "es"). / Response language: pass "en" for English output. |
No output schema declared.
No examples provided.
auditar_contexto §6: El contexto como activo estratégico ~155
Audita el contexto del agente: separación en 3 capas (identidad/dominio/referencia), relevancia y caducidad de datos, estrategia de integración (directo/snapshot/RAG/estático), gobernanza, least-privilege, trazabilidad y defensa anti prompt-injection. Recibe la definición del agente y devuelve un brief de evaluación estructurado.
| Name | Type | Req | Description |
|---|---|---|---|
| agent_definition | string | yes | Definición completa del agente a mejorar: system prompt, frontmatter, configuración, descripción de tools y cualquier doc de diseño. Cuanto más completa, mejor la evaluación. |
| language | string | — | Idioma de la respuesta (default "es"). / Response language: pass "en" for English output. |
No output schema declared.
No examples provided.
challenger_decision §5: Red-team de una decisión ~104
Genera el brief para cuestionar una decisión concreta del agente: 3 razones basadas en datos para NO hacerla. No bloquea, informa al humano.
| Name | Type | Req | Description |
|---|---|---|---|
| contexto | string | — | Contexto y datos relevantes a la decisión. |
| decision | string | yes | La decisión o recomendación del agente que se quiere cuestionar. |
| language | string | — | Idioma de la respuesta (default "es"). / Response language: pass "en" for English output. |
No output schema declared.
No examples provided.
checklist_nacimiento §11: Checklist de nacimiento (19 puntos) ~126
Corre el checklist completo de 19 puntos contra la definición del agente. Es el gate final antes de mergear: el agente debe NACER cumpliéndolo, no corregirse después. Devuelve veredicto punto por punto.
| Name | Type | Req | Description |
|---|---|---|---|
| agent_definition | string | yes | Definición completa del agente a mejorar: system prompt, frontmatter, configuración, descripción de tools y cualquier doc de diseño. Cuanto más completa, mejor la evaluación. |
| language | string | — | Idioma de la respuesta (default "es"). / Response language: pass "en" for English output. |
No output schema declared.
No examples provided.
construir_agente Construir la definición final del agente ~214
El paso de CIERRE del flujo: toma la definición iterada (tras pasar por las revisiones) y produce el artefacto final listo para usar: identity layer, tools con least-privilege, límites duros, schema de output, gates, contexto, evaluación. Llamala cuando checklist_nacimiento dé APTO.
| Name | Type | Req | Description |
|---|---|---|---|
| contrato | string | — | El contrato generado por generar_contrato, si existe. |
| definicion_final | string | yes | Definición completa del agente a mejorar: system prompt, frontmatter, configuración, descripción de tools y cualquier doc de diseño. Cuanto más completa, mejor la evaluación. |
| formato | string | — | Formato del artefacto: 'markdown' (doc de diseño completo, default), 'claude_skill' (SKILL.md con frontmatter), 'system_prompt' (prompt + config para cualquier framework). |
| language | string | — | Idioma de la respuesta (default "es"). / Response language: pass "en" for English output. |
No output schema declared.
No examples provided.
disenar_evaluacion §7: ¿Cómo saber si funciona bien? ~147
Evalúa el plan de evaluación del agente (o ayuda a crearlo): 3 dimensiones (capacidades/trayectoria/resultado), métricas (tasa de éxito, consistencia, coste por tarea, adopción), golden set de 20-50 tareas, monitoreo de drift y self-consistency para alto stake.
| Name | Type | Req | Description |
|---|---|---|---|
| agent_definition | string | yes | Definición completa del agente a mejorar: system prompt, frontmatter, configuración, descripción de tools y cualquier doc de diseño. Cuanto más completa, mejor la evaluación. |
| language | string | — | Idioma de la respuesta (default "es"). / Response language: pass "en" for English output. |
No output schema declared.
No examples provided.
evaluar_autonomia §3: Nivel de autonomía permitido ~143
Determina el nivel de autonomía adecuado (copiloto / ejecutor supervisado / autónomo con guardrails) y verifica los mecanismos de reducción de riesgo: sandbox/shadow mode, límites duros, frenos progresivos, override humano. Recibe la definición del agente y devuelve un brief de evaluación estructurado.
| Name | Type | Req | Description |
|---|---|---|---|
| agent_definition | string | yes | Definición completa del agente a mejorar: system prompt, frontmatter, configuración, descripción de tools y cualquier doc de diseño. Cuanto más completa, mejor la evaluación. |
| language | string | — | Idioma de la respuesta (default "es"). / Response language: pass "en" for English output. |
No output schema declared.
No examples provided.
evaluar_necesidad §0: ¿De verdad hace falta un agente? ~127
Evalúa si el problema justifica un agente o se resuelve con menos (prompt, workflow, skill). Detecta antipatrones: agente genérico, sobre-orquestación, agente sin contexto, autonomía total día 1. Usala ANTES de construir, o para cuestionar un agente existente.
| Name | Type | Req | Description |
|---|---|---|---|
| language | string | — | Idioma de la respuesta (default "es"). / Response language: pass "en" for English output. |
| problema | string | yes | Descripción del problema que se quiere resolver y, si existe, cómo lo resuelve el agente actual. |
No output schema declared.
No examples provided.
evaluar_sistema §9: De skills aisladas a sistema coherente ~145
Evalúa cómo el agente encaja en el sistema mayor: catálogo de skills, reutilización vs especialización, orquestación ligera vs orquestador, qué automatizar vs supervisar, y memorias separadas por agente.
| Name | Type | Req | Description |
|---|---|---|---|
| agent_definition | string | yes | Definición completa del agente a mejorar: system prompt, frontmatter, configuración, descripción de tools y cualquier doc de diseño. Cuanto más completa, mejor la evaluación. |
| ecosistema | string | — | Otros agentes/skills existentes en el sistema, si los hay. |
| language | string | — | Idioma de la respuesta (default "es"). / Response language: pass "en" for English output. |
No output schema declared.
No examples provided.
generar_contrato §8: Generar el contrato del agente ~235
Genera el contrato formal del agente (patrón contractor de §8) a partir de los campos provistos. Los campos faltantes quedan marcados como [PENDIENTE] para completar.
| Name | Type | Req | Description |
|---|---|---|---|
| autonomia | string | — | copiloto | supervisado | autónomo con guardrails (+ plan de progresión). |
| coste | string | — | Modelo + estimado de tokens/mes + tope. |
| evaluacion | string | — | Métricas de éxito (verde/amarillo/rojo) + cadencia de review. |
| inputs | string | — | Qué consume. |
| language | string | — | Idioma de la respuesta (default "es"). / Response language: pass "en" for English output. |
| no_puede | string | — | Acciones que requieren aprobación humana / fuera de alcance. |
| nombre | string | yes | Nombre del agente. |
| output | string | — | Schema exacto + formato legible para humano. |
| problema | string | — | Qué resuelve (1 frase) + qué deliberadamente NO toca. |
| puede | string | — | Acciones autónomas dentro de límites. |
No output schema declared.
No examples provided.
get_el_buen_agente Obtener la guía completa ~57
Devuelve la guía completa 'El Buen Agente (v2)' en Markdown. Usala para contexto general; para operar sobre un agente concreto usá las tools evaluar_*/revisar_*/auditar_*.
Input schema present but exposes no named parameters.
No output schema declared.
No examples provided.
plan_de_inicio §12: Cómo empezar ~115
Genera el plan de arranque correcto para un agente nuevo: un agente, un problema concreto, un humano revisando. Verifica que la tarea elegida sea recurrente, costosa en tiempo y con datos accesibles.
| Name | Type | Req | Description |
|---|---|---|---|
| equipo | string | — | Quiénes lo van a usar y revisar. |
| language | string | — | Idioma de la respuesta (default "es"). / Response language: pass "en" for English output. |
| problema | string | yes | El problema o tarea candidata para el primer agente. |
No output schema declared.
No examples provided.
plan_exposicion_mcp §10: Exponer el agente/skill vía MCP ~134
Evalúa qué partes del agente conviene exponer a la economía de agentes y cómo: qué modelar como tool (capacidad accionable), resource (doc/dato legible) o prompt (plantilla), y qué merece UI vs API/MCP.
| Name | Type | Req | Description |
|---|---|---|---|
| agent_definition | string | yes | Definición completa del agente a mejorar: system prompt, frontmatter, configuración, descripción de tools y cualquier doc de diseño. Cuanto más completa, mejor la evaluación. |
| language | string | — | Idioma de la respuesta (default "es"). / Response language: pass "en" for English output. |
No output schema declared.
No examples provided.
recomendar_flujo Recomendar el orden de las tools ~86
Devuelve el flujo recomendado de tools según la situación: agente nuevo (diseño desde cero) o agente existente (mejora). Llamala PRIMERO si no sabés por dónde empezar.
| Name | Type | Req | Description |
|---|---|---|---|
| situacion | string | yes | 'nuevo' = diseñar un agente desde cero; 'existente' = mejorar un agente que ya está definido o en producción. |
No output schema declared.
No examples provided.
revisar_frontera_ejecucion §4: Qué recomienda y qué ejecuta (frontera explícita desde el diseño) ~145
Verifica que la línea entre lo que el agente recomienda y lo que ejecuta esté definida en código (status + gates), no descubierta en producción: qué ejecuta directo, qué queda pendiente de aprobación, qué se rechaza. Recibe la definición del agente y devuelve un brief de evaluación estructurado.
| Name | Type | Req | Description |
|---|---|---|---|
| agent_definition | string | yes | Definición completa del agente a mejorar: system prompt, frontmatter, configuración, descripción de tools y cualquier doc de diseño. Cuanto más completa, mejor la evaluación. |
| language | string | — | Idioma de la respuesta (default "es"). / Response language: pass "en" for English output. |
No output schema declared.
No examples provided.
revisar_outputs §2: Inputs esperados y outputs útiles para humanos ~132
Verifica que los outputs sean accionables: schema estricto (JSON Schema/Pydantic), resumen legible para humanos separado del razonamiento, y exposición de qué gates pasó/falló. Recibe la definición del agente y devuelve un brief de evaluación estructurado.
| Name | Type | Req | Description |
|---|---|---|---|
| agent_definition | string | yes | Definición completa del agente a mejorar: system prompt, frontmatter, configuración, descripción de tools y cualquier doc de diseño. Cuanto más completa, mejor la evaluación. |
| language | string | — | Idioma de la respuesta (default "es"). / Response language: pass "en" for English output. |
No output schema declared.
No examples provided.
revisar_rol_y_frontera §1: Rol claro y frontera de responsabilidad ~135
Verifica que el agente tenga rol claro, dominio acotado y frontera explícita de qué NO es su responsabilidad, escrita en el identity layer y enforced en código donde se pueda. Recibe la definición del agente y devuelve un brief de evaluación estructurado.
| Name | Type | Req | Description |
|---|---|---|---|
| agent_definition | string | yes | Definición completa del agente a mejorar: system prompt, frontmatter, configuración, descripción de tools y cualquier doc de diseño. Cuanto más completa, mejor la evaluación. |
| language | string | — | Idioma de la respuesta (default "es"). / Response language: pass "en" for English output. |
No output schema declared.
No examples provided.
validar_veredicto Validar el veredicto del checklist (para CI) ~105
Cierra el ciclo de checklist_nacimiento con un contrato a nivel protocolo. Pasale los 19 puntos con su estado (ok|parcial|falta) y devuelve structuredContent validado: conteos y veredicto normalizado (apto solo si faltas === 0). Pensada para consumo programático / gates de CI: no depende de parsear texto.
| Name | Type | Req | Description |
|---|---|---|---|
| puntos | array | yes | Los 19 puntos del checklist con su estado evaluado. |
| Name | Type | Req | Description |
|---|---|---|---|
| aptos | integer | yes | — |
| completo | boolean | yes | true si están los 19 puntos sin números repetidos. |
| faltas | integer | yes | — |
| parciales | integer | yes | — |
| veredicto | string | yes | — |
No examples provided.