Skip to content
verify mcp Beta VerifyMCP is currently in beta. If you notice any issues, get in touch and we’ll put it right.

BananaBanana Image, Video & Speech Generation

REMOTE · BANANABANANA.PRO · SCANNED SEP 20

Images, video & speech: Nano Banana, GPT Image, Veo, Omni, Wan, Grok, Gemini TTS. Pay as you go.

Available components

0 this week 38 Trust /100
Trust breakdown (7 categories)

How this component scores in each security and reliability category. Every signal is checked automatically against the live server, and we only credit what we can confirm. How we score → Why this is hard to score →

Endpoint Security94
Transport & Reachability0
Schema Quality & AI Usability0
  • Schema blocked by authentication: the endpoint requires auth we don't have to read it. See how to fix → Unverified
Stability & Change Management0
  • Stability not yet verified: not enough scan history yet (needs a 30-day window).Unverified
Tool Coverage0
  • Tool coverage blocked by authentication: the endpoint requires auth we don't have to read its tools.Unverified
Tool Safety0
  • Tool safety blocked by authentication: the endpoint requires auth we don't have to read its tools.Unverified
Capabilities0
  • Capabilities blocked by authentication: the endpoint requires auth we don't have to read them. See how to fix → Unverified

Unverified: 6 categories

Categories scored 0 because we could not verify them: authentication we do not have, an unreachable endpoint, or not enough scan history. We only credit what we can confirm. Claim this server and supply a read-only token to verify it and lift the score.

Install

How do I install the BananaBanana Image, Video & Speech Generation MCP server?

BananaBanana Image, Video & Speech Generation is a hosted endpoint at https://bananabanana.pro/api/mcp, so there is nothing to install locally. Ready-made configuration for Claude, Cursor, VS Code, Codex and 5 more is on this page, copied from each client's own documentation.

remote · bananabanana.pro

# add to Claude Code
claude mcp add --transport http pro-bananabanana-image-video 'https://bananabanana.pro/api/mcp'
// .cursor/mcp.json
{
  "mcpServers": {
    "pro-bananabanana-image-video": {
      "url": "https://bananabanana.pro/api/mcp"
    }
  }
}
// .vscode/mcp.json
{
  "servers": {
    "pro-bananabanana-image-video": {
      "type": "http",
      "url": "https://bananabanana.pro/api/mcp"
    }
  }
}
# ~/.codex/config.toml
[mcp_servers.pro-bananabanana-image-video]
url = "https://bananabanana.pro/api/mcp"
// opencode.json
{
  "$schema": "https://opencode.ai/config.json",
  "mcp": {
    "pro-bananabanana-image-video": {
      "type": "remote",
      "url": "https://bananabanana.pro/api/mcp",
      "enabled": true
    }
  }
}
# add to OpenClaw
openclaw mcp add pro-bananabanana-image-video --url 'https://bananabanana.pro/api/mcp' --transport streamable-http
# ~/.hermes/config.yaml
mcp_servers:
  pro-bananabanana-image-video:
    url: "https://bananabanana.pro/api/mcp"
// ~/.netclaw/config/netclaw.json
{
  "McpServers": {
    "pro-bananabanana-image-video": {
      "Transport": "http",
      "Url": "https://bananabanana.pro/api/mcp"
    }
  }
}
# add to Vellum
assistant mcp add pro-bananabanana-image-video -t streamable-http -u 'https://bananabanana.pro/api/mcp'
// mcp.json
{
  "mcpServers": {
    "pro-bananabanana-image-video": {
      "type": "http",
      "url": "https://bananabanana.pro/api/mcp"
    }
  }
}

The mcpServers block is a cross-client convention. Remote transports vary, so check your client's docs.

Changelog

Every change we have recorded for this component, newest first. Security-relevant changes are always shown. ▲ marks a change for the better, ▼ a change for the worse; unmarked changes are neutral.

  • 26 Aug 26 0
    • We updated how we score, so this day's move reflects our rubric, not a change to the server See what changed → functional
  • 22 Aug 26 38
    • Endpoint reachability: reachable → behind authorisation security
    • Stability: 0.87 → unverified security
    • Transport: pass → unverified security
    • Authorization: The endpoint enforces authorisation, advertised via RFC 9728 protected-resource metadata. security
    • Capabilities: fail → unverified functional
    • Tool coverage: 100 → unverified functional
    • First check of Schema quality: unverified functional
  • 21 Aug 26 0
    • Tool “generate_image” rewrote its description, which is the text the model reads security
    • Schema quality: excellent → good functional
    • Server version: 1.1.0 → 1.0.8 functional
    • New tool “top_up” functional
    • “generate_image” reworded the description of “model” cosmetic
    • “generate_image” reworded the description of “resolution” cosmetic
  • 19 Aug 26 0
    • Tool “edit_image” rewrote its description, which is the text the model reads security
  • 16 Aug 26 0
    • Tool “generate_image” rewrote its description, which is the text the model reads security
    • Tool “generate_video” rewrote its description, which is the text the model reads security
    • Schema quality: 297 → 384 functional
    • Schema quality: 297 → 382 functional
    • Schema quality: excellent → good functional
    • “edit_image” added an optional parameter “relaxed_filter” cosmetic
    • “generate_image” added an optional parameter “relaxed_filter” cosmetic
    • “generate_video” added an optional parameter “relaxed_filter” cosmetic
    • “edit_image” reworded the description of “relaxed_filter” cosmetic
    • “generate_image” reworded the description of “relaxed_filter” cosmetic
  • 11 Aug 26 0
    • We updated how we score, so this day's move reflects our rubric, not a change to the server See what changed → functional
  • 8 Aug 26 0
    • Tool “generate_image” rewrote its description, which is the text the model reads security
    • Tool “generate_video” rewrote its description, which is the text the model reads security
    • “generate_image” added an optional parameter “reference_images” cosmetic
    • “generate_video” reworded the description of “first_frame” cosmetic
    • “generate_video” reworded the description of “reference_images” cosmetic
  • 6 Aug 26 0
    • The server rewrote its instructions, which are the text every model session reads security
    • Tool “list_models” rewrote its description, which is the text the model reads security
    • Schema quality: 2167 → 2475 functional
    • Server version: 1.0.0 → 1.1.0 functional
    • New tool “generate_speech” functional
Diagnostics

Diagnostic detail from the automated scan of this channel: what the scanner observed at each step, so you can see exactly where a check passed or failed. It is informational only and never changes the trust score.

Captured 20 Sept 2026 · Probed https://bananabanana.pro/api/mcp

TLS valid

Negotiated TLS 1.3 with TLS_AES_128_GCM_SHA256 .

Subject Issuer Valid from Valid until Key Signature Serial
CN=bananabanana.pro CN=YE2,O=Let's Encrypt,C=US 15 Sept 2026 14 Dec 2026 ECDSA 256 ECDSA-SHA384 5177b05796f3c80c0ac37b611484bc320d3
SANs: *.bananabanana.pro, bananabanana.pro
CN=YE2,O=Let's Encrypt,C=US (CA) CN=Root YE,O=ISRG,C=US 3 Sept 2025 2 Sept 2028 ECDSA 384 ECDSA-SHA384 4df3b15dd6c0784c507cd37b58e6f115
CN=Root YE,O=ISRG,C=US (CA) CN=ISRG Root X2,O=Internet Security Research Group,C=US 13 May 2026 2 Sept 2032 ECDSA 384 ECDSA-SHA384 872165fc34b6e5fba8add5b3705fb53a
CN=ISRG Root X2,O=Internet Security Research Group,C=US (CA) CN=ISRG Root X1,O=Internet Security Research Group,C=US 13 May 2026 2 Sept 2032 ECDSA 384 SHA256-RSA 6c8f1dc727c7117f7baf853ac980f9cd

Background: What to check on a remote MCP endpoint →

DNSSEC insecure

Validation of bananabanana.pro. Not signed

Zone DS Keys Algorithms Outcome
. trust_anchor 20326, 38696 8, 8 Verified
pro. present 42154 8 Verified
bananabanana.pro. absent Unsigned (proven) parent-signed NSEC/NSEC3 proves an unsigned delegation
Authentication Enforced and verified

The endpoint asked for a token and published valid RFC 9728 metadata describing how to get one.

Result Enforced and verified
Enforced On connection
HTTP status 401

WWW-Authenticate challenge Bearer realm="bananabanana", error_description="Authentication required. Connect this server with OAuth, or create an API key at https://bananabanana.pro/profile", resource_metadata="https://bananabanana.pro/.well-known/oauth-protected-resource/api/mcp", scope="mcp"

Bearer realm="bananabanana", error_description="Authentication required. Connect this server with OAuth, or create an API key at https://bananabanana.pro/profile", resource_metadata="https://bananabanana.pro/.well-known/oauth-protected-resource/api/mcp", scope="mcp"
Header Value
strict-transport-security max-age=31536000; includeSubDomains
x-content-type-options nosniff
x-frame-options DENY
referrer-policy strict-origin-when-cross-origin
permissions-policy camera=(), microphone=(), geolocation=()
www-authenticate Bearer realm="bananabanana", error_description="Authentication required. Connect this server with OAuth, or create an API key at https://bananabanana.pro/profile", resource_metadata="https://bananabanana.pro/.well-known/oauth-protected-resource/api/mcp", scope="mcp"

Protected resource metadata

Document https://bananabanana.pro/.well-known/oauth-protected-resource/api/mcp
Retrieved Yes
Resource https://bananabanana.pro/api/mcp
Authorisation server https://bananabanana.pro

Background: How OAuth 2.1 works in the 2026 MCP spec →

Transports 2 probes
Transport URL Outcome Status Location
streamable-http https://bananabanana.pro/api/mcp Auth required 401
http (plaintext) http://bananabanana.pro/api/mcp HTTPS enforced 301 https://bananabanana.pro/api/mcp
MCP tools · 10 exposed · ~3,504 tokens

The tools this component advertises to a client, with an estimated token cost for each. Expand a tool to see its parameters and schema. The per-tool counts are indicative and are not scored directly; the schema's total context footprint is one signal in Schema Quality & AI Usability. A tool's description is untrusted text the model reads on every call, which is what makes this list a security surface and not just an inventory: how tool poisoning works →

Tool Tokens
edit_image ~457

Edit / refine a previously generated image with a text instruction on Nano Banana 2 Lite / 2 / Pro (multi-turn editing: change colors, remove objects, restyle, etc.). This tool modifies an image that already exists — use generate_image to create a new image from a prompt. Pass the job_id of a COMPLETED image generation as source_generation_id. Charged like a single image of the chosen model/resolution; auto-refund on failure. Example: {"source_generation_id": "cmxyz...", "prompt": "make the background pure white and add soft shadow"}

NameTypeReqDescription
aspect_ratiostring
idempotency_keystring
modelstring
output_formatstring
promptstringyesThe edit instruction.
relaxed_filterbooleanSwitch this call to Google's own permissive presets: safety thresholds OFF and person generation allowed for adults (safetySettings and personGeneration=ALLOW_ADULT — documented Vertex AI parameters,…
resolutionstring
seedinteger
source_generation_idstringyesjob_id of a completed image generation owned by this account.

No output schema declared.

No examples provided.

edit_video ~474

Edit an EXISTING video with Gemini Omni Flash (video-to-video): restyle it, replace or add objects, relight the scene, change the mood — motion and composition of the source clip are preserved. Billed by output length at $0.10/s; cost confirmation is mandatory (first call returns the quote and charges nothing). Source: either source_generation_id (a completed video from this account — see list_generations) or video_url (public http(s) link, max 200 MB). The source is normalised to MP4 720p and the FIRST 10 SECONDS (model limit); output is 720p with sound, aspect ratio follows the source. OUTPUT LENGTH ALWAYS EQUALS SOURCE LENGTH (the model cannot stretch or shorten a clip), so duration only works downwards: it trims the source to the first N seconds. Omit duration to edit the whole clip — the quote tells you the resolved length. Returns a job_id; poll get_result. Failed edits are auto-refunded. Example: {"prompt": "make the whole scene look like a pencil sketch, keep the motion identical", "source_generation_id": "clx…", "confirm_cost": 1}

NameTypeReqDescription
audio_promptstringDescribe the desired sound — Omni Flash always generates audio.
confirm_costnumberThe quoted USD cost you accept. Omit on the first call to get the quote.
durationintegerOptional: trim the source to the first N seconds and edit only that part ($0.10/s). Values above the source length are ignored — the clip cannot be made longer.
idempotency_keystringOptional unique key; retries with the same key never double-charge.
promptstringyesWhat to change in the video.
source_generation_idstringjob_id of a completed video generation on this account.
source_refstringReturned by the quote when video_url is used: pass it back with confirm_cost to reuse the already downloaded clip instead of downloading it again.
video_urlstringPublic http(s) URL of the source video (mp4/mov/webm/mkv/avi/wmv/flv/3gpp, max 200 MB). Use instead of source_generation_id.

No output schema declared.

No examples provided.

generate_image ~812

Start an AI image generation (Google Nano Banana family). Charges the account balance immediately and returns a job_id — poll get_result for the finished image URLs. Typical completion: 10–60 seconds. The default nano-banana-2-lite model is the cheapest option and produces 1024px images; choose nano-banana-2 explicitly for 512px, 2048px or 4096px output. Optional reference_images provide the model with the actual subject, product, character or style pixels; Nano Banana Pro supports up to 14 references. Costs $0.03–$0.20 per image depending on model and resolution (see list_models). Failed generations are automatically refunded. If a legitimate prompt is rejected by Google's content filter, retry with relaxed_filter: true. Generating several images at once (number_of_images > 1) is a batch: the first call returns a price quote and charges nothing — repeat the call with confirm_cost set to the quoted amount to start. Example: {"prompt": "studio photo of a ceramic mug on linen, soft daylight", "model": "nano-banana-2-lite", "aspect_ratio": "4:5", "resolution": "1024"}

NameTypeReqDescription
aspect_ratiostring
confirm_costnumberRequired for batches (number_of_images > 1): the quoted total USD cost you accept.
idempotency_keystringOptional unique key; retries with the same key never double-charge.
modelstringnano-banana-2-lite: cheapest default, 1024 only. nano-banana-2: choose for 512, 2048 or 4096 output. nano-banana-pro: top quality, up to 4K.
negative_promptstring
number_of_imagesintegerVariants per call. >1 requires confirm_cost.
output_formatstring
promptstringyesWhat to generate. English works best.
reference_imagesarrayActual visual references for the generated image. Each item is a job_id of a completed image generation on this account, a public http(s) image URL, or an inline data:image/png|jpeg|webp;base64,... U…
relaxed_filterbooleanSwitch this call to Google's own permissive presets: safety thresholds OFF and person generation allowed for adults (safetySettings and personGeneration=ALLOW_ADULT — documented Vertex AI parameters,…
resolutionstringLite requires 1024. Choose nano-banana-2 for 512, 2048 or 4096; 4096 costs the most.
seedintegerFor reproducible results.

No output schema declared.

No examples provided.

generate_speech ~300

Generate natural speech with Gemini 3.1 Flash TTS Preview. This synchronous tool returns a hosted WAV URL directly (no get_result polling). Supports one voice or an exactly two-speaker dialogue, automatic language detection or a BCP-47 language_code, natural-language direction for accent/tone/pace, and inline performance tags such as [whispers], [laughs], [very slow] and [excited]. Price is $0.01 per started 200 transcript characters; the account is charged only after Google has returned valid audio. Example: {"text":"[cheerfully] Welcome to BananaBanana!","voice":"Kore","style":"Warm product announcement, medium pace."}

NameTypeReqDescription
idempotency_keystringOptional unique key; retries with the same key never double-charge.
language_codestringOptional language/locale such as en-US, ru-RU, ja-JP or es-MX. Omit for automatic detection.
speakersarrayExactly two dialogue speakers. Their names must prefix the turns in text.
stylestringOptional overall direction: persona, scene, emotion, accent, pace, pronunciation and delivery notes.
textstringyesExact transcript to speak. For dialogue, prefix every turn with the matching speaker name, e.g. Sam: Hello.
voicestringSingle-speaker voice. Ignored when speakers is provided.

No output schema declared.

No examples provided.

generate_video ~1,070

Start an AI video generation (Google Veo 3.1 family or Gemini Omni Flash). EXPENSIVE: $0.10–$4.40 per clip. Cost confirmation is mandatory: the first call always returns a USD quote and charges nothing — repeat the call with confirm_cost set to the quoted amount to actually start. Returns a job_id; poll get_result (videos take 1–10+ minutes). Failed generations are auto-refunded. Models: veo-3.1-fast (default, good quality/price), veo-3.1 (best Veo quality), veo-3.1-lite (cheapest, 720p/1080p), omni-flash (always has sound, any duration from 3 to 10 s at $0.10/s, supports conversational editing via edit_from_generation_id). Content filtering differs by model: omni-flash is by far the strictest, so a clip rejected as SAFETY_FILTERED on omni-flash is often produced by veo-3.1-fast without changing a word. Video has no configurable safety settings on Google's side (relaxed_filter is accepted for Veo but only pins its default personGeneration=allow_adult), and refused clips are refunded, so a retry is cheap in money and expensive only in time. Image inputs: first_frame animates a still picture, reference_images keep a subject/style consistent — both accept a job_id of a completed image generation on this account, a public image URL, or inline base64 image data, and on omni-flash they can be combined (up to 10 images total). Example: {"prompt": "drone shot over a misty pine forest at sunrise", "model": "veo-3.1-fast", "duration": 8, "resolution": "720p", "confirm_cost": 0.70}

NameTypeReqDescription
aspect_ratiostring
audio_promptstringDescribe the desired sound (used when audio is on).
confirm_costnumberThe quoted USD cost you accept. Omit on the first call to get the quote.
durationintegerClip length in seconds. Veo accepts only 4, 6 or 8; omni-flash accepts any value from 3 to 10.
edit_from_generation_idstringomni-flash only: job_id of a completed omni video to refine conversationally; prompt describes the changes. Duration, aspect ratio and the scene are inherited from that clip — duration is ignored her…
first_framestringStart the clip from a still image, which is then animated. Either a job_id of a completed image generation on this account (see list_generations), a public http(s) image URL, or an inline data:image/…
idempotency_keystringOptional unique key; retries with the same key never double-charge.
modelstring
negative_promptstringWhat to avoid. On omni-flash it is appended to the prompt as plain text (the model has no separate negative field).
promptstringyes
reference_imagesarrayReference images that keep a subject, character or style consistent (they are NOT used as literal frames). Each item is a job_id of a completed image generation on this account, a public http(s) imag…
relaxed_filterbooleanVeo only (omni-flash ignores it), and much weaker than its image counterpart: Google's video API exposes no configurable safety settings at all, so this flag can do only one thing — pin personGenerat…
resolutionstring4k only on veo-3.1 / veo-3.1-fast; omni-flash is 720p only.
seedintegerVeo only.
with_audiobooleanNative audio for Veo models (costs more). omni-flash always has audio.

No output schema declared.

No examples provided.

get_account ~50

Get the current account balance (USD), this API key's name, optional daily spend cap and how much of it is used today. Free, no charge. Use it to check affordability before starting expensive generations.

Input schema present but exposes no named parameters.

No output schema declared.

No examples provided.

get_result ~124

Get the status and result of a generation job started with generate_image / edit_image / generate_video / edit_video. Waits up to wait_seconds for completion before returning (long-poll). On success returns hosted media URLs (valid 24 h — call again for fresh links), cost_charged_usd and balance_remaining_usd, plus a small inline preview for images. Free, no charge. Poll roughly every 10–15 s for videos.

NameTypeReqDescription
job_idstringyes
wait_secondsintegerHow long to wait server-side before answering.

No output schema declared.

No examples provided.

list_generations ~98

List this account's recent generations (both MCP and website) — id, type, model, status, cost and prompt preview. Use it to find a job_id to re-download results or to pick a source for edit_image / edit_video / generate_video edit_from_generation_id. Free, no charge.

NameTypeReqDescription
limitinteger
statusstringFilter by status.
typestringFilter by media type.

No output schema declared.

No examples provided.

list_models ~60

List all available image, video and speech generation models with current per-unit USD prices, supported resolutions, durations and constraints. Prices come from the same source as the website — call this before quoting costs to a user or choosing a model. Free, no charge.

Input schema present but exposes no named parameters.

No output schema declared.

No examples provided.

top_up ~59

Get a balance top-up link. OAuth connections receive a one-time deposit-only link valid for 30 minutes; opening it cannot expose API keys, profile data or generation history. Bearer API-key users receive the normal profile link. Free, no charge.

Input schema present but exposes no named parameters.

No output schema declared.

No examples provided.

Common questions

What is the BananaBanana Image, Video & Speech Generation MCP server?

BananaBanana Image, Video & Speech Generation is an MCP server listed in the public MCP registry as pro.bananabanana/image-video. Images, video & speech: Nano Banana, GPT Image, Veo, Omni, Wan, Grok, Gemini TTS. Pay as you go. This page covers its hosted endpoint (https://bananabanana.pro/api/mcp).

Is the BananaBanana Image, Video & Speech Generation MCP server safe to use?

BananaBanana Image, Video & Speech Generation scores 38 out of 100 on VerifyMCP. That is a record of what we were able to check automatically, not an endorsement. The category breakdown on this page shows every signal behind the number, including the ones we could not confirm.

What tools does the BananaBanana Image, Video & Speech Generation MCP server expose?

BananaBanana Image, Video & Speech Generation exposes 10 tools: list_models, get_account, top_up, generate_image, edit_image, and 5 more. Their descriptions and schemas cost roughly 3,504 tokens of context every time the server is loaded.

Does the BananaBanana Image, Video & Speech Generation MCP server require authentication?

Yes. BananaBanana Image, Video & Speech Generation asked us for credentials when we connected, so you will need to authorise it in your MCP client before it can do anything.

Is the BananaBanana Image, Video & Speech Generation MCP server still maintained?

BananaBanana Image, Video & Speech Generation is still listed as active in the MCP registry. We last reached this channel on 20 September 2026. Those dates come from our own scans of the registry and the channel itself, not from anything the publisher announced.