Skip to content
verify mcp Beta VerifyMCP is currently in beta. If you notice any issues, email [email protected] and we’ll put it right.

Speech AI - Pronunciation, STT & TTS

REMOTE · APIM-AI-APIS.AZURE-API.NET · SCANNED AUG 3

Pronunciation scoring, speech-to-text, and text-to-speech for language learning

+7 this week 73 Trust /100
Trust breakdown (6 categories)

How this component scores in each security and reliability category. Every signal is checked automatically against the live server, and we only credit what we can confirm. How we score →

Endpoint Security77
Transport & Reachability100
Schema Quality & AI Usability64
  • AI-judged instruction clarity (excellent).Pass
  • Context-footprint check failed: tool/resource definitions use about 2017 tokens (~201/item across 10 items; 10 tools + 0 resources), over budget; trim descriptions and params. See how to fix → Fail
  • Usage-examples check failed: none of the tools include examples. See how to fix → Fail
Stability & Change Management27
  • Stability observed for 8 of 30 days with no destabilising changes; credit accrues until the full window elapses.Partial
Tool Coverage100
  • 100% of tools have a non-trivial description (not blank, and not just the tool's name).Pass
  • 100% of tool parameters carry a description.Pass
Capabilities100
  • Implements a supported MCP spec version (2025-11-25); the latest is 2026-07-28.Pass
Install

Add this component to your MCP client. Where a client-specific snippet is available, pick your client below and copy it straight into your config; otherwise use the connection detail shown.

remote · apim-ai-apis.azure-api.net

# add to Claude Code
claude mcp add --transport http fasuizu-br-speech-ai https://apim-ai-apis.azure-api.net/mcp/pronunciation/mcp
# ~/.codex/config.toml
[mcp_servers.fasuizu-br-speech-ai]
url = "https://apim-ai-apis.azure-api.net/mcp/pronunciation/mcp"
// opencode.json
{
  "$schema": "https://opencode.ai/config.json",
  "mcp": {
    "fasuizu-br-speech-ai": {
      "type": "remote",
      "url": "https://apim-ai-apis.azure-api.net/mcp/pronunciation/mcp",
      "enabled": true
    }
  }
}
# add to OpenClaw
openclaw mcp add fasuizu-br-speech-ai --url https://apim-ai-apis.azure-api.net/mcp/pronunciation/mcp --transport streamable-http
# ~/.hermes/config.yaml
mcp_servers:
  fasuizu-br-speech-ai:
    url: "https://apim-ai-apis.azure-api.net/mcp/pronunciation/mcp"
// mcp.json
{
  "mcpServers": {
    "fasuizu-br-speech-ai": {
      "type": "http",
      "url": "https://apim-ai-apis.azure-api.net/mcp/pronunciation/mcp"
    }
  }
}

The mcpServers block is a cross-client convention. Remote transports vary, so check your client's docs.

Changelog

Every change we have recorded for this component, newest first. Security-relevant changes are always shown. ▲ marks a change for the better, ▼ a change for the worse; unmarked changes are neutral.

  • 3 Aug 26 +1

    No change was recorded against any check on this day. Stability & Change Management went from 23 to 27. That category is still filling its 30-day observation window: 7 days of observed history at the previous scan, 8 at this one. The score rises as the window fills, whether or not the server changes.

  • 1 Aug 26 +1

    No change was recorded against any check on this day. Stability & Change Management went from 17 to 20. That category is still filling its 30-day observation window: 5 days of observed history at the previous scan, 6 at this one. The score rises as the window fills, whether or not the server changes.

  • 31 Jul 26 +3
    • We updated how we score, so this day's move reflects our rubric, not a change to the server See what changed → functional
  • 30 Jul 26 +1
    • We updated how we score, so this day's move reflects our rubric, not a change to the server See what changed → functional
  • 28 Jul 26 +1

    No change was recorded against any check on this day. Stability & Change Management went from 3 to 7. That category is still filling its 30-day observation window: 1 days of observed history at the previous scan, 2 at this one. The score rises as the window fills, whether or not the server changes.

  • 27 Jul 26 +1
    • We updated how we score, so this day's move reflects our rubric, not a change to the server See what changed → functional
  • 26 Jul 26 65

    First indexed and scored.

Diagnostics

Diagnostic detail from the automated scan of this channel: what the scanner observed at each step, so you can see exactly where a check passed or failed. It is informational only and never changes the trust score.

Captured 3 Aug 2026 · Probed https://apim-ai-apis.azure-api.net/mcp/pronunciation/mcp

TLS valid

Negotiated TLS 1.3 with TLS_AES_256_GCM_SHA384 .

Subject Issuer Valid from Valid until Key Signature Serial
CN=*.azure-api.net,O=Microsoft Corporation,L=Redmond,ST=WA,C=US CN=Microsoft TLS G2 RSA CA OCSP 02,O=Microsoft Corporation,C=US 7 Jun 2026 4 Dec 2026 RSA 2048 SHA384-RSA 41004cfee075a86fa3cd4ae1450000004cfee0
SANs: *.azure-api.net, *.portal.azure-api.net, *.management.azure-api.net, *.scm.azure-api.net, *.configuration.azure-api.net, *.regional.azure-api.net, *.developer.azure-api.net, *.data.azure-api.net, *.portal-editor.azure-api.net, *.unique.azure-api.net, *.unique.portal.azure-api.net, *.unique.management.azure-api.net and 6 more
CN=Microsoft TLS G2 RSA CA OCSP 02,O=Microsoft Corporation,C=US (CA) CN=Microsoft TLS RSA Root G2,O=Microsoft Corporation,C=US 1 Aug 2025 3 Jun 2029 RSA 4096 SHA384-RSA 330000000c4964a16f44203b2200000000000c
CN=Microsoft TLS RSA Root G2,O=Microsoft Corporation,C=US (CA) CN=DigiCert Global Root G2,OU=www.digicert.com,O=DigiCert Inc,C=US 21 May 2025 19 Jun 2029 RSA 4096 SHA384-RSA b0c6b2c466917b04773c647d4afc0c8
DNSSEC insecure

Validation of apim-ai-apis.azure-api.net. Not signed

Zone DS Keys Algorithms Outcome
. trust_anchor 20326, 38696 8, 8 Verified
net. present 37331 13 Verified
azure-api.net. absent Unsigned (proven) parent-signed NSEC/NSEC3 proves an unsigned delegation
Authentication Challenged, unverified

The endpoint asked for a token, but we could not retrieve and validate the RFC 9728 metadata that tells a client how to obtain one.

Result Challenged, unverified
Enforced On tool calls
HTTP status 200
Header Value
strict-transport-security max-age=31536000; includeSubDomains
x-content-type-options nosniff
x-frame-options DENY
referrer-policy strict-origin-when-cross-origin
permissions-policy camera=(), microphone=(), geolocation=()

Protected resource metadata

Retrieved No
Problem no_resource_metadata
Transports 2 probes
Transport URL Outcome Status Location
streamable-http https://apim-ai-apis.azure-api.net/mcp/pronunciation/mcp Verified 200
http (plaintext) http://apim-ai-apis.azure-api.net/mcp/pronunciation/mcp Inconclusive 404
MCP tools — 10 exposed · ~1,868 tokens

The tools this component advertises to a client, with an estimated token cost for each. Expand a tool to see its parameters and schema. The per-tool counts are indicative and are not scored directly; the schema's total context footprint is one signal in Schema Quality & AI Usability.

Tool Tokens
assess_pronunciation ~379

Assess English pronunciation quality from audio. Scores pronunciation at four levels: overall, sentence, word, and phoneme. Each score is 0-100. Phonemes are returned in both IPA and ARPAbet notation. Sub-300ms inference latency. Args: audio_base64: Base64-encoded audio data. Supports WAV, MP3, OGG, and WebM formats. text: The reference English text that the speaker was expected to read aloud. audio_format: Audio format hint — one of 'wav', 'mp3', 'ogg', 'webm'. Defaults to 'wav'. Returns: dict with keys: - overallScore (int 0-100): Overall pronunciation quality - sentenceScore (int 0-100): Sentence-level fluency and accuracy - words (list): Per-word scores, each containing: - word (str): The word - score (int 0-100): Word pronunciation score - phonemes (list): Per-phoneme scores with IPA/ARPAbet notation - decodedTranscript (str): What the model heard (ASR transcript) - transcript (str): Reference text - confidence (float 0-1): Scoring confidence - warnings (list[str]): Quality warnings if any - audioQuality (dict): Audio metrics (SNR, peak/RMS dB, etc.)

NameTypeReqDescription
audio_base64stringyesBase64-encoded audio data. Supports WAV, MP3, OGG, and WebM formats.
audio_formatstringAudio format hint — one of 'wav', 'mp3', 'ogg', 'webm'.
textstringyesThe reference English text that the speaker was expected to read aloud.

No output schema declared.

No examples provided.

check_pronunciation_service ~65

Check if the pronunciation assessment service is healthy and ready. Returns: dict with keys: - status (str): 'healthy' or error state - modelLoaded (bool): Whether the scoring model is loaded - version (str): API version

Input schema present but exposes no named parameters.

No output schema declared.

No examples provided.

check_stt_service ~66

Check if the speech-to-text service is healthy and ready. Returns: dict with keys: - status (str): 'healthy' or error state - modelLoaded (bool): Whether the STT model is loaded - version (str): API version

Input schema present but exposes no named parameters.

No output schema declared.

No examples provided.

check_tts_service ~67

Check if the text-to-speech service is healthy and ready. Returns: dict with keys: - status (str): 'healthy' or error state - modelLoaded (bool): Whether the TTS model is loaded - version (str): API version

Input schema present but exposes no named parameters.

No output schema declared.

No examples provided.

check_whisper_service ~103

Check if the Whisper STT Pro service is healthy and ready. Returns: dict with keys: - status (str): 'healthy' or error state - modelLoaded (bool): Whether the Whisper model is loaded - diarizeLoaded (bool): Whether the diarization pipeline is loaded - version (str): API version - modelName (str): Whisper model name (e.g. 'large-v3-turbo')

Input schema present but exposes no named parameters.

No output schema declared.

No examples provided.

get_phoneme_inventory ~138

Get the full phoneme inventory supported by the pronunciation scorer. Returns a list of all English phonemes the engine can assess, including ARPAbet symbol, IPA equivalent, example word, and phoneme category (vowel, consonant, diphthong). Returns: list of dicts, each with keys: - arpabet (str): ARPAbet symbol (e.g. 'AA', 'TH') - ipa (str): IPA notation - example (str): Example word containing the phoneme - category (str): vowel, consonant, or diphthong

Input schema present but exposes no named parameters.

No output schema declared.

No examples provided.

list_tts_voices ~61

List all available text-to-speech voices with metadata. Returns: dict with keys: - voices (list): Available voices, each with id, name, gender, accent, grade - defaultVoice (str): Default voice ID

Input schema present but exposes no named parameters.

No output schema declared.

No examples provided.

synthesize_speech ~323

Generate natural speech audio from English text. Produces high-quality speech with 12 English voices. Returns base64-encoded WAV audio (16-bit PCM, 24kHz mono) along with metadata. Available voices: - af_heart (default), af_bella, af_nicole, af_sarah, af_sky (American female) - am_adam, am_michael (American male) - bf_emma, bf_isabella (British female) - bm_george, bm_lewis, bm_daniel (British male) Args: text: English text to synthesize (1-5000 characters). voice: Voice ID. See list above. Defaults to 'af_heart'. speed: Speed multiplier from 0.5 to 2.0 (default: 1.0). Returns: dict with keys: - audio_base64 (str): Base64-encoded WAV audio (16-bit PCM, 24kHz) - duration_ms (str): Audio duration in milliseconds - voice (str): Voice ID used - text_length (str): Input text character count - processing_ms (str): Synthesis time in milliseconds

NameTypeReqDescription
speednumberSpeech speed multiplier (0.5 = half speed, 2.0 = double).
textstringyesEnglish text to convert to speech. Max 5000 characters.
voiceVoice ID (e.g. 'af_heart', 'am_adam'). Uses default if omitted.

No output schema declared.

No examples provided.

transcribe_audio ~313

Transcribe audio to text with word-level timestamps. Converts spoken English audio into text with optional word-level timestamps and per-word confidence scores. Args: audio_base64: Base64-encoded audio data (WAV, MP3, OGG, FLAC, WebM). audio_format: Audio format hint. Auto-detected from magic bytes if omitted. include_timestamps: Whether to include word-level timing (default: true). Returns: dict with keys: - text (str): Full decoded transcript - words (list): Per-word results with timestamps, each containing: - word (str): The transcribed word - start (float): Start time in seconds - end (float): End time in seconds - confidence (float 0-1): Word-level confidence - audioDurationMs (int): Audio duration in milliseconds - metadata (dict): Processing time, audio length, model version - audioQuality (dict): Audio metrics (SNR, peak/RMS dB, etc.)

NameTypeReqDescription
audio_base64stringyesBase64-encoded audio data. Supports WAV, MP3, OGG, FLAC, and WebM formats.
audio_formatAudio format hint — 'wav', 'mp3', 'ogg', 'flac', 'webm'. Auto-detected if omitted.
include_timestampsbooleanIf true, include word-level start/end times and confidence.

No output schema declared.

No examples provided.

transcribe_audio_pro ~353

Transcribe audio with Whisper Large V3 Turbo — multilingual STT. Supports 99 languages with automatic language detection, word-level timestamps, per-word confidence scores, and optional speaker diarization (identifies who spoke each word). Best-in-class WER (~2%). Args: audio_base64: Base64-encoded audio (WAV, MP3, OGG, FLAC, WebM). language: Language code. Auto-detected if omitted. Supports 99 languages. diarize: Enable speaker diarization (default: false). When true, each word includes a speaker label (e.g. SPEAKER_00, SPEAKER_01). Returns: dict with keys: - text (str): Full decoded transcript - words (list): Per-word results with timestamps, each containing: - word (str), start (float), end (float), confidence (float 0-1) - speaker (str|null): Speaker label when diarize=true - speakers (dict|null): Speaker info with count and labels - audioDurationMs (int): Audio duration in milliseconds - metadata (dict): Processing time, language, languageProbability - audioQuality (dict): Audio metrics (SNR, peak/RMS dB, etc.)

NameTypeReqDescription
audio_base64stringyesBase64-encoded audio data. Supports WAV, MP3, OGG, FLAC, and WebM formats.
diarizebooleanEnable speaker diarization to identify who spoke each word.
languageLanguage code (e.g. 'en', 'es', 'zh'). Auto-detected when omitted.

No output schema declared.

No examples provided.