* fix(mcp): drop unsupported FastMCP kwargs (mcp SDK >= 1.10)
The MCP server passes `version=` and `description=` to FastMCP(), but
neither kwarg exists on mcp >= 1.10 — the protocol version is now
managed internally and `description` was renamed to `instructions`.
Symptom on a fresh install (uv sync && pip install 'mcp[cli]'):
TypeError: FastMCP.__init__() got an unexpected keyword argument 'version'
Tested locally end-to-end:
- create_mcp_server() now constructs cleanly
- All 5 tools register and are listable via FastMCP.list_tools()
- generate_speech round-trip returns base64 WAV; ~24s server-side
for 4.2s of audio at steps=16 on Apple Silicon MPS
- pytest backend/ -x -q: 45 passed
* feat: bundle Claude Code agent skill at .claude/skills/omnivoice/
CLAUDE.md already invites contributions at .claude/skills/:
"No project skills found. Add skills to any of: .claude/skills/,
.agents/skills/, .cursor/skills/, .github/skills/, or .codex/skills/
with a SKILL.md index file."
But the existing .gitignore blanket-ignored .claude/ (line 41), making
the invited path un-trackable. This commit narrows the ignore so ad-hoc
Claude state stays out while deliberate skill bundles are tracked:
-.claude/
+.claude/*
+!.claude/skills/
+!.claude/skills/**
Once merged, any compatible agent client running
`npx skills add debpalash/OmniVoice-Studio` gets immediate context on:
- What the MCP server exposes (5 tools + 2 resources)
- When to pick OmniVoice vs other engines
- How to wire the stdio MCP server into a client config
- Backend lifecycle: start / health / stop scripts
- Common failure modes + fixes (port collision, model download stall,
missing HF_TOKEN, MPS fallback, voice-profile-not-found, etc.)
Conforms to Anthropic skill-creator conventions: frontmatter
description under 1024-char limit, body under 500 lines, references/
for detail, scripts/ for deterministic ops, no README/CHANGELOG
inside the skill, validates clean against quick_validate.py.
Verified locally that `npx skills list` discovers the bundled skill
automatically once cloned. End-to-end tested through MCP:
- generate_speech (English, demo voice, steps=16) -> 4.2 s WAV
- generate_speech (voice design via instruct only, steps=8) -> 6.3 s WAV
- generate_speech (Spanish, demo voice, steps=16) -> 2.8 s WAV
Depends on #112 (FastMCP API fix). Without it, every MCP tool call
fails with TypeError at server construction.
* feat(skill): add voice-clone end-to-end recipe + record-reference.sh helper
Two additions to the bundled skill, closing the gap where agents had no
procedural knowledge for creating a voice profile (the previous SKILL.md
said "use the UI or POST /profiles" but didn't include the recording +
trimming + verification workflow).
1. scripts/record-reference.sh — macOS-only helper that records a clean
reference clip with **audible** countdown + start/stop cues via
`say` + /System/Library/Sounds/Ping.aiff. Solves the buffering bug
where text-mode "speak now" prompts arrive after recording starts.
Captures a longer raw window then trims to ~10 sec of speech via
silenceremove + atrim. Plays back for verification. Prints the
next-step `curl` command for POST /profiles.
2. SKILL.md "Voice clone — end-to-end recipe" section (replaces the
stub one-liner). Covers:
- Path A: the bundled helper (one command, audible cues)
- Path B: manual ffmpeg flow if the helper doesn't fit
- POST /profiles multipart/form-data fields (required: name +
ref_audio; optional: ref_text, language, instruct, seed, personality)
- Reference clip quality factors that materially affect output
(single speaker, natural prosody, 3-10 sec sweet spot, ref_text
alignment, language correctness, loudness ≥ -15 dB peak)
Tested locally: recorded a 10-sec Spanish reference + 3-sec English
reference, created two profiles via the helper + curl flow, generated
14.1 sec of Spanish + 10.2 sec of English audio in the user's cloned
voice. Round-trip works end-to-end at steps=16 on Apple Silicon MPS.
Frontmatter description unchanged (860 chars, under the 1024 limit).
Body grew from ~120 to 169 lines (still well under the 500-line skill
ceiling).
* fix(skill): address P20 cross-review findings on PR #113
Adversarial multi-agent review (code + comment + silent-failure analyzers
on parallel reviewers) surfaced one blocker, one critical silent-failure
class, two medium-severity bugs, and two minor doc inaccuracies. All
addressed in this commit.
Blocker (cited 3x by both code-reviewer and comment-analyzer):
- SKILL.md linked references/engines-comparison.md three times (lines 44,
153, 160) but the file was never copied into the upstream skill tree.
+ Added the file (engine decision tree across OmniVoice / kokoro /
Voicebox / Edge TTS / ElevenLabs / cloud APIs).
Critical — record-reference.sh (was 4/10):
- Mic-permission silent failure: macOS denies the mic by sending a silent
stream; ffmpeg exits 0 with a valid silent WAV. The script printed
"✓ raw captured" and produced a degenerate reference clip that would
train a broken voice profile.
+ Parse mean_volume from volumedetect; exit 3 with a diagnostic
pointing the user to System Settings → Privacy → Microphone if
the recording is below -50 dB.
- afplay backgrounded with no exit check; if /System/Library/Sounds/*.aiff
is missing the user gets no audible cue.
+ beep() helper falls back to printf '\a' (terminal bell) when the
system sound file is missing.
- silenceremove silent corruption: silent input → near-empty output WAV,
exit 0.
+ ffprobe duration check after trim; exit 4 if < 2.0 sec.
- trap only covered EXIT; Ctrl-C / SIGTERM mid-recording leaked tmp file.
+ trap '...' EXIT INT TERM HUP.
- macOS guard ran after mktemp + trap.
+ Moved guard to first executable line.
- afplay verification swallowed stderr.
+ Drop 2>/dev/null; surface failure as a warning.
- Documented exit codes in header (0/2/3/4).
Medium — start-backend.sh (was 6/10):
- TOCTOU race: lsof check → uvicorn start could lose the port to another
process; only signal was a 60s health timeout.
+ Added `kill -0 $PID` check inside the probe loop; immediate exit 5
with log tail if uvicorn died.
- lsof check couldn't tell "stale us" from "third party" — same exit 3
for both.
+ ps -o command attribution; the message now tells the user whether
it's a stale uvicorn (suggest stop-backend.sh) or unknown process.
- Documented exit codes (0/2/3/4/5).
Medium — stop-backend.sh (was 7/10):
- No post-SIGKILL verification — script exited 0 even if process still
bound.
+ Added current_pids() helper; re-query after SIGKILL; exit 1 if still
bound, with lsof dump for diagnostics.
- 2>/dev/null || true on kill swallowed EPERM silently.
+ Capture stderr; classify EPERM vs ESRCH; exit 2 on EPERM with
actionable hint (try sudo).
- Documented exit codes (0/1/2).
Minor docs (comment-analyzer):
- SKILL.md line 120 claimed profiles persist as `<id>.wav`. Actual
backend (profiles.py:48-50) preserves uploaded extension.
+ Reworded to `<id>.<ext>` with explanation.
- mcp-setup.md line 68 cited HF cache path as Linux/macOS only.
Windows redirects via backend/core/config.py:38 to
%LOCALAPPDATA%\OmniVoice\hf_cache.
+ Added Windows row + reference to config.py.
Re-validated: all 6 files compile under set -euo pipefail; SKILL.md
frontmatter description stays at 860 chars (under 1024 cap); skill body
under 500 lines.
Diff: 6 files changed, ~+269/-47.
214 lines
7.3 KiB
Python
214 lines
7.3 KiB
Python
"""
|
||
OmniVoice MCP Server — expose voice synthesis as AI-agent tools.
|
||
|
||
Run standalone:
|
||
python -m backend.mcp_server # stdio transport (Claude Desktop)
|
||
python -m backend.mcp_server --sse # SSE transport (remote agents)
|
||
|
||
Tools exposed:
|
||
generate_speech — text → WAV audio (voice clone or design)
|
||
list_voices — enumerate saved voice profiles
|
||
list_languages — available TTS languages
|
||
list_personalities — voice personality presets
|
||
|
||
Resources exposed:
|
||
voice://{profile_id} — voice profile metadata
|
||
history://recent — last 20 generated audio items
|
||
"""
|
||
from __future__ import annotations
|
||
|
||
import argparse
|
||
import base64
|
||
import logging
|
||
import os
|
||
import sys
|
||
|
||
logger = logging.getLogger("omnivoice.mcp")
|
||
|
||
# ── Lazy imports — keeps startup fast when not using MCP ────────────────
|
||
|
||
|
||
def _ensure_mcp():
|
||
"""Import `mcp` SDK lazily so the rest of the backend doesn't pay
|
||
for the import unless the MCP server is actually started."""
|
||
try:
|
||
from mcp.server.fastmcp import FastMCP # noqa: F811
|
||
return FastMCP
|
||
except ImportError:
|
||
logger.error(
|
||
"MCP SDK not installed. Install with:\n"
|
||
" pip install 'mcp[cli]'\n"
|
||
"Then re-run this module."
|
||
)
|
||
sys.exit(1)
|
||
|
||
|
||
def create_mcp_server():
|
||
"""Build and return the FastMCP server instance."""
|
||
FastMCP = _ensure_mcp()
|
||
mcp = FastMCP(
|
||
"OmniVoice Studio",
|
||
instructions=(
|
||
"AI-agent interface for OmniVoice Studio — voice cloning, "
|
||
"voice design, and video dubbing in 646 languages."
|
||
),
|
||
)
|
||
|
||
# ── Helpers ─────────────────────────────────────────────────────────
|
||
|
||
def _api_base() -> str:
|
||
return os.environ.get("OMNIVOICE_API_URL", "http://localhost:3900")
|
||
|
||
async def _api_get(path: str):
|
||
import httpx
|
||
async with httpx.AsyncClient(base_url=_api_base(), timeout=30) as c:
|
||
r = await c.get(path)
|
||
r.raise_for_status()
|
||
return r.json()
|
||
|
||
async def _api_post_form(path: str, data: dict, files: dict | None = None):
|
||
import httpx
|
||
async with httpx.AsyncClient(base_url=_api_base(), timeout=120) as c:
|
||
r = await c.post(path, data=data, files=files or {})
|
||
r.raise_for_status()
|
||
return r
|
||
|
||
# ── Tools ───────────────────────────────────────────────────────────
|
||
|
||
@mcp.tool()
|
||
async def generate_speech(
|
||
text: str,
|
||
language: str = "Auto",
|
||
profile_id: str | None = None,
|
||
instruct: str | None = None,
|
||
speed: float = 1.0,
|
||
steps: int = 16,
|
||
) -> str:
|
||
"""Generate speech audio from text.
|
||
|
||
Args:
|
||
text: The text to synthesize into speech.
|
||
language: Target language (ISO code or 'Auto'). 646 languages supported.
|
||
profile_id: ID of a saved voice profile to clone. Omit for voice design mode.
|
||
instruct: Style instruction (e.g. 'whisper', 'excited', 'narrator').
|
||
speed: Speech speed multiplier (0.5–2.0, default 1.0).
|
||
steps: Diffusion steps (8=fast/draft, 16=balanced, 32=quality).
|
||
|
||
Returns:
|
||
JSON with audio_id, generation_time, audio_duration, and
|
||
base64-encoded WAV data.
|
||
"""
|
||
form = {
|
||
"text": text,
|
||
"language": language,
|
||
"speed": str(speed),
|
||
"num_step": str(steps),
|
||
}
|
||
if profile_id:
|
||
form["profile_id"] = profile_id
|
||
if instruct:
|
||
form["instruct"] = instruct
|
||
|
||
r = await _api_post_form("/generate", data=form)
|
||
|
||
audio_id = r.headers.get("X-Audio-Id", "unknown")
|
||
gen_time = r.headers.get("X-Gen-Time", "?")
|
||
duration = r.headers.get("X-Audio-Duration", "?")
|
||
|
||
wav_b64 = base64.b64encode(r.content).decode("ascii")
|
||
|
||
return (
|
||
f'{{"audio_id":"{audio_id}",'
|
||
f'"generation_time_s":{gen_time},'
|
||
f'"audio_duration_s":{duration},'
|
||
f'"format":"wav",'
|
||
f'"wav_base64":"{wav_b64}"}}'
|
||
)
|
||
|
||
@mcp.tool()
|
||
async def list_voices() -> str:
|
||
"""List all saved voice profiles.
|
||
|
||
Returns a JSON array of voice profiles with id, name, type (clone/design),
|
||
and personality.
|
||
"""
|
||
profiles = await _api_get("/profiles")
|
||
return str(profiles)
|
||
|
||
@mcp.tool()
|
||
async def list_personalities() -> str:
|
||
"""List available voice personality presets.
|
||
|
||
Returns presets like Narrator, Casual, News Anchor, etc. with their
|
||
instruct text. Use the instruct text with generate_speech.
|
||
"""
|
||
presets = await _api_get("/personalities")
|
||
return str(presets)
|
||
|
||
@mcp.tool()
|
||
async def list_languages() -> str:
|
||
"""List a sample of supported TTS languages.
|
||
|
||
OmniVoice supports 646 languages. This returns the most popular ones
|
||
plus a note about the full count.
|
||
"""
|
||
return (
|
||
'{"total":646,"popular":['
|
||
'"en","es","fr","de","it","pt","ru","ja","ko","zh",'
|
||
'"ar","hi","tr","nl","pl","sv","da","fi","no","el"'
|
||
'],"note":"Pass any ISO 639 code or set language=Auto for detection."}'
|
||
)
|
||
|
||
@mcp.tool()
|
||
async def check_health() -> str:
|
||
"""Check if the OmniVoice backend is running and what GPU device is active."""
|
||
info = await _api_get("/health")
|
||
return str(info)
|
||
|
||
# ── Resources ───────────────────────────────────────────────────────
|
||
|
||
@mcp.resource("voice://{profile_id}")
|
||
async def get_voice(profile_id: str) -> str:
|
||
"""Get details of a specific voice profile."""
|
||
profiles = await _api_get("/profiles")
|
||
for p in profiles:
|
||
if p.get("id") == profile_id:
|
||
return str(p)
|
||
return f'{{"error":"Voice profile {profile_id} not found"}}'
|
||
|
||
@mcp.resource("history://recent")
|
||
async def get_recent_history() -> str:
|
||
"""Get the 20 most recent generation history items."""
|
||
history = await _api_get("/history")
|
||
return str(history[:20])
|
||
|
||
return mcp
|
||
|
||
|
||
# ── CLI entrypoint ──────────────────────────────────────────────────────
|
||
|
||
def main():
|
||
parser = argparse.ArgumentParser(description="OmniVoice MCP Server")
|
||
parser.add_argument(
|
||
"--sse", action="store_true",
|
||
help="Use SSE transport instead of stdio (for remote agents)",
|
||
)
|
||
parser.add_argument(
|
||
"--port", type=int, default=8765,
|
||
help="Port for SSE transport (default: 8765)",
|
||
)
|
||
args = parser.parse_args()
|
||
|
||
mcp = create_mcp_server()
|
||
|
||
if args.sse:
|
||
logger.info("Starting MCP server on SSE transport, port %d", args.port)
|
||
mcp.run(transport="sse", port=args.port)
|
||
else:
|
||
logger.info("Starting MCP server on stdio transport")
|
||
mcp.run(transport="stdio")
|
||
|
||
|
||
if __name__ == "__main__":
|
||
main()
|