Files
VoiceStudio/docs/features.yaml
T
acd36badee feat(engines): PocketTTS CPU-only sidecar shape (#1306) (#1328)
* feat(engines): PocketTTS CPU-only sidecar shape (#1306)

Sidecar SHAPE for review, mirroring omnivoice-subprocess: PocketTTSBackend(SubprocessBackend) (CPU-only, parent interpreter, optional-dep gate) plus a stdio sidecar (ready/ping/synthesize/shutdown, lazy TTSModel.load_model, per-ref voice cache, generate_audio to int16 PCM). Registered in services/tts_backend.py. Batch protocol; streaming raised as a follow-up. CI smoke, gated-weights preflight, 4-platform install, licence-accept gate deferred to on-top after shape review.

* feat(engines): PocketTTS sidecar handles 6 languages (en/fr/de/pt/it/es)

load_model(language=...) per language (cached), maps OmniVoice's language value to a pocket-tts model language, and picks the default preset voice per language when no ref clip is given. Represents PocketTTS accurately: it is multilingual, not english-only. The HF model card's 'English only' line is stale, confirmed by the GitHub README and pocket-tts 2.1.0.

* fix(engines): list pockettts in docs inventory; drop unused logger

docs/features.yaml tts_engines now includes pockettts, clearing the docs-drift test that failed CI (every registered engine must be in the inventory). Removed the unused logger line CodeQL flagged. No readme/doc entry yet, matching opt-in engines like supertonic3 and omnivoice-gguf; a doc page can land with the rest of the integration.

* fix(engines): address PocketTTS sidecar review findings

- Cold-load watchdog: heartbeat progress frames during the gated weights download so the parent does not kill a healthy sidecar mid-load, plus a 600s recv timeout on the backend.
- Unsupported language: raise a clear error instead of silently falling back to English and mispronouncing.
- Voice-state cache: LRU-bounded to 8 entries so a long session cannot leak memory.
- ref_audio SSRF: reject URLs (local file paths only) to preserve local-first.

Addresses the 3 Greptile P1 + 1 CodeRabbit Major on #1328.

* fix(engines): invalidate voice cache on ref-file change; reject non-finite recv timeout

- Voice-state cache key now folds the ref_audio file mtime+size, so a file replaced at the same path no longer returns a stale voice from the previous contents (Greptile P1).
- recv_timeout_s rejects inf/nan env values via math.isfinite and falls back to 600s, so the deadline can't be silently disabled (CodeRabbit Major).

* fix(engines): nanosecond mtime in voice cache fingerprint

int(st.st_mtime) lost sub-second precision, so a file replaced at the same path within one second with the same size kept the old key and returned a stale voice. Use st.st_mtime_ns for full resolution (Greptile P1 on the follow-up fix commit).

* fix(engines): raise on multi-channel audio instead of unsafe downmix

The defensive mean(axis=0) assumed channels-first; on channels-last (N,2) it averaged across time, producing garbage. The engine returns mono, so the branch is unreachable in practice. Raise on ndim>1 so an upstream shape change surfaces as a loud error frame instead of silent noise. (debpalash review on #1328)

* fix(engines): include import error in pockettts is_available message

CodeRabbit Minor on #1328: the exception was caught as 'e' but never shown.

* fix(engines): lock _send to prevent concurrent-write framing corruption

Greptile P1 on #1328: the cold-load heartbeat thread and the main loop both call _send (stdout write). The stop+join serializes the normal case, but a join timeout leaves a window where both threads write length+body segments concurrently, interleaving the wire framing. Add a threading.Lock around the write so concurrent _send calls are serialized regardless.

* test(engines): cover the PocketTTS sidecar's silent failure modes

The four review findings fixed on this PR are all silent by construction:
an unsupported language rendered fluent, confident, wrong audio; the
channels-last downmix produced noise; interleaved frames desynchronized
the pipe permanently; a re-recorded clip kept serving the old voice. None
of them raise, and none would be caught by an end-to-end smoke test that
only asserts audio came back.

49 tests over the sidecar's pure logic — language selection, PCM
conversion, wire framing, the LRU voice cache — plus the backend surface
(recv-timeout guards, CPU-only declaration, sample-rate lockstep with the
sidecar, lazy registration). The model is mocked and the sidecar is
stdlib-only at import time, so none of it needs the optional pocket-tts
wheel or a child process.

Verified fail-before/pass-after by reverting the lock and the multi-channel
guard: the framing test fails with a length header decoded from inside
another frame's body, which is the corruption itself rather than a proxy
for it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: debpalash <nizam4103@gmail.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-05 00:34:30 +05:30

93 lines
2.7 KiB
YAML

# Canonical feature inventory — the single source of truth that the daily
# docs-drift job (.github/workflows/docs-drift.yml) diffs against README.md,
# docs/, and the engine registries via scripts/check-docs-drift.py.
#
# When a PR adds or removes an engine or user-facing feature, update this
# file in the same PR — otherwise the nightly job opens/updates the rolling
# `docs-drift` issue. Spec: docs/competitive-analysis.md Spec 9a /
# docs/specs/2026-06-12-elevenlabs-parity-program.md Wave 0.1.
# Each name must appear verbatim in README.md (the Features grid).
features:
- Voice Cloning
- Voice Design
- Video Dubbing
- Dictation Widget
- Vocal Isolation
- Speaker Diarization
- Batch Queue
- MCP Server
- AI Watermark
- 100% Local
- GPU Auto-Detect
- Extensible
# id: must exactly match the registry keys in backend/services/tts_backend.py
# (_REGISTRY eager entries + _LAZY_REGISTRY).
# readme (optional): a string that must appear in README.md (engine table row).
# doc (optional): a repo-relative doc file that must exist.
tts_engines:
- id: omnivoice
readme: "**OmniVoice** (default)"
- id: omnivoice-subprocess
doc: docs/engines/omnivoice-subprocess.md
- id: cosyvoice
readme: CosyVoice 3
doc: docs/engines/cosyvoice.md
- id: kittentts
readme: KittenTTS
- id: mlx-audio
readme: MLX-Audio
- id: voxcpm2
readme: VoxCPM2
- id: moss-tts-nano
readme: MOSS-TTS-Nano
- id: gpt-sovits
- id: sherpa-onnx
- id: indextts2
doc: docs/engines/indextts.md
- id: omnivoice-gguf
- id: supertonic3
- id: moss-tts-v15
readme: "**MOSS-TTS-v1.5**"
doc: docs/engines/moss-tts-v15.md
- id: dots-tts
readme: "**dots.tts**"
doc: docs/engines/dots-tts.md
- id: confucius4-tts
readme: "**Confucius4-TTS**"
doc: docs/engines/confucius4-tts.md
- id: pockettts
# Same contract against backend/services/asr_backend.py _REGISTRY.
asr_engines:
- id: whisperx
readme: "**WhisperX** (default)"
- id: faster-whisper
readme: Faster-Whisper
- id: mlx-whisper
readme: MLX Whisper
- id: pytorch-whisper
readme: PyTorch Whisper
- id: nemo-parakeet
readme: Parakeet TDT
- id: parakeet-mlx
readme: Parakeet TDT v3 (MLX)
- id: moonshine
readme: Moonshine
- id: funasr
readme: FunASR
- id: sherpa-onnx-asr
readme: "**sherpa-onnx** (live dictation)"
- id: openai-compat-asr
readme: "**OpenAI-compatible** ⚠️ remote"
# Doc files that must exist (the install path users are sent to).
docs:
- docs/install/macos.md
- docs/install/windows.md
- docs/install/linux.md
- docs/install/docker.md
- docs/install/troubleshooting.md
- docs/migration/real-time-voice-cloning.md