* feat(engines): Confucius4-TTS scaffold (opt-in, needs hardware validation) (#590) Plumbing for netease-youdao's Confucius4-TTS — LLM-based 14-language cross-lingual zero-shot voice cloning, Apache-2.0 — mirroring the opt-in subprocess-venv pattern of dots.tts / MOSS-TTS-v1.5: - engines/confucius4/__init__.py: Confucius4Backend(SubprocessBackend), CUDA-only (gpu_compat=("cuda",)), language passthrough, ref_audio→prompt_wav. is_available reports a clear reason and stays unavailable without a clone. - bootstrap.py: dedicated Python 3.10 venv resolution (user clone-level venv → package venv → uv bootstrap), import-probed on `confuciustts`. - main.py: sidecar speaking the same length-prefixed JSON-over-stdio protocol as the other engines, calling ConfuciusTTS(config_path, device).generate(text, lang, prompt_wav). - Registered lazily in _LAZY_REGISTRY; docs/engines/confucius4-tts.md. Gated behind OMNIVOICE_CONFUCIUS4_TTS_DIR — inert on every default install, never imports the upstream package unless opted in. The sidecar's synthesis API is derived from the upstream README and is NOT yet validated on a CUDA box; the module, docs, and CHANGELOG all flag this. 4 tests pin registration + inert-by-default. No version bump. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(#590): register Confucius4 in install-hints + docs inventory (CI gates) Registering the engine tripped two completeness gates: every backend needs an install_hint (test_issue_fixes) and every registry engine must appear in the tts_engines docs inventory + README (check-docs-drift). Add the install_hint, the docs/features.yaml entry, and the README engine-table row (with the scaffold caveat). Docs-drift clean; gates pass. No version bump. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(confucius4): finalize — validate API vs upstream, add 22 sidecar unit tests, document external deps (Amphion/w2v-bert/weights) The synthesis API (ConfuciusTTS(config_path, device) → generate(text, lang, prompt_wav) → tensor, model.sample_rate) is confirmed against the netease-youdao/Confucius4-TTS repo. Added runnable unit tests for the sidecar's pure logic (language norm, tensor→PCM mono/stereo/clip, config resolution, wire framing, synthesize dispatch with the model mocked) — 22 cases, all green. Docs now list the external deps (Amphion/MaskGCT codec, facebook/w2v-bert-2.0, ~2-4GB HF checkpoint) and CUDA 12.6. Softened the scaffold warnings to reflect API-validated + unit-tested status; a one-time CUDA GPU run is still needed to confirm live inference + true sample rate. --------- Co-authored-by: mergetest <test@local> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
85 lines
2.4 KiB
YAML
85 lines
2.4 KiB
YAML
# Canonical feature inventory — the single source of truth that the daily
|
|
# docs-drift job (.github/workflows/docs-drift.yml) diffs against README.md,
|
|
# docs/, and the engine registries via scripts/check-docs-drift.py.
|
|
#
|
|
# When a PR adds or removes an engine or user-facing feature, update this
|
|
# file in the same PR — otherwise the nightly job opens/updates the rolling
|
|
# `docs-drift` issue. Spec: docs/competitive-analysis.md Spec 9a /
|
|
# docs/specs/2026-06-12-elevenlabs-parity-program.md Wave 0.1.
|
|
|
|
# Each name must appear verbatim in README.md (the Features grid).
|
|
features:
|
|
- Voice Cloning
|
|
- Voice Design
|
|
- Video Dubbing
|
|
- Dictation Widget
|
|
- Vocal Isolation
|
|
- Speaker Diarization
|
|
- Batch Queue
|
|
- MCP Server
|
|
- AI Watermark
|
|
- 100% Local
|
|
- GPU Auto-Detect
|
|
- Extensible
|
|
|
|
# id: must exactly match the registry keys in backend/services/tts_backend.py
|
|
# (_REGISTRY eager entries + _LAZY_REGISTRY).
|
|
# readme (optional): a string that must appear in README.md (engine table row).
|
|
# doc (optional): a repo-relative doc file that must exist.
|
|
tts_engines:
|
|
- id: omnivoice
|
|
readme: "**OmniVoice** (default)"
|
|
- id: cosyvoice
|
|
readme: CosyVoice 3
|
|
doc: docs/engines/cosyvoice.md
|
|
- id: kittentts
|
|
readme: KittenTTS
|
|
- id: mlx-audio
|
|
readme: MLX-Audio
|
|
- id: voxcpm2
|
|
readme: VoxCPM2
|
|
- id: moss-tts-nano
|
|
readme: MOSS-TTS-Nano
|
|
- id: gpt-sovits
|
|
- id: sherpa-onnx
|
|
- id: indextts2
|
|
doc: docs/engines/indextts.md
|
|
- id: omnivoice-gguf
|
|
- id: supertonic3
|
|
- id: moss-tts-v15
|
|
readme: "**MOSS-TTS-v1.5**"
|
|
doc: docs/engines/moss-tts-v15.md
|
|
- id: dots-tts
|
|
readme: "**dots.tts**"
|
|
doc: docs/engines/dots-tts.md
|
|
- id: confucius4-tts
|
|
readme: "**Confucius4-TTS**"
|
|
doc: docs/engines/confucius4-tts.md
|
|
|
|
# Same contract against backend/services/asr_backend.py _REGISTRY.
|
|
asr_engines:
|
|
- id: whisperx
|
|
readme: "**WhisperX** (default)"
|
|
- id: faster-whisper
|
|
readme: Faster-Whisper
|
|
- id: mlx-whisper
|
|
readme: MLX Whisper
|
|
- id: pytorch-whisper
|
|
readme: PyTorch Whisper
|
|
- id: nemo-parakeet
|
|
readme: Parakeet TDT
|
|
- id: moonshine
|
|
readme: Moonshine
|
|
- id: funasr
|
|
readme: FunASR
|
|
- id: sherpa-onnx-asr
|
|
readme: "**sherpa-onnx** (live dictation)"
|
|
|
|
# Doc files that must exist (the install path users are sent to).
|
|
docs:
|
|
- docs/install/macos.md
|
|
- docs/install/windows.md
|
|
- docs/install/linux.md
|
|
- docs/install/docker.md
|
|
- docs/install/troubleshooting.md
|