Files
VoiceStudio/docs/features.yaml
T
3d0705fdb7 feat(engines): Confucius4-TTS — finalized (API-validated + unit-tested; opt-in, GPU run pending) (#590) (#637)
* feat(engines): Confucius4-TTS scaffold (opt-in, needs hardware validation) (#590)

Plumbing for netease-youdao's Confucius4-TTS — LLM-based 14-language
cross-lingual zero-shot voice cloning, Apache-2.0 — mirroring the opt-in
subprocess-venv pattern of dots.tts / MOSS-TTS-v1.5:

- engines/confucius4/__init__.py: Confucius4Backend(SubprocessBackend), CUDA-only
  (gpu_compat=("cuda",)), language passthrough, ref_audio→prompt_wav. is_available
  reports a clear reason and stays unavailable without a clone.
- bootstrap.py: dedicated Python 3.10 venv resolution (user clone-level venv →
  package venv → uv bootstrap), import-probed on `confuciustts`.
- main.py: sidecar speaking the same length-prefixed JSON-over-stdio protocol as
  the other engines, calling ConfuciusTTS(config_path, device).generate(text,
  lang, prompt_wav).
- Registered lazily in _LAZY_REGISTRY; docs/engines/confucius4-tts.md.

Gated behind OMNIVOICE_CONFUCIUS4_TTS_DIR — inert on every default install, never
imports the upstream package unless opted in. The sidecar's synthesis API is
derived from the upstream README and is NOT yet validated on a CUDA box; the
module, docs, and CHANGELOG all flag this. 4 tests pin registration +
inert-by-default. No version bump.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(#590): register Confucius4 in install-hints + docs inventory (CI gates)

Registering the engine tripped two completeness gates: every backend needs an
install_hint (test_issue_fixes) and every registry engine must appear in the
tts_engines docs inventory + README (check-docs-drift). Add the install_hint,
the docs/features.yaml entry, and the README engine-table row (with the scaffold
caveat). Docs-drift clean; gates pass. No version bump.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(confucius4): finalize — validate API vs upstream, add 22 sidecar unit tests, document external deps (Amphion/w2v-bert/weights)

The synthesis API (ConfuciusTTS(config_path, device) → generate(text, lang,
prompt_wav) → tensor, model.sample_rate) is confirmed against the
netease-youdao/Confucius4-TTS repo. Added runnable unit tests for the sidecar's
pure logic (language norm, tensor→PCM mono/stereo/clip, config resolution, wire
framing, synthesize dispatch with the model mocked) — 22 cases, all green.
Docs now list the external deps (Amphion/MaskGCT codec, facebook/w2v-bert-2.0,
~2-4GB HF checkpoint) and CUDA 12.6. Softened the scaffold warnings to reflect
API-validated + unit-tested status; a one-time CUDA GPU run is still needed to
confirm live inference + true sample rate.

---------

Co-authored-by: mergetest <test@local>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-01 23:37:40 +05:30

85 lines
2.4 KiB
YAML

# Canonical feature inventory — the single source of truth that the daily
# docs-drift job (.github/workflows/docs-drift.yml) diffs against README.md,
# docs/, and the engine registries via scripts/check-docs-drift.py.
#
# When a PR adds or removes an engine or user-facing feature, update this
# file in the same PR — otherwise the nightly job opens/updates the rolling
# `docs-drift` issue. Spec: docs/competitive-analysis.md Spec 9a /
# docs/specs/2026-06-12-elevenlabs-parity-program.md Wave 0.1.
# Each name must appear verbatim in README.md (the Features grid).
features:
- Voice Cloning
- Voice Design
- Video Dubbing
- Dictation Widget
- Vocal Isolation
- Speaker Diarization
- Batch Queue
- MCP Server
- AI Watermark
- 100% Local
- GPU Auto-Detect
- Extensible
# id: must exactly match the registry keys in backend/services/tts_backend.py
# (_REGISTRY eager entries + _LAZY_REGISTRY).
# readme (optional): a string that must appear in README.md (engine table row).
# doc (optional): a repo-relative doc file that must exist.
tts_engines:
- id: omnivoice
readme: "**OmniVoice** (default)"
- id: cosyvoice
readme: CosyVoice 3
doc: docs/engines/cosyvoice.md
- id: kittentts
readme: KittenTTS
- id: mlx-audio
readme: MLX-Audio
- id: voxcpm2
readme: VoxCPM2
- id: moss-tts-nano
readme: MOSS-TTS-Nano
- id: gpt-sovits
- id: sherpa-onnx
- id: indextts2
doc: docs/engines/indextts.md
- id: omnivoice-gguf
- id: supertonic3
- id: moss-tts-v15
readme: "**MOSS-TTS-v1.5**"
doc: docs/engines/moss-tts-v15.md
- id: dots-tts
readme: "**dots.tts**"
doc: docs/engines/dots-tts.md
- id: confucius4-tts
readme: "**Confucius4-TTS**"
doc: docs/engines/confucius4-tts.md
# Same contract against backend/services/asr_backend.py _REGISTRY.
asr_engines:
- id: whisperx
readme: "**WhisperX** (default)"
- id: faster-whisper
readme: Faster-Whisper
- id: mlx-whisper
readme: MLX Whisper
- id: pytorch-whisper
readme: PyTorch Whisper
- id: nemo-parakeet
readme: Parakeet TDT
- id: moonshine
readme: Moonshine
- id: funasr
readme: FunASR
- id: sherpa-onnx-asr
readme: "**sherpa-onnx** (live dictation)"
# Doc files that must exist (the install path users are sent to).
docs:
- docs/install/macos.md
- docs/install/windows.md
- docs/install/linux.md
- docs/install/docker.md
- docs/install/troubleshooting.md