* feat(dictation): live local dictation via sherpa-onnx + Voice settings panel Add a sherpa-onnx ASR engine alongside the existing Whisper/NeMo dictation path, powering a genuinely live experience: as you speak, words type straight into the focused field (streaming partials via a new simulate_type command, self-correcting with backspaces) and commit per pause. Backend: - SherpaDictationBackend + sherpa_dictation registry of the 7 models (Parakeet TDT v3/v2, streaming Zipformer EN/ZH/bilingual, Paraformer bilingual, Whisper Tiny) from csukuangfj/* int8 HF repos; CPU provider for cross-platform parity. - /dictation/models + /dictation/prefs router; get_capture_asr_backend() honors the selected dictation model. get_active_asr_backend() (dub transcription) and the legacy WebM/Opus capture path are untouched. - True streaming over /ws/transcribe (OnlineRecognizer: live partials + per-endpoint finals); offline models surface partials via short re-decode. Frontend: - New "Voice" settings panel (enable, Toggle/Hold mode, model picker with offline/streaming/recommended badges + per-model download/delete). - Live word-by-word typing via simulate_type (enigo) with prefix-diff delta and backspace correction; paste fallback retained, no double-insertion. Deps: sherpa-onnx>=1.13.3 (+ sherpa-onnx-core); uv.lock regenerated, Docker frozen-install verified. API route-inventory snapshot updated. 40+ new tests. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs(dictation): register sherpa-onnx-asr engine in README + features inventory Fixes the docs-drift CI guard: the new sherpa-onnx-asr ASR engine existed in the registry but not in docs/features.yaml or README. Adds the live-dictation engine row to the ASR Engines table, bumps the engine counts (8→9), and adds the inventory entry. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs(changelog): fold live-dictation into the [0.3.8] section main is 0.3.8 (untagged), so the dictation feature belongs in that release, not a separate [Unreleased] block. Merge the two Added lists under one [0.3.8], refresh the headline to lead with live dictation, and correct the capture description to reflect live word-by-word typing (not paste-on-pause). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: mergetest <test@local> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
82 lines
2.3 KiB
YAML
82 lines
2.3 KiB
YAML
# Canonical feature inventory — the single source of truth that the daily
|
|
# docs-drift job (.github/workflows/docs-drift.yml) diffs against README.md,
|
|
# docs/, and the engine registries via scripts/check-docs-drift.py.
|
|
#
|
|
# When a PR adds or removes an engine or user-facing feature, update this
|
|
# file in the same PR — otherwise the nightly job opens/updates the rolling
|
|
# `docs-drift` issue. Spec: docs/competitive-analysis.md Spec 9a /
|
|
# docs/specs/2026-06-12-elevenlabs-parity-program.md Wave 0.1.
|
|
|
|
# Each name must appear verbatim in README.md (the Features grid).
|
|
features:
|
|
- Voice Cloning
|
|
- Voice Design
|
|
- Video Dubbing
|
|
- Dictation Widget
|
|
- Vocal Isolation
|
|
- Speaker Diarization
|
|
- Batch Queue
|
|
- MCP Server
|
|
- AI Watermark
|
|
- 100% Local
|
|
- GPU Auto-Detect
|
|
- Extensible
|
|
|
|
# id: must exactly match the registry keys in backend/services/tts_backend.py
|
|
# (_REGISTRY eager entries + _LAZY_REGISTRY).
|
|
# readme (optional): a string that must appear in README.md (engine table row).
|
|
# doc (optional): a repo-relative doc file that must exist.
|
|
tts_engines:
|
|
- id: omnivoice
|
|
readme: "**OmniVoice** (default)"
|
|
- id: cosyvoice
|
|
readme: CosyVoice 3
|
|
doc: docs/engines/cosyvoice.md
|
|
- id: kittentts
|
|
readme: KittenTTS
|
|
- id: mlx-audio
|
|
readme: MLX-Audio
|
|
- id: voxcpm2
|
|
readme: VoxCPM2
|
|
- id: moss-tts-nano
|
|
readme: MOSS-TTS-Nano
|
|
- id: gpt-sovits
|
|
- id: sherpa-onnx
|
|
- id: indextts2
|
|
doc: docs/engines/indextts.md
|
|
- id: omnivoice-gguf
|
|
- id: supertonic3
|
|
- id: moss-tts-v15
|
|
readme: "**MOSS-TTS-v1.5**"
|
|
doc: docs/engines/moss-tts-v15.md
|
|
- id: dots-tts
|
|
readme: "**dots.tts**"
|
|
doc: docs/engines/dots-tts.md
|
|
|
|
# Same contract against backend/services/asr_backend.py _REGISTRY.
|
|
asr_engines:
|
|
- id: whisperx
|
|
readme: "**WhisperX** (default)"
|
|
- id: faster-whisper
|
|
readme: Faster-Whisper
|
|
- id: mlx-whisper
|
|
readme: MLX Whisper
|
|
- id: pytorch-whisper
|
|
readme: PyTorch Whisper
|
|
- id: nemo-parakeet
|
|
readme: Parakeet TDT
|
|
- id: moonshine
|
|
readme: Moonshine
|
|
- id: funasr
|
|
readme: FunASR
|
|
- id: sherpa-onnx-asr
|
|
readme: "**sherpa-onnx** (live dictation)"
|
|
|
|
# Doc files that must exist (the install path users are sent to).
|
|
docs:
|
|
- docs/install/macos.md
|
|
- docs/install/windows.md
|
|
- docs/install/linux.md
|
|
- docs/install/docker.md
|
|
- docs/install/troubleshooting.md
|