First slice of the community's two-track proposal for #877: a generic OpenAI-compatible ASR backend that works TODAY, without waiting on transformers to ship a direct Qwen3-ASR integration (tracked separately, still blocked upstream). Points OmniVoice's transcription at any server exposing POST /v1/audio/transcriptions — a self-hosted Qwen3-ASR/ FunASR/SenseVoice server, or OpenAI's own API. - New OpenAICompatASRBackend (backend/services/asr_backend.py): a pure network client, no local model, no install. Prefers response_format=verbose_json for real per-segment timestamps, degrades to plain text (matching MoonshineASRBackend's shape) when a minimal server rejects that format. Never leaks a raw SDK/httpx exception to the caller (#977 convention) — wraps network/auth failures in a clean, actionable RuntimeError naming the server. - Settings persist via the same encrypted-secret convention as services/llm_providers.py (settings_store.set_secret for the API key — Fernet-encrypted, never a .env row, never echoed back; get_text/ set_text for base_url/model). New GET/PUT /api/settings/ asr-openai-compat, loopback-gated like every other settings route. - Frontend: a small settings panel (Settings → Models) mirroring HFMirrorPanel's exact structure. No ASR engine picker exists yet for ANY ASR backend (only TTS has one) — activating this engine still needs OMNIVOICE_ASR_BACKEND=openai-compat-asr; documented plainly rather than pretending otherwise. - README's ASR Engines table (9 → 10 engines) and docs/features.yaml's drift-checker inventory updated; the '9 engines, all fully local' claim corrected since this one genuinely isn't. - docs/engines/openai-compatible-asr.md: setup steps + an explicit privacy note (unlike every other ASR engine, audio leaves the machine to whatever server is configured). Regression tests: tests/test_asr_openai_compat_877.py (12 tests) — is_available() gating, verbose_json + plain-text response adaptation, network-failure error hygiene, SDK retry disabling, and the settings endpoints' persist/mask/clear-vs-unchanged semantics. Fixed two real full-suite-only failures found during verification (not brushed aside): the API route inventory snapshot needed regenerating for the two new routes, and this file's own tests had a module- staleness bug — a collection-time settings_store import went stale relative to a test-time-fresh fixture when another test elsewhere in the ~2400-test suite reimports the module — fixed by making settings_store itself a fixture resolved at test-run time, same lesson already applied to tests/test_mm2_lifecycle.py earlier this session. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
87 lines
2.5 KiB
YAML
87 lines
2.5 KiB
YAML
# Canonical feature inventory — the single source of truth that the daily
|
|
# docs-drift job (.github/workflows/docs-drift.yml) diffs against README.md,
|
|
# docs/, and the engine registries via scripts/check-docs-drift.py.
|
|
#
|
|
# When a PR adds or removes an engine or user-facing feature, update this
|
|
# file in the same PR — otherwise the nightly job opens/updates the rolling
|
|
# `docs-drift` issue. Spec: docs/competitive-analysis.md Spec 9a /
|
|
# docs/specs/2026-06-12-elevenlabs-parity-program.md Wave 0.1.
|
|
|
|
# Each name must appear verbatim in README.md (the Features grid).
|
|
features:
|
|
- Voice Cloning
|
|
- Voice Design
|
|
- Video Dubbing
|
|
- Dictation Widget
|
|
- Vocal Isolation
|
|
- Speaker Diarization
|
|
- Batch Queue
|
|
- MCP Server
|
|
- AI Watermark
|
|
- 100% Local
|
|
- GPU Auto-Detect
|
|
- Extensible
|
|
|
|
# id: must exactly match the registry keys in backend/services/tts_backend.py
|
|
# (_REGISTRY eager entries + _LAZY_REGISTRY).
|
|
# readme (optional): a string that must appear in README.md (engine table row).
|
|
# doc (optional): a repo-relative doc file that must exist.
|
|
tts_engines:
|
|
- id: omnivoice
|
|
readme: "**OmniVoice** (default)"
|
|
- id: cosyvoice
|
|
readme: CosyVoice 3
|
|
doc: docs/engines/cosyvoice.md
|
|
- id: kittentts
|
|
readme: KittenTTS
|
|
- id: mlx-audio
|
|
readme: MLX-Audio
|
|
- id: voxcpm2
|
|
readme: VoxCPM2
|
|
- id: moss-tts-nano
|
|
readme: MOSS-TTS-Nano
|
|
- id: gpt-sovits
|
|
- id: sherpa-onnx
|
|
- id: indextts2
|
|
doc: docs/engines/indextts.md
|
|
- id: omnivoice-gguf
|
|
- id: supertonic3
|
|
- id: moss-tts-v15
|
|
readme: "**MOSS-TTS-v1.5**"
|
|
doc: docs/engines/moss-tts-v15.md
|
|
- id: dots-tts
|
|
readme: "**dots.tts**"
|
|
doc: docs/engines/dots-tts.md
|
|
- id: confucius4-tts
|
|
readme: "**Confucius4-TTS**"
|
|
doc: docs/engines/confucius4-tts.md
|
|
|
|
# Same contract against backend/services/asr_backend.py _REGISTRY.
|
|
asr_engines:
|
|
- id: whisperx
|
|
readme: "**WhisperX** (default)"
|
|
- id: faster-whisper
|
|
readme: Faster-Whisper
|
|
- id: mlx-whisper
|
|
readme: MLX Whisper
|
|
- id: pytorch-whisper
|
|
readme: PyTorch Whisper
|
|
- id: nemo-parakeet
|
|
readme: Parakeet TDT
|
|
- id: moonshine
|
|
readme: Moonshine
|
|
- id: funasr
|
|
readme: FunASR
|
|
- id: sherpa-onnx-asr
|
|
readme: "**sherpa-onnx** (live dictation)"
|
|
- id: openai-compat-asr
|
|
readme: "**OpenAI-compatible** ⚠️ remote"
|
|
|
|
# Doc files that must exist (the install path users are sent to).
|
|
docs:
|
|
- docs/install/macos.md
|
|
- docs/install/windows.md
|
|
- docs/install/linux.md
|
|
- docs/install/docker.md
|
|
- docs/install/troubleshooting.md
|