Files
VoiceStudio/docs/features.yaml
T
5a7d9cc05c feat(asr): generic OpenAI-compatible transcription backend (#877) (#1003)
First slice of the community's two-track proposal for #877: a generic
OpenAI-compatible ASR backend that works TODAY, without waiting on
transformers to ship a direct Qwen3-ASR integration (tracked separately,
still blocked upstream). Points OmniVoice's transcription at any server
exposing POST /v1/audio/transcriptions — a self-hosted Qwen3-ASR/
FunASR/SenseVoice server, or OpenAI's own API.

- New OpenAICompatASRBackend (backend/services/asr_backend.py): a pure
  network client, no local model, no install. Prefers
  response_format=verbose_json for real per-segment timestamps,
  degrades to plain text (matching MoonshineASRBackend's shape) when a
  minimal server rejects that format. Never leaks a raw SDK/httpx
  exception to the caller (#977 convention) — wraps network/auth
  failures in a clean, actionable RuntimeError naming the server.
- Settings persist via the same encrypted-secret convention as
  services/llm_providers.py (settings_store.set_secret for the API key
  — Fernet-encrypted, never a .env row, never echoed back; get_text/
  set_text for base_url/model). New GET/PUT /api/settings/
  asr-openai-compat, loopback-gated like every other settings route.
- Frontend: a small settings panel (Settings → Models) mirroring
  HFMirrorPanel's exact structure. No ASR engine picker exists yet for
  ANY ASR backend (only TTS has one) — activating this engine still
  needs OMNIVOICE_ASR_BACKEND=openai-compat-asr; documented plainly
  rather than pretending otherwise.
- README's ASR Engines table (9 → 10 engines) and docs/features.yaml's
  drift-checker inventory updated; the '9 engines, all fully local'
  claim corrected since this one genuinely isn't.
- docs/engines/openai-compatible-asr.md: setup steps + an explicit
  privacy note (unlike every other ASR engine, audio leaves the
  machine to whatever server is configured).

Regression tests: tests/test_asr_openai_compat_877.py (12 tests) —
is_available() gating, verbose_json + plain-text response adaptation,
network-failure error hygiene, SDK retry disabling, and the settings
endpoints' persist/mask/clear-vs-unchanged semantics.

Fixed two real full-suite-only failures found during verification (not
brushed aside): the API route inventory snapshot needed regenerating
for the two new routes, and this file's own tests had a module-
staleness bug — a collection-time settings_store import went stale
relative to a test-time-fresh fixture when another test elsewhere in
the ~2400-test suite reimports the module — fixed by making
settings_store itself a fixture resolved at test-run time, same
lesson already applied to tests/test_mm2_lifecycle.py earlier this
session.

Co-authored-by: mergetest <test@local>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-08 08:33:03 +05:30

87 lines
2.5 KiB
YAML

# Canonical feature inventory — the single source of truth that the daily
# docs-drift job (.github/workflows/docs-drift.yml) diffs against README.md,
# docs/, and the engine registries via scripts/check-docs-drift.py.
#
# When a PR adds or removes an engine or user-facing feature, update this
# file in the same PR — otherwise the nightly job opens/updates the rolling
# `docs-drift` issue. Spec: docs/competitive-analysis.md Spec 9a /
# docs/specs/2026-06-12-elevenlabs-parity-program.md Wave 0.1.
# Each name must appear verbatim in README.md (the Features grid).
features:
- Voice Cloning
- Voice Design
- Video Dubbing
- Dictation Widget
- Vocal Isolation
- Speaker Diarization
- Batch Queue
- MCP Server
- AI Watermark
- 100% Local
- GPU Auto-Detect
- Extensible
# id: must exactly match the registry keys in backend/services/tts_backend.py
# (_REGISTRY eager entries + _LAZY_REGISTRY).
# readme (optional): a string that must appear in README.md (engine table row).
# doc (optional): a repo-relative doc file that must exist.
tts_engines:
- id: omnivoice
readme: "**OmniVoice** (default)"
- id: cosyvoice
readme: CosyVoice 3
doc: docs/engines/cosyvoice.md
- id: kittentts
readme: KittenTTS
- id: mlx-audio
readme: MLX-Audio
- id: voxcpm2
readme: VoxCPM2
- id: moss-tts-nano
readme: MOSS-TTS-Nano
- id: gpt-sovits
- id: sherpa-onnx
- id: indextts2
doc: docs/engines/indextts.md
- id: omnivoice-gguf
- id: supertonic3
- id: moss-tts-v15
readme: "**MOSS-TTS-v1.5**"
doc: docs/engines/moss-tts-v15.md
- id: dots-tts
readme: "**dots.tts**"
doc: docs/engines/dots-tts.md
- id: confucius4-tts
readme: "**Confucius4-TTS**"
doc: docs/engines/confucius4-tts.md
# Same contract against backend/services/asr_backend.py _REGISTRY.
asr_engines:
- id: whisperx
readme: "**WhisperX** (default)"
- id: faster-whisper
readme: Faster-Whisper
- id: mlx-whisper
readme: MLX Whisper
- id: pytorch-whisper
readme: PyTorch Whisper
- id: nemo-parakeet
readme: Parakeet TDT
- id: moonshine
readme: Moonshine
- id: funasr
readme: FunASR
- id: sherpa-onnx-asr
readme: "**sherpa-onnx** (live dictation)"
- id: openai-compat-asr
readme: "**OpenAI-compatible** ⚠️ remote"
# Doc files that must exist (the install path users are sent to).
docs:
- docs/install/macos.md
- docs/install/windows.md
- docs/install/linux.md
- docs/install/docker.md
- docs/install/troubleshooting.md