* fix: eliminate DB connection leaks, race conditions, and deprecated asyncio API ## DB Connection Leaks (P0) - Convert 38 raw get_db() calls to db_conn() context manager across 14 router files - Connections are now guaranteed to close even when exceptions are raised - profiles.py create_profile: clean up orphaned audio file if DB insert fails - profiles.py lock_profile: consolidate 3 separate conn.close() error paths ## Race Condition (P1) - Add _dub_jobs_lock (threading.Lock) to protect _dub_jobs dict in dub_pipeline.py - get_job/put_job now thread-safe for concurrent dub sessions ## asyncio Deprecation (P2) - Replace 23 asyncio.get_event_loop() calls with asyncio.get_running_loop() - Prevents DeprecationWarning on Python 3.12+ and future breakage on 3.14 ## Quick Fixes - gallery.py preview_voice: remove filesystem path from error response (P2) - dub_pipeline.py parse_vtt_segments: remove redundant `import re` inside loop (P3) - gallery.py _init_gallery_db: use db_conn() context manager (P2) * refactor: extract hooks, centralize isTauri, add pytest-cov ## Frontend - Extract useTTS hook (150 LOC) — TTS generation, streaming, audio ingestion - Extract useProfiles hook (219 LOC) — voice profile CRUD, lock/unlock, preview - Centralize isTauri detection: dialog.js, VoiceGallery.jsx, Settings.jsx now import from utils/media.js instead of 4 different detection patterns ## Backend - Add pytest-cov to dev dependencies - Baseline coverage: 39% across backend/ (214 tests pass) - Add .coverage to .gitignore * feat: add Vitest + checkJs, extract useDubWorkflow + useAppData hooks ## Frontend Testing (new) - Set up Vitest with jsdom environment + @testing-library/react - 11 tests: utils (isTauri, formatTime, constants) + Zustand store (mode, text, dubStep, pill) - Scripts: 'test' (vitest run), 'test:watch' (vitest), 'test:legacy' (node runner) ## App.jsx Decomposition (continued) - Extract useDubWorkflow hook (387 LOC) — upload, ingest, transcribe SSE, translate, generate SSE, abort, stop, cleanup - Extract useAppData hook (181 LOC) — data loading, localStorage persistence, WebSocket real-time updates, model-status pill management ## TypeScript checkJs - Enable checkJs: true in tsconfig.json for IDE-level type checking - 947 existing errors (informational, not blocking builds) - noImplicitAny remains false to avoid blocking * ci: add Vitest step, fix useProfiles duplicate state ## CI - Add 'Run Vitest (frontend)' step — runs 11 unit tests - Override --checkJs false in CI typecheck to avoid 947 pre-existing errors - Rename legacy test step for clarity ## Hooks - Fix useProfiles to accept loadProfiles from parent (useAppData) instead of managing its own duplicate profiles array * refactor: wire hooks into App.jsx — 2067 → 1129 LOC (-45%) App.jsx now delegates to extracted hooks instead of inline logic: - useAppData: data loading, localStorage, WebSocket, model pill - useProfiles: voice profile CRUD, lock/unlock, preview - useTTS: generation, streaming, audio ingestion - useDubWorkflow: upload, transcribe SSE, translate, generate SSE 988 lines removed. All handler logic lives in focused, independently testable hooks. Store selectors and render JSX stay in App.jsx as the shell. Verified: vite build clean, 11 frontend + 214 backend tests pass. * feat: show real-time percentage on model loading pill Backend: register hf_progress listener during _load_model_sync() so download/weight-loading tqdm events update _loading_detail with a progress percentage (0-99%). get_model_status() now includes a 'progress' field that the frontend polls. Frontend: useAppData reads msQuery.data.progress and calls setPillProgress() — the FloatingPill already renders the percentage text and progress bar width from this value. * fix: prevent FileNotFoundError in desktop bundle during model init transformers >=4.52 calls _can_set_experts_implementation() and _can_set_attn_implementation() during PreTrainedModel.__init__, which open the class source file via open(class_file). In a Tauri desktop bundle, module.__file__ points to a path that doesn't exist on disk, causing: FileNotFoundError: .../omnivoice/models/omnivoice.py Override both classmethods on OmniVoice to return static values without filesystem access. OmniVoice doesn't use MoE experts (return False), but does support flex/flash attn (return True). * fix: sync source dirs on every bootstrap, not just first run The Tauri bootstrap previously only copied omnivoice/ and backend/ to Application Support on the first run. Subsequent app updates kept using stale source files, preventing bug fixes from landing. Now ensure_venv_ready() always syncs both directories from the bundle resources before returning, even when the venv is healthy. This fixes the FileNotFoundError crash where the old omnivoice.py lacked the _can_set_experts_implementation override. * ui: premium setup wizard polish - Primary button: solid gradient fill with hover glow + lift + press - Stepper nav: connected pills with glow ring on active step - Welcome cards: glassmorphism with stagger-in animations, lucide icons, left-border accent strip, hover translate - Preflight panel: colored icon pill backgrounds, stagger-slide entrance - Step transitions: fade+slide animation via keyed wrapper - Footnote: shortened paths (~/ notation), Reveal in Finder button - Recommendation banner: gradient background with accent glow - Compact spacing throughout for denser, professional layout * fix: kill zombie backend on clean+retry bootstrap When clean_and_retry_bootstrap removes the project dir, any old uvicorn process still running from the deleted paths remains alive on port 3900. The subsequent retry_bootstrap sees the port is healthy and attaches to the zombie instead of re-bootstrapping. Now explicitly kill any process on the backend port after cleaning, before calling retry_bootstrap. * feat: integrate speaker clones into dubbing interface, sanitize system environment variables for subprocesses, and improve FFMPEG binary path resolution. * fix: restore docker compose default + drop dead setSeed call - deploy/docker-compose.yml: remove profiles: ["cpu"] from the default service so `docker compose up` matches the comment on line 5. With the profile present, no service auto-started. - frontend/src/App.jsx: drop the setSeed call in restoreHistory. The selector was never reintroduced after the App.jsx hooks split, and there is no seed state in the store — seeds are generated fresh per call in useTTS and only read from history items for display. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix: address CodeRabbit review — async detection, dub stream, bootstrap fail-fast - backend/services/tts_backend.py: invert async-context detection in _ensure_loaded. The previous code unconditionally caught its own diagnostic RuntimeError and then called asyncio.run() inside a running loop, masking the intended error message. - frontend/src/hooks/useDubWorkflow.js: require a terminal `done` event before reporting dub success. Without this, a dropped stream after partial progress would flip the UI to `done`, refresh history, and play the completion ping as if generation finished. - frontend/src/hooks/useDubWorkflow.js: restore the previous step when tasksCancel() fails. The UI was getting stuck in `stopping` forever on cancel errors. - frontend/src-tauri/src/bootstrap.rs: fail-fast when source sync fails after the existing directory has already been removed. The previous warn-and-continue path could leave the install with no backend/ or omnivoice/ sources and defer the failure to backend startup with a cryptic error. - backend/api/routers/generation.py: add `from e` to the ValueError → HTTPException re-raise (Ruff B904). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix: preserve % suffix in TTS generation timer The 100ms timer in useTTS was rewriting generationTime to a plain elapsed-seconds string, which immediately wiped the "(xx%)" download suffix written on the next iteration of the response-body loop. The real-time percentage was flickering on/off as a result. Read the previous value inside the setter and reattach any existing percent suffix. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
127 lines
4.4 KiB
Python
127 lines
4.4 KiB
Python
"""
|
|
Standalone transcription endpoint for the Capture / Dictation feature.
|
|
|
|
Unlike /dub/transcribe/{job_id}, this endpoint is job-free — callers POST
|
|
raw audio bytes and get back transcribed text immediately. Used by:
|
|
|
|
• The frontend "Capture" (global hotkey dictation) mode
|
|
• The MCP server's future `transcribe_audio` tool
|
|
• CLI consumers that just want speech-to-text
|
|
|
|
The ASR engine is whatever `get_active_asr_backend()` returns — WhisperX
|
|
by default, or MLX Whisper on Apple Silicon when configured.
|
|
"""
|
|
from __future__ import annotations
|
|
|
|
import io
|
|
import logging
|
|
import os
|
|
import tempfile
|
|
import time
|
|
|
|
from fastapi import APIRouter, File, Form, HTTPException, UploadFile
|
|
from typing import Optional
|
|
|
|
router = APIRouter()
|
|
logger = logging.getLogger("omnivoice.capture")
|
|
|
|
|
|
@router.post("/transcribe")
|
|
async def transcribe_audio(
|
|
audio: UploadFile = File(...),
|
|
language: Optional[str] = Form(None),
|
|
model: Optional[str] = Form(None),
|
|
mode: Optional[str] = Form(None),
|
|
):
|
|
"""Transcribe an audio file to text.
|
|
|
|
Args:
|
|
audio: The audio file to transcribe.
|
|
language: Optional language hint (not currently used; auto-detected).
|
|
model: Whisper model size (legacy; ignored in dual-mode architecture).
|
|
mode: 'fast' (default) uses MLX Turbo for speed; 'accurate' uses
|
|
WhisperX with forced alignment for word-level timing.
|
|
|
|
Returns:
|
|
{
|
|
"text": "full transcription",
|
|
"segments": [ {"start": 0.0, "end": 1.5, "text": "..."}, ... ],
|
|
"language": "en",
|
|
"duration_s": 4.2,
|
|
"transcription_time_s": 0.8,
|
|
"engine": "mlx-whisper"
|
|
}
|
|
"""
|
|
import asyncio
|
|
|
|
# Save upload to a temp file (all backends need a file path)
|
|
ext = os.path.splitext(audio.filename or "audio.wav")[1] or ".wav"
|
|
tmp = tempfile.NamedTemporaryFile(delete=False, suffix=ext)
|
|
try:
|
|
content = await audio.read()
|
|
tmp.write(content)
|
|
tmp.close()
|
|
|
|
use_accurate = (mode or "").strip().lower() == "accurate"
|
|
|
|
def _run():
|
|
if use_accurate:
|
|
# Accurate mode: full WhisperX with forced alignment —
|
|
# for when the user explicitly wants word-level timing.
|
|
from services.asr_backend import get_active_asr_backend
|
|
backend = get_active_asr_backend()
|
|
result = backend.transcribe(tmp.name, word_timestamps=True)
|
|
else:
|
|
# Fast mode (default): use the fastest available engine
|
|
# (MLX Turbo on Apple Silicon). Skip word_timestamps for
|
|
# ~30% latency reduction — dictation doesn't need them.
|
|
from services.asr_backend import get_capture_asr_backend
|
|
backend = get_capture_asr_backend()
|
|
result = backend.transcribe(tmp.name, word_timestamps=False)
|
|
return result, backend.id
|
|
|
|
from services.model_manager import _gpu_pool
|
|
loop = asyncio.get_running_loop()
|
|
t0 = time.perf_counter()
|
|
result, engine_id = await loop.run_in_executor(_gpu_pool, _run)
|
|
elapsed = round(time.perf_counter() - t0, 2)
|
|
|
|
# Normalize result shape
|
|
segments = result.get("segments", [])
|
|
full_text = result.get("text", "")
|
|
if not full_text and segments:
|
|
full_text = " ".join(s.get("text", "") for s in segments).strip()
|
|
|
|
# Calculate audio duration from segments if available
|
|
duration = 0.0
|
|
if segments:
|
|
duration = max(s.get("end", 0) for s in segments)
|
|
|
|
detected_lang = result.get("language", language or "unknown")
|
|
|
|
logger.info(
|
|
"Capture transcription done: engine=%s, elapsed=%.2fs, duration=%.1fs, mode=%s",
|
|
engine_id, elapsed, duration, "accurate" if use_accurate else "fast",
|
|
)
|
|
|
|
return {
|
|
"text": full_text,
|
|
"segments": [
|
|
{
|
|
"start": round(s.get("start", 0), 2),
|
|
"end": round(s.get("end", 0), 2),
|
|
"text": s.get("text", "").strip(),
|
|
}
|
|
for s in segments
|
|
],
|
|
"language": detected_lang,
|
|
"duration_s": round(duration, 2),
|
|
"transcription_time_s": elapsed,
|
|
"engine": engine_id,
|
|
}
|
|
finally:
|
|
try:
|
|
os.unlink(tmp.name)
|
|
except OSError:
|
|
pass
|