diff --git a/.github/CONTRIBUTING.md b/.github/CONTRIBUTING.md
index 36e3a745..db1bfa1c 100644
--- a/.github/CONTRIBUTING.md
+++ b/.github/CONTRIBUTING.md
@@ -28,8 +28,24 @@ Read it before opening a proposal; the licence check in particular ends most of
- [Bun](https://bun.sh/) (frontend package manager)
- [uv](https://docs.astral.sh/uv/) (Python environment manager)
- [ffmpeg](https://ffmpeg.org/) (audio/video processing)
+- [Rust / Cargo](https://rustup.rs/) (desktop shell only)
- Python 3.10+ (managed automatically by `uv`)
+Linux desktop development also needs WebKitGTK/GTK development libraries. On
+Debian or Ubuntu, install the same packages used by CI:
+
+```bash
+sudo apt-get update
+sudo apt-get install -y \
+ libwebkit2gtk-4.1-dev libgtk-3-dev libpango1.0-dev libcairo2-dev \
+ libsoup-3.0-dev libgdk-pixbuf-2.0-dev \
+ libayatana-appindicator3-dev librsvg2-dev libssl-dev libxdo-dev \
+ libasound2-dev build-essential curl wget file
+```
+
+See the [Linux source-build guide](../docs/install/linux.md#building-from-source)
+for Fedora and Arch packages.
+
### Clone & Run
```bash
@@ -68,6 +84,18 @@ names: there is no `desktop=prod` (note the **hyphen** in `desktop-prod`).
Requires [Rust](https://rustup.rs/) and platform-specific Tauri dependencies — see the [Tauri prerequisites](https://v2.tauri.app/start/prerequisites/).
+After installing Rust with rustup on macOS/Linux, either open a new terminal or
+load Cargo into the current one before starting the desktop app:
+
+```bash
+source "$HOME/.cargo/env"
+bun desktop
+```
+
+On Linux, errors such as `Package gdk-3.0 was not found`, `pango.pc` missing,
+or `javascriptcoregtk-4.1` missing mean the native packages above were not
+installed; changing `PKG_CONFIG_PATH` does not fix libraries that are absent.
+
If the app opens but stays on the **setup splash with no buttons**, the Python
backend didn't finish starting — the splash surfaces the stall reason, a log
panel, and a **Retry** button (and Settings → Logs → Backend has the full trace).
diff --git a/CHANGELOG.md b/CHANGELOG.md
index 067b2094..e8e56ebb 100644
--- a/CHANGELOG.md
+++ b/CHANGELOG.md
@@ -30,6 +30,7 @@ The bundled TTS model package (`pyproject.toml`) is versioned independently.
### Added
+- Voice recording now offers microphone and channel selection with a live input-level meter on every desktop platform. (#1481)
- Settings → Appearance → **Navigation style** switches the workspace switcher between the icon rail down the window edge and browser-style tabs across the title bar. Both offer the same workspaces; the choice sticks across launches, and the rail stays the default. Tab labels fold down to icons when the title bar runs out of room — the workspace you're in keeps its name. (#1412)
- Portable mode lets you choose the folder — press **Change…** on the first-run setup screen and put the whole install on an external drive. It also stops being greyed out after a default Program Files install. (#766)
- Settings → Privacy now has an **Invisible watermark** toggle. On by default, available to everyone, and it only affects audio generated after the change. (#1308)
@@ -42,6 +43,9 @@ The bundled TTS model package (`pyproject.toml`) is versioned independently.
### Fixed
+- Basic Dubbing translation remains available without an LLM; Cinematic and Autofit now degrade through the existing Fast translation path instead of blocking the quality choice. (#1481)
+- Linux microphone recording now falls back to WAV when WebKit cannot encode MediaRecorder audio, and desktop scaling/titlebar controls remain responsive at every UI scale. (#1481)
+- Dubbing can install a missing ASR model and retry the same job, navigate back through completed stages, and finish transcription under low GPU memory without producing an empty transcript. (#1481)
- Filenames and other outside data can no longer forge extra lines or terminal commands in backend and frontend diagnostic logs. (#1457)
- Backend journal, dictation reset, voice-catalog, and crash-notification failures are now visible and retryable instead of being silently ignored. (#1459)
- Backend failures keep raw tracebacks, local paths and credentials in the local log instead of returning them in API responses. (#1454)
diff --git a/README.md b/README.md
index 00acb7fa..5526337f 100644
--- a/README.md
+++ b/README.md
@@ -295,6 +295,8 @@ another machine, set `OMNIVOICE_GPTSOVITS_URL` to its credential-free
> Whisper-family engines cover ~100 languages; **FunASR / SenseVoice** adds an all-in-one multilingual path with built-in voice-activity detection and inline speaker diarization. **sherpa-onnx** powers the live dictation model picker — you talk and text appears as you speak. Every engine runs on-device — no API keys, no cloud.
+> If Dubbing needs an ASR model that is not installed yet, it offers the recommended download in place, shows its progress, and retries transcription on the same job when the model is ready.
+
> **GPU without efficient float16?** On older NVIDIA GPUs (Maxwell/Pascal, GTX 16xx) or after a CTranslate2/cuDNN mismatch, the CTranslate2 ASR engines (WhisperX, Faster-Whisper) can't run `float16` and VoiceStudio automatically retries on `int8` — no config needed. If transcription still fails, pin the compute type with the `ASR_COMPUTE_TYPE` env var (escape hatch): `ASR_COMPUTE_TYPE=int8` (or `float32` for CPU). Set it to `int8` and restart the backend.
diff --git a/backend/api/routers/capture_ws.py b/backend/api/routers/capture_ws.py
index 0bcc15d6..2bc1be24 100644
--- a/backend/api/routers/capture_ws.py
+++ b/backend/api/routers/capture_ws.py
@@ -9,7 +9,9 @@ Protocol:
→ Client sends binary audio frames (16-bit PCM or WebM/Opus blobs)
← Server sends JSON messages:
- Opt-in AEC mode (``?aec=1[&sr=16000]``, parity Action 8b): for dictating
+ Raw PCM mode (``?pcm=1&sr=16000``) is the container-free fallback for
+ WebViews without MediaRecorder. Opt-in AEC mode
+ (``?aec=1[&sr=16000]``, parity Action 8b): for dictating
while the app plays audio. Frames must be raw int16 mono PCM, each tagged
with a 1-byte prefix — 0x00 = microphone, 0x01 = playback reference. The
server runs an NLMS echo canceller, cleaning the mic against the reference
@@ -68,6 +70,19 @@ _AEC_NEAR = 0x00 # microphone frame (clean it, then buffer for ASR)
_AEC_FAR = 0x01 # playback reference frame (feed the echo model only)
+def _requested_pcm_sample_rate(query_params) -> int | None:
+ """Return a bounded PCM rate for ``?pcm=1``/``?aec=1`` sessions."""
+ raw_pcm = query_params.get("pcm") in ("1", "true", "on")
+ aec = query_params.get("aec") in ("1", "true", "on")
+ if not raw_pcm and not aec:
+ return None
+ try:
+ sample_rate = int(query_params.get("sr", "16000"))
+ except (TypeError, ValueError):
+ return 16000
+ return sample_rate if 8000 <= sample_rate <= 96000 else 16000
+
+
def _demux_aec_frame(data: bytes) -> tuple[str, bytes]:
"""Split a prefixed AEC binary frame into ``(kind, pcm)``.
@@ -210,12 +225,11 @@ async def ws_transcribe(websocket: WebSocket):
# identical legacy behaviour. When on, frames are 1-byte-tagged raw PCM
# and the cleaned mic stream is muxed via stdlib wave (not ffmpeg).
aec = None
- pcm_sr: int | None = None
+ pcm_sr = _requested_pcm_sample_rate(websocket.query_params)
if websocket.query_params.get("aec") in ("1", "true", "on"):
try:
- pcm_sr = int(websocket.query_params.get("sr", "16000"))
from services.aec import NlmsEchoCanceller
- aec = NlmsEchoCanceller(sample_rate=pcm_sr)
+ aec = NlmsEchoCanceller(sample_rate=pcm_sr or 16000)
logger.info("AEC enabled for dictation session (sr=%d)", pcm_sr)
except Exception as e:
# Bad sr or import failure → fall back to plain dictation.
diff --git a/backend/api/routers/dub_core.py b/backend/api/routers/dub_core.py
index 7113bab8..92509479 100644
--- a/backend/api/routers/dub_core.py
+++ b/backend/api/routers/dub_core.py
@@ -769,6 +769,16 @@ async def dub_transcribe_stream(
preflight_payload = _missing
if _missing is None:
try:
+ # Free recoverable TTS VRAM before ASR chooses its
+ # device. Probing first falsely routed Whisper to
+ # CPU even when this offload made CUDA viable.
+ try:
+ await asyncio.get_running_loop().run_in_executor(
+ _cpu_pool, offload_tts_for_asr
+ )
+ _tts_offloaded["v"] = True
+ except Exception as e:
+ logger.warning("offload_tts_for_asr failed (continuing): %s", e)
# The PyTorch-Whisper backend lazily builds its own pipeline
# when no preloaded `_asr_pipe` is present (issue #255), so it
# no longer needs OMNIVOICE_PRELOAD_TTS_ASR=1.
@@ -898,17 +908,6 @@ async def dub_transcribe_stream(
"chunk_s": transcribe_chunk_s,
})
- # Free VRAM: move TTS model to CPU so WhisperX + VAD can fit.
- # Only offloads when free GPU memory is < 4 GB (e.g. laptop GPUs).
- # Non-fatal: an offload failure must not drop the stream (#255) —
- # transcription can still proceed (it just has less headroom).
- try:
- await loop.run_in_executor(_cpu_pool, offload_tts_for_asr)
- # Restore is now owed on every exit path, not just success (#1191).
- _tts_offloaded["v"] = True
- except Exception as e:
- logger.warning("offload_tts_for_asr failed (continuing): %s", e)
-
all_segments: list[dict] = []
# Words (global-timeline) retained so diarization can re-split a segment
# that spans two speakers' turns at the word boundary (#486).
@@ -916,6 +915,7 @@ async def dub_transcribe_stream(
detected_lang = None
next_seg_id = 0
chunk_errors: list[str] = []
+ chunk_error_codes: list[str] = []
# Speaker turns from an ASR backend that diarizes inline (FunASR cam++).
# When present, _diarize() uses them and skips pyannote (Phase 2, #182).
asr_speaker_turns: list[dict] = []
@@ -960,10 +960,20 @@ async def dub_transcribe_stream(
continue
turns.append({"start": s0 + offset, "end": s1 + offset, "speaker": spk})
return {"chunks": shifted, "language": r.get("language"), "speaker_turns": turns}
- except Exception:
- logger.error("Chunk transcription failed (backend=%s)", _asr_backend.id)
+ except Exception as exc:
+ # Keep diagnostics local and fixed-shape. In particular,
+ # CUDA OOM is a distinct, actionable recovery class rather
+ # than the generic "no segments" dead end.
+ is_memory = isinstance(exc, torch.OutOfMemoryError)
+ logger.error(
+ "Chunk transcription failed (backend=%s; class=%s; details withheld)",
+ _asr_backend.id,
+ type(exc).__name__,
+ )
from core.public_errors import stream_failure
- failure = stream_failure("transcription_failed")
+ failure = stream_failure(
+ "transcription_memory" if is_memory else "transcription_failed"
+ )
return {
"chunks": [],
"language": None,
@@ -986,7 +996,6 @@ async def dub_transcribe_stream(
# worker, and raises the actionable ASRTimeoutError. Run it as
# a task and poll so we can keep yielding pings — the
# EventSource connection drops without them.
- pool_reset_by_guard = False
task = asyncio.ensure_future(run_transcribe_guarded(
_gpu_pool, _transcribe_chunk,
what=f"Dub chunk {i + 1}/{chunks_n}",
@@ -1004,7 +1013,6 @@ async def dub_transcribe_stream(
# The guard already reset the pool; keep the actionable
# message (it names the durable fixes, and — after repeated
# timeouts — the crash-isolated engine escape hatch).
- pool_reset_by_guard = True
logger.error(
"Transcribe chunk %d/%d timed out after %.0fs (attempt %d/%d, job=%s)",
i + 1, chunks_n, transcribe_timeout_s, _attempt,
@@ -1028,11 +1036,14 @@ async def dub_transcribe_stream(
"Retrying transcribe chunk %d/%d after failure/timeout (next attempt %d/%d, job=%s)",
i + 1, chunks_n, _attempt + 1, _CHUNK_TRANSCRIBE_ATTEMPTS, log_safe(job_id),
)
- if not pool_reset_by_guard:
- reset_pool_after_wedge(
- _gpu_pool, what=f"Dub chunk {i + 1}/{chunks_n}")
+ # A completed exception did not wedge the worker. Resetting
+ # the pool here leaked a healthy executor on every ordinary
+ # decode failure; run_transcribe_guarded already resets the
+ # pool on the only case that needs it: a real timeout.
if part.get("error"):
chunk_errors.append(part["error"])
+ if part.get("error_code"):
+ chunk_error_codes.append(part["error_code"])
logger.warning("Chunk %d/%d error: %s", i + 1, chunks_n, log_safe(part["error"]))
if detected_lang is None and part.get("language"):
detected_lang = part["language"]
@@ -1097,7 +1108,9 @@ async def dub_transcribe_stream(
seen.add(s)
uniq.append(s)
if uniq:
- detail = "Transcription produced no segments. " + " | ".join(uniq[:3])
+ # Chunk failures already carry a complete recovery message.
+ # Do not prepend another generic sentence to it.
+ detail = " | ".join(uniq[:3])
# Add the actionable hint for a recognized failure class
# (e.g. pkg_resources missing → install setuptools).
hint = build_failure(" ".join(uniq), stage="transcribe", include_diagnostic=False).get("hint")
@@ -1110,7 +1123,10 @@ async def dub_transcribe_stream(
"check that the source has an audible speech track."
)
logger.error("transcribe yielded 0 segments (job=%s): %s", log_safe(job_id), log_safe(detail))
- yield _sse_event("error", {"detail": detail, "retryable": True})
+ payload = {"detail": detail, "retryable": True}
+ if chunk_error_codes:
+ payload["code"] = chunk_error_codes[0]
+ yield _sse_event("error", payload)
yield _sse_event("done", {})
return
diff --git a/backend/api/routers/setup/download.py b/backend/api/routers/setup/download.py
index ed10f778..fd463210 100644
--- a/backend/api/routers/setup/download.py
+++ b/backend/api/routers/setup/download.py
@@ -13,6 +13,7 @@ import json
import logging
import os
import sys
+import threading
from fastapi import APIRouter, HTTPException
from fastapi.responses import StreamingResponse
@@ -69,6 +70,13 @@ def clear_install_cooldowns() -> None:
# cancelled, and clears the cooldown so a cancel isn't rate-limited.
_cancelled: set[str] = set()
+# One worker per repo. Repeated clicks and feature-level recovery can converge
+# on the same install; starting a second snapshot_download against the same HF
+# cache is wasteful and can corrupt the user-visible progress stream.
+_active_installs: set[str] = set()
+_active_installs_lock = threading.Lock()
+_install_tasks: set[asyncio.Task] = set()
+
def _download_max_workers() -> int:
"""Parallel-FILES worker count for snapshot_download (FDL-02). Default 8 —
@@ -394,6 +402,10 @@ async def install_model(req: InstallModelRequest):
f"Retry in {remaining}s or check your network."
),
)
+ with _active_installs_lock:
+ if req.repo_id in _active_installs:
+ return {"status": "already_running", "repo_id": req.repo_id}
+ _active_installs.add(req.repo_id)
loop = asyncio.get_running_loop()
def _do():
@@ -652,8 +664,17 @@ async def install_model(req: InstallModelRequest):
_cancelled.discard(req.repo_id)
download_aggregator.finish(req.repo_id)
hf_progress.current_repo_id.reset(token)
+ with _active_installs_lock:
+ _active_installs.discard(req.repo_id)
- loop.create_task(asyncio.to_thread(_do))
+ try:
+ task = loop.create_task(asyncio.to_thread(_do))
+ _install_tasks.add(task)
+ task.add_done_callback(_install_tasks.discard)
+ except Exception:
+ with _active_installs_lock:
+ _active_installs.discard(req.repo_id)
+ raise
return {"status": "install_started", "repo_id": req.repo_id}
diff --git a/backend/core/public_errors.py b/backend/core/public_errors.py
index 63ab172b..53d32769 100644
--- a/backend/core/public_errors.py
+++ b/backend/core/public_errors.py
@@ -43,6 +43,15 @@ def stream_failure(code: str) -> dict[str, object]:
"detail": "Transcription failed. Check the selected ASR engine and try again.",
"retryable": True,
},
+ "transcription_memory": {
+ "code": "transcription_memory",
+ "detail": (
+ "Transcription ran out of GPU memory. Close other GPU apps or "
+ "Flush models, then try again; VoiceStudio will use CPU when "
+ "the remaining GPU memory is too low."
+ ),
+ "retryable": True,
+ },
"transcription_timeout": {
"code": "transcription_timeout",
"detail": (
diff --git a/backend/services/asr_backend.py b/backend/services/asr_backend.py
index 80babb54..4ef2153d 100644
--- a/backend/services/asr_backend.py
+++ b/backend/services/asr_backend.py
@@ -1232,6 +1232,11 @@ class PyTorchWhisperBackend(ASRBackend):
# Reuses the `_asr_pipe` attached to the TTS model when available.
self._pipe = asr_pipe
+ # whisper-large-v3-turbo occupies roughly 3.2 GiB before generation adds
+ # its encoder/decoder workspace. Loading it onto a nearly full card works,
+ # then the first transcribe fails with a CUDA OOM and yields zero segments.
+ _CUDA_VRAM_BUDGET_GB = 5.0
+
@classmethod
def is_available(cls) -> tuple[bool, str]:
try:
@@ -1240,6 +1245,40 @@ class PyTorchWhisperBackend(ASRBackend):
except ImportError as e:
return False, f"transformers not installed: {e}"
+ @classmethod
+ def _pick_device(cls) -> str:
+ from services.model_manager import get_best_device
+
+ device = str(get_best_device())
+ if not device.startswith("cuda") or os.environ.get(
+ "OMNIVOICE_ASR_VRAM_PREFLIGHT", "1"
+ ).strip().lower() in ("0", "false", "no"):
+ return device
+ try:
+ import torch
+
+ free, _total = torch.cuda.mem_get_info()
+ free_gb = free / 1024**3
+ except Exception: # noqa: BLE001 — an unavailable probe must not block ASR
+ return device
+ if free_gb >= cls._CUDA_VRAM_BUDGET_GB:
+ return device
+ logger.warning(
+ "PyTorch Whisper VRAM preflight: %.1f GB free < %.1f GB needed "
+ "for reliable CUDA transcription — using CPU instead. Close other "
+ "GPU apps or Flush models to restore GPU-speed ASR.",
+ free_gb,
+ cls._CUDA_VRAM_BUDGET_GB,
+ )
+ return "cpu"
+
+ def ensure_loaded(self) -> None:
+ # Unlike the CTranslate2 backends, this fallback used to inherit the
+ # protocol's no-op loader. Import/model failures therefore appeared on
+ # every chunk as the misleading "produced no segments" result. Load it
+ # once during the stream preflight so the real failure is reported once.
+ self._ensure_pipe()
+
def _ensure_pipe(self):
if self._pipe is not None:
return
@@ -1252,12 +1291,10 @@ class PyTorchWhisperBackend(ASRBackend):
# constructor and this path is skipped.
import torch
from transformers import pipeline as hf_pipeline
- from services.model_manager import get_best_device
-
model_name = os.environ.get(
"OMNIVOICE_PYTORCH_ASR_MODEL", "openai/whisper-large-v3-turbo"
)
- device = get_best_device()
+ device = self._pick_device()
asr_dtype = torch.float16 if str(device).startswith("cuda") else torch.float32
logger.info(
"PyTorchWhisperBackend: loading standalone ASR pipeline %s on %s",
@@ -1268,7 +1305,11 @@ class PyTorchWhisperBackend(ASRBackend):
"automatic-speech-recognition",
model=model_name,
dtype=asr_dtype,
- device_map=device,
+ # `device_map="cpu"` only controls weight placement; the
+ # pipeline can still choose CUDA as its execution device.
+ # `device` is the pipeline-level contract and keeps the
+ # low-VRAM fallback entirely on CPU.
+ device=device,
)
except Exception as e:
# #549: an incomplete transformers install fails to build the ASR
diff --git a/backend/tests/test_asr_oom_fallback.py b/backend/tests/test_asr_oom_fallback.py
index b5d513fb..c5911103 100644
--- a/backend/tests/test_asr_oom_fallback.py
+++ b/backend/tests/test_asr_oom_fallback.py
@@ -85,6 +85,7 @@ def test_float16_unsupported_falls_back_to_int8(monkeypatch):
return object() # cuda int8 succeeds
monkeypatch.setattr(whisperx, "load_model", fake_load_model)
+ monkeypatch.setattr(WhisperXBackend, "_free_vram_gb", staticmethod(lambda: 10.0))
be = WhisperXBackend()
be._device, be._compute_type = "cuda", "float16"
diff --git a/docs/generation-parameters.md b/docs/generation-parameters.md
index 69821441..cdcd2ca9 100644
--- a/docs/generation-parameters.md
+++ b/docs/generation-parameters.md
@@ -62,6 +62,8 @@ Priority: `duration` > `speed`.
>
> **Tip — reference-clip quality transfers.** Zero-shot cloning mirrors the acoustics of the reference clip, not just the voice: a clip recorded in an echoey room clones echoey. Record dry and close-mic for clean output. No effect preset adds reverb unless you choose one that declares it (Cinematic, Warm).
+For an in-app recording, choose the microphone and Auto, Mono, or Stereo in the Voice panel. While recording, the input meter confirms whether VoiceStudio is receiving a usable signal; monitoring is visual and never plays the microphone through the speakers.
+
## Long-Form Generation
To support stable long-form speech generation with low VRAM consumption, the text is automatically split into smaller segments when the estimated duration of the generated speech exceeds `audio_chunk_duration`, with each segment producing approximately `audio_chunk_duration` seconds of audio. This approach allows the model to accept arbitrarily long text and generate arbitrarily long speech with near-constant VRAM consumption.
diff --git a/docs/install/linux.md b/docs/install/linux.md
index aa22e8dc..5522cca1 100644
--- a/docs/install/linux.md
+++ b/docs/install/linux.md
@@ -33,7 +33,12 @@ Everything above, plus the toolchain:
```bash
# Debian / Ubuntu
- sudo apt install libwebkit2gtk-4.1-dev libayatana-appindicator3-dev librsvg2-dev libssl-dev libxdo-dev build-essential
+ sudo apt-get update
+ sudo apt-get install -y \
+ libwebkit2gtk-4.1-dev libgtk-3-dev libpango1.0-dev libcairo2-dev \
+ libsoup-3.0-dev libgdk-pixbuf-2.0-dev \
+ libayatana-appindicator3-dev librsvg2-dev libssl-dev libxdo-dev \
+ libasound2-dev build-essential curl wget file
# Fedora
sudo dnf install webkit2gtk4.1-devel libappindicator-gtk3-devel librsvg2-devel openssl-devel
@@ -51,11 +56,30 @@ Everything above, plus the toolchain:
git clone https://github.com/debpalash/VoiceStudio.git
cd VoiceStudio
bun install
-bun run desktop-prod
+source "$HOME/.cargo/env" # only needed in a shell opened before rustup finished
+bun desktop # development build with hot reload
```
-The first launch creates the Python venv via `uv`, syncs deps, and downloads
-model weights (~2.4 GB). Subsequent launches start in seconds.
+Use `bun run desktop-prod` instead when you need to build and launch the
+production bundle. Both commands create the Python environment via `uv`, sync
+dependencies, and start the backend automatically; do not start the backend in
+a second terminal.
+
+The first Rust build takes longer because Cargo compiles the Tauri shell. If it
+fails with `Package gdk-3.0 was not found`, `pango.pc` missing,
+`libsoup-3.0` missing, or `javascriptcoregtk-4.1` missing, install the complete
+Debian/Ubuntu package block above. Those messages mean the development
+libraries are absent, not that `PKG_CONFIG_PATH` needs changing. Verify them
+with:
+
+```bash
+pkg-config --exists \
+ gdk-3.0 pango cairo libsoup-3.0 javascriptcoregtk-4.1 gdk-pixbuf-2.0 \
+ && echo "Tauri system libraries are ready"
+```
+
+The first app launch downloads model weights on demand. Subsequent launches
+reuse the Rust build, Python environment, and installed models.
## Install (AppImage)
diff --git a/docs/install/troubleshooting.md b/docs/install/troubleshooting.md
index 04a6d23b..8092839f 100644
--- a/docs/install/troubleshooting.md
+++ b/docs/install/troubleshooting.md
@@ -681,6 +681,14 @@ up, the shell now prints the exit code and where to look (the cargo/tauri
output above it, plus `omnivoice.log` and `backend_err.log` in your VoiceStudio
data folder) instead of exiting silently.
+If Cargo stops before the window is built with `Package gdk-3.0 was not found`,
+`pango.pc` missing, `libsoup-3.0` missing, or `javascriptcoregtk-4.1` missing,
+the Ubuntu/Debian WebKitGTK development packages are absent. Install the full
+package block in the [Linux source-build guide](linux.md#building-from-source),
+then rerun `source "$HOME/.cargo/env"` and `bun desktop`. Do not set a custom
+`PKG_CONFIG_PATH` unless the libraries were deliberately installed outside the
+system package manager.
+
**Still stuck?** Open the details, copy the output, and file it with **Report**
— that output is the thing that makes the failure diagnosable.
diff --git a/frontend/public/aec-worklet.js b/frontend/public/aec-worklet.js
index 4a34bdc5..f722c439 100644
--- a/frontend/public/aec-worklet.js
+++ b/frontend/public/aec-worklet.js
@@ -1,5 +1,5 @@
// AudioWorklet processor for the opt-in dictate-over-playback AEC (parity
-// Action 8). Accumulates mono Float32 input into fixed-size frames and posts
+// Action 8). Accumulates Float32 input into fixed-size frames and posts
// them to the main thread, which converts them to tagged int16 PCM. The same
// processor serves both the microphone capture and the playback-reference tap
// — only the wiring on the main thread differs.
@@ -12,21 +12,27 @@ class AecFrameEmitter extends AudioWorkletProcessor {
super();
const frame =
(options && options.processorOptions && options.processorOptions.frameSize) || 320;
- this._frameSize = frame; // 320 samples = 20 ms @ 16 kHz
- this._buf = new Float32Array(frame);
+ this._frameSize = frame; // 320 frames = 20 ms @ 16 kHz
+ this._channels = Math.max(
+ 1,
+ Math.min(2, (options && options.processorOptions && options.processorOptions.channels) || 1),
+ );
+ this._buf = new Float32Array(frame * this._channels);
this._n = 0;
}
process(inputs) {
const input = inputs[0];
- // input[0] is the first (mono) channel; absent when upstream is idle.
+ // Interleave requested channels. A missing secondary channel mirrors the
+ // first so the WAV layout stays valid on single-channel devices.
if (input && input[0]) {
- const ch = input[0];
- for (let i = 0; i < ch.length; i++) {
- this._buf[this._n++] = ch[i];
- if (this._n >= this._frameSize) {
+ for (let i = 0; i < input[0].length; i++) {
+ for (let channel = 0; channel < this._channels; channel++) {
+ this._buf[this._n++] = (input[channel] || input[0])[i];
+ }
+ if (this._n >= this._buf.length) {
// Copy out — the buffer is reused for the next frame.
- this.port.postMessage(this._buf.slice(0, this._frameSize));
+ this.port.postMessage(this._buf.slice());
this._n = 0;
}
}
diff --git a/frontend/src-tauri/capabilities/default.json b/frontend/src-tauri/capabilities/default.json
index ae938b97..12962e8b 100644
--- a/frontend/src-tauri/capabilities/default.json
+++ b/frontend/src-tauri/capabilities/default.json
@@ -13,6 +13,7 @@
"core:window:allow-set-fullscreen",
"core:window:allow-minimize",
"core:window:allow-close",
+ "core:webview:allow-set-webview-zoom",
"dialog:allow-save",
"dialog:allow-open",
"dialog:allow-message",
diff --git a/frontend/src-tauri/tauri.conf.json b/frontend/src-tauri/tauri.conf.json
index 568ef369..be50e40b 100644
--- a/frontend/src-tauri/tauri.conf.json
+++ b/frontend/src-tauri/tauri.conf.json
@@ -22,6 +22,7 @@
"minHeight": 600,
"resizable": true,
"fullscreen": false,
+ "decorations": false,
"titleBarStyle": "Overlay",
"hiddenTitle": true,
"maximized": true,
diff --git a/frontend/src/App.jsx b/frontend/src/App.jsx
index a4c0a241..57b75107 100644
--- a/frontend/src/App.jsx
+++ b/frontend/src/App.jsx
@@ -72,6 +72,7 @@ import { toastErrorWithReport } from './utils/errorToast';
import { listenDictationNotice, showDictationNotice } from './utils/dictationNotice';
import { addBreadcrumb } from './utils/breadcrumbs';
import { appShellClasses } from './utils/appShellClasses';
+import { applyUiScale } from './utils/uiScaleEngine';
import { recordValueMoment } from './utils/donationMoments';
import {
POPULAR_LANGS,
@@ -145,26 +146,13 @@ function App() {
return () => ro.disconnect();
}, []);
- // Engine capability probe (#523/#524): does this WebView honor `zoom` as a
- // LAYOUT transform? Chromium (WebView2 / macOS WebKit) and modern WebKitGTK
- // do; older WebKitGTK (Linux) treats it as a no-op. The .app-container sizing
- // branches on the result (index.css) so the shell fills the window on BOTH —
- // no black band on WebKitGTK, no clipped Generate/Settings CTAs on Chromium.
- // Measuring a real zoomed element is robust where @supports(zoom)/UA-sniffing
- // aren't (both report "supported" on WebKitGTK even when zoom doesn't lay out).
+ // Desktop UI scale belongs at the webview boundary. A CSS `zoom` probe can
+ // report the expected bounding box on WebKitGTK even when the painted shell
+ // still occupies only the upper-left of the window. Tauri's native zoom keeps
+ // layout and paint in agreement; browser/dev sessions retain the CSS path.
useLayoutEffect(() => {
- let honored = true;
- try {
- const probe = document.createElement('div');
- probe.style.cssText = 'position:absolute;left:-9999px;top:0;width:100px;height:100px;zoom:2';
- document.body.appendChild(probe);
- honored = Math.round(probe.getBoundingClientRect().width) >= 150;
- probe.remove();
- } catch {
- honored = true;
- } // safe default: the existing zoom path
- document.documentElement.dataset.zoomLayout = honored ? 'on' : 'off';
- }, []);
+ void applyUiScale(uiScale);
+ }, [uiScale]);
const shellSizeClass =
shellWidth <= 600 ? 'shell-mini' : shellWidth <= 1100 ? 'shell-narrow' : '';
const theme = useAppStore((s) => s.theme);
@@ -433,8 +421,19 @@ function App() {
const [compareProgress, setCompareProgress] = useState('');
// ═══ MIC RECORDING ═══
- const { isRecording, isCleaning, recordingTime, startRecording, stopRecording } =
- useRecording(ingestRefAudio);
+ const {
+ isRecording,
+ isCleaning,
+ recordingTime,
+ audioInputs,
+ selectedAudioInputId,
+ setSelectedAudioInputId,
+ channelMode,
+ setChannelMode,
+ inputLevelStore,
+ startRecording,
+ stopRecording,
+ } = useRecording(ingestRefAudio);
// ═══ DUB STATE ═══
const dubJobId = useAppStore((s) => s.dubJobId);
@@ -525,10 +524,12 @@ function App() {
setPreviewAudios,
transcribeElapsed,
transcribeProgress,
+ asrInstall,
handleDubUpload: _handleDubUpload,
handleDubIngestUrl,
handleDubAbort,
handleDubRetryTranscribe,
+ handleInstallMissingAsr,
handleDubStop,
handleDubGenerate,
handleCleanupSegments,
@@ -1176,7 +1177,7 @@ function App() {
// awaiting_setup stage would never get to render.
if (bootstrapStage === 'awaiting_setup') {
return (
-
+
);
@@ -1189,7 +1190,7 @@ function App() {
// flash the empty studio before the wizard has a chance to mount.
if (!setupChecked) {
return (
-
+
);
@@ -1472,6 +1473,7 @@ function App() {
dubLocalBlobUrl={dubLocalBlobUrl}
transcribeElapsed={transcribeElapsed}
transcribeProgress={transcribeProgress}
+ asrInstall={asrInstall}
translateProvider={translateProvider}
setTranslateProvider={setTranslateProvider}
onGlossaryChange={setGlossaryTerms}
@@ -1487,6 +1489,7 @@ function App() {
handleDubUpload={handleDubUpload}
handleDubIngestUrl={handleDubIngestUrl}
handleDubRetryTranscribe={handleDubRetryTranscribe}
+ handleInstallMissingAsr={handleInstallMissingAsr}
handleDubStop={handleDubStop}
handleDubGenerate={handleDubGenerate}
handleDubDownload={handleDubDownload}
@@ -1596,6 +1599,12 @@ function App() {
isRecording={isRecording}
isCleaning={isCleaning}
recordingTime={recordingTime}
+ audioInputs={audioInputs}
+ selectedAudioInputId={selectedAudioInputId}
+ setSelectedAudioInputId={setSelectedAudioInputId}
+ channelMode={channelMode}
+ setChannelMode={setChannelMode}
+ inputLevelStore={inputLevelStore}
vdStates={vdStates}
setVdStates={setVdStates}
isGenerating={isGenerating}
diff --git a/frontend/src/components/CaptureWidget.jsx b/frontend/src/components/CaptureWidget.jsx
index 96fbdbfd..742a3d00 100644
--- a/frontend/src/components/CaptureWidget.jsx
+++ b/frontend/src/components/CaptureWidget.jsx
@@ -13,6 +13,7 @@ import { showMicDeniedGuide } from '../utils/micDeniedToast';
import { asrMissingPayload, toastAsrModelMissing } from '../utils/asrModelMissing';
import { createWaveform } from './captureWaveform';
import { emitDictationNotice } from '../utils/dictationNotice';
+import { audioFormatForMimeType, startSupportedMediaRecorder } from '../utils/mediaRecorder';
// True inside the Tauri shell (desktop app / widget window); false in the
// browser webui / Docker, where the native commands don't exist. Gating on
@@ -101,8 +102,8 @@ const IDLE_VISIBLE_POLL_MS = 600;
// A dictation model id is a sherpa-onnx live model when it carries the
// `sherpa-` prefix the backend assigns (see services/sherpa_dictation.py). Only
-// then do we open the low-latency raw-PCM streaming path; anything else (or no
-// selection) falls through to the legacy MediaRecorder/WebM path unchanged.
+// then do we open the low-latency raw-PCM streaming path. Other models use a
+// supported MediaRecorder container when available, or raw PCM on WebKitGTK.
export function isSherpaModel(id) {
return typeof id === 'string' && id.startsWith('sherpa-');
}
@@ -366,6 +367,7 @@ export default function CaptureWidget({ onDismiss }) {
const mediaRecorderRef = useRef(null);
const chunksRef = useRef([]);
+ const recordingFormatRef = useRef({ mimeType: 'audio/webm', extension: 'webm' });
const streamRef = useRef(null);
const timerRef = useRef(null);
const wsRef = useRef(null);
@@ -383,6 +385,7 @@ export default function CaptureWidget({ onDismiss }) {
// raw PCM via an AudioWorklet and tag mic/far-end frames instead of using
// MediaRecorder. All AEC state lives in refs so the default path is inert.
const aecModeRef = useRef(false);
+ const pcmModeRef = useRef(false);
const aecStopRef = useRef(null); // async teardown of the mic worklet graph
const farEndUnsubRef = useRef(null); // unsubscribe from the far-end bus
@@ -401,6 +404,7 @@ export default function CaptureWidget({ onDismiss }) {
console.warn('mic worklet teardown failed:', err);
}
aecModeRef.current = false;
+ pcmModeRef.current = false;
}, []);
// Hydrate dictation prefs (enabled / mode / model) from the backend once. The
@@ -564,7 +568,7 @@ export default function CaptureWidget({ onDismiss }) {
clearTimeout(dismissTimerRef.current);
dismissTimerRef.current = null;
}
- if (aecModeRef.current || sherpaModeRef.current) teardownAec();
+ if (aecModeRef.current || sherpaModeRef.current || pcmModeRef.current) teardownAec();
setState('idle');
setTranscript('');
setPartialText('');
@@ -595,7 +599,7 @@ export default function CaptureWidget({ onDismiss }) {
if (mediaRecorderRef.current && mediaRecorderRef.current.state !== 'inactive') {
mediaRecorderRef.current.stop();
}
- if (aecModeRef.current || sherpaModeRef.current) {
+ if (aecModeRef.current || sherpaModeRef.current || pcmModeRef.current) {
teardownAec();
}
if (streamRef.current) {
@@ -782,6 +786,7 @@ export default function CaptureWidget({ onDismiss }) {
});
streamRef.current = stream;
chunksRef.current = [];
+ recordingFormatRef.current = { mimeType: 'audio/webm', extension: 'webm' };
wsPendingRef.current = [];
wsHadFinalRef.current = false;
committedRef.current = [];
@@ -804,10 +809,6 @@ export default function CaptureWidget({ onDismiss }) {
dismissTimerRef.current = null;
}
- const mimeType = MediaRecorder.isTypeSupported('audio/webm;codecs=opus')
- ? 'audio/webm;codecs=opus'
- : 'audio/webm';
-
// Read prefs at start time (avoids stale closures). AEC is opt-in; the
// sherpa live engine is selected when the persisted dictation model is a
// sherpa-onnx model — that path streams raw int16 PCM and emits live
@@ -815,10 +816,29 @@ export default function CaptureWidget({ onDismiss }) {
const aecOn = useAppStore.getState().aecEnabled === true;
const modelId = useAppStore.getState().dictationModelId;
const sherpaOn = isSherpaModel(modelId);
+ const supportedRecorder =
+ aecOn || sherpaOn
+ ? null
+ : startSupportedMediaRecorder(stream, {
+ onData: (e) => {
+ if (e.data.size === 0) return;
+ if (e.data.type) recordingFormatRef.current = audioFormatForMimeType(e.data.type);
+ chunksRef.current.push(e.data);
+ void e.data.arrayBuffer().then((buf) => {
+ const ws = wsRef.current;
+ if (ws && ws.readyState === WebSocket.OPEN) ws.send(buf);
+ else wsPendingRef.current.push(buf);
+ });
+ },
+ onStop: () => {},
+ });
+ const pcmFallback = !aecOn && !sherpaOn && supportedRecorder === null;
+ if (supportedRecorder) mediaRecorderRef.current = supportedRecorder.recorder;
aecModeRef.current = aecOn;
sherpaModeRef.current = sherpaOn;
+ pcmModeRef.current = pcmFallback;
// Raw-PCM transport is used whenever AEC or the sherpa live engine is on.
- const pcmMode = aecOn || sherpaOn;
+ const pcmMode = aecOn || sherpaOn || pcmFallback;
// Open WebSocket BEFORE starting capture.
try {
@@ -827,14 +847,31 @@ export default function CaptureWidget({ onDismiss }) {
// • sherpa → ?model=&sr=16000 (raw int16 PCM, live partials)
// • AEC → ?aec=1&sr=16000 (tagged raw PCM, NLMS canceller)
// • both → ?model=&aec=1&sr=16000
- // • neither → /ws/transcribe (legacy MediaRecorder/WebM)
+ // • no recorder → ?pcm=1&sr=16000 (WebKitGTK fallback)
+ // • otherwise → /ws/transcribe (negotiated media container)
const params = [];
if (sherpaOn) params.push(`model=${encodeURIComponent(modelId)}`);
if (aecOn) params.push('aec=1');
+ if (pcmFallback) params.push('pcm=1');
if (pcmMode) params.push('sr=16000');
const wsPath = params.length ? `/ws/transcribe?${params.join('&')}` : '/ws/transcribe';
const ws = new WebSocket(buildWsUrl(wsPath));
ws.binaryType = 'arraybuffer';
+ const failRawPcmSession = () => {
+ if (
+ wsHadFinalRef.current ||
+ !(sherpaModeRef.current || aecModeRef.current || pcmModeRef.current)
+ ) {
+ return false;
+ }
+ wsHadFinalRef.current = true;
+ stopCaptureGraph();
+ setTrayRecording(false);
+ setModelStatus(null);
+ setErrorInfo({ kind: 'server', message: '' });
+ setState('error');
+ return true;
+ };
ws.onopen = () => {
for (const buf of wsPendingRef.current) {
try {
@@ -975,7 +1012,7 @@ export default function CaptureWidget({ onDismiss }) {
toastAsrModelMissing(asrMissingPayload(msg));
setErrorInfo({ kind: 'transcription', message: t('asr_missing.message') });
setState('error');
- } else if (sherpaModeRef.current || aecModeRef.current) {
+ } else if (sherpaModeRef.current || aecModeRef.current || pcmModeRef.current) {
// Raw-PCM paths have no WebM blob to re-POST — surface the
// backend's error instead of leaving the pill wedged in
// "Transcribing…" forever.
@@ -992,6 +1029,7 @@ export default function CaptureWidget({ onDismiss }) {
};
ws.onerror = () => {
wsRef.current = null;
+ failRawPcmSession();
};
ws.onclose = () => {
wsRef.current = null;
@@ -1005,6 +1043,7 @@ export default function CaptureWidget({ onDismiss }) {
}
return;
}
+ if (failRawPcmSession()) return;
if (
!wsHadFinalRef.current &&
mediaRecorderRef.current &&
@@ -1091,22 +1130,8 @@ export default function CaptureWidget({ onDismiss }) {
}
mediaRecorderRef.current = null;
} else {
- const recorder = new MediaRecorder(stream, { mimeType });
- recorder.ondataavailable = (e) => {
- if (e.data.size > 0) {
- chunksRef.current.push(e.data);
- e.data.arrayBuffer().then((buf) => {
- const ws = wsRef.current;
- if (ws && ws.readyState === WebSocket.OPEN) {
- ws.send(buf);
- } else {
- wsPendingRef.current.push(buf);
- }
- });
- }
- };
- recorder.onstop = () => {};
- recorder.start(250);
+ const { recorder, mimeType, extension } = supportedRecorder;
+ recordingFormatRef.current = { mimeType, extension };
mediaRecorderRef.current = recorder;
}
// The session may already have RESOLVED while the mic graph was being
@@ -1141,6 +1166,7 @@ export default function CaptureWidget({ onDismiss }) {
stopCaptureGraph();
return;
}
+ stopCaptureGraph();
// Distinguish "permission denied" (→ per-OS settings hint) from
// "no device" / "device busy" / anything else (#323).
toast.error(micErrorMessage(t, err), { duration: 6000 });
@@ -1193,13 +1219,14 @@ export default function CaptureWidget({ onDismiss }) {
const sendForTranscription = useCallback(async () => {
if (wsHadFinalRef.current) return;
- // No WebM blob exists on any raw-PCM path (AEC or sherpa live) — the WS is
- // the only result channel there.
- if (aecModeRef.current || sherpaModeRef.current) return;
+ // No encoded blob exists on a raw-PCM path — the WS is the only result
+ // channel there.
+ if (aecModeRef.current || sherpaModeRef.current || pcmModeRef.current) return;
- const blob = new Blob(chunksRef.current, { type: 'audio/webm' });
+ const { mimeType, extension } = recordingFormatRef.current;
+ const blob = new Blob(chunksRef.current, { type: mimeType });
const formData = new FormData();
- formData.append('audio', blob, 'capture.webm');
+ formData.append('audio', blob, `capture.${extension}`);
formData.append('mode', captureMode);
try {
diff --git a/frontend/src/components/CaptureWidget.test.jsx b/frontend/src/components/CaptureWidget.test.jsx
index 26f263bd..528fedc9 100644
--- a/frontend/src/components/CaptureWidget.test.jsx
+++ b/frontend/src/components/CaptureWidget.test.jsx
@@ -132,6 +132,7 @@ describe('CaptureWidget', () => {
mocks.holder.paste = async () => undefined;
mocks.holder.calls = [];
mocks.holder.onFrame = null;
+ mocks.state.dictationModelId = 'sherpa-parakeet-tdt-v3';
FakeWebSocket.instances = [];
global.WebSocket = FakeWebSocket;
global.MediaRecorder = FakeMediaRecorder;
@@ -151,6 +152,19 @@ describe('CaptureWidget', () => {
const pasteCalls = () => mocks.holder.calls.filter(([c]) => c === 'simulate_paste');
const typeCalls = () => mocks.holder.calls.filter(([c]) => c === 'simulate_type');
+ it('falls back to raw PCM when MediaRecorder cannot be constructed', async () => {
+ mocks.state.dictationModelId = null;
+ delete global.MediaRecorder;
+ render(withI18n());
+
+ const ws = await startSession();
+ expect(ws.url).toContain('/ws/transcribe?pcm=1&sr=16000');
+ expect(mocks.holder.onFrame).toBeTypeOf('function');
+
+ act(() => mocks.holder.onFrame(new Float32Array([0.25, -0.25])));
+ expect(ws.sent.some((value) => value instanceof ArrayBuffer)).toBe(true);
+ });
+
it('renders truthful model status from {type:"status"} frames', async () => {
render(withI18n());
const ws = await startSession();
diff --git a/frontend/src/components/Header.jsx b/frontend/src/components/Header.jsx
index f67c99b3..557db85f 100644
--- a/frontend/src/components/Header.jsx
+++ b/frontend/src/components/Header.jsx
@@ -16,6 +16,9 @@ import {
Library,
FileText,
Trash2,
+ Minus,
+ Square,
+ X,
} from 'lucide-react';
import { Button, Badge } from '../ui';
import NotificationPanel from './NotificationPanel';
@@ -112,6 +115,18 @@ function WaveBars({ color = '#f3a5b6', active }) {
);
}
+async function runWindowAction(action) {
+ try {
+ const { getCurrentWindow } = await import('@tauri-apps/api/window');
+ const appWindow = getCurrentWindow();
+ if (action === 'minimize') await appWindow.minimize();
+ else if (action === 'maximize') await appWindow.toggleMaximize();
+ else if (action === 'close') await appWindow.close();
+ } catch {
+ console.warn('Window control action failed');
+ }
+}
+
export default function Header({
mode,
setMode,
@@ -125,6 +140,7 @@ export default function Header({
// breadcrumb + wordmark normally sit — the tabs already say where you are,
// and two answers to that question in one bar is one too many.
const tabsInTitlebar = navStyle === 'tabs';
+ const showWindowControls = typeof window !== 'undefined' && '__TAURI_INTERNALS__' in window;
const { t } = useTranslation();
// Sysinfo is subscribed here (not in App via useAppData) so the 5s poll
// only re-renders the header chrome, not the whole App tree.
@@ -233,7 +249,6 @@ export default function Header({
) : (
/* Left: view title + breadcrumb */
-
- OmniVoice
+ VoiceStudio
)}
@@ -455,6 +468,37 @@ export default function Header({
)}