Files
VoiceStudio/scripts/render_demos_omnivoice.py
T
Palash DebnathandClaude Opus 4.8 8b00dc1f4f feat: onboarding demos, opt-in bug reporting, error-docs deeplinks + issue triage (#133)
* feat: onboarding demos, opt-in bug reporting, error-docs deeplinks + issue triage

Working-tree snapshot bundling several in-flight workstreams (v0.3.0):

- Onboarding/demo system: DemoPresetGrid, DictationDemo, DubbingDemo components
  + tests, render scripts (render_demos_omnivoice.py, build_demos.sh,
  build_dub_demo.sh), personalities preview URLs, alembic 0002 voice-profile
  demo fields.
- Opt-in bug reporting: ReportBugButton (prefilled GitHub-issue URL path).
- Error transparency UX: errorDocsMap deeplinks + BootstrapSplash/error wiring.
- Dub workspace: DubSegmentRow/Table, WaveformTimeline, dubSlice tweaks.
- Issue triage: .planning/issue-clusters/ (plan-01..05 root-cause masters,
  GH #128-#132).
- CLAUDE.md: hard rule — everything ships on v0.3.0, no version bumps.

KNOWN GAP (why this is a draft): the generated demo audio assets are NOT in
this tree, and backend/assets/samples/demo_voice.wav is deleted. onboarding.py
guards the missing file (skips seeding the demo profile with a warning), so no
crash — but first-run Launchpad will be empty and /demo_audio/ preview URLs
404 until assets are regenerated via scripts/build_demos.sh. Do not merge
before regenerating + committing the demo assets.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(dub): timing strategies — kill audio compression, add Concise + Stretch Video

Replaces the current audio time-compression default (atempo squeeze to fit
slot) that produced chipmunk/alien output on high-density target languages
like Bengali. Two new user-selectable modes; legacy behaviour kept behind
an explicit "Strict slot" choice.

New `DubRequest.timing_strategy` enum (default "concise"):
  - "concise"        Translator trims text to fit at natural rate; if it
                     still overflows, hard-trim at slot with a fade so we
                     never overlap the next speaker. Surface overflow_s
                     per segment so the user can shorten the text.
  - "stretch_video"  Audio plays at natural 1.0× rate. Backend computes a
                     per-segment new timeline; persists a video_stretch_plan
                     on the job. Mux step (dub_export) builds an ffmpeg
                     trim+setpts+concat filter graph that stretches each
                     segment's video portion to match the natural-rate dub
                     audio. Gaps/pre-roll/tail pass through at 1.0×.
                     Sub burn under stretch_video is skipped in one pass
                     (cues would drift).
  - "strict_slot"    Legacy atempo squeeze. Retained for back-compat.

Director rate-bias side-effect (seg_speed *= bias) now gated on strict_slot
only, so "urgent"/"slow" direction tokens keep their instruct effect in
the new modes without chipmunking.

Per-segment fit_status emitted in the SSE done event:
  {status: "fits" | "overflows" | "video_stretched", overflow_s?, stretch_ratio?}
DubSegmentRow's "Sync: 100%" badge (which was lying — sync_ratio was always
~1.0 because the TTS loop pre-trimmed to slot) is replaced with a truthful
"Fits / Overflows +Ns / Video 1.18×" label.

Frontend:
  - prefsSlice.timingStrategy (persisted, store v3→v4 with safe migrate).
  - DubTab footer Segmented control: "Concise · Stretch Video · Strict slot".
  - useDubWorkflow passes timing_strategy on /dub/generate; consumes fit_status.

Tests: tests/test_dub_timing_strategy.py — 13 cases covering schema
defaults/validation, _build_video_stretch_filter_graph (pre-roll, gap,
tail, empty-plan early return, post-subtitle chain-in), and
_video_stretch_plan_for guards. 30/30 existing dub tests still pass.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(waveform): surface missing source as "Source media missing" instead of code-4 black box

When a project's underlying media file is gone (moved or deleted between
save and reload) the <video> element fires MediaError code 4 and the
companion audio fetch returns HTTP 404 — both were silently warned to
the console while the user stared at an unresponsive black panel and an
empty waveform.

- WaveformTimeline now flips loadError when the video element rejects
  code 3 (decode) or 4 (src not supported), and tracks `sourceMissing`
  separately so the error UI can name the actual problem.
- The audio decode fallback chain catches HTTP 404 specifically and
  treats it as source-missing instead of loading silent empty peaks —
  an empty waveform on a deleted source is more confusing than a clear
  "Re-upload the video to continue" message.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(tray): "Show OmniVoice" reloads when the webview is blank

When the dev Vite server restarts (or the main window is created before
the backend is ready), the webview load fails and the window is left
with `<body></body>` plus a "Could not connect to the server" console
error. Clicking "Show OmniVoice" from the tray menu just re-showed the
broken window — there was no recovery path short of quit+relaunch.

Now the show handler runs a tiny eval after `show()`/`set_focus()` that
calls `location.reload()` only when `document.body.childElementCount === 0`.
A healthy window doesn't blink (body is non-empty); a blank one
self-recovers as soon as the user clicks Show.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(#133): bug-report diagnostics field mapping + drop unused imports

Address PR #133 review:
- ReportBugButton: /system/info exposes `platform` + `device`, not
  `os`/`torch_device`/`gpu` — those reads silently dropped OS/GPU from every
  bug report. Map to the real fields (CodeRabbit). Also remove the dead
  `home` local in stripHome (CodeQL unused-variable).
- DictationDemo: drop unused `Loader` import (CodeQL unused-import).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-29 17:24:25 +05:30

221 lines
8.2 KiB
Python
Executable File

#!/usr/bin/env python3
"""Re-render the demo bundle using the real OmniVoice TTS engine.
This is the production-quality counterpart to scripts/build_demos.sh, which
uses macOS `say` to bootstrap the demo bundle. Run this once on a machine
with OmniVoice model weights cached (typically your dev box) to replace the
`say`-rendered placeholders with engine output. Commit the resulting WAVs.
Prerequisites:
* The project's .venv exists and is activated (`uv sync`).
* OmniVoice model weights cached under $HF_HUB_CACHE (the first
`model.generate()` call will download them otherwise — ~5 GB).
* Run from the repo root: `python3 scripts/render_demos_omnivoice.py`.
What it produces:
* backend/assets/samples/demo_voice.wav (clone reference)
* backend/assets/samples/demo_clone_output.wav (clone pre-rendered)
* backend/assets/samples/voice_design/demo_voice_design_<slug>.wav (7)
* Updated manifest with rendered_by="omnivoice@<git_sha>"
Not regenerated by this script:
* backend/assets/samples/dictation/*.wav — those need to be human speech
or human-quality TTS for the WhisperX replay path to demonstrate real
transcription. `say` output is fine; engine TTS is overkill and slow.
* backend/assets/demo/dubbing/*.mp4 — see scripts/build_dub_demo.sh.
Reproducibility:
* seed=42 fixed across all renders so re-running the script regenerates
the same audio byte-for-byte (modulo torch nondeterminism).
"""
from __future__ import annotations
import argparse
import json
import os
import subprocess
import sys
import time
from pathlib import Path
REPO_ROOT = Path(__file__).resolve().parent.parent
BACKEND_DIR = REPO_ROOT / "backend"
SAMPLES_DIR = BACKEND_DIR / "assets" / "samples"
VOICE_DESIGN_DIR = SAMPLES_DIR / "voice_design"
# Make `backend/` importable so we can pull personalities + the engine.
sys.path.insert(0, str(BACKEND_DIR))
# Cloning demo — must match scripts/build_demos.sh exactly so the manifest
# stays in sync with what the bootstrap script produced.
CLONE_REF_TEXT = (
"Hi, I'm the OmniVoice demo voice. Everything you hear me say from now on "
"was synthesized on your own machine. No cloud, no account, just you and "
"the model."
)
CLONE_OUTPUT_TEXT = (
"Welcome aboard. I was just a three-second clip a moment ago. Now I can "
"say anything you'd like, in your voice or mine."
)
def _git_sha() -> str:
try:
return subprocess.check_output(
["git", "rev-parse", "--short", "HEAD"],
cwd=REPO_ROOT, text=True,
).strip()
except Exception:
return "unknown"
def _save_wav(audio_tensor, sample_rate: int, out_path: Path):
"""Save a torch tensor (C, T) or (T,) to a 16-bit PCM WAV."""
import torch
import torchaudio
if audio_tensor.dim() == 1:
audio_tensor = audio_tensor.unsqueeze(0)
# Ensure mono — most OmniVoice outputs are mono already.
if audio_tensor.shape[0] > 1:
audio_tensor = audio_tensor.mean(dim=0, keepdim=True)
out_path.parent.mkdir(parents=True, exist_ok=True)
# Normalize to safe headroom and clip — matches the `say` output level.
peak = audio_tensor.abs().max().item()
if peak > 0:
audio_tensor = audio_tensor / peak * 0.97
torchaudio.save(
str(out_path),
audio_tensor.to(torch.float32),
sample_rate,
encoding="PCM_S", bits_per_sample=16,
)
def render_cloning(model, args):
"""Render the cloning demo: reference clip + pre-rendered output.
The reference clip is itself synthesized — chicken-and-egg, but the
OmniVoice engine in non-zero-shot mode (no ref_audio) accepts a plain
`instruct=` taxonomy string and produces a clean voice.
"""
print("── Cloning demo ─────────────────────────────────────")
sr = getattr(model, "sampling_rate", 24000)
# 1) Reference clip — synthesized with a "neutral female narrator"
# instruct so the timbre is the same across re-renders.
out_ref = SAMPLES_DIR / "demo_voice.wav"
if args.skip_existing and out_ref.exists():
print(f" · skip (exists): {out_ref.name}")
else:
audios = model.generate(
text=CLONE_REF_TEXT,
instruct="female, middle-aged, moderate pitch, american accent",
num_step=24,
)
_save_wav(audios[0], sr, out_ref)
print(f" ✓ {out_ref.name} ({sr} Hz, omnivoice)")
# 2) Pre-rendered clone output — same voice, different text. Use the
# reference clip we just rendered as the speaker reference so this
# actually demonstrates cloning rather than independent synthesis.
out_clone = SAMPLES_DIR / "demo_clone_output.wav"
if args.skip_existing and out_clone.exists():
print(f" · skip (exists): {out_clone.name}")
else:
audios = model.generate(
text=CLONE_OUTPUT_TEXT,
ref_audio=str(out_ref),
ref_text=CLONE_REF_TEXT,
num_step=24,
)
_save_wav(audios[0], sr, out_clone)
print(f" ✓ {out_clone.name} ({sr} Hz, cloned)")
def render_voice_design(model, args):
"""Re-render the 7 voice-design preset previews."""
print("\n── Voice design presets ─────────────────────────────")
from core.personalities import PERSONALITIES
sr = getattr(model, "sampling_rate", 24000)
demos = [p for p in PERSONALITIES if p.get("is_demo")]
if not demos:
print(" No is_demo presets found in personalities.py")
return
for preset in demos:
slug = preset["id"]
out = VOICE_DESIGN_DIR / f"demo_voice_design_{slug}.wav"
if args.skip_existing and out.exists():
print(f" · skip (exists): {out.name}")
continue
audios = model.generate(
text=preset["script"],
instruct=preset["instruct"],
language=preset.get("language"),
num_step=24,
)
_save_wav(audios[0], sr, out)
print(f" ✓ {out.name} ({preset['name']})")
def update_manifest(args):
"""Update the existing manifest with rendered_by + rendered_at."""
print("\n── Manifest ─────────────────────────────────────────")
mpath = SAMPLES_DIR / "demo" / "manifest.json"
if not mpath.exists():
print(f" ! manifest not found at {mpath} — run scripts/build_demos.sh first")
return
data = json.loads(mpath.read_text())
data["rendered_by"] = f"omnivoice@{_git_sha()}"
data["rendered_at"] = time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime())
mpath.write_text(json.dumps(data, indent=2, ensure_ascii=False))
print(f" ✓ {mpath.name}")
def main():
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument(
"--skip-existing", action="store_true",
help="Don't re-render files that already exist on disk.",
)
parser.add_argument(
"--only", choices=["cloning", "design", "manifest"],
help="Render only a subset (default: all).",
)
args = parser.parse_args()
print("Loading OmniVoice engine (this can take 30-60 s on first run)…")
try:
import asyncio
from services.model_manager import get_model
try:
asyncio.get_running_loop()
raise RuntimeError("Run this script outside an async context.")
except RuntimeError:
model = asyncio.run(get_model())
except Exception as e:
print(f"\nERROR: Could not load OmniVoice engine: {e}\n")
print("Check that:")
print(" 1. You're running inside the project venv (uv sync first).")
print(" 2. The omnivoice package is importable: `python -c 'import omnivoice'`.")
print(" 3. Model weights are downloaded (~5 GB on first synthesis).")
sys.exit(1)
print("Engine loaded.\n")
if args.only in (None, "cloning"):
render_cloning(model, args)
if args.only in (None, "design"):
render_voice_design(model, args)
if args.only in (None, "manifest"):
update_manifest(args)
print("\nDone. Re-run scripts/build_demos.sh to regenerate dictation samples.")
if __name__ == "__main__":
main()