Files
VoiceStudio/scripts/build_dub_demo.sh
T
Palash DebnathandClaude Opus 5 41c098e009 feat(demos): ship the demo audio and video the app already advertises (#1517)
* feat(demos): ship the demo audio and video the app already advertises

Every demo asset in the app was a dead link on anything but a Mac.

`personalities.py` has carried a `preview_url` for each of the seven
voice-design presets since they were added; DictationDemo.jsx posts three
bundled WAVs to /transcribe so the feature can be shown without microphone
permission; the Dub workspace reads a manifest and plays a source video plus
four dubbed languages. None of those files were committed, because the tooling
that renders them (scripts/build_demos.sh, scripts/build_dub_demo.sh) hard-
requires macOS `say` — it even carries a `TODO: add espeak-ng path for Linux
contributors`. So the presets returned 404, the replay buttons did nothing, and
the dubbing demo never loaded.

Rendered with VoiceStudio's own engine, which runs wherever the app does:

- 7 voice-design previews (2.2 MB)
- 3 dictation replay clips (1.1 MB) — verified by transcribing them back:
  the conversational and French clips round-trip exactly
- dubbing demo: source + 4 dubbed videos with subtitles and manifest (9.6 MB)

Tooling fixes this turned up:

- build_dub_demo.sh wrote to backend/assets/demo/dubbing, but main.py mounts
  backend/assets/samples at /demo_audio — so the frontend's
  /demo_audio/demo/dubbing/manifest.json could never have resolved even after
  a successful Mac build. Output moved under the mount.
- `say` is now the fallback rather than the requirement: the new
  scripts/render_dub_demo_audio.py renders the five tracks with the engine and
  the shell script picks them up.
- The five demo paragraphs lived in two files. They are now one JSON both read
  — two copies is one edit away from a video whose subtitles disagree with it.
- render_demos_omnivoice.py peak-normalized, which a single-sample transient
  defeats: the Helpdesk preset landed at -30 dB RMS against -17 dB for its
  neighbours, so the preview row played at wildly different volumes. Now EBU
  R128 at -18 LUFS with a -1.5 dBTP ceiling.
- …and pinning the output rate, because loudnorm resamples to 192 kHz
  internally and writes there unless told otherwise, which turned 2.1 MB of
  previews into 17.5 MB of identical-sounding audio.
- update_manifest() looked for a manifest at a path nothing writes, so it
  always printed "not found" and did nothing.
- Dictation is rendered here now too. It was excluded on the grounds that
  `say` was good enough and engine TTS was overkill — true only on macOS.

tests/test_demo_assets_exist.py resolves every advertised URL against the
directory main.py actually mounts, and checks each dubbing subtitle matches the
script its manifest entry claims. A missing static file is not an import error
and not a failing request; nothing would have caught this otherwise.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(changelog): stamp the demo-asset entries with their PR ref

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(demos): watermark rendered demo audio, and harden the render scripts

Review findings on #1517:

- Greptile P1: the renderers wrote engine output straight to disk, so a
  re-render shipped demo audio with no provenance mark. These clips play
  back to users as VoiceStudio output — they are synthetic audio leaving
  the app like any other, and now go through mark_synthetic (#1169), the
  one chokepoint every producing route uses. It runs on the file AFTER
  loudnorm, since loudnorm re-encodes what it is handed, and says so
  loudly when marking is unavailable rather than committing an unmarked
  asset. The dubbing renderer shares the same helper.
- CodeRabbit: build_dub_demo.sh checked only source.src.wav before
  deciding it could run without macOS `say`, so a Linux or Windows run
  with four of five tracks present reached a missing one, called `say`,
  and left a half-built bundle. It now requires all five.
- CodeRabbit: shutil.move over an existing path delegates to os.rename,
  which raises FileExistsError on Windows — os.replace overwrites
  atomically everywhere.
- CodeRabbit: the preview test discovered presets in a parametrize
  argument, importing app code at collection time and leaving
  core.personalities in sys.modules for later tests. Discovery moved into
  the test body.

CI: the rendered dub bundle's zh/ja subtitles, its manifest and the
script source are dubbing CONTENT, not UI strings — allowlisted in
test_no_hardcoded_cjk.py with that justification.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(demos): a render that cannot be watermarked fails instead of warning

CodeRabbit and Greptile, #1517: mark_synthetic degrades rather than
raising — correct for generation, wrong for a render script, whose whole
job is to produce files a human then commits. A printed warning on a
scrolling console is not a gate, so both scripts exited 0 with unmarked
assets sitting on disk ready to commit. They now raise, with the reason
and the fix; OMNIVOICE_DEMO_ALLOW_UNMARKED=1 stays for a local listen.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* ci: stop a flaky dependency fetch from failing green runs

en-core-web-sm resolves to a direct GitHub release URL, and github.com
intermittently answers `http2 error: refused stream before processing
any application logic`. uv's own three retries all land within the same
few seconds and fail together, so the whole job dies on a dependency
that has nothing to do with the change under test — it cost #1518 and
#1517 an otherwise-green run tonight.

Two changes: back off between whole `uv sync` attempts, which is what
actually clears it, and pass --no-sync to the pytest steps. `uv run`
re-resolves the environment before running, so every test step was a
fresh chance to hit the same fetch even though the install step had
already synced — that is exactly how #1518 failed, in the isolated
backend/tests step, with all 5467 tests already passed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* ci: one retry seam for every uv sync, not just the job that failed last

en-core-web-sm resolves to a direct GitHub *release* URL rather than a
package index, and github.com intermittently answers `http2 error:
refused stream before processing any application logic`. uv's own
retries all land inside the same ~10 seconds and fail together, so a job
dies on a dependency unrelated to the change under test. Tonight that
cost four otherwise-green runs across #1515, #1517 and #1518 — and the
first fix only covered the Tests job, so the next failure simply moved
to Smoke (Linux), which syncs separately.

The fetch is per-job, so the fix has to be per-job: scripts/uv-sync-retry.sh
backs off between whole attempts (15s, 45s, 90s) and every workflow that
syncs now goes through it — ci.yml (tests + the platform matrix),
release.yml, security.yml, evals.yml. It still fails loudly after four
attempts, so a genuinely broken lockfile is not disguised as a flake.

The Tests job also lacked the UV_HTTP_TIMEOUT / UV_HTTP_RETRIES the smoke
matrix has always set, which is part of why it was the one that kept
dying; it has them now.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(ci): pin the Intel-Mac contract by intent, not by command spelling

test_ci_verifies_intel_mac_as_the_documented_remote_only_host asserted
the literal line `run: uv sync --extra pockettts`, so routing every sync
through scripts/uv-sync-retry.sh read as a broken Intel-Mac contract. The
contract it exists to protect is that the pockettts extra installs ONLY
on backend_supported legs — which the regex now pins, while leaving how
the sync is invoked free to change.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* ci: keep every uv run out of the resolver, and bound the retry budget

CodeRabbit, #1517:

- `uv run` re-resolves before running, so the smoke suite, the
  worker-artifact tests, the release test run and the eval run were each
  a fresh chance to hit the flaky direct-URL fetch outside the retry
  loop. All of them pass --no-sync now; the environment is already
  synced by the step that owns the retries. security.yml's
  `uv run --with pip-audit` is deliberately left alone — it layers an
  ephemeral package rather than running the project's own tests.
- The retry count multiplied uv's own budget (UV_HTTP_RETRIES=5 with a
  120 s timeout on the smoke matrix). Three attempts and 60 s of total
  backoff outlast the refusals actually observed while staying well
  inside the jobs' timeout-minutes.
- The Intel-Mac contract test pinned the smoke command literally too, so
  --no-sync tripped it exactly like the sync line did. Same fix: assert
  the contract (smoke runs only on backend_supported legs), not its
  spelling.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-12 23:29:19 +00:00

187 lines
7.6 KiB
Bash
Executable File
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
#!/bin/bash
# Build the synthetic dubbing demo: one English source video + 4 dubbed
# variants. Each video is a 720p H.264 file with the showwaves audio
# visualizer over a dark gradient background — no copyrighted footage,
# no third-party voice IP.
#
# Output layout (under backend/assets/demo/dubbing/):
# source.mp4 English audio + visualizer
# dubbed_es.mp4 Spanish (Mónica)
# dubbed_fr.mp4 French (Thomas)
# dubbed_zh.mp4 Mandarin (Tingting)
# dubbed_ja.mp4 Japanese (Kyoko)
# *.srt transcripts for each
# manifest.json single source of truth the frontend reads
#
# Bundle target: ~5 MB per mp4 × 5 files = ~25 MB. Well under the 45 MB
# cap from the dubbing-demo design spec.
set -e
REPO_ROOT="$(cd "$(dirname "$0")/.." && pwd)"
OUT_DIR="${REPO_ROOT}/backend/assets/samples/demo/dubbing"
mkdir -p "$OUT_DIR"
# Compatibility: must run on macOS default bash 3.2 (no associative arrays,
# no ${var@Q}). We sidestep both by passing scripts as env vars to python3
# below. Just guard the basics.
if ! command -v ffmpeg >/dev/null; then
echo "ERROR: ffmpeg not found." >&2
exit 1
fi
# `say` is macOS-only, which used to make this script macOS-only and left the
# Dub workspace's demo player pointing at files no other platform could build.
# scripts/render_dub_demo_audio.py renders the same five tracks with the app's
# own engine, anywhere; `say` is now the fallback, not the requirement.
HAVE_SAY=0
command -v say >/dev/null && HAVE_SAY=1
# Check EVERY track, not just the source. Checking one let a run start with
# four of five present, reach a missing dubbed track, call `say` — which does
# not exist off macOS — and leave a half-built bundle behind.
if [ "$HAVE_SAY" = 0 ]; then
MISSING=""
for stem in source dubbed_es dubbed_fr dubbed_zh dubbed_ja; do
[ -f "${OUT_DIR}/${stem}.src.wav" ] || MISSING="${MISSING} ${stem}.src.wav"
done
if [ -n "$MISSING" ]; then
echo "ERROR: no macOS 'say', and these pre-rendered tracks are missing:${MISSING}" >&2
echo "Run: python3 scripts/render_dub_demo_audio.py" >&2
exit 1
fi
fi
if ! command -v python3 >/dev/null; then
echo "ERROR: python3 not found — needed to emit manifest.json." >&2
exit 1
fi
# ── Scripts: source + 4 translations ─────────────────────────────────────
# Read from scripts/dub_demo_scripts.json, which render_dub_demo_audio.py reads
# too — two copies of the same five paragraphs is one edit away from a source
# video whose subtitles say something else.
SCRIPTS_JSON="${REPO_ROOT}/scripts/dub_demo_scripts.json"
read_script() {
SCRIPTS_JSON="$SCRIPTS_JSON" CODE="$1" python3 -c "
import json, os, sys
spec = json.load(open(os.environ['SCRIPTS_JSON'], encoding='utf-8'))
code = os.environ['CODE']
entry = spec['source'] if code == 'en' else next(
e for e in spec['dubbed'] if e['code'] == code)
sys.stdout.write(entry['script'])
"
}
EN_SCRIPT="$(read_script en)"
ES_SCRIPT="$(read_script es)"
FR_SCRIPT="$(read_script fr)"
ZH_SCRIPT="$(read_script zh)"
JA_SCRIPT="$(read_script ja)"
# ── render_lang(code, voice, text) ──────────────────────────────────────
# Produces: $OUT_DIR/{source|dubbed_$code}.mp4 + matching .srt
# Uses showwaves to render a styled audio visualizer over a dark backdrop.
# Resolution 1280x720 keeps each output around 4-6 MB at CRF 28.
render_lang() {
local code="$1" voice="$2" text="$3"
local stem
if [ "$code" = "en" ]; then stem="source"; else stem="dubbed_${code}"; fi
local aiff="${OUT_DIR}/${stem}.aiff"
local wav="${OUT_DIR}/${stem}.wav"
local mp4="${OUT_DIR}/${stem}.mp4"
local srt="${OUT_DIR}/${stem}.srt"
local prerendered="${OUT_DIR}/${stem}.src.wav"
AUDIO_SOURCE="$voice"
if [ -f "$prerendered" ]; then
AUDIO_SOURCE="omnivoice"
ffmpeg -y -loglevel error -i "$prerendered" -ar 44100 -ac 1 "$wav"
else
say -v "$voice" -o "$aiff" "$text"
ffmpeg -y -loglevel error -i "$aiff" -ar 44100 -ac 1 "$wav"
rm -f "$aiff"
fi
# Visual: showwaves p2p mode over a dark gradient with a colored line.
# `nullsrc` + `geq` would let us make a static gradient backdrop, but
# simpler is `color` with overlay'd waveform.
ffmpeg -y -loglevel error \
-i "$wav" \
-filter_complex "[0:a]showwaves=s=1280x720:mode=p2p:rate=30:colors=0xf3a5b6[wave]; \
color=c=0x1d2021:s=1280x720:d=60[bg]; \
[bg][wave]overlay=format=auto,format=yuv420p,trim=duration=$(ffprobe -v error -show_entries format=duration -of default=nw=1:nk=1 "$wav")[v]" \
-map "[v]" -map 0:a \
-c:v libx264 -crf 28 -preset medium \
-c:a aac -b:a 96k -ac 1 \
-shortest \
"$mp4"
rm -f "$wav"
# Single-cue SRT — entire utterance as one block, "good enough" for the
# demo player which shows it as a caption strip rather than karaoke.
local dur
dur=$(ffprobe -v error -show_entries format=duration -of default=nw=1:nk=1 "$mp4")
local end_ts
end_ts=$(awk -v d="$dur" 'BEGIN { h=int(d/3600); m=int((d%3600)/60); s=d-int(d/60)*60; printf "%02d:%02d:%06.3f", h, m, s }' | tr '.' ',')
cat > "$srt" <<EOF
1
00:00:00,000 --> ${end_ts}
${text}
EOF
local size
size=$(du -h "$mp4" | awk '{print $1}')
echo " ✓ ${stem}.mp4 ($size, ${AUDIO_SOURCE})"
}
if [ -f "${OUT_DIR}/source.src.wav" ]; then
RENDERED_BY="omnivoice engine + ffmpeg showwaves"
else
RENDERED_BY="macOS say + ffmpeg showwaves"
fi
echo "── Source video (English) ────────────────────────────────"
render_lang en Samantha "$EN_SCRIPT"
echo ""
echo "── Dubbed videos (4 languages) ───────────────────────────"
render_lang es Mónica "$ES_SCRIPT"
render_lang fr Thomas "$FR_SCRIPT"
render_lang zh Tingting "$ZH_SCRIPT"
render_lang ja Kyoko "$JA_SCRIPT"
echo ""
echo "── Manifest ──────────────────────────────────────────────"
# bash 3.2 (default on macOS) lacks ${var@Q}; pass scripts as env vars to
# Python so escaping of unicode + quotes is handled correctly.
OUT_DIR="$OUT_DIR" RENDERED_BY="$RENDERED_BY" \
EN_SCRIPT="$EN_SCRIPT" ES_SCRIPT="$ES_SCRIPT" FR_SCRIPT="$FR_SCRIPT" \
ZH_SCRIPT="$ZH_SCRIPT" JA_SCRIPT="$JA_SCRIPT" \
python3 - <<'PY'
import json, datetime, os
out = os.path.join(os.environ["OUT_DIR"], "manifest.json")
manifest = {
"version": "0.3.0",
"rendered_by": os.environ.get("RENDERED_BY", "macOS say + ffmpeg showwaves"),
"rendered_at": datetime.datetime.now(datetime.timezone.utc).strftime("%Y-%m-%dT%H:%M:%SZ"),
"license": "MIT (synthetic, no third-party IP)",
"source": {
"code": "en", "label": "English",
"video": "source.mp4", "srt": "source.srt",
"script": os.environ["EN_SCRIPT"],
},
"dubbed": [
{"code":"es","label":"Español","video":"dubbed_es.mp4","srt":"dubbed_es.srt","dir":"ltr","script":os.environ["ES_SCRIPT"]},
{"code":"fr","label":"Français","video":"dubbed_fr.mp4","srt":"dubbed_fr.srt","dir":"ltr","script":os.environ["FR_SCRIPT"]},
{"code":"zh","label":"中文","video":"dubbed_zh.mp4","srt":"dubbed_zh.srt","dir":"ltr","script":os.environ["ZH_SCRIPT"]},
{"code":"ja","label":"日本語","video":"dubbed_ja.mp4","srt":"dubbed_ja.srt","dir":"ltr","script":os.environ["JA_SCRIPT"]},
],
}
with open(out, "w", encoding="utf-8") as f:
json.dump(manifest, f, ensure_ascii=False, indent=2)
PY
echo " ✓ manifest.json"
echo ""
du -sh "$OUT_DIR" | awk '{print " Bundle: " $1}'
echo "Done."