Files
VoiceStudio/backend/core/personalities.py
T
Palash DebnathandClaude Opus 4.8 8b00dc1f4f feat: onboarding demos, opt-in bug reporting, error-docs deeplinks + issue triage (#133)
* feat: onboarding demos, opt-in bug reporting, error-docs deeplinks + issue triage

Working-tree snapshot bundling several in-flight workstreams (v0.3.0):

- Onboarding/demo system: DemoPresetGrid, DictationDemo, DubbingDemo components
  + tests, render scripts (render_demos_omnivoice.py, build_demos.sh,
  build_dub_demo.sh), personalities preview URLs, alembic 0002 voice-profile
  demo fields.
- Opt-in bug reporting: ReportBugButton (prefilled GitHub-issue URL path).
- Error transparency UX: errorDocsMap deeplinks + BootstrapSplash/error wiring.
- Dub workspace: DubSegmentRow/Table, WaveformTimeline, dubSlice tweaks.
- Issue triage: .planning/issue-clusters/ (plan-01..05 root-cause masters,
  GH #128-#132).
- CLAUDE.md: hard rule — everything ships on v0.3.0, no version bumps.

KNOWN GAP (why this is a draft): the generated demo audio assets are NOT in
this tree, and backend/assets/samples/demo_voice.wav is deleted. onboarding.py
guards the missing file (skips seeding the demo profile with a warning), so no
crash — but first-run Launchpad will be empty and /demo_audio/ preview URLs
404 until assets are regenerated via scripts/build_demos.sh. Do not merge
before regenerating + committing the demo assets.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(dub): timing strategies — kill audio compression, add Concise + Stretch Video

Replaces the current audio time-compression default (atempo squeeze to fit
slot) that produced chipmunk/alien output on high-density target languages
like Bengali. Two new user-selectable modes; legacy behaviour kept behind
an explicit "Strict slot" choice.

New `DubRequest.timing_strategy` enum (default "concise"):
  - "concise"        Translator trims text to fit at natural rate; if it
                     still overflows, hard-trim at slot with a fade so we
                     never overlap the next speaker. Surface overflow_s
                     per segment so the user can shorten the text.
  - "stretch_video"  Audio plays at natural 1.0× rate. Backend computes a
                     per-segment new timeline; persists a video_stretch_plan
                     on the job. Mux step (dub_export) builds an ffmpeg
                     trim+setpts+concat filter graph that stretches each
                     segment's video portion to match the natural-rate dub
                     audio. Gaps/pre-roll/tail pass through at 1.0×.
                     Sub burn under stretch_video is skipped in one pass
                     (cues would drift).
  - "strict_slot"    Legacy atempo squeeze. Retained for back-compat.

Director rate-bias side-effect (seg_speed *= bias) now gated on strict_slot
only, so "urgent"/"slow" direction tokens keep their instruct effect in
the new modes without chipmunking.

Per-segment fit_status emitted in the SSE done event:
  {status: "fits" | "overflows" | "video_stretched", overflow_s?, stretch_ratio?}
DubSegmentRow's "Sync: 100%" badge (which was lying — sync_ratio was always
~1.0 because the TTS loop pre-trimmed to slot) is replaced with a truthful
"Fits / Overflows +Ns / Video 1.18×" label.

Frontend:
  - prefsSlice.timingStrategy (persisted, store v3→v4 with safe migrate).
  - DubTab footer Segmented control: "Concise · Stretch Video · Strict slot".
  - useDubWorkflow passes timing_strategy on /dub/generate; consumes fit_status.

Tests: tests/test_dub_timing_strategy.py — 13 cases covering schema
defaults/validation, _build_video_stretch_filter_graph (pre-roll, gap,
tail, empty-plan early return, post-subtitle chain-in), and
_video_stretch_plan_for guards. 30/30 existing dub tests still pass.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(waveform): surface missing source as "Source media missing" instead of code-4 black box

When a project's underlying media file is gone (moved or deleted between
save and reload) the <video> element fires MediaError code 4 and the
companion audio fetch returns HTTP 404 — both were silently warned to
the console while the user stared at an unresponsive black panel and an
empty waveform.

- WaveformTimeline now flips loadError when the video element rejects
  code 3 (decode) or 4 (src not supported), and tracks `sourceMissing`
  separately so the error UI can name the actual problem.
- The audio decode fallback chain catches HTTP 404 specifically and
  treats it as source-missing instead of loading silent empty peaks —
  an empty waveform on a deleted source is more confusing than a clear
  "Re-upload the video to continue" message.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(tray): "Show OmniVoice" reloads when the webview is blank

When the dev Vite server restarts (or the main window is created before
the backend is ready), the webview load fails and the window is left
with `<body></body>` plus a "Could not connect to the server" console
error. Clicking "Show OmniVoice" from the tray menu just re-showed the
broken window — there was no recovery path short of quit+relaunch.

Now the show handler runs a tiny eval after `show()`/`set_focus()` that
calls `location.reload()` only when `document.body.childElementCount === 0`.
A healthy window doesn't blink (body is non-empty); a blank one
self-recovers as soon as the user clicks Show.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(#133): bug-report diagnostics field mapping + drop unused imports

Address PR #133 review:
- ReportBugButton: /system/info exposes `platform` + `device`, not
  `os`/`torch_device`/`gpu` — those reads silently dropped OS/GPU from every
  bug report. Map to the real fields (CodeRabbit). Also remove the dead
  `home` local in stripHome (CodeQL unused-variable).
- DictationDemo: drop unused `Loader` import (CodeQL unused-import).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-29 17:24:25 +05:30

231 lines
9.6 KiB
Python

"""Built-in voice personality presets.
Each personality is a named bundle of TTS-instruct taxonomy tokens (gender,
age, pitch, accent, dialect, style — see ``omnivoice.utils.voice_design``)
that users can pick from a strip in the Voice Design tab.
Why taxonomy tokens and not prose
=================================
OmniVoice's ``model.generate(instruct=...)`` runs every instruct string
through ``_resolve_instruct``, which splits on commas and validates each
item against a fixed vocabulary (e.g. ``male``, ``middle-aged``,
``moderate pitch``, ``british accent``). Anything outside that vocabulary
raises ``ValueError`` and surfaces to the user as a generation failure.
Earlier versions of this file shipped prose like
``"Speak clearly and professionally like a television news presenter"``
which always failed the validator — clicking *any* personality button in
the Design tab made the next Synthesize call crash (issue #89). Each
personality below maps to a comma-separated list of valid taxonomy
tokens so the picked instruct string is accepted by the model.
The ``description`` field is the human-readable explanation kept for
parity with the old field and for any future UI tooltips; the frontend
today only renders ``name`` and ``icon``. ``instruct`` is what flows into
``setInstruct`` on click.
"""
PERSONALITIES = [
{
"id": "narrator",
"name": "Narrator",
# Calm documentary-narrator vibe: settled adult, mid-low pitch.
"instruct": "middle-aged, low pitch",
"description": "Calm, authoritative documentary narrator with measured pacing",
"icon": "📖",
},
{
"id": "casual",
"name": "Casual",
# Relaxed, conversational — younger speaker, neutral pitch.
"instruct": "young adult, moderate pitch",
"description": "Relaxed, conversational tone like talking to a friend",
"icon": "😊",
},
{
"id": "news_anchor",
"name": "News Anchor",
# Clear, professional broadcaster — adult voice with US accent.
"instruct": "middle-aged, moderate pitch, american accent",
"description": "Clear, professional television news presenter",
"icon": "📺",
},
{
"id": "storyteller",
"name": "Storyteller",
# Dramatic bedtime-story flair — British accent reads well here.
"instruct": "middle-aged, moderate pitch, british accent",
"description": "Dramatic flair and engaging pacing like reading a bedtime story",
"icon": "🧙",
},
{
"id": "corporate",
"name": "Corporate",
# Polished business-presentation register.
"instruct": "middle-aged, moderate pitch",
"description": "Polished, professional tone suitable for business presentations",
"icon": "💼",
},
{
"id": "energetic",
"name": "Energetic",
# High-energy podcast host — younger speaker, higher pitch.
"instruct": "young adult, high pitch",
"description": "High energy and enthusiasm like a podcast host",
"icon": "⚡",
},
# ─────────────────────────────────────────────────────────────────────
# Demo presets — render as a 7-card grid in the empty Design tab.
# Each preset bundles taxonomy `attrs` (drives the CATEGORIES sliders),
# a sample `script` (pre-fills the textarea), and a `preview_url`
# pointing at a pre-rendered WAV under /demo_audio.
#
# `is_demo: True` is the marker the frontend filters on; the legacy 6
# entries above stay as chips, the entries below render as full cards.
# WAVs are generated by scripts/build_demos.sh — keep slugs in sync.
# ─────────────────────────────────────────────────────────────────────
{
"id": "audiobook_uk_narrator",
"name": "The Librarian",
"icon": "📚",
"description": "Warm UK audiobook narrator — measured, atmospheric.",
"instruct": "female, middle-aged, low pitch, british accent",
"attrs": {
"Gender": "female", "Age": "middle-aged", "Pitch": "low pitch",
"Style": "Auto", "EnglishAccent": "british accent",
"ChineseDialect": "Auto",
},
"script": (
"The clock tower struck thirteen, and for the first time in her "
"life, Eleanor wondered if she had been counting wrong all along."
),
"preview_url": "/demo_audio/voice_design/demo_voice_design_audiobook_uk_narrator.wav",
"language": "English",
"is_demo": True,
},
{
"id": "us_news_anchor",
"name": "The Anchor",
"icon": "📺",
"description": "Clear American broadcaster — primetime evening news.",
"instruct": "male, middle-aged, moderate pitch, american accent",
"attrs": {
"Gender": "male", "Age": "middle-aged", "Pitch": "moderate pitch",
"Style": "Auto", "EnglishAccent": "american accent",
"ChineseDialect": "Auto",
},
"script": (
"Good evening. Topping our broadcast tonight: scientists at the "
"coastal observatory have confirmed the signal is, in fact, repeating."
),
"preview_url": "/demo_audio/voice_design/demo_voice_design_us_news_anchor.wav",
"language": "English",
"is_demo": True,
},
{
"id": "indian_support_agent",
"name": "The Helpdesk",
"icon": "🎧",
"description": "Patient Indian-English customer-service voice.",
"instruct": "female, young adult, moderate pitch, indian accent",
"attrs": {
"Gender": "female", "Age": "young adult", "Pitch": "moderate pitch",
"Style": "Auto", "EnglishAccent": "indian accent",
"ChineseDialect": "Auto",
},
"script": (
"Thank you for calling OmniVoice support. I can see your account "
"here. Let's get this sorted out together."
),
"preview_url": "/demo_audio/voice_design/demo_voice_design_indian_support_agent.wav",
"language": "English",
"is_demo": True,
},
{
"id": "gravelly_villain",
"name": "Captain Crusty",
"icon": "☠️",
"description": "Gravelly old sailor — cartoon villain energy, original character.",
"instruct": "male, elderly, very low pitch",
"attrs": {
"Gender": "male", "Age": "elderly", "Pitch": "very low pitch",
"Style": "Auto", "EnglishAccent": "Auto",
"ChineseDialect": "Auto",
},
"script": (
"You came a long way for an answer you already had. Sit. The "
"fire is warm, and the truth is not."
),
"preview_url": "/demo_audio/voice_design/demo_voice_design_gravelly_villain.wav",
"language": "English",
"is_demo": True,
},
{
"id": "aussie_podcaster",
"name": "The Podcaster",
"icon": "🎙️",
"description": "Aussie explainer-show host — quick, punchy, friendly.",
"instruct": "female, young adult, high pitch, australian accent",
"attrs": {
"Gender": "female", "Age": "young adult", "Pitch": "high pitch",
"Style": "Auto", "EnglishAccent": "australian accent",
"ChineseDialect": "Auto",
},
"script": (
"Right, so here's the wild bit. Nobody told the engineers the "
"satellite was supposed to be in orbit by Tuesday. Tuesday came and went."
),
"preview_url": "/demo_audio/voice_design/demo_voice_design_aussie_podcaster.wav",
"language": "English",
"is_demo": True,
},
{
"id": "bedtime_storyteller",
"name": "Junior Quacks",
"icon": "🦆",
"description": "Anxious squawky sidekick — cartoon nephew energy, original character.",
"instruct": "young adult, high pitch",
"attrs": {
"Gender": "Auto", "Age": "young adult", "Pitch": "high pitch",
"Style": "Auto", "EnglishAccent": "Auto",
"ChineseDialect": "Auto",
},
"script": (
"Once, in a town where every street was named after a kind of "
"bread, a small fox decided she was going to learn to play the cello."
),
"preview_url": "/demo_audio/voice_design/demo_voice_design_bedtime_storyteller.wav",
"language": "English",
"is_demo": True,
},
{
"id": "mandarin_sichuan",
"name": "The Sichuan Friend",
"icon": "🌶️",
"description": "四川话 — non-English showcase, dialect-aware design.",
"instruct": "female, young adult, moderate pitch, 四川话",
"attrs": {
"Gender": "female", "Age": "young adult", "Pitch": "moderate pitch",
"Style": "Auto", "EnglishAccent": "Auto",
"ChineseDialect": "四川话",
},
"script": "今天天气巴适得很,我们去吃火锅嘛!记得多加点豆芽。",
"preview_url": "/demo_audio/voice_design/demo_voice_design_mandarin_sichuan.wav",
"language": "Chinese",
"is_demo": True,
},
]
def get_personalities():
"""Return the full list of built-in personality presets."""
return PERSONALITIES
def get_personality(personality_id: str):
"""Look up a single personality by ID, or None."""
for p in PERSONALITIES:
if p["id"] == personality_id:
return p
return None