Files
VoiceStudio/docs/specs/longform/25-inline-create-voice.md
T
Palash Debnath 5cab8e0149 feat: rename the product to VoiceStudio (previously OmniVoice-Studio)
Renames what users see. The app, the installers, the window title, the
docs and all 21 locales now say VoiceStudio, with "(previously
OmniVoice-Studio)" noted near the title of each doc surface so people
recognise it.

Deliberately NOT renamed, because renaming any of them silently breaks
an existing install — there is no legacy-path fallback anywhere in this
codebase:

  - bundle identifier com.debpalash.omnivoice-studio (MSI UpgradeCode,
    macOS TCC grants, managed venv, WebView localStorage, the
    single-instance lock)
  - data directories OmniVoice / .omnivoice and omnivoice.db
  - the ~150 OMNIVOICE_* environment variables
  - the X-OmniVoice-* HTTP headers (a wire protocol)
  - the published Docker image paths
  - the OmniVoice ENGINE, which is a model name and not this product

tests/test_identity_paths_survive_the_rename.py pins every one of those
so a future well-meaning sweep cannot orphan a user's library.

Linux .deb users install a new package name and should apt remove
omnivoice-studio; that note is in the changelog.
2026-08-07 01:30:58 +05:30

93 KiB
Raw Blame History

Implementation Spec — TASK #25: Inline "Create Voice" in Stories cast + Audiobook narrator

Lens applied (round 1/10): Codebase grounding. Every file/symbol/line reference below was verified against the working tree at feat/stories-shared-render. Corrections from the prior draft are called out inline as [grounding] notes so reviewers can see what changed and why. The biggest five: (1) buildDesignInstruct() returns an object { instruct, unsupported, duplicates }, not a string — the design save path must destructure .instruct; (2) the ui/Dialog shell has no header slot — its title prop renders the header and the method Segmented must live in the body; (3) the realtime refresh event is emitted from backend/api/routers/profiles.py:124 (not events.py:7) and wired in the frontend via useAppData.js:111-117, not a direct import; (4) no test files exist for CloneDesignTab/StoriesEditor/AudiobookTab today, so the extraction is not guarded by existing tests — they must be written from scratch; (5) POST /profiles returns only {id, name, kind} at runtime, so onCreated may rely only on profile.id.

Lens applied (round 2/10): Completeness. Every state, empty path, and failure path the modal and its triggers must handle is now enumerated explicitly — see the new § State machine & edge cases section and the per-anchor edge-case notes. Verified against code: useRecording.js has a too-short guard (<1000 bytes → toast, no ingest) and a denoise-failure fallback (raw-webm ingest with a "denoising unavailable" toast) that the modal inherits for free (useRecording.js:38-41,56-60); voice_profiles has no UNIQUE constraint on name (db.py:39-59), so duplicate-name creation silently succeeds and must be handled as a product decision, not an error; ApiError carries .status and .detail (api/client.ts:76-85), so failure-path toasts can branch on HTTP status (422 vs 503 vs network); ingestRefAudio sets pendingTrimFile and toasts but does not block — the studio keeps the too-long file pending, which the modal must instead reject (no AudioTrimmer in v1). All eight discrete states and every "and then…" are spelled out below.

Lens applied (round 3/10): Project constraints. The spec now explicitly maps every relevant VoiceStudio hard rule to how this feature satisfies it — see the rewritten § Constraints section. Verified-against-code refinements this lens added: (a) i18n fallback is fallbackLng: 'en' (frontend/src/i18n/index.ts:64) and there are 21 locale files under frontend/src/i18n/locales/, so the correct mechanism is "add keys to en.json; the 20 other locales resolve missing keys via the en fallback chain" — not per-t() defaultValue strings; (b) the recording.too_short / recording.loaded_raw and tts_errors.trim_hint keys already exist in en.json (:1859-1860, :1693), so only the genuinely-new createVoice.* namespace is net-new; (c) the only user-input regex the modal's free-text feeds is String(freeText).split(/[,,]/) (voiceInstruct.js:53) — a linear-time character class, not a CodeQL py/polynomial-redos (or JS-ReDoS) hazard, and the fullwidth-comma CJK char is functional text-processing already covered by the allowlist; (d) voiceInstruct.js and constants.js are already in _ALLOWED_FILES of tests/test_no_hardcoded_cjk.py (:57, :64), so the ChineseDialect picker values surfaced by the modal trip no CJK gate; (e) the version is 0.3.6 across all three lockstep files (pyproject.toml, frontend/src-tauri/tauri.conf.json, frontend/src-tauri/Cargo.toml) — no bump; (f) the describe-debounce effect already has a cancelled race guard (CloneDesignTab.jsx:96,114), correcting the round-2 hedge.

Lens applied (round 4/10): API + data shapes. Every wire shape this feature touches is now pinned to the byte: the exact POST /profiles FormData field list with types and required-ness, the literal 2xx/422/503 response bodies ({"id","name","kind"} and {"detail": "..."}), the ApiError field shape thrown by api/client.ts, the POST /clean-audio response (a binary FileResponse + X-Clean-Filename header — not JSON), the POST /design/describe request/response JSON (note matched is a list of objects and the ignored-bucket field is unmatched, not duplicates), the realtime WebSocket event frame ({"type":"profiles","data":{"action":"created","id":"…"}} and the handler-signature mismatch — the frontend handler ignores the payload), the new pure-helper and component function signatures (isRefAudioTooLong, CreateVoiceModal, the three extracted sub-components, MicButton), and the SQLite voice_profiles row shape that backs the response. Corrections this lens added: (1) createProfile is typed Promise<Profile> but the route returns a partial {id,name,kind} — onCreated must treat name/kind as the only other reliable fields and never read ref_audio/created_at/etc. off the create response; (2) /clean-audio returns audio/wav bytes, consumed via res.blob() + res.headers.get("X-Clean-Filename") (useRecording.js:50-52) — it is not an apiJson call; (3) the describe response's matched field is [{category,token,phrase}] and the UI only reads .length (CloneDesignTab.jsx:103), and the dropped-text field is unmatched: string[] — distinct from buildDesignInstruct's unsupported/duplicates buckets, which are a different local computation; (4) the realtime handler is registered as profiles: () => loadProfiles() (useAppData.js:114) — zero-arg, so the {action,id} payload on the wire is intentionally unused and the modal must not depend on reading it.

Lens applied (round 5/10): Test plan. The § Test plan section below is rewritten to be executable, not aspirational: it pins the test layer for every assertion (pure-helper harness vs. fetch-mocked component render vs. store-direct), states the explicit strategy for keeping the suite import-light so it never pulls main+torch/GPU (the front-end suite is pure jsdom and never imports the Python backend; the heavy parts of the modal — createProfile, cleanAudio, /design/describe — are mocked at the lowest seam: global.fetch returning {ok,status,json,text}, exactly as ApiKeysPanel.test.jsx:32-52 does, so no real network and no apiPost/apiFetch rewrite is needed), names the exact mocking seams verified in the existing harness (global.fetch; vi.mock('react-hot-toast', …) per EngineCompatibilityMatrix.test.jsx:7-10; import '../i18n' is auto-loaded by setup.js:6; localStorage is stubbed by setup.js:8-28), and the concrete CI gates with file:line (ci.yml:107 bunx vitest run; ci.yml:67 uv run pytest tests/ which transitively runs tests/test_no_hardcoded_cjk.py and tests/probe/test_probe_i18n.py; security.yml:65-108 CodeQL Python+JS/TS). Corrections this lens added: (i) the pure-helper tests (refAudio.test.js) and the store-action assertion (setCharacterVoice) must not render React at all — follow the storiesSlice.test.ts harness pattern (:4-10) and the storyCast.test.js pure-import pattern, which is also the mechanism that keeps the suite off the Python/torch path; (ii) toast assertions go through the vi.mock('react-hot-toast') spy and assert the i18n key or interpolated message passed to toast.error/toast.success, not rendered DOM text (toast renders into a portal the component test doesn't mount); (iii) useRecording is exercised by mocking global.fetch for /clean-audio (binary Blob response with the X-Clean-Filename header) rather than mocking the hook, so the too-short/denoise-fallback branches are covered through the real hook code; (iv) the i18n key-completeness check is not a bespoke script — orphan/coverage findings are advisory (non-gating) per tests/probe/test_probe_i18n.py:1-3, so the hard requirement is only "add keys to en.json" (the fallbackLng:'en' chain makes the other 20 locales resolve them), and no new locale file needs editing to stay green; (v) there is no node:test legacy gate for these files — package.json:16 test:legacy targets ../tests/frontend and is not the vitest path; the only frontend gate is bunx vitest run.

TL;DR

Add a reusable <CreateVoiceModal> that wraps the two existing voice-definition paths (audio-clone with mic/upload + by-design with sliders/describe) into a self-contained Radix Dialog. On save it calls the existing createProfile() API (frontend/src/api/profiles.ts:16), refreshes the profile list, and auto-assigns the new profile to the call site (a Stories cast member, or the Audiobook default-voice field). Surface a "+ New voice" trigger in StoriesEditor's Cast panel rows and in AudiobookTab's default-voice field. The modal reuses the design/clone control logic that today lives only in CloneDesignTab.jsx — extracted into self-contained sub-components so neither the modal nor the main studio carries duplicated logic.

The task names <VoiceSelector> (#22) as the ideal host. <VoiceSelector> does not exist yet (#22 is still pending). This spec is written to ship #25 standalone (wiring the modal directly into the two raw <select> sites), with an explicit "if #22 lands first" integration note so the modal becomes the selector's footer action rather than a separate trigger.

Problem

Both longform surfaces let users pick a voice only from the already-saved profile list:

  • Stories cast (frontend/src/components/StoriesEditor.jsx:550-558): each cast row has a <select> of profiles bound to setCharacterVoice(c.id, e.target.value || null); if the list is empty it shows t('stories.noProfiles') (:571) with no way to create one. [grounding] Note the empty option label is t('stories.defaultVoice') (:556), and the noProfiles hint renders in addition to the row(s) only when profiles.length === 0 (:571) — the cast itself always has at least a narrator row whose delete is locked (:563).
  • Audiobook narrator/default (frontend/src/pages/AudiobookTab.jsx:220-224): a single <select> of profiles bound to the local setDefaultVoice state setter, defaulting to "engine default" (empty value → t('audiobook.engine_default'), :222).

A user who lands in Stories/Audiobook with no profiles (or who wants a new character voice mid-project) must leave to the Voice studio, clone/design a voice, then navigate back and re-find their project. That is a hard wall directly contradicting the project's "first-run that actually works" core value. The clone (mic + upload) and design (sliders + "describe") flows already exist but are locked inside CloneDesignTab.jsx — a 792-line page component (frontend/src/pages/CloneDesignTab.jsx) driven entirely by props from App.jsx/useTTS/useProfiles, not reusable as-is. [grounding] The file is 792 lines (prior draft said 793).

Goal / Non-goals

Goals

  1. A <CreateVoiceModal> component reachable from Stories cast rows and the Audiobook default-voice field.
  2. Two methods inside the modal, matching the studio's defineMethod Segmented control: From audio (mic record + drag/upload reference, with backend denoise) and By design (category sliders + "describe your voice" free-text → attrs).
  3. On save: createProfile(formData) (clone or design kind), then loadProfiles() (the parent's refresh), then auto-assign the returned profile.id to the originating cast member / audiobook default.
  4. Extract the clone/design control bodies out of CloneDesignTab.jsx into shared, prop-driven sub-components so the modal and the studio render the same UI from one source.
  5. Full i18n; cross-platform-identical default behavior (mic permission errors mapped via the existing micErrorMessage helper, frontend/src/utils/micError.js, already used inside useRecording.js:78).
  6. Every failure and empty path produces a clear toast and never silently drops user input or leaves the modal in an indeterminate state — see § State machine & edge cases.

Non-goals

  • Building the <VoiceSelector> (#22) shared dropdown — out of scope; this spec degrades gracefully whether or not #22 lands first.
  • Production-overrides sliders (CFG, t_shift, etc., rendered at CloneDesignTab.jsx:624-665 — those are synthesis params, not profile params; not stored on a profile and not sent in the create FormData).
  • Consent-lock / verified-own-voice flow (separate feature; the verified_own_voice/consent_* fields exist on the Profile type at frontend/src/api/types.ts:119-122 and on the POST /profiles/{id}/consent route, but the modal does not gate on them or send them).
  • Changing the backend POST /profiles contract (it already supports both kinds — see § API / data shapes).
  • Dub CastingView inline-create (the task scopes Stories + Audiobook; CastingView is mentioned only as the existing assignment widget). Leave a follow-up note.
  • Duplicate-name detection / rename-on-collision. [grounding] voice_profiles.name has no UNIQUE constraint (db.py:41 — name TEXT NOT NULL, no index), so the backend will happily create a second profile with an identical name and a fresh id. v1 allows this (matches today's studio behavior, which never dedupes). The modal MUST NOT pre-validate names against the existing list — but the assign step keys on profile.id, never on name, so a collision is harmless to assignment. (Optional soft-warning UX is a documented follow-up.)
  • In-modal trimming of over-15s clips. v1 rejects + toasts; embedding AudioTrimmer is a follow-up (see Risk).
  • No new platform-only behavior. This feature is a default feature (no opt-in toggle), so by the strict cross-platform-parity rule it must behave identically on macOS/Windows/Linux — see § Constraints. There is no macOS-only or Windows-only path introduced here.

Design

New extracted sub-components (shared between modal + studio)

Create frontend/src/components/voice-create/ (does not exist yet) with three pieces, all fully controlled (state owned by the caller via props — the same contract CloneDesignTab already uses via the props it destructures at CloneDesignTab.jsx:25-54):

  1. CloneSourcePicker.jsx — the "From audio" body: the drop-zone <input type=file> + <MicButton> + ref-transcript/style inputs. Lifted from CloneDesignTab.jsx:359-452 (the defineMethod === 'audio' branch), minus the studio-only "using profile" banner (:399-413) and the inline save-as-profile row (:426-451) — the modal owns save. Props: { refAudio, ingestRefAudio, refText, setRefText, instruct, setInstruct, isRecording, isCleaning, recordingTime, startRecording, stopRecording }. [grounding] Note the JSX at :389-395 passes the recording handlers to <MicButton> as onStart={startRecording}/onStop={stopRecording} (MicButton's own prop names are onStart/onStop, see :768) — keep that mapping when extracting. [completeness] This component must render a visibly-distinct "recording in progress" state (isRecording, with recordingTime ticking, driven by useRecording.js:71-73) and a "cleaning/denoising" state (isCleaning, between stopRecording and ingestRefAudio, useRecording.js:44-62). Both states must disable the file-drop input and the method Segmented so the user can't start a second source mid-capture.

  2. DesignControls.jsx — the "By design" body: the describe-voice textarea + "Starting points" presets + identity recipe + category chip/select groups. Lifted from CloneDesignTab.jsx:453-613 (the else branch), minus the inline save row (:587-612). Props: { vdStates, setVdStates, instruct, setInstruct, language, setLanguage, describeText, setDescribeText, ...feedback }. The /design/describe debounce effect (CloneDesignTab.jsx:93-116), onDescribeChange (:83-91), applyPersonality (:125-139), applyPreset (passed in as a prop today, defined in useTTS.js:69-72), applyDemoPreset (:255-261), and onChipKeyDown (:228-235) move here too (they are self-contained); personalities is fetched via the existing useQuery(['personalities']) (CloneDesignTab.jsx:119-123). [grounding] applyDemoPreset and the DemoPresetGrid empty-state (:275-277) are tied to the studio's text/activePersonality state and the prefilled script; for the modal they are not needed (the modal has no script textarea). Keep them in the extracted component but gate the demo grid behind a prop so the modal can suppress it. [completeness] The /design/describe debounce effect fires an async fetch on every keystroke (debounced 450 ms). It has three terminal sub-states the modal inherits: pending (in-flight — show the same spinner/feedback the studio shows at :465-474), success (apply mergeDescribedAttrs(res.attrs) to vdStates, set describeUnmatched/describeMatchedAny feedback), and failure (network/5xx). On failure the component must not wipe the user's manual slider picks — the merge only runs on success; a failed describe leaves vdStates untouched (the catch at :109-112 is a no-op by design) and shows a non-fatal toast/inline hint. [grounding, round-3 correction] The existing effect already guards against a stale response overwriting a newer one on rapid typing: it declares let cancelled = false (:96) and its cleanup returns () => { cancelled = true; clearTimeout(id); } (:114), with the async body short-circuiting on if (cancelled) return; (:100). The extraction must preserve this guard verbatim (do not regress it) — no new abort logic is required.

  3. MicButton.jsx — already exists privately at CloneDesignTab.jsx:768-792; promote to its own file and import from both CloneDesignTab and CloneSourcePicker. Signature is MicButton({ isCleaning, isRecording, recordingTime, onStart, onStop }).

CloneDesignTab.jsx is then refactored to render <CloneSourcePicker .../> and <DesignControls .../> in place of the inlined JSX, threading the same props it already receives. Net behavior unchanged for the studio — this is a pure extraction. [grounding] There are currently NO tests for CloneDesignTab.jsx (confirmed: no CloneDesignTab.test.* exists anywhere under frontend/). The extraction is therefore not guarded by an existing suite — a regression test must be created (see Test plan), and a /verify pass on the studio synth flow is mandatory.

<CreateVoiceModal> (frontend/src/components/CreateVoiceModal.jsx)

A self-contained Dialog (imported from ../ui; see frontend/src/ui/index.js:18) that owns its own local voice-definition state (so it never touches the studio's useTTS/useAppStore generation state):

state: method ('audio'|'design'), refAudio, refText, instruct,
       vdStates (init: all 'Auto'), language ('Auto'), describeText + feedback,
       profileName, saving (bool), describePending (bool)
hooks: useRecording(localIngest)   // reuse the exact hook — mic + denoise + raw-fallback

[completeness] State reset on open/close. When open transitions false→true, the modal resets all local state to initial (method = defaultMethod, refAudio = null, vdStates all 'Auto', profileName = '', saving = false). When it closes (Cancel / ESC / backdrop / post-success), it must also stop any active recording: call stopRecording() if isRecording so the mic stream's tracks are released (useRecording.js:35 releases tracks only on onstop). Closing mid-saving is blocked — see the state machine. Closing mid-isCleaning is allowed but the in-flight apiCleanAudio/ingestRefAudio resolves into a now-unmounted component, so the modal must guard setState after unmount (mounted-ref or AbortController), or simply let useRecording's finally run harmlessly.

[grounding] The ui/Dialog component has no header slot. Its signature is Dialog({ open, onClose, title = null, footer = null, size = 'md', dismissable = true, children }) (frontend/src/ui/Dialog.jsx:20-28). It renders its own <header> containing the title node and a close button only when title || dismissable (:51-66), the children in .ui-dialog__body (:68), and the footer node in .ui-dialog__footer (:69). So:

  • title prop: t('createVoice.title').

  • Body (children), top: a Segmented From audio / By design. Reuse the studio's items/labels verbatim (CloneDesignTab.jsx:348-356, keys clone.define_from_audio / clone.define_by_design). Segmented is exported from ../ui (ui/index.js:22). [completeness] Disable the Segmented while isRecording || isCleaning || saving so a method switch can't strand an in-flight capture or save.

  • Body, middle: <CloneSourcePicker> or <DesignControls> by method, plus a required name Input (exported from ../ui, see ui/index.js:17) and a language selector. [grounding] The language picker is SearchableSelect, imported from ../components/SearchableSelect (frontend/src/components/SearchableSelect.jsx), NOT from ../ui — ui/ exposes a plain Select (from Input.jsx) but the studio uses SearchableSelect with options={ALL_LANGUAGES} / popular={POPULAR_LANGS} (CloneDesignTab.jsx:9, 671-677). Match that for parity.

  • footer prop: Cancel + a primary Create voice button (loading={saving}, disabled until valid). Button is exported from ../ui (ui/index.js:15); it supports loading and block props (used at CloneDesignTab.jsx:732-741). [completeness] The Cancel button must be disabled (or relabeled) while saving is true to enforce the "no close mid-save" rule (Dialog's own close button must also be suppressed — pass dismissable={!saving} so ESC/backdrop/X are all disabled during the request, Dialog.jsx:30,34,38,58). [grounding] When dismissable={false} and title is set, the header still renders (because of the title || dismissable guard at :51) but the X close button disappears (:58) — exactly the desired "locked-during-save" header.

  • size="lg" (Dialog accepts 'sm'|'md'|'lg'|'xl', Dialog.jsx:17).

  • Local ingestRefAudio: reuse the trim guard from useTTS.js:48-59 — probeAudioDuration (from utils/format.js) + CLONE_MAX_SECONDS (= 15, utils/constants.js:48). [grounding] The studio guard at useTTS.js:50-56 sets pendingTrimFile and calls setSelectedProfile(null), then toasts t('tts_errors.trim_hint', { duration, max }) — crucially it does not clear/replace refAudio and leaves the over-long file "pending" for the AudioTrimmer UI. The modal has no selectedProfile and no trimmer. So extract just the duration-probe + threshold check into frontend/src/utils/refAudio.js as a pure helper (signature below) and let each caller decide. [completeness] For modal v1 the local ingestRefAudio must: (a) if !file → set refAudio = null and return (matches useTTS.js:49); (b) probe duration; (c) if tooLong → reject (do NOT set refAudio), toast t('tts_errors.trim_hint', { duration, max }), and leave any previously-good refAudio untouched; (d) if probeAudioDuration returns falsy/NaN (corrupt or unprobeable file) → accept the file (do not block on an unknown duration — backend /clean-audio and POST /profiles will reject a truly invalid clip), matching useTTS.js:51 which only blocks when dur && dur > MAX; (e) otherwise set refAudio = file.

  • Save handler builds the same FormData the existing handlers build (see § API / data shapes for the exact field list) and calls createProfile, then onCreated(profile). See the state machine for the full success/failure branching.

State machine & edge cases (modal)

The modal is a small explicit FSM. Every transition's "and then…" is spelled out.

States: IDLE → (user edits) → READY (valid) ⇄ READY? (invalid) → SAVING → SUCCESS (closes) | ERROR (stays open, returns to READY). Plus orthogonal sub-states on the audio method: RECORDING, CLEANING, and on the design method: DESCRIBING.

  1. IDLE (just opened). Name empty, no source. Create voice disabled. Cancel/ESC/backdrop all dismiss freely (dismissable=true).
  2. Validity (READY vs invalid). Create is enabled iff profileName.trim() is non-empty AND one of:
    • method 'audio': refAudio != null. (Mirrors useProfiles.js:42 need_name_audio.)
    • method 'design': buildDesignInstruct(vdStates, describeText/instruct).instruct is non-empty — i.e. at least one non-Auto category or one valid free-text tag. [completeness] All-Auto sliders + empty/all-unsupported free-text → empty .instruct → Create disabled, with an inline hint (t('createVoice.design_needs_attr')) explaining why, because the backend returns 422 {"detail": "design profiles require instruct"} (profiles.py:75-76) and we must never let the user click into a guaranteed 422.
  3. Switching method (audio↔design). Allowed only when not RECORDING/CLEANING/SAVING. Switching does not clear the other method's state (so a user who fills audio, peeks at design, and switches back keeps their clip). Validity is re-evaluated for the now-active method.
  4. RECORDING. Mic granted → MicButton shows stop + recordingTime. Segmented, file-drop, and Create are disabled. Stop → CLEANING.
    • Mic permission denied / no device / device busy: useRecording catches and toasts micErrorMessage(t, e) (useRecording.js:75-79); state never enters RECORDING; modal stays in IDLE/READY. No crash. [grounding] micError.js maps the DOMException.name to an i18n key and, for the permission case, attaches a per-OS hint (capture.mic_hint_mac|_windows|_linux, chosen by detectPlatform()) — this is implementation-level OS detection, not a divergence in user-visible default behavior (every platform gets the same "here's how to re-enable mic access" UX, just with the correct path). See § Constraints.
    • Recording too short (<1000 bytes): useRecording.js:38-41 toasts recording.too_short and does not ingest — refAudio stays as it was. User can re-record.
  5. CLEANING. Blob sent to POST /clean-audio. On success → ingestRefAudio(cleanFile) (then trim guard, step §Design ingestRefAudio).
    • Denoise failure (backend down / 5xx / network): useRecording.js:56-60 falls back to the raw webm file and ingests it, toasting recording.loaded_raw ("denoising unavailable"). The modal treats this as a valid refAudio (a clone profile can be built from raw audio). No error state.
    • Cleaned clip > 15s: the local trim guard rejects it (toast trim_hint), so refAudio is not set even though recording "succeeded." User must re-record shorter. (Acceptable v1 — documented.)
  6. DESCRIBING (design method). Debounced POST /design/describe in-flight. Show pending feedback (CloneDesignTab.jsx:465-474). Success → merge attrs (mergeDescribedAttrs(res.attrs)). Failure → leave vdStates untouched, non-fatal toast. Create remains usable from the manual sliders regardless of describe state.
  7. SAVING. createProfile(formData) in flight. Create shows loading, Cancel disabled, dismissable=false. Exactly one of:
    • SUCCESS (2xx): response body is exactly { "id": string, "name": string, "kind": "clone"|"design" } (see § API). Then: (a) await onCreated(profile) — caller does await onProfilesChanged() then assign by profile.id; (b) onClose(); (c) reset local state. The realtime profiles event (profiles.py:124) is a belt-and-suspenders refresh; assignment does not depend on it.
    • ERROR (422): validation rejected server-side (should be unreachable for a correctly-built FormData, but possible if e.g. instruct collapsed to empty due to a tag/whitelist drift). createProfile rejects with ApiError whose .status === 422 and .detail is the FastAPI detail string (api/client.ts:120-121). Toast e.detail || e.message. Modal stays open, inputs preserved, returns to READY. Do NOT close.
    • ERROR (503, design only): deterministic-sample render failed (no TTS engine ready) — profiles.py:103-107 returns 503 with {"detail": "could not render the design sample: <reason>"}. Toast a friendly message keyed off e.status === 503 (t('createVoice.render_unavailable'), falling back to e.detail). Modal stays open, inputs preserved. User can retry (e.g. after the engine warms up) or switch to the audio method.
    • ERROR (network / 5xx / timeout): apiPost throws ApiError (with .status from the HTTP code, or a raw fetch TypeError on a network drop where .status is undefined). Toast e.message. Modal stays open, inputs preserved.
    • In every ERROR branch: saving resets to false, re-enabling Create/Cancel/dismiss. User input (name, clip, sliders, describe text) is never cleared on failure — only on SUCCESS.
  8. Concurrency guard. A second Create click while saving is true is a no-op (button disabled). A close attempt while saving is true is blocked (dismissable=false). If the user closes during CLEANING/DESCRIBING, the in-flight async resolves into an unmounted tree — guard setState (mounted ref) to avoid React warnings.

Trigger + auto-assign wiring

Stories (StoriesEditor.jsx): in each cast row (:541-570), after the voice <select> (:550-558), add a small icon button (<UserPlus> from lucide-react — already used in frontend/src/pages/VoiceGallery.jsx:11,442) → opens <CreateVoiceModal> with assignTarget = { type:'cast', castId: c.id }. Also add a "+ New voice" affordance to the cast-panel header (:535-540, alongside the existing "+ Add character" button at :537-539) and a "Create one" CTA replacing the bare t('stories.noProfiles') hint (:571). onCreated(profile) → await onProfilesChanged() (passed down — see below) then setCharacterVoice(castId, profile.id). [grounding] setCharacterVoice(castId: string, profileId: string | null) is a real store action (frontend/src/store/storiesSlice.ts:45 declaration, :82-83 impl: set((s) => ({ cast: s.cast.map((c) => (c.id === castId ? { ...c, profileId } : c)) }))) already consumed in StoriesEditor.jsx:121,553.

  • [completeness] Cast row deleted while modal open. A user can delete the originating cast member (removeCastMember/deleteCharacter, StoriesEditor.jsx:562) after opening the modal but before save. On onCreated, setCharacterVoice on a now-missing castId is a harmless no-op against the store (the .map simply finds no matching c.id, storiesSlice.ts:83), but the new profile still lands in the list and is selectable manually. Either accept (simplest) or capture castId at open-time and skip the assign if the row is gone — document the chosen behavior. The narrator row cannot be deleted (:563), so the common case is safe.
  • [completeness] Header / empty-state trigger has no castId (it's a "create + add to library" action, not "assign to row X"). For these, onCreated should await onProfilesChanged() only (no assign) — the new voice appears in every row's <select> for manual pick. The empty-state CTA replaces the dead-end noProfiles hint with an actionable button.

Audiobook (AudiobookTab.jsx): next to the default-voice <select> (:218-225), add a "+ New voice" button (Plus icon already imported, AudiobookTab.jsx:3) → opens the modal with assignTarget = { type:'audiobookDefault' }. onCreated → await onProfilesChanged() then setDefaultVoice(profile.id). [grounding] setDefaultVoice here is the component-local useState setter (AudiobookTab.jsx:22), NOT a store action — there is no audiobook slice. Calling it just updates the local defaultVoice string bound to the <select> at :220.

Profile refresh plumbing. Today profiles is passed into both components (App.jsx:1074 <StoriesEditor profiles={profiles} />, :1080 <AudiobookTab profiles={profiles} />) but loadProfiles is not. Two options (pick the lighter):

  • (A) Pass onProfilesChanged={loadProfiles} from App.jsx into StoriesEditor/AudiobookTab. Preferred — explicit, matches the existing profiles={profiles} prop style. [grounding] loadProfiles is in scope in App.jsx: it is destructured from useAppData() at App.jsx:214 (defined in frontend/src/hooks/useAppData.js:105 as const loadProfiles = useCallback(async () => { try { setProfiles(await listProfiles()); } catch (e) {} }, []) — signature: () => Promise<void>, swallows its own errors, returned at useAppData.js:192). Prior draft cited "returned :229/used :266" — :229 is actually the useProfiles({...}) destructure call and :266 is one of several existing loadProfiles() call sites; the authoritative scope binding is App.jsx:214.
  • (B) Have <CreateVoiceModal> import useAppData's loadProfiles. Rejected — loadProfiles lives in useAppData/App scope, not the store; threading the prop is cleaner.

[completeness] onProfilesChanged failure. loadProfiles() swallows its own errors (useAppData.js:105 empty catch), so it never rejects — the await onProfilesChanged() in onCreated resolves even if the backend is momentarily unreachable (the list just stays stale until the realtime profiles event reconciles it on socket reconnect, profiles.py:124). The onCreated handler should still await onProfilesChanged() and then assign; the assign uses the valid profile.id regardless, so a stale list shows a transient value-not-in-options until the refresh lands. Do not block the assign on the refresh's content.

After createProfile, the backend emits a profiles realtime event. [grounding] The emit is event_bus.emit("profiles", {"action": "created", "id": profile_id}) at backend/api/routers/profiles.py:124 (event_bus.emit defined in backend/core/event_bus.py:44). It is delivered over the WebSocket at backend/api/routers/events.py (/ws/events) as a frame {"type": "profiles", "data": {"action": "created", "id": "<id>"}} and handled on the frontend by useRealtimeEvents({ profiles: () => loadProfiles(), ... }) in frontend/src/hooks/useAppData.js:111-117. [grounding, round-4] The registered handler is zero-arg (profiles: () => loadProfiles(), :114) — the {action, id} payload is intentionally discarded; the modal must NOT depend on reading the realtime payload. So the list eventually updates even without an explicit refresh — but we still await onProfilesChanged() before assigning so the new id is present in profiles when the cast/audiobook <select> re-renders (avoids a flash of "value not in options").

If #22 (<VoiceSelector>) lands first

Make <CreateVoiceModal> self-contained and trigger-agnostic (it takes open, onClose, onCreated, defaultMethod?). Then #22's selector simply renders a "+ Create new voice…" item in its dropdown that sets open=true; the cast/audiobook call sites pass their existing assignment callback as onCreated. No modal rework needed — only the trigger moves.

API / data shapes

No backend changes. Every endpoint below already exists; this section pins their exact wire shapes so the modal can be implemented without guessing.

POST /profiles — create a voice profile (backend/api/routers/profiles.py:41-125)

Content-Type: multipart/form-data (built as a FormData object; apiPost sets the body to the FormData and lets the browser set the boundary — no manual Content-Type header, see api/client.ts:137-138).

Handler signature (profiles.py:42-52):

async def create_profile(
    name:       str            = Form(...),        # REQUIRED
    ref_audio:  Optional[UploadFile] = File(None),
    ref_text:   str            = Form(""),
    instruct:   str            = Form(""),
    language:   str            = Form("Auto"),
    seed:       Optional[int]  = Form(None),
    personality:str            = Form(""),
    kind:       str            = Form("clone"),
    vd_states:  Optional[str]  = Form(None),       # JSON string
)

Clone-method FormData (mirrors useProfiles.js:43-50 exactly):

field type required value the modal sends
name string yes (.trim() non-empty) profileName
ref_audio Blob yes (kind=clone) reconstructed: new Blob([await refAudio.arrayBuffer()], { type: refAudio.type }), appended as ("ref_audio", safeBlob, refAudio.name || "profile.wav") (useProfiles.js:45-47)
ref_text string no optional transcript (default "")
instruct string no optional style string (default "")
language string no the SearchableSelect value (default "Auto")
kind string no "clone" (may be omitted — backend defaults to "clone")

Design-method FormData (mirrors useProfiles.js:88-93 exactly):

field type required value the modal sends
name string yes profileName
kind string yes the literal string "design"
vd_states string yes (must JSON-parse to an object) JSON.stringify(vdStates || {}) — a { [category]: value } map, e.g. {"Gender":"female","Age":"elderly","Pitch":"Auto",...}
instruct string yes (non-empty) buildDesignInstruct(vdStates, freeInstruct).instruct — the .instruct field of the object, a comma-joined token string like "female, elderly, low pitch"
language string no "Auto" default

[grounding, round-4] CRITICAL — the design instruct is a string, not the whole object. buildDesignInstruct returns { instruct: string, unsupported: string[], duplicates: string[] } (voiceInstruct.js:66). Append .instruct. Reference correct usage at useTTS.js:121 (const { instruct: finalInstruct, unsupported, duplicates } = buildDesignInstruct(...)). Note that the studio's save path at CloneDesignTab.jsx:606 passes the whole object into handleSaveDesignProfile, which would serialize as [object Object] — that is a latent pre-existing bug in the studio save row; the modal must NOT replicate it. Do not "fix" the studio bug as part of this task unless the owner asks (out of scope); just avoid it in the modal.

The modal never sends seed or personality. seed is Form(None) → backend defaults design profiles to _DESIGN_SEED = 42 (profiles.py:38,108); the modal omits it. personality is empty by default. Production-override synthesis params (cfg/t_shift/etc.) are explicitly not profile fields and are never appended.

Validation order (each a discrete 422 the modal must avoid or surface) (profiles.py:61-76):

  1. kind not in ("clone","design") → 422 {"detail": "kind must be 'clone' or 'design'"}. (Modal only ever sends one of the two; unreachable.)
  2. kind=="clone" and ref_audio is None → 422 {"detail": "clone profiles require ref_audio"}. (Modal disables Create until refAudio set; unreachable unless the file is dropped between validity-check and submit.)
  3. kind=="design" and vd_states empty/whitespace → 422 {"detail": "design profiles require vd_states"}.
  4. kind=="design" and vd_states does not json.loads to a dict → 422 {"detail": "vd_states must be a JSON object"}. (Modal always sends JSON.stringify(obj), so unreachable.)
  5. kind=="design" and instruct.strip() empty → 422 {"detail": "design profiles require instruct"}. (Modal disables Create when .instruct is empty; the FSM §2 hint guards this.)

Success response — 200 (profiles.py:125, the literal body):

{ "id": "a1b2c3d4", "name": "Grandma Vera", "kind": "design" }
  • id is an 8-char hex slug (str(uuid.uuid4())[:8], profiles.py:78), not a full UUID.
  • [grounding, round-4] The route returns only these three keys even though createProfile() is typed Promise<Profile> (profiles.ts:16) and the full Profile interface (frontend/src/api/types.ts:109-123) declares language_code, ref_audio, ref_text, description, created_at, is_locked, verified_own_voice, consent_text, consent_recorded_at. onCreated must read only profile.id (always present) and may read name/kind; it must NOT read any other Profile field off the create response (they are undefined). The full record is fetched separately via GET /profiles/{id} or arrives in the refreshed list from GET /profiles.

Design-render-failure response — 503 (profiles.py:103-107):

{ "detail": "could not render the design sample: <python exception text>" }

Raised when _render_archetype_wav throws (no TTS engine ready). Surface via e.status === 503.

DB-insert-failure (profiles.py:119-123): if the INSERT raises, the handler removes the orphaned audio file and re-raises the original exception → FastAPI turns it into a 500 with a generic body. Surfaces to the modal as the generic network/5xx ERROR branch. Rare.

Realtime side-effect (profiles.py:124): on success, event_bus.emit("profiles", {"action": "created", "id": profile_id}). Wire frame: {"type": "profiles", "data": {"action": "created", "id": "<id>"}}. Frontend handler ignores the payload (zero-arg () => loadProfiles()).

POST /clean-audio — mic-recording denoise (backend/api/routers/system.py:912-972)

Used only inside useRecording.js:48; the modal inherits it via the hook and never calls it directly.

  • Request: multipart/form-data, one field audio (the raw webm Blob, appended as ("audio", blob, "recording.webm"), useRecording.js:46-47).
  • Response: NOT JSON. It is a binary FileResponse (media_type="audio/wav") with header X-Clean-Filename: mic_<id>.wav (system.py:971-972). The frontend wrapper cleanAudio(formData) returns the raw Response (api/system.ts:95-98) precisely so the caller can read await res.blob() + res.headers.get("X-Clean-Filename") (useRecording.js:50-52). Do not apiJson this endpoint.
  • Failure: any non-2xx (or network drop) makes useRecording's try throw → its catch falls back to ingesting the raw webm with the recording.loaded_raw toast (useRecording.js:56-60). No JSON error shape matters to the modal.

POST /design/describe — free-text → design attrs (backend/api/routers/describe_voice.py:22-35)

Used inside the DesignControls describe-debounce effect (extracted from CloneDesignTab.jsx:99); the modal inherits it.

  • Request (JSON): { "description": string } (max_length=2000, describe_voice.py:18-19). Sent via apiPost('/design/describe', { description: q }) (CloneDesignTab.jsx:99) → Content-Type: application/json.
  • Response (JSON, exact shape, describe_voice.py:26-33):
    {
      "attrs":     { "Gender": "female", "Age": "elderly", "Pitch": "Auto", "Style": "Auto", "EnglishAccent": "british", "ChineseDialect": "Auto" },
      "instruct":  "female, elderly, british accent",
      "matched":   [ { "category": "Age", "token": "elderly", "phrase": "elderly" } ],
      "unmatched": [ "slightly raspy" ]
    }
    
    • [grounding, round-4] attrs is a complete { [category]: token|"Auto" } map over the six categories. The effect feeds it through mergeDescribedAttrs(res.attrs) (voiceInstruct.js:82-89), which re-validates each token against CATEGORIES and resets anything unrecognized to "Auto" — so a drifted/older backend can never inject an out-of-taxonomy value.
    • [grounding, round-4] matched is a list of objects {category, token, phrase}; the UI reads only (res.matched || []).length > 0 to set describeMatchedAny (CloneDesignTab.jsx:103). Do not treat matched as a string array.
    • [grounding, round-4] The dropped-text field is unmatched: string[] (CloneDesignTab.jsx:102 → setDescribeUnmatched(res.unmatched || [])). This is the describe endpoint's "I couldn't map these words" bucket — distinct from buildDesignInstruct's unsupported/duplicates buckets (which are a separate client-side computation over the manual free-text + sliders at save time). Keep the two feedback sources clearly separated in DesignControls.
  • Failure: the debounce effect's try/catch swallows any error (CloneDesignTab.jsx:109-112) — controls untouched, next keystroke retries. No error shape surfaces. (This endpoint is pure-CPU/stdlib, so a failure here is effectively only a network drop.)

ApiError — the rejection shape every failed call throws (frontend/src/api/client.ts:76-85,114-122)

class ApiError extends Error {
  name = 'ApiError';
  status?: number;   // the HTTP status (422, 503, 500, …); undefined on a raw fetch network failure
  detail?: unknown;  // the parsed `detail`/`error` field from the body, else the raw text (readError, :92-100)
  // .message is `"<status> <statusText>: <detail>"` (:121)
}

createProfile/apiPost/apiFetch reject with this on any non-2xx. Failure-path toasts branch on e.status (422 / 503 / other) and may show e.detail (a string for FastAPI HTTPException) or fall back to e.message. Note: a true network drop (server gone) throws a fetch TypeError, not ApiError, so e.status is undefined there — the generic branch must handle both (e?.status optional-chained).

[test seam, round-5] How tests reproduce each ApiError branch. ApiError is a plain exported class (client.ts:76) constructed as new ApiError(message, { status, detail }). The lowest mockable seam is global.fetch (everything funnels through apiFetch → fetch(apiUrl(path), …), client.ts:114). To make createProfile reject with a given status, a test mocks global.fetch to resolve { ok:false, status:422, text: async () => JSON.stringify({detail:'design profiles require instruct'}) } — readError (client.ts:92-100) parses .detail out and apiFetch throws the ApiError with the right .status/.detail (client.ts:120-121). For a network drop, mock global.fetch to reject with a TypeError — that propagates raw (no .status), exercising the generic branch. This means the modal's 422/503/network toast branches are tested through the real client.ts code, not by hand-throwing — closer to production and guards against readError/apiFetch regressions.

voice_profiles table — the row that backs the response (backend/core/db.py:39-59)

CREATE TABLE IF NOT EXISTS voice_profiles (
    id TEXT PRIMARY KEY,            -- the 8-char slug returned as "id"
    name TEXT NOT NULL,             -- NO UNIQUE constraint → duplicate names allowed
    ref_audio_path TEXT,            -- "<id>.<ext>" (clone) or "<id>.wav" (design sample)
    ref_text TEXT DEFAULT '',
    instruct TEXT DEFAULT '',
    language TEXT DEFAULT 'Auto',
    locked_audio_path TEXT DEFAULT '',
    seed INTEGER DEFAULT NULL,      -- 42 for design (the deterministic sample), NULL/clone otherwise
    is_locked INTEGER DEFAULT 0,
    personality TEXT DEFAULT '',
    description TEXT DEFAULT '',
    is_demo INTEGER DEFAULT 0,
    verified_own_voice INTEGER DEFAULT 0,
    consent_text TEXT DEFAULT '',
    consent_audio_path TEXT DEFAULT '',
    consent_recorded_at REAL DEFAULT NULL,
    kind TEXT DEFAULT 'clone',      -- 'clone' | 'design'
    vd_states TEXT DEFAULT NULL,    -- the design JSON string, verbatim
    created_at REAL
);

[grounding, round-4] No column is added or altered → no alembic migration, no localStorage/store schema change → existing omnivoice_data/ is untouched (see § Constraints). id is the PK and the only assignment key; name has no UNIQUE index, so the modal can never produce a name-collision error (only an optional, follow-up soft-warning).

Function signatures introduced by this task

// frontend/src/utils/refAudio.ts (or .js) — pure, no React.
// Returns whether `file` exceeds CLONE_MAX_SECONDS; `tooLong:false` for an
// unprobeable/NaN duration (accept rather than block — mirrors useTTS.js:51).
export async function isRefAudioTooLong(
  file: File,
): Promise<{ tooLong: boolean; duration: number }>;
// CLONE_MAX_SECONDS = 15 (utils/constants.js:48); duration via probeAudioDuration (utils/format.js).
// frontend/src/components/CreateVoiceModal.jsx
type ProfileKind = 'clone' | 'design';
interface CreateVoiceModalProps {
  open: boolean;
  onClose: () => void;
  // Receives ONLY the create-route's partial response — read id/name/kind only.
  onCreated: (profile: { id: string; name: string; kind: ProfileKind })
    => void | Promise<void>;   // caller does (await) onProfilesChanged + assign
  defaultMethod?: 'audio' | 'design';   // default 'audio'
}
// frontend/src/components/voice-create/CloneSourcePicker.jsx
function CloneSourcePicker({
  refAudio, ingestRefAudio, refText, setRefText, instruct, setInstruct,
  isRecording, isCleaning, recordingTime, startRecording, stopRecording,
});
// frontend/src/components/voice-create/DesignControls.jsx
function DesignControls({
  vdStates, setVdStates, instruct, setInstruct, language, setLanguage,
  describeText, setDescribeText,
  describeUnmatched, describeMatchedAny,   // describe-feedback state
  showDemoGrid = true,                     // modal passes false (no script state)
});
// frontend/src/components/voice-create/MicButton.jsx (promoted from CloneDesignTab:768)
function MicButton({ isCleaning, isRecording, recordingTime, onStart, onStop });

Assign callbacks the triggers wire as onCreated:

  • Stories cast row: async (p) => { await onProfilesChanged(); setCharacterVoice(castId, p.id); } (setCharacterVoice: (castId: string, profileId: string|null) => void, store action).
  • Stories header / empty-state: async () => { await onProfilesChanged(); } (refresh only, no assign).
  • Audiobook default: async (p) => { await onProfilesChanged(); setDefaultVoice(p.id); } (setDefaultVoice = component-local useState setter, AudiobookTab.jsx:22).

[completeness] Design-method describe + feedback surfacing at save time. When buildDesignInstruct(...)'s unsupported.length or duplicates.length is non-empty (free-text tags dropped/collided), the modal should surface the same non-fatal tts_errors.ignored_unsupported / tts_errors.ignored_duplicate toasts the studio shows at synth time (useTTS.js:122-127) — but at save time, so the user understands why a tag they typed isn't reflected. This is advisory, not blocking, and is separate from the describe endpoint's unmatched feedback (§ /design/describe).

Integration points (file:line)

[grounding] All anchors below re-verified against the working tree.

  • frontend/src/pages/CloneDesignTab.jsx:359-452 — clone-source JSX (defineMethod === 'audio') to extract into CloneSourcePicker.jsx (drop the studio-only banner :399-413 and save row :426-451). [completeness] Preserve the recording-in-progress and cleaning states (driven by isRecording/isCleaning/recordingTime).
  • frontend/src/pages/CloneDesignTab.jsx:453-613 — design-controls JSX (else branch) → DesignControls.jsx. Co-located logic to move: describe effect :93-116 (including its cancelled race guard at :96,100,114 — preserve verbatim; the effect sets vdStates/describeUnmatched/describeMatchedAny/activePersonality/instruct on success per :101-108); onDescribeChange :83-91; applyPersonality :125-139; applyDemoPreset :255-261; onChipKeyDown :228-235; personalities useQuery :119-123. (applyPreset/insertTag are props from useTTS.js:61-72.) [completeness] Keep the describe-feedback render (:465-474) so the modal shows pending/success/failure of the describe call; gate DemoPresetGrid (:275-277) behind a prop (modal hides it — no script state).
  • frontend/src/pages/CloneDesignTab.jsx:768-792 — MicButton({ isCleaning, isRecording, recordingTime, onStart, onStop }) to promote into its own file.
  • frontend/src/pages/CloneDesignTab.jsx:348-356 — the Segmented items/labels (clone.define_from_audio / clone.define_by_design) reused by the modal header.
  • frontend/src/components/StoriesEditor.jsx:535-540 — cast-panel header (add "+ New voice" affordance). :541-570 — cast rows. :550-558 — cast-row <select> (empty option = t('stories.defaultVoice'), :556). :563 — narrator delete is locked; other rows can be deleted while the modal is open (handle the stale-castId assign). :571 — t('stories.noProfiles') hint (renders only when profiles.length === 0) to replace with a CTA. Assign via setCharacterVoice (store action, imported :121).
  • frontend/src/pages/AudiobookTab.jsx:218-225 — default-voice field; add trigger. :220-224 — the <select> (empty option = t('audiobook.engine_default'), :222). Assign via local setDefaultVoice (:22) — not persisted across nav (see Risk). Plus icon already imported (:3).
  • frontend/src/App.jsx:1074 (<StoriesEditor profiles={profiles} />) and :1080 (<AudiobookTab profiles={profiles} />) — add onProfilesChanged={loadProfiles}. loadProfiles (() => Promise<void>, error-swallowing) destructured at App.jsx:214 from useAppData().
  • frontend/src/hooks/useProfiles.js:41-57 — handleSaveProfile(refAudio, refText, instruct, language) (clone FormData reference shape: name guard at :42 toasts profiles.need_name_audio when !profileName.trim() || !refAudio; append block :43-50, Blob reconstruction via arrayBuffer() at :45-47; await createProfile(formData) then await loadProfiles() at :52,55; on error toasts e.message at :56). frontend/src/hooks/useProfiles.js:86-101 — handleSaveDesignProfile(vdStates, instruct, language) (design FormData; name guard at :87; append block :88-93; error toast :100). [grounding] The two relevant functions are :41-57 and :86-101 (with handleDeleteProfile/handleSelectProfile in between). [completeness] Both existing handlers await loadProfiles() after a successful create and toast e.message on failure without clearing inputs — the modal mirrors exactly this failure discipline (preserve inputs, toast, stay open).
  • frontend/src/hooks/useRecording.js:12-96 — useRecording(ingestRefAudio) reused verbatim by the modal. Returns { isRecording, isCleaning, recordingTime, startRecording, stopRecording } (:89-95). Internals the modal inherits: mic errors mapped via micErrorMessage(t, e) (:78); too-short guard (blob.size < 1000 → toast recording.too_short, no ingest, :38-41); /clean-audio call + X-Clean-Filename read (:48-52); denoise-failure fallback to raw webm with toast recording.loaded_raw (:56-60); tracks released on onstop (:35). [grounding, round-3] All three i18n keys (recording.too_short :1860, recording.loaded_raw :1859) already exist in en.json; the only net-new namespace is createVoice.*.
  • frontend/src/api/system.ts:95-98 — cleanAudio(formData): Promise<Response> (returns the raw Response so the caller reads blob() + X-Clean-Filename; not apiJson).
  • frontend/src/utils/micError.js — micErrorMessage(t, err) (+ detectPlatform() returning 'mac'|'windows'|'linux'); pure module. The permission case attaches capture.mic_hint_mac|_windows|_linux. OS-detection implementation the parity rule permits; behavior identical across platforms. [test, round-5] This module is already fully unit-tested at frontend/src/utils/micError.test.js (covers describeMicError, micErrorMessage, detectPlatform, micHintKey with a fake t); the modal tests do not re-test mic-error mapping — they only assert that the modal surfaces a toast on the denied path (the content is already covered). See § Constraints.
  • frontend/src/hooks/useTTS.js:48-59 — ingestRefAudio + CLONE_MAX_SECONDS/probeAudioDuration trim guard (logic :50-56) to factor into utils/refAudio.js. [completeness] The studio only blocks when dur && dur > MAX (:51) — an unprobeable file (dur falsy/NaN) is accepted. Mirror that in isRefAudioTooLong.
  • frontend/src/api/profiles.ts:16-18 — createProfile(formData: FormData): Promise<Profile> (typed Profile but the route returns the partial {id,name,kind} — see § API). Rejects with ApiError carrying .status/.detail.
  • frontend/src/utils/voiceInstruct.js:33-67 (buildDesignInstruct(vdStates = {}, freeText = '')) — returns { instruct: string, unsupported: string[], duplicates: string[] } (:66). The design save path uses .instruct. [completeness] When every picked value is 'Auto' and free-text is empty/unsupported, .instruct is the empty string (:66 joins an empty byCategory) — the disable-Create condition in state §2. A dropdown value not in CATEGORIES is silently skipped with a console.warn (:46), so a drifted picker can't inject a 422-bound tag. [grounding, round-3] The free-text split is String(freeText||'').split(/[,,]/) (:53) — a two-char character class (ASCII comma + fullwidth CJK comma), linear-time, no ReDoS surface. frontend/src/utils/voiceInstruct.js:82-89 (mergeDescribedAttrs(attrs = {})) — used by the /design/describe effect; returns a complete { [category]: token|'Auto' } map, resetting unmatched/out-of-taxonomy categories to 'Auto' (:86). [test, round-5] buildDesignInstruct is already covered by frontend/src/utils/voiceInstruct.test.js (six cases incl. the object-shape {instruct, unsupported, duplicates} return, the all-Auto empty-instruct case :47-50, and the full-width-comma normalisation :42-45); the modal tests reuse the behavior but do not re-test the helper — they assert the modal sends .instruct correctly (case d below).
  • frontend/src/utils/constants.js:17-31 (CATEGORIES), :10-15 (TAGS), :33-46 (PRESETS), :48 (CLONE_MAX_SECONDS = 15) — design control source data + trim threshold. [grounding] TAGS is :10-15 (separate from PRESETS). [grounding, round-3] CATEGORIES.ChineseDialect (:27-30) holds CJK picker values; constants.js is already in _ALLOWED_FILES (tests/test_no_hardcoded_cjk.py:64).
  • frontend/src/components/SearchableSelect.jsx — language picker (import from ../components/SearchableSelect). frontend/src/languages.json (ALL_LANGUAGES) + POPULAR_LANGS (utils/constants.js:1-4) feed it.
  • frontend/src/ui/Dialog.jsx:20-74 — modal shell. Signature Dialog({ open, onClose, title=null, footer=null, size='md', dismissable=true, children }). dismissable controls ESC (:34), backdrop (:38), and the close-button presence (:58); pass dismissable={!saving} to lock the modal during a create request. The header renders when title || dismissable (:51); body is children (:68); footer is the footer node (:69). frontend/src/ui/index.js:18 exports Dialog; :15 Button, :17 Input, :22 Segmented.
  • frontend/src/api/client.ts:76-85 — ApiError (.status, .detail, .message); :92-100 — readError (parses detail/error from the body); :114-122 — where non-2xx throws it (this is the seam tests drive via global.fetch); :137-138 — apiPost sends FormData as the body untouched (browser sets multipart boundary).
  • frontend/src/i18n/index.ts:64 — [grounding, round-3] fallbackLng: 'en' is the resolution mechanism for missing keys in the other 20 locales. New createVoice.* keys go in en.json; the other locales resolve them through this fallback. (No per-t() defaultValue strings.)
  • backend/api/routers/profiles.py:41-125 — POST /profiles contract (no change). Validation :61-76; success body :125; 503 :103-107; DB-failure cleanup+re-raise :119-123; realtime emit :124.
  • backend/api/routers/system.py:912-972 — POST /clean-audio; binary FileResponse + X-Clean-Filename header (:971-972).
  • backend/api/routers/describe_voice.py:22-35 — POST /design/describe; request {description} (:18-19), response {attrs, instruct, matched, unmatched} (:26-33).
  • backend/core/db.py:39-59 — voice_profiles schema. name TEXT NOT NULL with no UNIQUE index → duplicate names allowed; id is the PK and the only assignment key. No column added/altered → no alembic migration.

Test plan

Strategy — keep the suite import-light so it never touches main+torch/GPU. This is a front-end-only feature, and the front-end vitest suite runs in jsdom (frontend/vite.config.js:30-36), never importing the Python backend, torch, CUDA/MPS, or any heavy native module. That is the structural reason these tests are safe to run locally (it sidesteps the documented tests/ pytest segfault from importing main+torch/Triton): create_profile is unchanged, so this task adds zero new Python tests — backend behavior is covered by the existing CI pytest job, not by anything new here. Within the front-end suite, the import-light discipline is enforced two ways:

  1. Pure helpers and store actions are tested with no React render at all. refAudio.test.js imports only isRefAudioTooLong (pure async, no JSX) — the storyCast.test.js pattern. The setCharacterVoice auto-assign contract is asserted against the store slice directly via the storiesSlice.test.ts harness (:4-10: a hand-rolled set/get over createStoriesSlice), not by mounting StoriesEditor — so that assertion never pulls in the whole editor tree.
  2. The heavy seams of the modal are mocked at the lowest layer: global.fetch. createProfile, cleanAudio, and /design/describe all funnel through apiFetch → fetch(apiUrl(path), …) (client.ts:114). Tests replace global.fetch with a vi.fn() returning the { ok, status, json, text } shape used by ApiKeysPanel.test.jsx:32-52 (and the Blob-body + X-Clean-Filename-header shape for /clean-audio). This exercises the real client.ts/useRecording code (so ApiError construction, readError, the denoise raw-fallback, and the too-short guard are all covered through production code paths) without any network and without rewriting apiPost. No module under test imports anything outside the front-end source tree.

Harness facts (verified, so test files don't re-discover them): setup.js (referenced at vite.config.js:33) auto-loads ../i18n (setup.js:6) — so import '../i18n' is not needed per-file and t('…') resolves real English strings; it also stubs window.localStorage (setup.js:8-28). Toast is a portal side-effect, so tests must mock it: vi.mock('react-hot-toast', () => ({ default: { error: vi.fn(), success: vi.fn() }, toast: { error: vi.fn(), success: vi.fn() } })) (the exact form in EngineCompatibilityMatrix.test.jsx:7-10) and assert on the arguments passed to toast.error/toast.success (the resolved i18n string), never on rendered DOM toast text. Components that fetch personalities via useQuery must be wrapped in a QueryClientProvider (or the query mocked); design-method tests that don't touch personalities can omit it if the component tolerates an absent provider — prefer wrapping to match production. Fake timers (vi.useFakeTimers()) are used to flush the 450 ms describe debounce.

Unit / component (vitest + RTL, gated by ci.yml:107 bunx vitest run)

frontend/src/utils/refAudio.test.js (new — pure helper, no React):

  • it('reports tooLong:true with the duration for a >15s clip') — mock probeAudioDuration (via vi.mock('./format')) to resolve 20 → { tooLong: true, duration: 20 }.
  • it('reports tooLong:false for a clip at or under 15s') — probeAudioDuration → 15 → { tooLong: false, duration: 15 } (boundary, > MAX not >= MAX, mirrors useTTS.js:51).
  • it('accepts (tooLong:false) an unprobeable clip whose duration is NaN/undefined') — probeAudioDuration → NaN and → undefined both yield { tooLong: false }.

frontend/src/components/CreateVoiceModal.test.jsx (new — fetch-mocked render). Each case beforeEach(() => vi.restoreAllMocks()); mock react-hot-toast once at module top:

  • (a) it('renders both methods and toggles via the Segmented control') — both labels (clone.define_from_audio/clone.define_by_design) render; clicking switches the visible body between CloneSourcePicker and DesignControls.
  • (b) it('disables Create until the form is valid for the active method') — name empty → Create disabled; name set, audio method, no refAudio → disabled; design method all-Auto + empty free-text (so buildDesignInstruct(...).instruct === '') → disabled and the createVoice.design_needs_attr hint is present; design method with one non-Auto category → enabled. (Assert via getByRole('button', { name: /create voice/i }).disabled.)
  • (c) it('builds clone-method FormData and fires onCreated on success') — mock global.fetch to resolve { ok:true, status:200, json: async () => ({ id:'a1b2c3d4', name:'X', kind:'clone' }) }; ingest a File; submit; assert the captured fetch init body instanceof FormData, body.get('name') is the string, body.get('ref_audio') instanceof Blob with the file's type, and body.get('kind') is 'clone' or absent; assert onCreated was called with { id:'a1b2c3d4', name:'X', kind:'clone' } and onClose fired.
  • (d) it('sends design instruct as a non-empty string, never [object Object]') — design method, set Gender→female; submit; assert body.get('kind') === 'design', JSON.parse(body.get('vd_states')) is an object (not array/string), and body.get('instruct') is a non-empty string and expect(body.get('instruct')).not.toBe('[object Object]') (the round-4 latent-bug guard).
  • (e) it('keeps the modal open and preserves inputs on a 422') — global.fetch → { ok:false, status:422, text: async () => JSON.stringify({ detail:'design profiles require instruct' }) }; submit; assert toast.error was called with a string containing the detail, onClose not called, and the name <input> still holds its value.
  • (f) it('shows a friendly render-unavailable toast on a 503 (design)') — global.fetch → { ok:false, status:503, text: async () => JSON.stringify({ detail:'could not render the design sample: no engine' }) }; assert toast.error arg equals the resolved createVoice.render_unavailable string (or contains e.detail), modal stays open.
  • (g) it('shows e.message and stays open on a network drop') — global.fetch set to vi.fn().mockRejectedValue(new TypeError('Failed to fetch')); assert a toast fired, onClose not called (covers the .status===undefined branch).
  • (h) it('closes and resets to IDLE on success') — after a successful create + close, re-open with open toggled and assert Create is disabled again and no refAudio/name carried over.
  • (i) it('blocks close while saving and re-enables after the request resolves') — make global.fetch return a never-yet-resolved promise; while pending, fire ESC / click the backdrop / click Cancel and assert onClose not called and dismissable is false (the X is absent); then resolve and assert dismiss works.
  • (j) it('toasts recording.too_short and leaves refAudio unset on a <1000-byte capture') — drive the real useRecording by mocking MediaRecorder/getUserMedia (jsdom) to emit a <1000-byte Blob on stop; assert toast.error got the recording.too_short string and Create stays disabled (no refAudio).
  • (k) it('falls back to the raw webm and toasts recording.loaded_raw when /clean-audio fails') — global.fetch for /clean-audio → { ok:false, status:500 }; assert the raw file is ingested (Create becomes enabled with a name) and toast got the recording.loaded_raw string.
  • (l) it('rejects an over-15s upload with the trim_hint toast and no refAudio') — vi.mock('../utils/format', …) so probeAudioDuration → 20; ingest a file; assert toast.error got the interpolated tts_errors.trim_hint string and refAudio is unset (Create disabled).
  • (m) it('accepts an unprobeable upload (NaN duration)') — probeAudioDuration → NaN; ingest; assert refAudio is set (Create enabled with a name).
  • (n) it('reads only id/name/kind off the create response') — global.fetch resolves exactly { id:'a1b2c3d4', name:'X', kind:'design' } (no other keys); assert onCreated receives that object verbatim and the test spy verifies the assign callback was invoked with .id; (defensive) assert no thrown access on absent ref_audio/created_at.

frontend/src/components/StoriesEditor.test.jsx (new — fetch-mocked render; setCharacterVoice assertion done separately at the slice level per the strategy note):

  • it('opens CreateVoiceModal from a cast row trigger') — click the row's UserPlus button → modal visible.
  • it('auto-assigns the new profile to the row on onCreated') — simulate onCreated({ id:'p_new', name:'X', kind:'clone' }) and assert the passed onProfilesChanged was awaited and setCharacterVoice (a spy injected as a prop / via a mocked store) was called with (rowId, 'p_new').
  • it('replaces the noProfiles dead-end with an actionable CTA when profiles is empty') — render with profiles={[]} and assert a button (not just the stories.noProfiles text) is present.
  • it('header trigger refreshes but does not assign') — header onCreated calls onProfilesChanged and not setCharacterVoice.
  • it('does not throw when the originating cast row was deleted before onCreated resolves') — delete the row, then fire onCreated; assert no throw (the slice .map no-ops on a missing id).

frontend/src/components/StoriesEditor.cast.test.ts (new — pure store, the import-light assertion): using the storiesSlice.test.ts harness (createStoriesSlice + hand-rolled set/get), it('setCharacterVoice on a missing castId is a no-op') and it('setCharacterVoice(rowId, newId) updates only that row'). This is where the auto-assign contract is really pinned, without mounting any React tree.

frontend/src/pages/AudiobookTab.test.jsx (new — fetch-mocked render):

  • it('opens CreateVoiceModal from the + New voice button').
  • it('sets the default-voice select to the new id on onCreated') — after onCreated({ id:'p_ab', … }), assert the <select> value updates to 'p_ab' (component-local setDefaultVoice).

frontend/src/pages/CloneDesignTab.test.jsx (new — extraction regression, the one that proves the pure refactor didn't drift; no prior suite existed):

  • it('renders CloneSourcePicker for the audio method and DesignControls for design') — toggle the Segmented and assert each body mounts.
  • it('shows the recording-in-progress and cleaning states with distinct UI') — drive isRecording/isCleaning (via the real useRecording with mocked MediaRecorder) and assert the MicButton/cleaning indicator and disabled Segmented.
  • it('applies mergeDescribedAttrs and sets describe feedback on a /design/describe success') — vi.useFakeTimers(); global.fetch for /design/describe → the documented { attrs, instruct, matched, unmatched } shape; type into the describe box, advance 450 ms; assert the merged attrs land in the controls, describeMatchedAny reflects matched.length>0, and describeUnmatched equals unmatched.
  • it('preserves the describe race guard — the later edit wins on out-of-order resolution') — fire two rapid describe edits, resolve the first fetch after the second; assert the second edit's attrs are the ones applied (the cancelled guard at CloneDesignTab.jsx:96,100,114 discards the stale response). This is the single highest-value regression assertion for PR slice 1.
  • it('leaves vdStates untouched and does not crash on a /design/describe failure') — global.fetch rejects; assert manual slider picks survive (the catch no-op).

i18n / CJK (no bespoke gate; covered transitively by ci.yml:67 uv run pytest tests/)

  • No allowlist edit needed. The only CJK the modal surfaces comes from voiceInstruct.js (split delimiter) and constants.js (ChineseDialect picker) — both already in _ALLOWED_FILES of tests/test_no_hardcoded_cjk.py (:57, :64). The new component files (CreateVoiceModal.jsx, voice-create/*.jsx) introduce no hardcoded CJK, so tests/test_no_hardcoded_cjk.py stays green untouched.
  • Add the net-new keys to frontend/src/i18n/locales/en.json only. No createVoice.* namespace exists yet (verified: 0 matches for createVoice/design_needs_attr/render_unavailable). Add at minimum: createVoice.title, createVoice.design_needs_attr, createVoice.render_unavailable, plus a Cancel / Create-voice label pair. Do not edit the other 20 locale files — fallbackLng: 'en' (i18n/index.ts:64) resolves missing keys, so the suite stays green and no defaultValue literals enter JSX. The inherited keys (recording.too_short :1860, recording.loaded_raw :1859, tts_errors.trim_hint :1693, tts_errors.ignored_unsupported :1696, tts_errors.ignored_duplicate :1697) already exist; reused clone.* keys are unchanged.
  • The locale-parity probe is advisory, not gating. tests/probe/test_probe_i18n.py:1-3 states orphan-key + coverage findings are "reported, not gating," so adding en-only keys does not fail CI. (If a future hardening pass turns this gating, run scripts/translate_all.py to backfill the 20 locales — out of scope here.)

Manual / /verify (run the app — covers the GPU/engine paths the unit suite deliberately mocks out)

The unit suite never warms a real TTS engine (design 503 is mocked). The /verify pass is where the actual engine/GPU path is exercised:

  • Stories with zero profiles → cast "+ New voice" → record 3s via mic → denoise → name → Create → cast row auto-selects it; preview the line.
  • Audiobook → "+ New voice" → By design → set Gender/Age controls → name → Create → default-voice selects it; "Preview plan" then a chapter preview uses it.
  • Cancel/ESC/backdrop dismiss (Dialog dismissable default true, Dialog.jsx:18,29-39) leaves profiles + assignments untouched.
  • Mic-denied path shows the per-OS micErrorMessage hint (no crash) on macOS / Windows / Linux — confirm the hint text is the correct per-OS path on each (parity check, see § Constraints).
  • Failure paths: kill the backend mid-create → modal stays open, error toast (e.message), inputs preserved; restart backend → retry succeeds. Design create with no engine warmed → real 503 toast (createVoice.render_unavailable), modal open. Upload a >15s clip → trim_hint toast, clip rejected. Record <1s → too_short toast.
  • No-close-mid-save: spam ESC during a slow create → modal does not close until the request resolves.
  • Studio regression: the main Voice studio (CloneDesignTab) clone + design + describe + personalities + recording/cleaning states all behave identically post-extraction.

CI gates (exact jobs that must be green before merge)

  • bunx vitest run — .github/workflows/ci.yml:107 (the only front-end test gate; package.json:14 "test": "vitest run"). Run locally first per MEMORY's merge-discipline rule. There is no node:test legacy gate for these files — package.json:16 test:legacy targets ../tests/frontend, a different path not exercised by this task.
  • uv run pytest tests/ -q — ci.yml:67. Transitively runs tests/test_no_hardcoded_cjk.py (CJK gate — passes with no allowlist edit) and tests/probe/test_probe_i18n.py (i18n parity — advisory, non-gating). No new Python test is added; create_profile is unchanged.
  • uv run pytest backend/tests/ -q — ci.yml:81 (isolated backend session). Unaffected (no backend change).
  • CodeQL (Python + JS/TS) — .github/workflows/security.yml:65-108. This change adds no new server-side or client-side regex over user input; the one user-input regex the modal feeds is the pre-existing linear-time character class voiceInstruct.js:53 split(/[,,]/) — no py/polynomial-redos / JS-ReDoS surface (see § Constraints). CodeQL stays green.
  • Local-loop note (MEMORY): run bunx vitest run before every push; never merge before all PR checks report green (gh pr checks monitor). The documented local tests/ pytest segfault (importing main+torch/Triton) does not affect this PR — it adds no Python test and the front-end suite is jsdom-only.

Constraints

This is a default feature (no opt-in toggle), so the strict cross-platform-parity rule applies in full. Each relevant VoiceStudio hard rule and how this spec satisfies it:

  • Cross-platform parity (P0, strict rule 2026-05-20). The modal's user-visible behavior is identical on macOS, Windows, and Linux. All capture goes through the already-shipping useRecording/getUserMedia path (useRecording.js:21-80), which is platform-agnostic in the WebView. The only OS-aware code touched is micError.js's detectPlatform() → capture.mic_hint_{mac,windows,linux} — this is permitted implementation code for an OS API (the privacy-settings path differs per OS), and the behavior (an actionable mic-denied hint) is uniform across platforms; it is already unit-tested across all three platforms in micError.test.js:20-24. No macOS-only/Windows-only feature is introduced, so nothing here needs an opt-in toggle. The too-short guard, denoise raw-fallback, trim rejection, and the full save/422/503/network FSM are all pure JS with no platform branches. The /verify matrix above explicitly checks the per-OS mic hint on all three.
  • Local-first guarantee (no cloud/accounts/telemetry). The modal makes only same-origin localhost calls to the existing FastAPI backend: POST /profiles (createProfile, multipart), POST /clean-audio (apiCleanAudio, multipart, used in useRecording.js:48), and the optional POST /design/describe (JSON, already used by the design controls). No third-party endpoint, no account, no API key, no telemetry. The app remains fully functional with no network beyond 127.0.0.1. No HF/GitHub/Sentry calls are added by this feature.
  • Backward-compatible project data (alembic for DB, lazy migration for localStorage). No DB schema change — the feature reuses the existing voice_profiles table (db.py:39-59) and the unchanged POST /profiles contract, so there is no alembic migration to write and existing omnivoice_data/ (voices, projects, settings) keeps working untouched. No persisted client schema change — the modal owns ephemeral local React state only; the Stories assignment writes through the existing setCharacterVoice store action (whose persisted CastMember.profileId shape is unchanged), and the Audiobook assignment is component-local useState (not persisted at all). Therefore no localStorage/store lazy-migration is required, and no version key needs bumping. Duplicate profile names remain accepted (no UNIQUE constraint) — that's existing behavior, not a data-format change.
  • CodeQL py/polynomial-redos (and JS ReDoS) on user-input regex. This feature adds no new regex that consumes user input on either the Python backend (unchanged) or the new frontend components. The modal's design free-text does flow into one existing regex — String(freeText||'').split(/[,,]/) at voiceInstruct.js:53 — but that is a two-element character class (ASCII , + fullwidth CJK ,), which is linear-time and not a polynomial-ReDoS pattern (no nested/overlapping quantifiers, no [^x]* with ambiguous delimiters, no \s*/.+ overlap). It is unchanged by this task. So no CodeQL ReDoS gate is implicated (security.yml:65-108 stays green); if the modal later needs to parse free-text differently, keep delimiters as a character class and avoid backtracking-prone quantifiers per the project's ReDoS guidance.
  • Localization (hard rule — no hardcoded non-English UI text; all UI via i18n t() keys). All new user-facing strings go through t('createVoice.*') keys added to frontend/src/i18n/locales/en.json; the 20 other locale files resolve missing keys via fallbackLng: 'en' (i18n/index.ts:64) until translated. No English (or any) literal user-facing string is placed in component code, and no per-t() defaultValue fallbacks are used (that pattern isn't used in this layer and would smuggle English literals into JSX). The only CJK the feature renders is functional (the ChineseDialect picker values in constants.js:27-30 and the fullwidth-comma split delimiter in voiceInstruct.js:53); both files are already allowlisted in tests/test_no_hardcoded_cjk.py (:64, :57), so the CI CJK gate passes without any allowlist edit. The new CreateVoiceModal.jsx / CloneSourcePicker.jsx / DesignControls.jsx / MicButton.jsx files contribute no hardcoded CJK.
  • Versioning (continuous-to-main patch; no RCs, no minor/major bump). main is 0.3.6 (verified in lockstep across pyproject.toml, frontend/src-tauri/tauri.conf.json, frontend/src-tauri/Cargo.toml). This feature does not touch any version file — it ships continuous-to-main as part of the open v0.3.x line; the owner tags a patch when the state is worth cutting. No -rc tag, no codename, no "defer to v0.4."
  • Docs-sync (hard rule). If any docs/** or README flow describes "how to create a voice" (voice-cloning / voice-design walkthrough), it must be updated in the same PR that adds inline create — a small note that voices can now be created inline from Stories cast rows and the Audiobook default-voice field, without leaving the editor. Stale docs are treated as bugs; the docs delta lands with PR slice 3/4 (whichever first ships a user-visible trigger), not as backlog.
  • GSD workflow. Implement via a GSD command (/gsd-quick for the small slices or /gsd-execute-phase for the planned extraction), not raw edits, so planning artifacts stay in sync.

Dependencies

  • Hard: none new. Reuses @radix-ui/react-dialog (already behind ui/Dialog.jsx), react-hot-toast, lucide-react (UserPlus/Plus), react-i18next, @tanstack/react-query (personalities fetch at CloneDesignTab.jsx:119), and the existing createProfile/cleanAudio APIs.
  • Test-only: @testing-library/react + vitest + jsdom (already the harness — see ApiKeysPanel.test.jsx, EngineCompatibilityMatrix.test.jsx, storiesSlice.test.ts); no new test dependency. A QueryClientProvider wrapper (from the already-installed @tanstack/react-query) is needed in the DesignControls/CloneDesignTab tests that exercise the personalities query.
  • Soft / sequencing: <VoiceSelector> (#22) is the preferred host. Recommend landing #25 standalone now (direct wiring into the two <select> sites) and, if #22 ships, moving only the trigger into the selector (see "If #22 lands first"). This unblocks the user-visible value without waiting on #22.
  • Internal refactor dependency: the CloneDesignTab extraction (PR slice 1) must land before/with the modal so there is a single source of clone/design UI.

Risk

  • Extraction regression in the main studio (medium): moving ~250 lines out of CloneDesignTab.jsx risks subtle prop/state drift (e.g. the describe-debounce effect :93-116 and its cancelled race guard :96,114, and the applyPersonality reset logic :125-139 that fixed issue #114). Mitigation: keep sub-components fully controlled with identical prop contracts; preserve the cancelled guard verbatim; ship extraction as its own reviewed PR slice with the new CloneDesignTab.test.jsx (including the out-of-order describe-response assertion that directly exercises the cancelled guard); do a /verify pass on the studio synth flow, including the recording/cleaning states.
  • Design instruct shape bug (medium): buildDesignInstruct returns an object; passing it raw to FormData yields "[object Object]" → likely a 422 (profiles.py:75) or a garbage profile. Mitigation: always append .instruct; the unit test (case d) asserts the appended instruct is a non-empty string and .not.toBe('[object Object]'). Disable Create when .instruct would be empty (all-Auto + no valid free text), with a hint (case b).
  • Partial create-response misuse (medium): createProfile is typed Promise<Profile> but the route returns only {id,name,kind} (profiles.py:125). Code that reads profile.ref_audio/created_at/etc. off the create response will silently get undefined. Mitigation: onCreated reads only .id (and optionally .name/.kind); unit case (n) asserts the assign uses .id and nothing else; the full record arrives via the onProfilesChanged() refresh / GET /profiles.
  • Design render 503 (low): the deterministic sample render (profiles.py:91-102) can fail if no TTS engine is ready → 503 (:103-107). Mitigation: surface the toast (branch on e.status === 503, unit case f), keep the modal open, don't clear inputs; the real-engine path is checked in /verify.
  • Profile-list staleness (low): assigning before onProfilesChanged() resolves shows a transient empty <select> value. Mitigation: await onProfilesChanged() before calling the assign setter (it never rejects — error-swallowing useAppData.js:105); the realtime profiles event reconciles the list on socket reconnect regardless.
  • Too-long reference clip (low): modal v1 rejects clips over CLONE_MAX_SECONDS (=15s) with a tts_errors.trim_hint toast instead of embedding AudioTrimmer. Mitigation: documented follow-up; consistent with the trim guard messaging in useTTS.js:54; unit cases (l)/(m) cover rejection and the unprobeable-accept boundary. Note: a cleaned recording that comes back >15s is also rejected — the user must re-record shorter.
  • Recording too short / denoise unavailable (low): useRecording already handles both — a <1000-byte blob toasts recording.too_short and skips ingest (:38-41); a denoise failure ingests the raw webm with recording.loaded_raw (:56-60). Mitigation: surface the existing i18n keys (both already present in en.json); unit cases (j)/(k) cover both through the real hook (mocked MediaRecorder/getUserMedia + global.fetch), not by stubbing the hook.
  • Close mid-save / unmount during async (low): closing while createProfile is in flight could lose track of the result or leave the list half-refreshed; closing during isCleaning/describe resolves into an unmounted tree. Mitigation: dismissable={!saving} blocks close during save (unit case i); a mounted-ref guard prevents post-unmount setState; stopRecording() on close releases the mic stream.
  • Duplicate profile names (low): no UNIQUE constraint (db.py:41) → two voices can share a name. Mitigation: assignment keys on id, so this is cosmetic; v1 allows it. Optional soft-warning is a follow-up.
  • Cast row deleted while modal open (low): the originating cast member can be removed before save; setCharacterVoice on a missing id is a no-op (storiesSlice.ts:83 .map). Mitigation: accept (new voice still lands in the library) or capture castId at open and skip the assign if the row is gone — document which; the pure-store test (StoriesEditor.cast.test.ts) pins the no-op and the StoriesEditor render test asserts no throw.
  • Audiobook default is local-only state (low): setDefaultVoice is component-local (AudiobookTab.jsx:22), so the assignment is not persisted across navigation away from the tab. That matches today's behavior (the <select> is already local); no regression, but note it so reviewers don't expect persistence.

PR slices

  1. Refactor: extract clone/design UI — create components/voice-create/{MicButton,CloneSourcePicker,DesignControls}.jsx, refactor CloneDesignTab.jsx to consume them; extract utils/refAudio.js (isRefAudioTooLong, accepts unprobeable durations). No behavior change; preserve the describe-effect cancelled race guard verbatim. Tests: new CloneDesignTab.test.jsx (incl. the out-of-order describe-response race assertion + the documented /design/describe response shape) + refAudio.test.js + the existing voiceInstruct.test.js stays green. Gate: bunx vitest run.
  2. CreateVoiceModal — new modal component owning local state + useRecording, both methods, the full state machine (validity, recording/cleaning/describe sub-states, save success/422/503/network branches, no-close-mid-save), save → createProfile (exact FormData per § API) → onCreated. New createVoice.* i18n keys in en.json (no defaultValue literals). Tests: CreateVoiceModal.test.jsx (new), cases (a)–(n), all via global.fetch mock + react-hot-toast spy.
  3. Wire into Stories — cast-row + header + empty-state triggers; onProfilesChanged prop from App.jsx:1074; auto-assign via setCharacterVoice (row trigger) / refresh-only (header + empty-state). Handle deleted-row + refresh-failure. Docs-sync: update any "create a voice" doc flow in this PR. Tests: StoriesEditor.test.jsx (new) + StoriesEditor.cast.test.ts (pure-store, the import-light setCharacterVoice contract).
  4. Wire into Audiobook — default-voice trigger; onProfilesChanged prop from App.jsx:1080; auto-assign via local setDefaultVoice. Tests: AudiobookTab.test.jsx (new). Docs-sync note in the same PR if not already covered by slice 3.
  5. (optional, only if #22 ships first) — collapse triggers into <VoiceSelector>'s "+ Create new voice…" item.

Acceptance criteria

  • From Stories cast (including the zero-profiles state, where the dead-end stories.noProfiles hint is replaced by an actionable CTA) a user can open a Create Voice modal, define a voice by audio (mic or upload) or by design (sliders/describe), save it, and the originating cast member is auto-assigned the new voice (setCharacterVoice(castId, profile.id)) without leaving the editor.
  • From Audiobook, the same modal creates a voice and auto-sets it as the default voice (local setDefaultVoice(profile.id)).
  • Saving calls createProfile() with the correct kind-specific FormData (clone: name+ref_audio Blob; design: name+kind=design+vd_states JSON-object string+instruct string from buildDesignInstruct(...).instruct, asserted not "[object Object]"); the success body is {id,name,kind} only and onCreated reads no other field; the profile list refreshes via onProfilesChanged and the new voice appears in the relevant <select>.
  • The main Voice studio (CloneDesignTab) renders and behaves identically post-extraction (clone + design + describe + personalities + recording/cleaning states all unchanged, including the describe race guard and the documented /design/describe response handling); the new CloneDesignTab.test.jsx passes (no prior suite existed) and its out-of-order describe-response case proves the cancelled guard survived the extraction.
  • All new UI text is i18n keys under createVoice.* in en.json; no hardcoded user-facing strings; the 20 non-en locales fall back via fallbackLng: 'en'; the CI CJK check (tests/test_no_hardcoded_cjk.py, via ci.yml:67) passes with no allowlist edit (functional CJK lives only in already-allowlisted voiceInstruct.js/constants.js).
  • Default-feature parity: the modal behaves identically on macOS/Windows/Linux; the only OS-aware code is the per-OS mic-denied hint, whose behavior is uniform; no opt-in toggle is required because no platform-only feature is introduced.
  • Local-first: the feature makes only same-origin localhost calls (POST /profiles multipart, POST /clean-audio multipart, optional POST /design/describe JSON); no cloud, accounts, keys, or telemetry; app remains functional offline-except-localhost.
  • No data-format change: no alembic migration, no localStorage/store schema change, no version-file bump (main stays 0.3.6); existing profiles/projects continue to work.
  • Every failure/empty path is handled: mic-permission-denied, recording-too-short, denoise-unavailable (raw fallback), oversized clip, unprobeable clip, design .instruct-empty (Create disabled), design-render-503, 422, and network failure each show a clear toast (or disabled control) and never crash, never silently lose user input, and (for save failures) keep the modal open with inputs preserved. Each path has a named unit test (CreateVoiceModal cases b–n).
  • No close mid-save: while createProfile is in flight, ESC/backdrop/Cancel/close-button do not dismiss; they re-enable after the request resolves (unit case i).
  • Duplicate names are accepted without error (no UNIQUE constraint); assignment keys on id, so a name collision never breaks assignment.
  • Any "create a voice" doc flow (docs/**/README) is updated in the same PR (docs-sync).
  • CI green: bunx vitest run (ci.yml:107), uv run pytest tests/ and backend/tests/ (ci.yml:67,81), and CodeQL (security.yml:65-108 — no new ReDoS surface) all pass; PR is not merged before all checks report green (MEMORY merge-discipline).