Renames what users see. The app, the installers, the window title, the
docs and all 21 locales now say VoiceStudio, with "(previously
OmniVoice-Studio)" noted near the title of each doc surface so people
recognise it.
Deliberately NOT renamed, because renaming any of them silently breaks
an existing install — there is no legacy-path fallback anywhere in this
codebase:
- bundle identifier com.debpalash.omnivoice-studio (MSI UpgradeCode,
macOS TCC grants, managed venv, WebView localStorage, the
single-instance lock)
- data directories OmniVoice / .omnivoice and omnivoice.db
- the ~150 OMNIVOICE_* environment variables
- the X-OmniVoice-* HTTP headers (a wire protocol)
- the published Docker image paths
- the OmniVoice ENGINE, which is a model name and not this product
tests/test_identity_paths_survive_the_rename.py pins every one of those
so a future well-meaning sweep cannot orphan a user's library.
Linux .deb users install a new package name and should apt remove
omnivoice-studio; that note is in the changelog.
93 KiB
Implementation Spec — TASK #25: Inline "Create Voice" in Stories cast + Audiobook narrator
Lens applied (round 1/10): Codebase grounding. Every file/symbol/line reference below was verified against the working tree at
feat/stories-shared-render. Corrections from the prior draft are called out inline as [grounding] notes so reviewers can see what changed and why. The biggest five: (1)buildDesignInstruct()returns an object{ instruct, unsupported, duplicates }, not a string — the design save path must destructure.instruct; (2) theui/Dialogshell has noheaderslot — itstitleprop renders the header and the methodSegmentedmust live in the body; (3) the realtime refresh event is emitted frombackend/api/routers/profiles.py:124(notevents.py:7) and wired in the frontend viauseAppData.js:111-117, not a direct import; (4) no test files exist forCloneDesignTab/StoriesEditor/AudiobookTabtoday, so the extraction is not guarded by existing tests — they must be written from scratch; (5)POST /profilesreturns only{id, name, kind}at runtime, soonCreatedmay rely only onprofile.id.Lens applied (round 2/10): Completeness. Every state, empty path, and failure path the modal and its triggers must handle is now enumerated explicitly — see the new § State machine & edge cases section and the per-anchor edge-case notes. Verified against code:
useRecording.jshas a too-short guard (<1000bytes → toast, no ingest) and a denoise-failure fallback (raw-webm ingest with a "denoising unavailable" toast) that the modal inherits for free (useRecording.js:38-41,56-60);voice_profileshas no UNIQUE constraint onname(db.py:39-59), so duplicate-name creation silently succeeds and must be handled as a product decision, not an error;ApiErrorcarries.statusand.detail(api/client.ts:76-85), so failure-path toasts can branch on HTTP status (422 vs 503 vs network);ingestRefAudiosetspendingTrimFileand toasts but does not block — the studio keeps the too-long file pending, which the modal must instead reject (no AudioTrimmer in v1). All eight discrete states and every "and then…" are spelled out below.Lens applied (round 3/10): Project constraints. The spec now explicitly maps every relevant VoiceStudio hard rule to how this feature satisfies it — see the rewritten § Constraints section. Verified-against-code refinements this lens added: (a) i18n fallback is
fallbackLng: 'en'(frontend/src/i18n/index.ts:64) and there are 21 locale files underfrontend/src/i18n/locales/, so the correct mechanism is "add keys toen.json; the 20 other locales resolve missing keys via theenfallback chain" — not per-t()defaultValuestrings; (b) therecording.too_short/recording.loaded_rawandtts_errors.trim_hintkeys already exist inen.json(:1859-1860,:1693), so only the genuinely-newcreateVoice.*namespace is net-new; (c) the only user-input regex the modal's free-text feeds isString(freeText).split(/[,,]/)(voiceInstruct.js:53) — a linear-time character class, not a CodeQLpy/polynomial-redos(or JS-ReDoS) hazard, and the fullwidth-comma CJK char is functional text-processing already covered by the allowlist; (d)voiceInstruct.jsandconstants.jsare already in_ALLOWED_FILESoftests/test_no_hardcoded_cjk.py(:57,:64), so theChineseDialectpicker values surfaced by the modal trip no CJK gate; (e) the version is0.3.6across all three lockstep files (pyproject.toml,frontend/src-tauri/tauri.conf.json,frontend/src-tauri/Cargo.toml) — no bump; (f) the describe-debounce effect already has acancelledrace guard (CloneDesignTab.jsx:96,114), correcting the round-2 hedge.Lens applied (round 4/10): API + data shapes. Every wire shape this feature touches is now pinned to the byte: the exact
POST /profilesFormDatafield list with types and required-ness, the literal 2xx/422/503 response bodies ({"id","name","kind"}and{"detail": "..."}), theApiErrorfield shape thrown byapi/client.ts, thePOST /clean-audioresponse (a binaryFileResponse+X-Clean-Filenameheader — not JSON), thePOST /design/describerequest/response JSON (notematchedis a list of objects and the ignored-bucket field isunmatched, notduplicates), the realtime WebSocket event frame ({"type":"profiles","data":{"action":"created","id":"…"}}and the handler-signature mismatch — the frontend handler ignores the payload), the new pure-helper and component function signatures (isRefAudioTooLong,CreateVoiceModal, the three extracted sub-components,MicButton), and the SQLitevoice_profilesrow shape that backs the response. Corrections this lens added: (1)createProfileis typedPromise<Profile>but the route returns a partial{id,name,kind}—onCreatedmust treatname/kindas the only other reliable fields and never readref_audio/created_at/etc. off the create response; (2)/clean-audioreturnsaudio/wavbytes, consumed viares.blob()+res.headers.get("X-Clean-Filename")(useRecording.js:50-52) — it is not anapiJsoncall; (3) the describe response'smatchedfield is[{category,token,phrase}]and the UI only reads.length(CloneDesignTab.jsx:103), and the dropped-text field isunmatched: string[]— distinct frombuildDesignInstruct'sunsupported/duplicatesbuckets, which are a different local computation; (4) the realtime handler is registered asprofiles: () => loadProfiles()(useAppData.js:114) — zero-arg, so the{action,id}payload on the wire is intentionally unused and the modal must not depend on reading it.Lens applied (round 5/10): Test plan. The § Test plan section below is rewritten to be executable, not aspirational: it pins the test layer for every assertion (pure-helper harness vs. fetch-mocked component render vs. store-direct), states the explicit strategy for keeping the suite import-light so it never pulls
main+torch/GPU (the front-end suite is pure jsdom and never imports the Python backend; the heavy parts of the modal —createProfile,cleanAudio,/design/describe— are mocked at the lowest seam:global.fetchreturning{ok,status,json,text}, exactly asApiKeysPanel.test.jsx:32-52does, so no real network and noapiPost/apiFetchrewrite is needed), names the exact mocking seams verified in the existing harness (global.fetch;vi.mock('react-hot-toast', …)perEngineCompatibilityMatrix.test.jsx:7-10;import '../i18n'is auto-loaded bysetup.js:6;localStorageis stubbed bysetup.js:8-28), and the concrete CI gates with file:line (ci.yml:107bunx vitest run;ci.yml:67uv run pytest tests/which transitively runstests/test_no_hardcoded_cjk.pyandtests/probe/test_probe_i18n.py;security.yml:65-108CodeQL Python+JS/TS). Corrections this lens added: (i) the pure-helper tests (refAudio.test.js) and the store-action assertion (setCharacterVoice) must not render React at all — follow thestoriesSlice.test.tsharness pattern (:4-10) and thestoryCast.test.jspure-import pattern, which is also the mechanism that keeps the suite off the Python/torch path; (ii) toast assertions go through thevi.mock('react-hot-toast')spy and assert the i18n key or interpolated message passed totoast.error/toast.success, not rendered DOM text (toast renders into a portal the component test doesn't mount); (iii)useRecordingis exercised by mockingglobal.fetchfor/clean-audio(binaryBlobresponse with theX-Clean-Filenameheader) rather than mocking the hook, so the too-short/denoise-fallback branches are covered through the real hook code; (iv) the i18n key-completeness check is not a bespoke script — orphan/coverage findings are advisory (non-gating) pertests/probe/test_probe_i18n.py:1-3, so the hard requirement is only "add keys toen.json" (thefallbackLng:'en'chain makes the other 20 locales resolve them), and no new locale file needs editing to stay green; (v) there is nonode:testlegacy gate for these files —package.json:16test:legacytargets../tests/frontendand is not the vitest path; the only frontend gate isbunx vitest run.
TL;DR
Add a reusable <CreateVoiceModal> that wraps the two existing voice-definition paths (audio-clone with mic/upload + by-design with sliders/describe) into a self-contained Radix Dialog. On save it calls the existing createProfile() API (frontend/src/api/profiles.ts:16), refreshes the profile list, and auto-assigns the new profile to the call site (a Stories cast member, or the Audiobook default-voice field). Surface a "+ New voice" trigger in StoriesEditor's Cast panel rows and in AudiobookTab's default-voice field. The modal reuses the design/clone control logic that today lives only in CloneDesignTab.jsx — extracted into self-contained sub-components so neither the modal nor the main studio carries duplicated logic.
The task names <VoiceSelector> (#22) as the ideal host. <VoiceSelector> does not exist yet (#22 is still pending). This spec is written to ship #25 standalone (wiring the modal directly into the two raw <select> sites), with an explicit "if #22 lands first" integration note so the modal becomes the selector's footer action rather than a separate trigger.
Problem
Both longform surfaces let users pick a voice only from the already-saved profile list:
- Stories cast (
frontend/src/components/StoriesEditor.jsx:550-558): each cast row has a<select>ofprofilesbound tosetCharacterVoice(c.id, e.target.value || null); if the list is empty it showst('stories.noProfiles')(:571) with no way to create one. [grounding] Note the empty option label ist('stories.defaultVoice')(:556), and thenoProfileshint renders in addition to the row(s) only whenprofiles.length === 0(:571) — the cast itself always has at least anarratorrow whose delete is locked (:563). - Audiobook narrator/default (
frontend/src/pages/AudiobookTab.jsx:220-224): a single<select>ofprofilesbound to the localsetDefaultVoicestate setter, defaulting to "engine default" (empty value →t('audiobook.engine_default'),:222).
A user who lands in Stories/Audiobook with no profiles (or who wants a new character voice mid-project) must leave to the Voice studio, clone/design a voice, then navigate back and re-find their project. That is a hard wall directly contradicting the project's "first-run that actually works" core value. The clone (mic + upload) and design (sliders + "describe") flows already exist but are locked inside CloneDesignTab.jsx — a 792-line page component (frontend/src/pages/CloneDesignTab.jsx) driven entirely by props from App.jsx/useTTS/useProfiles, not reusable as-is. [grounding] The file is 792 lines (prior draft said 793).
Goal / Non-goals
Goals
- A
<CreateVoiceModal>component reachable from Stories cast rows and the Audiobook default-voice field. - Two methods inside the modal, matching the studio's
defineMethodSegmented control: From audio (mic record + drag/upload reference, with backend denoise) and By design (category sliders + "describe your voice" free-text → attrs). - On save:
createProfile(formData)(clone or designkind), thenloadProfiles()(the parent's refresh), then auto-assign the returnedprofile.idto the originating cast member / audiobook default. - Extract the clone/design control bodies out of
CloneDesignTab.jsxinto shared, prop-driven sub-components so the modal and the studio render the same UI from one source. - Full i18n; cross-platform-identical default behavior (mic permission errors mapped via the existing
micErrorMessagehelper,frontend/src/utils/micError.js, already used insideuseRecording.js:78). - Every failure and empty path produces a clear toast and never silently drops user input or leaves the modal in an indeterminate state — see § State machine & edge cases.
Non-goals
- Building the
<VoiceSelector>(#22) shared dropdown — out of scope; this spec degrades gracefully whether or not #22 lands first. - Production-overrides sliders (CFG, t_shift, etc., rendered at
CloneDesignTab.jsx:624-665— those are synthesis params, not profile params; not stored on a profile and not sent in the create FormData). - Consent-lock / verified-own-voice flow (separate feature; the
verified_own_voice/consent_*fields exist on theProfiletype atfrontend/src/api/types.ts:119-122and on thePOST /profiles/{id}/consentroute, but the modal does not gate on them or send them). - Changing the backend
POST /profilescontract (it already supports both kinds — see § API / data shapes). - Dub
CastingViewinline-create (the task scopes Stories + Audiobook; CastingView is mentioned only as the existing assignment widget). Leave a follow-up note. - Duplicate-name detection / rename-on-collision. [grounding]
voice_profiles.namehas no UNIQUE constraint (db.py:41—name TEXT NOT NULL, no index), so the backend will happily create a second profile with an identical name and a freshid. v1 allows this (matches today's studio behavior, which never dedupes). The modal MUST NOT pre-validate names against the existing list — but the assign step keys onprofile.id, never on name, so a collision is harmless to assignment. (Optional soft-warning UX is a documented follow-up.) - In-modal trimming of over-15s clips. v1 rejects + toasts; embedding
AudioTrimmeris a follow-up (see Risk). - No new platform-only behavior. This feature is a default feature (no opt-in toggle), so by the strict cross-platform-parity rule it must behave identically on macOS/Windows/Linux — see § Constraints. There is no macOS-only or Windows-only path introduced here.
Design
New extracted sub-components (shared between modal + studio)
Create frontend/src/components/voice-create/ (does not exist yet) with three pieces, all fully controlled (state owned by the caller via props — the same contract CloneDesignTab already uses via the props it destructures at CloneDesignTab.jsx:25-54):
-
CloneSourcePicker.jsx— the "From audio" body: the drop-zone<input type=file>+<MicButton>+ ref-transcript/style inputs. Lifted fromCloneDesignTab.jsx:359-452(thedefineMethod === 'audio'branch), minus the studio-only "using profile" banner (:399-413) and the inline save-as-profile row (:426-451) — the modal owns save. Props:{ refAudio, ingestRefAudio, refText, setRefText, instruct, setInstruct, isRecording, isCleaning, recordingTime, startRecording, stopRecording }. [grounding] Note the JSX at:389-395passes the recording handlers to<MicButton>asonStart={startRecording}/onStop={stopRecording}(MicButton's own prop names areonStart/onStop, see:768) — keep that mapping when extracting. [completeness] This component must render a visibly-distinct "recording in progress" state (isRecording, withrecordingTimeticking, driven byuseRecording.js:71-73) and a "cleaning/denoising" state (isCleaning, betweenstopRecordingandingestRefAudio,useRecording.js:44-62). Both states must disable the file-drop input and the methodSegmentedso the user can't start a second source mid-capture. -
DesignControls.jsx— the "By design" body: the describe-voice textarea + "Starting points" presets + identity recipe + category chip/select groups. Lifted fromCloneDesignTab.jsx:453-613(theelsebranch), minus the inline save row (:587-612). Props:{ vdStates, setVdStates, instruct, setInstruct, language, setLanguage, describeText, setDescribeText, ...feedback }. The/design/describedebounce effect (CloneDesignTab.jsx:93-116),onDescribeChange(:83-91),applyPersonality(:125-139),applyPreset(passed in as a prop today, defined inuseTTS.js:69-72),applyDemoPreset(:255-261), andonChipKeyDown(:228-235) move here too (they are self-contained);personalitiesis fetched via the existinguseQuery(['personalities'])(CloneDesignTab.jsx:119-123). [grounding]applyDemoPresetand theDemoPresetGridempty-state (:275-277) are tied to the studio'stext/activePersonalitystate and the prefilled script; for the modal they are not needed (the modal has no script textarea). Keep them in the extracted component but gate the demo grid behind a prop so the modal can suppress it. [completeness] The/design/describedebounce effect fires an async fetch on every keystroke (debounced 450 ms). It has three terminal sub-states the modal inherits: pending (in-flight — show the same spinner/feedback the studio shows at:465-474), success (applymergeDescribedAttrs(res.attrs)tovdStates, setdescribeUnmatched/describeMatchedAnyfeedback), and failure (network/5xx). On failure the component must not wipe the user's manual slider picks — the merge only runs on success; a failed describe leavesvdStatesuntouched (thecatchat:109-112is a no-op by design) and shows a non-fatal toast/inline hint. [grounding, round-3 correction] The existing effect already guards against a stale response overwriting a newer one on rapid typing: it declareslet cancelled = false(:96) and its cleanup returns() => { cancelled = true; clearTimeout(id); }(:114), with the async body short-circuiting onif (cancelled) return;(:100). The extraction must preserve this guard verbatim (do not regress it) — no new abort logic is required. -
MicButton.jsx— already exists privately atCloneDesignTab.jsx:768-792; promote to its own file and import from bothCloneDesignTabandCloneSourcePicker. Signature isMicButton({ isCleaning, isRecording, recordingTime, onStart, onStop }).
CloneDesignTab.jsx is then refactored to render <CloneSourcePicker .../> and <DesignControls .../> in place of the inlined JSX, threading the same props it already receives. Net behavior unchanged for the studio — this is a pure extraction. [grounding] There are currently NO tests for CloneDesignTab.jsx (confirmed: no CloneDesignTab.test.* exists anywhere under frontend/). The extraction is therefore not guarded by an existing suite — a regression test must be created (see Test plan), and a /verify pass on the studio synth flow is mandatory.
<CreateVoiceModal> (frontend/src/components/CreateVoiceModal.jsx)
A self-contained Dialog (imported from ../ui; see frontend/src/ui/index.js:18) that owns its own local voice-definition state (so it never touches the studio's useTTS/useAppStore generation state):
state: method ('audio'|'design'), refAudio, refText, instruct,
vdStates (init: all 'Auto'), language ('Auto'), describeText + feedback,
profileName, saving (bool), describePending (bool)
hooks: useRecording(localIngest) // reuse the exact hook — mic + denoise + raw-fallback
[completeness] State reset on open/close. When open transitions false→true, the modal resets all local state to initial (method = defaultMethod, refAudio = null, vdStates all 'Auto', profileName = '', saving = false). When it closes (Cancel / ESC / backdrop / post-success), it must also stop any active recording: call stopRecording() if isRecording so the mic stream's tracks are released (useRecording.js:35 releases tracks only on onstop). Closing mid-saving is blocked — see the state machine. Closing mid-isCleaning is allowed but the in-flight apiCleanAudio/ingestRefAudio resolves into a now-unmounted component, so the modal must guard setState after unmount (mounted-ref or AbortController), or simply let useRecording's finally run harmlessly.
[grounding] The ui/Dialog component has no header slot. Its signature is Dialog({ open, onClose, title = null, footer = null, size = 'md', dismissable = true, children }) (frontend/src/ui/Dialog.jsx:20-28). It renders its own <header> containing the title node and a close button only when title || dismissable (:51-66), the children in .ui-dialog__body (:68), and the footer node in .ui-dialog__footer (:69). So:
-
titleprop:t('createVoice.title'). -
Body (children), top: a
SegmentedFrom audio / By design. Reuse the studio's items/labels verbatim (CloneDesignTab.jsx:348-356, keysclone.define_from_audio/clone.define_by_design).Segmentedis exported from../ui(ui/index.js:22). [completeness] Disable theSegmentedwhileisRecording || isCleaning || savingso a method switch can't strand an in-flight capture or save. -
Body, middle:
<CloneSourcePicker>or<DesignControls>by method, plus a required nameInput(exported from../ui, seeui/index.js:17) and a language selector. [grounding] The language picker isSearchableSelect, imported from../components/SearchableSelect(frontend/src/components/SearchableSelect.jsx), NOT from../ui—ui/exposes a plainSelect(fromInput.jsx) but the studio usesSearchableSelectwithoptions={ALL_LANGUAGES}/popular={POPULAR_LANGS}(CloneDesignTab.jsx:9, 671-677). Match that for parity. -
footerprop: Cancel + a primary Create voice button (loading={saving}, disabled until valid).Buttonis exported from../ui(ui/index.js:15); it supportsloadingandblockprops (used atCloneDesignTab.jsx:732-741). [completeness] The Cancel button must be disabled (or relabeled) whilesavingis true to enforce the "no close mid-save" rule (Dialog's own close button must also be suppressed — passdismissable={!saving}so ESC/backdrop/X are all disabled during the request,Dialog.jsx:30,34,38,58). [grounding] Whendismissable={false}andtitleis set, the header still renders (because of thetitle || dismissableguard at:51) but the X close button disappears (:58) — exactly the desired "locked-during-save" header. -
size="lg"(Dialogaccepts'sm'|'md'|'lg'|'xl',Dialog.jsx:17). -
Local
ingestRefAudio: reuse the trim guard fromuseTTS.js:48-59—probeAudioDuration(fromutils/format.js) +CLONE_MAX_SECONDS(=15,utils/constants.js:48). [grounding] The studio guard atuseTTS.js:50-56setspendingTrimFileand callssetSelectedProfile(null), then toastst('tts_errors.trim_hint', { duration, max })— crucially it does not clear/replacerefAudioand leaves the over-long file "pending" for the AudioTrimmer UI. The modal has noselectedProfileand no trimmer. So extract just the duration-probe + threshold check intofrontend/src/utils/refAudio.jsas a pure helper (signature below) and let each caller decide. [completeness] For modal v1 the localingestRefAudiomust: (a) if!file→ setrefAudio = nulland return (matchesuseTTS.js:49); (b) probe duration; (c) iftooLong→ reject (do NOT setrefAudio), toastt('tts_errors.trim_hint', { duration, max }), and leave any previously-goodrefAudiountouched; (d) ifprobeAudioDurationreturns falsy/NaN (corrupt or unprobeable file) → accept the file (do not block on an unknown duration — backend/clean-audioandPOST /profileswill reject a truly invalid clip), matchinguseTTS.js:51which only blocks whendur && dur > MAX; (e) otherwise setrefAudio = file. -
Save handler builds the same
FormDatathe existing handlers build (see § API / data shapes for the exact field list) and callscreateProfile, thenonCreated(profile). See the state machine for the full success/failure branching.
State machine & edge cases (modal)
The modal is a small explicit FSM. Every transition's "and then…" is spelled out.
States: IDLE → (user edits) → READY (valid) ⇄ READY? (invalid) → SAVING → SUCCESS (closes) | ERROR (stays open, returns to READY). Plus orthogonal sub-states on the audio method: RECORDING, CLEANING, and on the design method: DESCRIBING.
- IDLE (just opened). Name empty, no source. Create voice disabled. Cancel/ESC/backdrop all dismiss freely (
dismissable=true). - Validity (READY vs invalid). Create is enabled iff
profileName.trim()is non-empty AND one of:- method
'audio':refAudio != null. (MirrorsuseProfiles.js:42need_name_audio.) - method
'design':buildDesignInstruct(vdStates, describeText/instruct).instructis non-empty — i.e. at least one non-Auto category or one valid free-text tag. [completeness] All-Auto sliders + empty/all-unsupported free-text → empty.instruct→ Create disabled, with an inline hint (t('createVoice.design_needs_attr')) explaining why, because the backend returns 422{"detail": "design profiles require instruct"}(profiles.py:75-76) and we must never let the user click into a guaranteed 422.
- method
- Switching method (audio↔design). Allowed only when not
RECORDING/CLEANING/SAVING. Switching does not clear the other method's state (so a user who fills audio, peeks at design, and switches back keeps their clip). Validity is re-evaluated for the now-active method. - RECORDING. Mic granted → MicButton shows stop +
recordingTime.Segmented, file-drop, and Create are disabled. Stop →CLEANING.- Mic permission denied / no device / device busy:
useRecordingcatches and toastsmicErrorMessage(t, e)(useRecording.js:75-79); state never enters RECORDING; modal stays in IDLE/READY. No crash. [grounding]micError.jsmaps theDOMException.nameto an i18n key and, for the permission case, attaches a per-OS hint (capture.mic_hint_mac|_windows|_linux, chosen bydetectPlatform()) — this is implementation-level OS detection, not a divergence in user-visible default behavior (every platform gets the same "here's how to re-enable mic access" UX, just with the correct path). See § Constraints. - Recording too short (<1000 bytes):
useRecording.js:38-41toastsrecording.too_shortand does not ingest —refAudiostays as it was. User can re-record.
- Mic permission denied / no device / device busy:
- CLEANING. Blob sent to
POST /clean-audio. On success →ingestRefAudio(cleanFile)(then trim guard, step §DesigningestRefAudio).- Denoise failure (backend down / 5xx / network):
useRecording.js:56-60falls back to the raw webm file and ingests it, toastingrecording.loaded_raw("denoising unavailable"). The modal treats this as a validrefAudio(a clone profile can be built from raw audio). No error state. - Cleaned clip > 15s: the local trim guard rejects it (toast
trim_hint), sorefAudiois not set even though recording "succeeded." User must re-record shorter. (Acceptable v1 — documented.)
- Denoise failure (backend down / 5xx / network):
- DESCRIBING (design method). Debounced
POST /design/describein-flight. Show pending feedback (CloneDesignTab.jsx:465-474). Success → merge attrs (mergeDescribedAttrs(res.attrs)). Failure → leavevdStatesuntouched, non-fatal toast. Create remains usable from the manual sliders regardless of describe state. - SAVING.
createProfile(formData)in flight. Create showsloading, Cancel disabled,dismissable=false. Exactly one of:- SUCCESS (2xx): response body is exactly
{ "id": string, "name": string, "kind": "clone"|"design" }(see § API). Then: (a)await onCreated(profile)— caller doesawait onProfilesChanged()then assign byprofile.id; (b)onClose(); (c) reset local state. The realtimeprofilesevent (profiles.py:124) is a belt-and-suspenders refresh; assignment does not depend on it. - ERROR (422): validation rejected server-side (should be unreachable for a correctly-built FormData, but possible if e.g.
instructcollapsed to empty due to a tag/whitelist drift).createProfilerejects withApiErrorwhose.status === 422and.detailis the FastAPI detail string (api/client.ts:120-121). Toaste.detail || e.message. Modal stays open, inputs preserved, returns to READY. Do NOT close. - ERROR (503, design only): deterministic-sample render failed (no TTS engine ready) —
profiles.py:103-107returns 503 with{"detail": "could not render the design sample: <reason>"}. Toast a friendly message keyed offe.status === 503(t('createVoice.render_unavailable'), falling back toe.detail). Modal stays open, inputs preserved. User can retry (e.g. after the engine warms up) or switch to the audio method. - ERROR (network / 5xx / timeout):
apiPostthrowsApiError(with.statusfrom the HTTP code, or a raw fetchTypeErroron a network drop where.statusisundefined). Toaste.message. Modal stays open, inputs preserved. - In every ERROR branch:
savingresets tofalse, re-enabling Create/Cancel/dismiss. User input (name, clip, sliders, describe text) is never cleared on failure — only on SUCCESS.
- SUCCESS (2xx): response body is exactly
- Concurrency guard. A second Create click while
savingis true is a no-op (button disabled). Acloseattempt whilesavingis true is blocked (dismissable=false). If the user closes duringCLEANING/DESCRIBING, the in-flight async resolves into an unmounted tree — guardsetState(mounted ref) to avoid React warnings.
Trigger + auto-assign wiring
Stories (StoriesEditor.jsx): in each cast row (:541-570), after the voice <select> (:550-558), add a small icon button (<UserPlus> from lucide-react — already used in frontend/src/pages/VoiceGallery.jsx:11,442) → opens <CreateVoiceModal> with assignTarget = { type:'cast', castId: c.id }. Also add a "+ New voice" affordance to the cast-panel header (:535-540, alongside the existing "+ Add character" button at :537-539) and a "Create one" CTA replacing the bare t('stories.noProfiles') hint (:571). onCreated(profile) → await onProfilesChanged() (passed down — see below) then setCharacterVoice(castId, profile.id). [grounding] setCharacterVoice(castId: string, profileId: string | null) is a real store action (frontend/src/store/storiesSlice.ts:45 declaration, :82-83 impl: set((s) => ({ cast: s.cast.map((c) => (c.id === castId ? { ...c, profileId } : c)) }))) already consumed in StoriesEditor.jsx:121,553.
- [completeness] Cast row deleted while modal open. A user can delete the originating cast member (
removeCastMember/deleteCharacter,StoriesEditor.jsx:562) after opening the modal but before save. OnonCreated,setCharacterVoiceon a now-missingcastIdis a harmless no-op against the store (the.mapsimply finds no matchingc.id,storiesSlice.ts:83), but the new profile still lands in the list and is selectable manually. Either accept (simplest) or capturecastIdat open-time and skip the assign if the row is gone — document the chosen behavior. The narrator row cannot be deleted (:563), so the common case is safe. - [completeness] Header / empty-state trigger has no
castId(it's a "create + add to library" action, not "assign to row X"). For these,onCreatedshouldawait onProfilesChanged()only (no assign) — the new voice appears in every row's<select>for manual pick. The empty-state CTA replaces the dead-endnoProfileshint with an actionable button.
Audiobook (AudiobookTab.jsx): next to the default-voice <select> (:218-225), add a "+ New voice" button (Plus icon already imported, AudiobookTab.jsx:3) → opens the modal with assignTarget = { type:'audiobookDefault' }. onCreated → await onProfilesChanged() then setDefaultVoice(profile.id). [grounding] setDefaultVoice here is the component-local useState setter (AudiobookTab.jsx:22), NOT a store action — there is no audiobook slice. Calling it just updates the local defaultVoice string bound to the <select> at :220.
Profile refresh plumbing. Today profiles is passed into both components (App.jsx:1074 <StoriesEditor profiles={profiles} />, :1080 <AudiobookTab profiles={profiles} />) but loadProfiles is not. Two options (pick the lighter):
- (A) Pass
onProfilesChanged={loadProfiles}fromApp.jsxintoStoriesEditor/AudiobookTab. Preferred — explicit, matches the existingprofiles={profiles}prop style. [grounding]loadProfilesis in scope inApp.jsx: it is destructured fromuseAppData()atApp.jsx:214(defined infrontend/src/hooks/useAppData.js:105asconst loadProfiles = useCallback(async () => { try { setProfiles(await listProfiles()); } catch (e) {} }, [])— signature:() => Promise<void>, swallows its own errors, returned atuseAppData.js:192). Prior draft cited "returned:229/used:266" —:229is actually theuseProfiles({...})destructure call and:266is one of several existingloadProfiles()call sites; the authoritative scope binding isApp.jsx:214. - (B) Have
<CreateVoiceModal>importuseAppData'sloadProfiles. Rejected —loadProfileslives inuseAppData/Appscope, not the store; threading the prop is cleaner.
[completeness] onProfilesChanged failure. loadProfiles() swallows its own errors (useAppData.js:105 empty catch), so it never rejects — the await onProfilesChanged() in onCreated resolves even if the backend is momentarily unreachable (the list just stays stale until the realtime profiles event reconciles it on socket reconnect, profiles.py:124). The onCreated handler should still await onProfilesChanged() and then assign; the assign uses the valid profile.id regardless, so a stale list shows a transient value-not-in-options until the refresh lands. Do not block the assign on the refresh's content.
After createProfile, the backend emits a profiles realtime event. [grounding] The emit is event_bus.emit("profiles", {"action": "created", "id": profile_id}) at backend/api/routers/profiles.py:124 (event_bus.emit defined in backend/core/event_bus.py:44). It is delivered over the WebSocket at backend/api/routers/events.py (/ws/events) as a frame {"type": "profiles", "data": {"action": "created", "id": "<id>"}} and handled on the frontend by useRealtimeEvents({ profiles: () => loadProfiles(), ... }) in frontend/src/hooks/useAppData.js:111-117. [grounding, round-4] The registered handler is zero-arg (profiles: () => loadProfiles(), :114) — the {action, id} payload is intentionally discarded; the modal must NOT depend on reading the realtime payload. So the list eventually updates even without an explicit refresh — but we still await onProfilesChanged() before assigning so the new id is present in profiles when the cast/audiobook <select> re-renders (avoids a flash of "value not in options").
If #22 (<VoiceSelector>) lands first
Make <CreateVoiceModal> self-contained and trigger-agnostic (it takes open, onClose, onCreated, defaultMethod?). Then #22's selector simply renders a "+ Create new voice…" item in its dropdown that sets open=true; the cast/audiobook call sites pass their existing assignment callback as onCreated. No modal rework needed — only the trigger moves.
API / data shapes
No backend changes. Every endpoint below already exists; this section pins their exact wire shapes so the modal can be implemented without guessing.
POST /profiles — create a voice profile (backend/api/routers/profiles.py:41-125)
Content-Type: multipart/form-data (built as a FormData object; apiPost sets the body to the FormData and lets the browser set the boundary — no manual Content-Type header, see api/client.ts:137-138).
Handler signature (profiles.py:42-52):
async def create_profile(
name: str = Form(...), # REQUIRED
ref_audio: Optional[UploadFile] = File(None),
ref_text: str = Form(""),
instruct: str = Form(""),
language: str = Form("Auto"),
seed: Optional[int] = Form(None),
personality:str = Form(""),
kind: str = Form("clone"),
vd_states: Optional[str] = Form(None), # JSON string
)
Clone-method FormData (mirrors useProfiles.js:43-50 exactly):
| field | type | required | value the modal sends |
|---|---|---|---|
name |
string | yes (.trim() non-empty) |
profileName |
ref_audio |
Blob | yes (kind=clone) | reconstructed: new Blob([await refAudio.arrayBuffer()], { type: refAudio.type }), appended as ("ref_audio", safeBlob, refAudio.name || "profile.wav") (useProfiles.js:45-47) |
ref_text |
string | no | optional transcript (default "") |
instruct |
string | no | optional style string (default "") |
language |
string | no | the SearchableSelect value (default "Auto") |
kind |
string | no | "clone" (may be omitted — backend defaults to "clone") |
Design-method FormData (mirrors useProfiles.js:88-93 exactly):
| field | type | required | value the modal sends |
|---|---|---|---|
name |
string | yes | profileName |
kind |
string | yes | the literal string "design" |
vd_states |
string | yes (must JSON-parse to an object) | JSON.stringify(vdStates || {}) — a { [category]: value } map, e.g. {"Gender":"female","Age":"elderly","Pitch":"Auto",...} |
instruct |
string | yes (non-empty) | buildDesignInstruct(vdStates, freeInstruct).instruct — the .instruct field of the object, a comma-joined token string like "female, elderly, low pitch" |
language |
string | no | "Auto" default |
[grounding, round-4] CRITICAL — the design
instructis a string, not the whole object.buildDesignInstructreturns{ instruct: string, unsupported: string[], duplicates: string[] }(voiceInstruct.js:66). Append.instruct. Reference correct usage atuseTTS.js:121(const { instruct: finalInstruct, unsupported, duplicates } = buildDesignInstruct(...)). Note that the studio's save path atCloneDesignTab.jsx:606passes the whole object intohandleSaveDesignProfile, which would serialize as[object Object]— that is a latent pre-existing bug in the studio save row; the modal must NOT replicate it. Do not "fix" the studio bug as part of this task unless the owner asks (out of scope); just avoid it in the modal.The modal never sends
seedorpersonality.seedisForm(None)→ backend defaults design profiles to_DESIGN_SEED = 42(profiles.py:38,108); the modal omits it.personalityis empty by default. Production-override synthesis params (cfg/t_shift/etc.) are explicitly not profile fields and are never appended.
Validation order (each a discrete 422 the modal must avoid or surface) (profiles.py:61-76):
kind not in ("clone","design")→ 422{"detail": "kind must be 'clone' or 'design'"}. (Modal only ever sends one of the two; unreachable.)kind=="clone" and ref_audio is None→ 422{"detail": "clone profiles require ref_audio"}. (Modal disables Create untilrefAudioset; unreachable unless the file is dropped between validity-check and submit.)kind=="design"andvd_statesempty/whitespace → 422{"detail": "design profiles require vd_states"}.kind=="design"andvd_statesdoes notjson.loadsto adict→ 422{"detail": "vd_states must be a JSON object"}. (Modal always sendsJSON.stringify(obj), so unreachable.)kind=="design"andinstruct.strip()empty → 422{"detail": "design profiles require instruct"}. (Modal disables Create when.instructis empty; the FSM §2 hint guards this.)
Success response — 200 (profiles.py:125, the literal body):
{ "id": "a1b2c3d4", "name": "Grandma Vera", "kind": "design" }
idis an 8-char hex slug (str(uuid.uuid4())[:8],profiles.py:78), not a full UUID.- [grounding, round-4] The route returns only these three keys even though
createProfile()is typedPromise<Profile>(profiles.ts:16) and the fullProfileinterface (frontend/src/api/types.ts:109-123) declareslanguage_code,ref_audio,ref_text,description,created_at,is_locked,verified_own_voice,consent_text,consent_recorded_at.onCreatedmust read onlyprofile.id(always present) and may readname/kind; it must NOT read any otherProfilefield off the create response (they areundefined). The full record is fetched separately viaGET /profiles/{id}or arrives in the refreshed list fromGET /profiles.
Design-render-failure response — 503 (profiles.py:103-107):
{ "detail": "could not render the design sample: <python exception text>" }
Raised when _render_archetype_wav throws (no TTS engine ready). Surface via e.status === 503.
DB-insert-failure (profiles.py:119-123): if the INSERT raises, the handler removes the orphaned audio file and re-raises the original exception → FastAPI turns it into a 500 with a generic body. Surfaces to the modal as the generic network/5xx ERROR branch. Rare.
Realtime side-effect (profiles.py:124): on success, event_bus.emit("profiles", {"action": "created", "id": profile_id}). Wire frame: {"type": "profiles", "data": {"action": "created", "id": "<id>"}}. Frontend handler ignores the payload (zero-arg () => loadProfiles()).
POST /clean-audio — mic-recording denoise (backend/api/routers/system.py:912-972)
Used only inside useRecording.js:48; the modal inherits it via the hook and never calls it directly.
- Request:
multipart/form-data, one fieldaudio(the raw webmBlob, appended as("audio", blob, "recording.webm"),useRecording.js:46-47). - Response: NOT JSON. It is a binary
FileResponse(media_type="audio/wav") with headerX-Clean-Filename: mic_<id>.wav(system.py:971-972). The frontend wrappercleanAudio(formData)returns the rawResponse(api/system.ts:95-98) precisely so the caller can readawait res.blob()+res.headers.get("X-Clean-Filename")(useRecording.js:50-52). Do notapiJsonthis endpoint. - Failure: any non-2xx (or network drop) makes
useRecording'strythrow → itscatchfalls back to ingesting the raw webm with therecording.loaded_rawtoast (useRecording.js:56-60). No JSON error shape matters to the modal.
POST /design/describe — free-text → design attrs (backend/api/routers/describe_voice.py:22-35)
Used inside the DesignControls describe-debounce effect (extracted from CloneDesignTab.jsx:99); the modal inherits it.
- Request (JSON):
{ "description": string }(max_length=2000,describe_voice.py:18-19). Sent viaapiPost('/design/describe', { description: q })(CloneDesignTab.jsx:99) →Content-Type: application/json. - Response (JSON, exact shape,
describe_voice.py:26-33):{ "attrs": { "Gender": "female", "Age": "elderly", "Pitch": "Auto", "Style": "Auto", "EnglishAccent": "british", "ChineseDialect": "Auto" }, "instruct": "female, elderly, british accent", "matched": [ { "category": "Age", "token": "elderly", "phrase": "elderly" } ], "unmatched": [ "slightly raspy" ] }- [grounding, round-4]
attrsis a complete{ [category]: token|"Auto" }map over the six categories. The effect feeds it throughmergeDescribedAttrs(res.attrs)(voiceInstruct.js:82-89), which re-validates each token againstCATEGORIESand resets anything unrecognized to"Auto"— so a drifted/older backend can never inject an out-of-taxonomy value. - [grounding, round-4]
matchedis a list of objects{category, token, phrase}; the UI reads only(res.matched || []).length > 0to setdescribeMatchedAny(CloneDesignTab.jsx:103). Do not treatmatchedas a string array. - [grounding, round-4] The dropped-text field is
unmatched: string[](CloneDesignTab.jsx:102→setDescribeUnmatched(res.unmatched || [])). This is the describe endpoint's "I couldn't map these words" bucket — distinct frombuildDesignInstruct'sunsupported/duplicatesbuckets (which are a separate client-side computation over the manual free-text + sliders at save time). Keep the two feedback sources clearly separated inDesignControls.
- [grounding, round-4]
- Failure: the debounce effect's
try/catchswallows any error (CloneDesignTab.jsx:109-112) — controls untouched, next keystroke retries. No error shape surfaces. (This endpoint is pure-CPU/stdlib, so a failure here is effectively only a network drop.)
ApiError — the rejection shape every failed call throws (frontend/src/api/client.ts:76-85,114-122)
class ApiError extends Error {
name = 'ApiError';
status?: number; // the HTTP status (422, 503, 500, …); undefined on a raw fetch network failure
detail?: unknown; // the parsed `detail`/`error` field from the body, else the raw text (readError, :92-100)
// .message is `"<status> <statusText>: <detail>"` (:121)
}
createProfile/apiPost/apiFetch reject with this on any non-2xx. Failure-path toasts branch on e.status (422 / 503 / other) and may show e.detail (a string for FastAPI HTTPException) or fall back to e.message. Note: a true network drop (server gone) throws a fetch TypeError, not ApiError, so e.status is undefined there — the generic branch must handle both (e?.status optional-chained).
[test seam, round-5] How tests reproduce each
ApiErrorbranch.ApiErroris a plain exported class (client.ts:76) constructed asnew ApiError(message, { status, detail }). The lowest mockable seam isglobal.fetch(everything funnels throughapiFetch→fetch(apiUrl(path), …),client.ts:114). To makecreateProfilereject with a given status, a test mocksglobal.fetchto resolve{ ok:false, status:422, text: async () => JSON.stringify({detail:'design profiles require instruct'}) }—readError(client.ts:92-100) parses.detailout andapiFetchthrows theApiErrorwith the right.status/.detail(client.ts:120-121). For a network drop, mockglobal.fetchto reject with aTypeError— that propagates raw (no.status), exercising the generic branch. This means the modal's 422/503/network toast branches are tested through the realclient.tscode, not by hand-throwing — closer to production and guards againstreadError/apiFetchregressions.
voice_profiles table — the row that backs the response (backend/core/db.py:39-59)
CREATE TABLE IF NOT EXISTS voice_profiles (
id TEXT PRIMARY KEY, -- the 8-char slug returned as "id"
name TEXT NOT NULL, -- NO UNIQUE constraint → duplicate names allowed
ref_audio_path TEXT, -- "<id>.<ext>" (clone) or "<id>.wav" (design sample)
ref_text TEXT DEFAULT '',
instruct TEXT DEFAULT '',
language TEXT DEFAULT 'Auto',
locked_audio_path TEXT DEFAULT '',
seed INTEGER DEFAULT NULL, -- 42 for design (the deterministic sample), NULL/clone otherwise
is_locked INTEGER DEFAULT 0,
personality TEXT DEFAULT '',
description TEXT DEFAULT '',
is_demo INTEGER DEFAULT 0,
verified_own_voice INTEGER DEFAULT 0,
consent_text TEXT DEFAULT '',
consent_audio_path TEXT DEFAULT '',
consent_recorded_at REAL DEFAULT NULL,
kind TEXT DEFAULT 'clone', -- 'clone' | 'design'
vd_states TEXT DEFAULT NULL, -- the design JSON string, verbatim
created_at REAL
);
[grounding, round-4] No column is added or altered → no alembic migration, no localStorage/store schema change → existing omnivoice_data/ is untouched (see § Constraints). id is the PK and the only assignment key; name has no UNIQUE index, so the modal can never produce a name-collision error (only an optional, follow-up soft-warning).
Function signatures introduced by this task
// frontend/src/utils/refAudio.ts (or .js) — pure, no React.
// Returns whether `file` exceeds CLONE_MAX_SECONDS; `tooLong:false` for an
// unprobeable/NaN duration (accept rather than block — mirrors useTTS.js:51).
export async function isRefAudioTooLong(
file: File,
): Promise<{ tooLong: boolean; duration: number }>;
// CLONE_MAX_SECONDS = 15 (utils/constants.js:48); duration via probeAudioDuration (utils/format.js).
// frontend/src/components/CreateVoiceModal.jsx
type ProfileKind = 'clone' | 'design';
interface CreateVoiceModalProps {
open: boolean;
onClose: () => void;
// Receives ONLY the create-route's partial response — read id/name/kind only.
onCreated: (profile: { id: string; name: string; kind: ProfileKind })
=> void | Promise<void>; // caller does (await) onProfilesChanged + assign
defaultMethod?: 'audio' | 'design'; // default 'audio'
}
// frontend/src/components/voice-create/CloneSourcePicker.jsx
function CloneSourcePicker({
refAudio, ingestRefAudio, refText, setRefText, instruct, setInstruct,
isRecording, isCleaning, recordingTime, startRecording, stopRecording,
});
// frontend/src/components/voice-create/DesignControls.jsx
function DesignControls({
vdStates, setVdStates, instruct, setInstruct, language, setLanguage,
describeText, setDescribeText,
describeUnmatched, describeMatchedAny, // describe-feedback state
showDemoGrid = true, // modal passes false (no script state)
});
// frontend/src/components/voice-create/MicButton.jsx (promoted from CloneDesignTab:768)
function MicButton({ isCleaning, isRecording, recordingTime, onStart, onStop });
Assign callbacks the triggers wire as onCreated:
- Stories cast row:
async (p) => { await onProfilesChanged(); setCharacterVoice(castId, p.id); }(setCharacterVoice: (castId: string, profileId: string|null) => void, store action). - Stories header / empty-state:
async () => { await onProfilesChanged(); }(refresh only, no assign). - Audiobook default:
async (p) => { await onProfilesChanged(); setDefaultVoice(p.id); }(setDefaultVoice= component-localuseStatesetter,AudiobookTab.jsx:22).
[completeness] Design-method describe + feedback surfacing at save time. When buildDesignInstruct(...)'s unsupported.length or duplicates.length is non-empty (free-text tags dropped/collided), the modal should surface the same non-fatal tts_errors.ignored_unsupported / tts_errors.ignored_duplicate toasts the studio shows at synth time (useTTS.js:122-127) — but at save time, so the user understands why a tag they typed isn't reflected. This is advisory, not blocking, and is separate from the describe endpoint's unmatched feedback (§ /design/describe).
Integration points (file:line)
[grounding] All anchors below re-verified against the working tree.
frontend/src/pages/CloneDesignTab.jsx:359-452— clone-source JSX (defineMethod === 'audio') to extract intoCloneSourcePicker.jsx(drop the studio-only banner:399-413and save row:426-451). [completeness] Preserve the recording-in-progress and cleaning states (driven byisRecording/isCleaning/recordingTime).frontend/src/pages/CloneDesignTab.jsx:453-613— design-controls JSX (elsebranch) →DesignControls.jsx. Co-located logic to move: describe effect:93-116(including itscancelledrace guard at:96,100,114— preserve verbatim; the effect setsvdStates/describeUnmatched/describeMatchedAny/activePersonality/instructon success per:101-108);onDescribeChange:83-91;applyPersonality:125-139;applyDemoPreset:255-261;onChipKeyDown:228-235; personalitiesuseQuery:119-123. (applyPreset/insertTagare props fromuseTTS.js:61-72.) [completeness] Keep the describe-feedback render (:465-474) so the modal shows pending/success/failure of the describe call; gateDemoPresetGrid(:275-277) behind a prop (modal hides it — no script state).frontend/src/pages/CloneDesignTab.jsx:768-792—MicButton({ isCleaning, isRecording, recordingTime, onStart, onStop })to promote into its own file.frontend/src/pages/CloneDesignTab.jsx:348-356— theSegmenteditems/labels (clone.define_from_audio/clone.define_by_design) reused by the modal header.frontend/src/components/StoriesEditor.jsx:535-540— cast-panel header (add "+ New voice" affordance).:541-570— cast rows.:550-558— cast-row<select>(empty option =t('stories.defaultVoice'),:556).:563— narrator delete is locked; other rows can be deleted while the modal is open (handle the stale-castIdassign).:571—t('stories.noProfiles')hint (renders only whenprofiles.length === 0) to replace with a CTA. Assign viasetCharacterVoice(store action, imported:121).frontend/src/pages/AudiobookTab.jsx:218-225— default-voice field; add trigger.:220-224— the<select>(empty option =t('audiobook.engine_default'),:222). Assign via localsetDefaultVoice(:22) — not persisted across nav (see Risk).Plusicon already imported (:3).frontend/src/App.jsx:1074(<StoriesEditor profiles={profiles} />) and:1080(<AudiobookTab profiles={profiles} />) — addonProfilesChanged={loadProfiles}.loadProfiles(() => Promise<void>, error-swallowing) destructured atApp.jsx:214fromuseAppData().frontend/src/hooks/useProfiles.js:41-57—handleSaveProfile(refAudio, refText, instruct, language)(clone FormData reference shape: name guard at:42toastsprofiles.need_name_audiowhen!profileName.trim() || !refAudio; append block:43-50, Blob reconstruction viaarrayBuffer()at:45-47;await createProfile(formData)thenawait loadProfiles()at:52,55; on error toastse.messageat:56).frontend/src/hooks/useProfiles.js:86-101—handleSaveDesignProfile(vdStates, instruct, language)(design FormData; name guard at:87; append block:88-93; error toast:100). [grounding] The two relevant functions are:41-57and:86-101(withhandleDeleteProfile/handleSelectProfilein between). [completeness] Both existing handlersawait loadProfiles()after a successful create and toaste.messageon failure without clearing inputs — the modal mirrors exactly this failure discipline (preserve inputs, toast, stay open).frontend/src/hooks/useRecording.js:12-96—useRecording(ingestRefAudio)reused verbatim by the modal. Returns{ isRecording, isCleaning, recordingTime, startRecording, stopRecording }(:89-95). Internals the modal inherits: mic errors mapped viamicErrorMessage(t, e)(:78); too-short guard (blob.size < 1000→ toastrecording.too_short, no ingest,:38-41);/clean-audiocall +X-Clean-Filenameread (:48-52); denoise-failure fallback to raw webm with toastrecording.loaded_raw(:56-60); tracks released ononstop(:35). [grounding, round-3] All three i18n keys (recording.too_short:1860,recording.loaded_raw:1859) already exist inen.json; the only net-new namespace iscreateVoice.*.frontend/src/api/system.ts:95-98—cleanAudio(formData): Promise<Response>(returns the rawResponseso the caller readsblob()+X-Clean-Filename; notapiJson).frontend/src/utils/micError.js—micErrorMessage(t, err)(+detectPlatform()returning'mac'|'windows'|'linux'); pure module. The permission case attachescapture.mic_hint_mac|_windows|_linux. OS-detection implementation the parity rule permits; behavior identical across platforms. [test, round-5] This module is already fully unit-tested atfrontend/src/utils/micError.test.js(coversdescribeMicError,micErrorMessage,detectPlatform,micHintKeywith a faket); the modal tests do not re-test mic-error mapping — they only assert that the modal surfaces a toast on the denied path (the content is already covered). See § Constraints.frontend/src/hooks/useTTS.js:48-59—ingestRefAudio+CLONE_MAX_SECONDS/probeAudioDurationtrim guard (logic:50-56) to factor intoutils/refAudio.js. [completeness] The studio only blocks whendur && dur > MAX(:51) — an unprobeable file (durfalsy/NaN) is accepted. Mirror that inisRefAudioTooLong.frontend/src/api/profiles.ts:16-18—createProfile(formData: FormData): Promise<Profile>(typedProfilebut the route returns the partial{id,name,kind}— see § API). Rejects withApiErrorcarrying.status/.detail.frontend/src/utils/voiceInstruct.js:33-67(buildDesignInstruct(vdStates = {}, freeText = '')) — returns{ instruct: string, unsupported: string[], duplicates: string[] }(:66). The design save path uses.instruct. [completeness] When every picked value is'Auto'and free-text is empty/unsupported,.instructis the empty string (:66joins an emptybyCategory) — the disable-Create condition in state §2. A dropdown value not inCATEGORIESis silently skipped with aconsole.warn(:46), so a drifted picker can't inject a 422-bound tag. [grounding, round-3] The free-text split isString(freeText||'').split(/[,,]/)(:53) — a two-char character class (ASCII comma + fullwidth CJK comma), linear-time, no ReDoS surface.frontend/src/utils/voiceInstruct.js:82-89(mergeDescribedAttrs(attrs = {})) — used by the/design/describeeffect; returns a complete{ [category]: token|'Auto' }map, resetting unmatched/out-of-taxonomy categories to'Auto'(:86). [test, round-5]buildDesignInstructis already covered byfrontend/src/utils/voiceInstruct.test.js(six cases incl. the object-shape{instruct, unsupported, duplicates}return, the all-Auto empty-instruct case:47-50, and the full-width-comma normalisation:42-45); the modal tests reuse the behavior but do not re-test the helper — they assert the modal sends.instructcorrectly (case d below).frontend/src/utils/constants.js:17-31(CATEGORIES),:10-15(TAGS),:33-46(PRESETS),:48(CLONE_MAX_SECONDS = 15) — design control source data + trim threshold. [grounding]TAGSis:10-15(separate fromPRESETS). [grounding, round-3]CATEGORIES.ChineseDialect(:27-30) holds CJK picker values;constants.jsis already in_ALLOWED_FILES(tests/test_no_hardcoded_cjk.py:64).frontend/src/components/SearchableSelect.jsx— language picker (import from../components/SearchableSelect).frontend/src/languages.json(ALL_LANGUAGES) +POPULAR_LANGS(utils/constants.js:1-4) feed it.frontend/src/ui/Dialog.jsx:20-74— modal shell. SignatureDialog({ open, onClose, title=null, footer=null, size='md', dismissable=true, children }).dismissablecontrols ESC (:34), backdrop (:38), and the close-button presence (:58); passdismissable={!saving}to lock the modal during a create request. The header renders whentitle || dismissable(:51); body ischildren(:68); footer is thefooternode (:69).frontend/src/ui/index.js:18exportsDialog;:15Button,:17Input,:22Segmented.frontend/src/api/client.ts:76-85—ApiError(.status,.detail,.message);:92-100—readError(parsesdetail/errorfrom the body);:114-122— where non-2xx throws it (this is the seam tests drive viaglobal.fetch);:137-138—apiPostsendsFormDataas the body untouched (browser sets multipart boundary).frontend/src/i18n/index.ts:64— [grounding, round-3]fallbackLng: 'en'is the resolution mechanism for missing keys in the other 20 locales. NewcreateVoice.*keys go inen.json; the other locales resolve them through this fallback. (No per-t()defaultValuestrings.)backend/api/routers/profiles.py:41-125—POST /profilescontract (no change). Validation:61-76; success body:125; 503:103-107; DB-failure cleanup+re-raise:119-123; realtime emit:124.backend/api/routers/system.py:912-972—POST /clean-audio; binaryFileResponse+X-Clean-Filenameheader (:971-972).backend/api/routers/describe_voice.py:22-35—POST /design/describe; request{description}(:18-19), response{attrs, instruct, matched, unmatched}(:26-33).backend/core/db.py:39-59—voice_profilesschema.name TEXT NOT NULLwith no UNIQUE index → duplicate names allowed;idis the PK and the only assignment key. No column added/altered → no alembic migration.
Test plan
Strategy — keep the suite import-light so it never touches
main+torch/GPU. This is a front-end-only feature, and the front-end vitest suite runs in jsdom (frontend/vite.config.js:30-36), never importing the Python backend, torch, CUDA/MPS, or any heavy native module. That is the structural reason these tests are safe to run locally (it sidesteps the documentedtests/pytest segfault from importingmain+torch/Triton):create_profileis unchanged, so this task adds zero new Python tests — backend behavior is covered by the existing CI pytest job, not by anything new here. Within the front-end suite, the import-light discipline is enforced two ways:
- Pure helpers and store actions are tested with no React render at all.
refAudio.test.jsimports onlyisRefAudioTooLong(pure async, no JSX) — thestoryCast.test.jspattern. ThesetCharacterVoiceauto-assign contract is asserted against the store slice directly via thestoriesSlice.test.tsharness (:4-10: a hand-rolledset/getovercreateStoriesSlice), not by mountingStoriesEditor— so that assertion never pulls in the whole editor tree.- The heavy seams of the modal are mocked at the lowest layer:
global.fetch.createProfile,cleanAudio, and/design/describeall funnel throughapiFetch→fetch(apiUrl(path), …)(client.ts:114). Tests replaceglobal.fetchwith avi.fn()returning the{ ok, status, json, text }shape used byApiKeysPanel.test.jsx:32-52(and theBlob-body +X-Clean-Filename-header shape for/clean-audio). This exercises the realclient.ts/useRecordingcode (soApiErrorconstruction,readError, the denoise raw-fallback, and the too-short guard are all covered through production code paths) without any network and without rewritingapiPost. No module under test imports anything outside the front-end source tree.Harness facts (verified, so test files don't re-discover them):
setup.js(referenced atvite.config.js:33) auto-loads../i18n(setup.js:6) — soimport '../i18n'is not needed per-file andt('…')resolves real English strings; it also stubswindow.localStorage(setup.js:8-28). Toast is a portal side-effect, so tests must mock it:vi.mock('react-hot-toast', () => ({ default: { error: vi.fn(), success: vi.fn() }, toast: { error: vi.fn(), success: vi.fn() } }))(the exact form inEngineCompatibilityMatrix.test.jsx:7-10) and assert on the arguments passed totoast.error/toast.success(the resolved i18n string), never on rendered DOM toast text. Components that fetchpersonalitiesviauseQuerymust be wrapped in aQueryClientProvider(or the query mocked); design-method tests that don't touch personalities can omit it if the component tolerates an absent provider — prefer wrapping to match production. Fake timers (vi.useFakeTimers()) are used to flush the 450 ms describe debounce.
Unit / component (vitest + RTL, gated by ci.yml:107 bunx vitest run)
frontend/src/utils/refAudio.test.js (new — pure helper, no React):
it('reports tooLong:true with the duration for a >15s clip')— mockprobeAudioDuration(viavi.mock('./format')) to resolve20→{ tooLong: true, duration: 20 }.it('reports tooLong:false for a clip at or under 15s')—probeAudioDuration→15→{ tooLong: false, duration: 15 }(boundary,> MAXnot>= MAX, mirrorsuseTTS.js:51).it('accepts (tooLong:false) an unprobeable clip whose duration is NaN/undefined')—probeAudioDuration→NaNand →undefinedboth yield{ tooLong: false }.
frontend/src/components/CreateVoiceModal.test.jsx (new — fetch-mocked render). Each case beforeEach(() => vi.restoreAllMocks()); mock react-hot-toast once at module top:
- (a)
it('renders both methods and toggles via the Segmented control')— both labels (clone.define_from_audio/clone.define_by_design) render; clicking switches the visible body betweenCloneSourcePickerandDesignControls. - (b)
it('disables Create until the form is valid for the active method')— name empty → Createdisabled; name set, audio method, norefAudio→ disabled; design method all-Auto + empty free-text (sobuildDesignInstruct(...).instruct === '') → disabled and thecreateVoice.design_needs_attrhint is present; design method with one non-Auto category → enabled. (Assert viagetByRole('button', { name: /create voice/i }).disabled.) - (c)
it('builds clone-method FormData and fires onCreated on success')— mockglobal.fetchto resolve{ ok:true, status:200, json: async () => ({ id:'a1b2c3d4', name:'X', kind:'clone' }) }; ingest aFile; submit; assert the capturedfetchinitbody instanceof FormData,body.get('name')is the string,body.get('ref_audio') instanceof Blobwith the file'stype, andbody.get('kind')is'clone'or absent; assertonCreatedwas called with{ id:'a1b2c3d4', name:'X', kind:'clone' }andonClosefired. - (d)
it('sends design instruct as a non-empty string, never [object Object]')— design method, set Gender→female; submit; assertbody.get('kind') === 'design',JSON.parse(body.get('vd_states'))is an object (not array/string), andbody.get('instruct')is a non-empty string andexpect(body.get('instruct')).not.toBe('[object Object]')(the round-4 latent-bug guard). - (e)
it('keeps the modal open and preserves inputs on a 422')—global.fetch→{ ok:false, status:422, text: async () => JSON.stringify({ detail:'design profiles require instruct' }) }; submit; asserttoast.errorwas called with a string containing the detail,onClosenot called, and the name<input>still holds its value. - (f)
it('shows a friendly render-unavailable toast on a 503 (design)')—global.fetch→{ ok:false, status:503, text: async () => JSON.stringify({ detail:'could not render the design sample: no engine' }) }; asserttoast.errorarg equals the resolvedcreateVoice.render_unavailablestring (or containse.detail), modal stays open. - (g)
it('shows e.message and stays open on a network drop')—global.fetchset tovi.fn().mockRejectedValue(new TypeError('Failed to fetch')); assert a toast fired,onClosenot called (covers the.status===undefinedbranch). - (h)
it('closes and resets to IDLE on success')— after a successful create + close, re-open withopentoggled and assert Create is disabled again and norefAudio/name carried over. - (i)
it('blocks close while saving and re-enables after the request resolves')— makeglobal.fetchreturn a never-yet-resolved promise; while pending, fire ESC / click the backdrop / click Cancel and assertonClosenot called anddismissableis false (the X is absent); then resolve and assert dismiss works. - (j)
it('toasts recording.too_short and leaves refAudio unset on a <1000-byte capture')— drive the realuseRecordingby mockingMediaRecorder/getUserMedia(jsdom) to emit a<1000-byteBlobon stop; asserttoast.errorgot therecording.too_shortstring and Create stays disabled (norefAudio). - (k)
it('falls back to the raw webm and toasts recording.loaded_raw when /clean-audio fails')—global.fetchfor/clean-audio→{ ok:false, status:500 }; assert the raw file is ingested (Create becomes enabled with a name) andtoastgot therecording.loaded_rawstring. - (l)
it('rejects an over-15s upload with the trim_hint toast and no refAudio')—vi.mock('../utils/format', …)soprobeAudioDuration→20; ingest a file; asserttoast.errorgot the interpolatedtts_errors.trim_hintstring andrefAudiois unset (Create disabled). - (m)
it('accepts an unprobeable upload (NaN duration)')—probeAudioDuration→NaN; ingest; assertrefAudiois set (Create enabled with a name). - (n)
it('reads only id/name/kind off the create response')—global.fetchresolves exactly{ id:'a1b2c3d4', name:'X', kind:'design' }(no other keys); assertonCreatedreceives that object verbatim and the test spy verifies the assign callback was invoked with.id; (defensive) assert no thrown access on absentref_audio/created_at.
frontend/src/components/StoriesEditor.test.jsx (new — fetch-mocked render; setCharacterVoice assertion done separately at the slice level per the strategy note):
it('opens CreateVoiceModal from a cast row trigger')— click the row'sUserPlusbutton → modal visible.it('auto-assigns the new profile to the row on onCreated')— simulateonCreated({ id:'p_new', name:'X', kind:'clone' })and assert the passedonProfilesChangedwas awaited andsetCharacterVoice(a spy injected as a prop / via a mocked store) was called with(rowId, 'p_new').it('replaces the noProfiles dead-end with an actionable CTA when profiles is empty')— render withprofiles={[]}and assert a button (not just thestories.noProfilestext) is present.it('header trigger refreshes but does not assign')— headeronCreatedcallsonProfilesChangedand notsetCharacterVoice.it('does not throw when the originating cast row was deleted before onCreated resolves')— delete the row, then fireonCreated; assert no throw (the slice.mapno-ops on a missing id).
frontend/src/components/StoriesEditor.cast.test.ts (new — pure store, the import-light assertion): using the storiesSlice.test.ts harness (createStoriesSlice + hand-rolled set/get), it('setCharacterVoice on a missing castId is a no-op') and it('setCharacterVoice(rowId, newId) updates only that row'). This is where the auto-assign contract is really pinned, without mounting any React tree.
frontend/src/pages/AudiobookTab.test.jsx (new — fetch-mocked render):
it('opens CreateVoiceModal from the + New voice button').it('sets the default-voice select to the new id on onCreated')— afteronCreated({ id:'p_ab', … }), assert the<select>value updates to'p_ab'(component-localsetDefaultVoice).
frontend/src/pages/CloneDesignTab.test.jsx (new — extraction regression, the one that proves the pure refactor didn't drift; no prior suite existed):
it('renders CloneSourcePicker for the audio method and DesignControls for design')— toggle theSegmentedand assert each body mounts.it('shows the recording-in-progress and cleaning states with distinct UI')— driveisRecording/isCleaning(via the realuseRecordingwith mockedMediaRecorder) and assert the MicButton/cleaning indicator and disabledSegmented.it('applies mergeDescribedAttrs and sets describe feedback on a /design/describe success')—vi.useFakeTimers();global.fetchfor/design/describe→ the documented{ attrs, instruct, matched, unmatched }shape; type into the describe box, advance 450 ms; assert the merged attrs land in the controls,describeMatchedAnyreflectsmatched.length>0, anddescribeUnmatchedequalsunmatched.it('preserves the describe race guard — the later edit wins on out-of-order resolution')— fire two rapid describe edits, resolve the first fetch after the second; assert the second edit's attrs are the ones applied (thecancelledguard atCloneDesignTab.jsx:96,100,114discards the stale response). This is the single highest-value regression assertion for PR slice 1.it('leaves vdStates untouched and does not crash on a /design/describe failure')—global.fetchrejects; assert manual slider picks survive (thecatchno-op).
i18n / CJK (no bespoke gate; covered transitively by ci.yml:67 uv run pytest tests/)
- No allowlist edit needed. The only CJK the modal surfaces comes from
voiceInstruct.js(split delimiter) andconstants.js(ChineseDialectpicker) — both already in_ALLOWED_FILESoftests/test_no_hardcoded_cjk.py(:57,:64). The new component files (CreateVoiceModal.jsx,voice-create/*.jsx) introduce no hardcoded CJK, sotests/test_no_hardcoded_cjk.pystays green untouched. - Add the net-new keys to
frontend/src/i18n/locales/en.jsononly. NocreateVoice.*namespace exists yet (verified: 0 matches forcreateVoice/design_needs_attr/render_unavailable). Add at minimum:createVoice.title,createVoice.design_needs_attr,createVoice.render_unavailable, plus a Cancel / Create-voice label pair. Do not edit the other 20 locale files —fallbackLng: 'en'(i18n/index.ts:64) resolves missing keys, so the suite stays green and nodefaultValueliterals enter JSX. The inherited keys (recording.too_short:1860,recording.loaded_raw:1859,tts_errors.trim_hint:1693,tts_errors.ignored_unsupported:1696,tts_errors.ignored_duplicate:1697) already exist; reusedclone.*keys are unchanged. - The locale-parity probe is advisory, not gating.
tests/probe/test_probe_i18n.py:1-3states orphan-key + coverage findings are "reported, not gating," so addingen-only keys does not fail CI. (If a future hardening pass turns this gating, runscripts/translate_all.pyto backfill the 20 locales — out of scope here.)
Manual / /verify (run the app — covers the GPU/engine paths the unit suite deliberately mocks out)
The unit suite never warms a real TTS engine (design 503 is mocked). The /verify pass is where the actual engine/GPU path is exercised:
- Stories with zero profiles → cast "+ New voice" → record 3s via mic → denoise → name → Create → cast row auto-selects it; preview the line.
- Audiobook → "+ New voice" → By design → set Gender/Age controls → name → Create → default-voice selects it; "Preview plan" then a chapter preview uses it.
- Cancel/ESC/backdrop dismiss (Dialog
dismissabledefault true,Dialog.jsx:18,29-39) leaves profiles + assignments untouched. - Mic-denied path shows the per-OS
micErrorMessagehint (no crash) on macOS / Windows / Linux — confirm the hint text is the correct per-OS path on each (parity check, see § Constraints). - Failure paths: kill the backend mid-create → modal stays open, error toast (
e.message), inputs preserved; restart backend → retry succeeds. Design create with no engine warmed → real 503 toast (createVoice.render_unavailable), modal open. Upload a >15s clip →trim_hinttoast, clip rejected. Record <1s →too_shorttoast. - No-close-mid-save: spam ESC during a slow create → modal does not close until the request resolves.
- Studio regression: the main Voice studio (
CloneDesignTab) clone + design + describe + personalities + recording/cleaning states all behave identically post-extraction.
CI gates (exact jobs that must be green before merge)
bunx vitest run—.github/workflows/ci.yml:107(the only front-end test gate;package.json:14"test": "vitest run"). Run locally first per MEMORY's merge-discipline rule. There is nonode:testlegacy gate for these files —package.json:16test:legacytargets../tests/frontend, a different path not exercised by this task.uv run pytest tests/ -q—ci.yml:67. Transitively runstests/test_no_hardcoded_cjk.py(CJK gate — passes with no allowlist edit) andtests/probe/test_probe_i18n.py(i18n parity — advisory, non-gating). No new Python test is added;create_profileis unchanged.uv run pytest backend/tests/ -q—ci.yml:81(isolated backend session). Unaffected (no backend change).- CodeQL (Python + JS/TS) —
.github/workflows/security.yml:65-108. This change adds no new server-side or client-side regex over user input; the one user-input regex the modal feeds is the pre-existing linear-time character classvoiceInstruct.js:53split(/[,,]/)— nopy/polynomial-redos/ JS-ReDoS surface (see § Constraints). CodeQL stays green. - Local-loop note (MEMORY): run
bunx vitest runbefore every push; never merge before all PR checks report green (gh pr checksmonitor). The documented localtests/pytest segfault (importingmain+torch/Triton) does not affect this PR — it adds no Python test and the front-end suite is jsdom-only.
Constraints
This is a default feature (no opt-in toggle), so the strict cross-platform-parity rule applies in full. Each relevant VoiceStudio hard rule and how this spec satisfies it:
- Cross-platform parity (P0, strict rule 2026-05-20). The modal's user-visible behavior is identical on macOS, Windows, and Linux. All capture goes through the already-shipping
useRecording/getUserMediapath (useRecording.js:21-80), which is platform-agnostic in the WebView. The only OS-aware code touched ismicError.js'sdetectPlatform()→capture.mic_hint_{mac,windows,linux}— this is permitted implementation code for an OS API (the privacy-settings path differs per OS), and the behavior (an actionable mic-denied hint) is uniform across platforms; it is already unit-tested across all three platforms inmicError.test.js:20-24. No macOS-only/Windows-only feature is introduced, so nothing here needs an opt-in toggle. The too-short guard, denoise raw-fallback, trim rejection, and the full save/422/503/network FSM are all pure JS with no platform branches. The/verifymatrix above explicitly checks the per-OS mic hint on all three. - Local-first guarantee (no cloud/accounts/telemetry). The modal makes only same-origin localhost calls to the existing FastAPI backend:
POST /profiles(createProfile, multipart),POST /clean-audio(apiCleanAudio, multipart, used inuseRecording.js:48), and the optionalPOST /design/describe(JSON, already used by the design controls). No third-party endpoint, no account, no API key, no telemetry. The app remains fully functional with no network beyond127.0.0.1. No HF/GitHub/Sentry calls are added by this feature. - Backward-compatible project data (alembic for DB, lazy migration for localStorage). No DB schema change — the feature reuses the existing
voice_profilestable (db.py:39-59) and the unchangedPOST /profilescontract, so there is no alembic migration to write and existingomnivoice_data/(voices, projects, settings) keeps working untouched. No persisted client schema change — the modal owns ephemeral local React state only; the Stories assignment writes through the existingsetCharacterVoicestore action (whose persistedCastMember.profileIdshape is unchanged), and the Audiobook assignment is component-localuseState(not persisted at all). Therefore no localStorage/store lazy-migration is required, and no version key needs bumping. Duplicate profile names remain accepted (no UNIQUE constraint) — that's existing behavior, not a data-format change. - CodeQL
py/polynomial-redos(and JS ReDoS) on user-input regex. This feature adds no new regex that consumes user input on either the Python backend (unchanged) or the new frontend components. The modal's design free-text does flow into one existing regex —String(freeText||'').split(/[,,]/)atvoiceInstruct.js:53— but that is a two-element character class (ASCII,+ fullwidth CJK,), which is linear-time and not a polynomial-ReDoS pattern (no nested/overlapping quantifiers, no[^x]*with ambiguous delimiters, no\s*/.+overlap). It is unchanged by this task. So no CodeQL ReDoS gate is implicated (security.yml:65-108stays green); if the modal later needs to parse free-text differently, keep delimiters as a character class and avoid backtracking-prone quantifiers per the project's ReDoS guidance. - Localization (hard rule — no hardcoded non-English UI text; all UI via i18n
t()keys). All new user-facing strings go throught('createVoice.*')keys added tofrontend/src/i18n/locales/en.json; the 20 other locale files resolve missing keys viafallbackLng: 'en'(i18n/index.ts:64) until translated. No English (or any) literal user-facing string is placed in component code, and no per-t()defaultValuefallbacks are used (that pattern isn't used in this layer and would smuggle English literals into JSX). The only CJK the feature renders is functional (theChineseDialectpicker values inconstants.js:27-30and the fullwidth-comma split delimiter invoiceInstruct.js:53); both files are already allowlisted intests/test_no_hardcoded_cjk.py(:64,:57), so the CI CJK gate passes without any allowlist edit. The newCreateVoiceModal.jsx/CloneSourcePicker.jsx/DesignControls.jsx/MicButton.jsxfiles contribute no hardcoded CJK. - Versioning (continuous-to-main patch; no RCs, no minor/major bump). main is
0.3.6(verified in lockstep acrosspyproject.toml,frontend/src-tauri/tauri.conf.json,frontend/src-tauri/Cargo.toml). This feature does not touch any version file — it ships continuous-to-main as part of the open v0.3.x line; the owner tags a patch when the state is worth cutting. No-rctag, no codename, no "defer to v0.4." - Docs-sync (hard rule). If any
docs/**or README flow describes "how to create a voice" (voice-cloning / voice-design walkthrough), it must be updated in the same PR that adds inline create — a small note that voices can now be created inline from Stories cast rows and the Audiobook default-voice field, without leaving the editor. Stale docs are treated as bugs; the docs delta lands with PR slice 3/4 (whichever first ships a user-visible trigger), not as backlog. - GSD workflow. Implement via a GSD command (
/gsd-quickfor the small slices or/gsd-execute-phasefor the planned extraction), not raw edits, so planning artifacts stay in sync.
Dependencies
- Hard: none new. Reuses
@radix-ui/react-dialog(already behindui/Dialog.jsx),react-hot-toast,lucide-react(UserPlus/Plus),react-i18next,@tanstack/react-query(personalities fetch atCloneDesignTab.jsx:119), and the existingcreateProfile/cleanAudioAPIs. - Test-only:
@testing-library/react+vitest+ jsdom (already the harness — seeApiKeysPanel.test.jsx,EngineCompatibilityMatrix.test.jsx,storiesSlice.test.ts); no new test dependency. AQueryClientProviderwrapper (from the already-installed@tanstack/react-query) is needed in theDesignControls/CloneDesignTabtests that exercise the personalities query. - Soft / sequencing:
<VoiceSelector>(#22) is the preferred host. Recommend landing #25 standalone now (direct wiring into the two<select>sites) and, if #22 ships, moving only the trigger into the selector (see "If #22 lands first"). This unblocks the user-visible value without waiting on #22. - Internal refactor dependency: the
CloneDesignTabextraction (PR slice 1) must land before/with the modal so there is a single source of clone/design UI.
Risk
- Extraction regression in the main studio (medium): moving ~250 lines out of
CloneDesignTab.jsxrisks subtle prop/state drift (e.g. the describe-debounce effect:93-116and itscancelledrace guard:96,114, and theapplyPersonalityreset logic:125-139that fixed issue #114). Mitigation: keep sub-components fully controlled with identical prop contracts; preserve thecancelledguard verbatim; ship extraction as its own reviewed PR slice with the newCloneDesignTab.test.jsx(including the out-of-order describe-response assertion that directly exercises thecancelledguard); do a/verifypass on the studio synth flow, including the recording/cleaning states. - Design
instructshape bug (medium):buildDesignInstructreturns an object; passing it raw to FormData yields"[object Object]"→ likely a 422 (profiles.py:75) or a garbage profile. Mitigation: always append.instruct; the unit test (case d) asserts the appendedinstructis a non-empty string and.not.toBe('[object Object]'). Disable Create when.instructwould be empty (all-Auto + no valid free text), with a hint (case b). - Partial create-response misuse (medium):
createProfileis typedPromise<Profile>but the route returns only{id,name,kind}(profiles.py:125). Code that readsprofile.ref_audio/created_at/etc. off the create response will silently getundefined. Mitigation:onCreatedreads only.id(and optionally.name/.kind); unit case (n) asserts the assign uses.idand nothing else; the full record arrives via theonProfilesChanged()refresh /GET /profiles. - Design render 503 (low): the deterministic sample render (
profiles.py:91-102) can fail if no TTS engine is ready → 503 (:103-107). Mitigation: surface the toast (branch one.status === 503, unit case f), keep the modal open, don't clear inputs; the real-engine path is checked in/verify. - Profile-list staleness (low): assigning before
onProfilesChanged()resolves shows a transient empty<select>value. Mitigation:await onProfilesChanged()before calling the assign setter (it never rejects — error-swallowinguseAppData.js:105); the realtimeprofilesevent reconciles the list on socket reconnect regardless. - Too-long reference clip (low): modal v1 rejects clips over
CLONE_MAX_SECONDS(=15s) with atts_errors.trim_hinttoast instead of embeddingAudioTrimmer. Mitigation: documented follow-up; consistent with the trim guard messaging inuseTTS.js:54; unit cases (l)/(m) cover rejection and the unprobeable-accept boundary. Note: a cleaned recording that comes back >15s is also rejected — the user must re-record shorter. - Recording too short / denoise unavailable (low):
useRecordingalready handles both — a <1000-byte blob toastsrecording.too_shortand skips ingest (:38-41); a denoise failure ingests the raw webm withrecording.loaded_raw(:56-60). Mitigation: surface the existing i18n keys (both already present inen.json); unit cases (j)/(k) cover both through the real hook (mockedMediaRecorder/getUserMedia+global.fetch), not by stubbing the hook. - Close mid-save / unmount during async (low): closing while
createProfileis in flight could lose track of the result or leave the list half-refreshed; closing duringisCleaning/describe resolves into an unmounted tree. Mitigation:dismissable={!saving}blocks close during save (unit case i); a mounted-ref guard prevents post-unmountsetState;stopRecording()on close releases the mic stream. - Duplicate profile names (low): no UNIQUE constraint (
db.py:41) → two voices can share a name. Mitigation: assignment keys onid, so this is cosmetic; v1 allows it. Optional soft-warning is a follow-up. - Cast row deleted while modal open (low): the originating cast member can be removed before save;
setCharacterVoiceon a missing id is a no-op (storiesSlice.ts:83.map). Mitigation: accept (new voice still lands in the library) or capturecastIdat open and skip the assign if the row is gone — document which; the pure-store test (StoriesEditor.cast.test.ts) pins the no-op and the StoriesEditor render test asserts no throw. - Audiobook default is local-only state (low):
setDefaultVoiceis component-local (AudiobookTab.jsx:22), so the assignment is not persisted across navigation away from the tab. That matches today's behavior (the<select>is already local); no regression, but note it so reviewers don't expect persistence.
PR slices
- Refactor: extract clone/design UI — create
components/voice-create/{MicButton,CloneSourcePicker,DesignControls}.jsx, refactorCloneDesignTab.jsxto consume them; extractutils/refAudio.js(isRefAudioTooLong, accepts unprobeable durations). No behavior change; preserve the describe-effectcancelledrace guard verbatim. Tests: newCloneDesignTab.test.jsx(incl. the out-of-order describe-response race assertion + the documented/design/describeresponse shape) +refAudio.test.js+ the existingvoiceInstruct.test.jsstays green. Gate:bunx vitest run. CreateVoiceModal— new modal component owning local state +useRecording, both methods, the full state machine (validity, recording/cleaning/describe sub-states, save success/422/503/network branches, no-close-mid-save), save →createProfile(exact FormData per § API) →onCreated. NewcreateVoice.*i18n keys inen.json(nodefaultValueliterals). Tests:CreateVoiceModal.test.jsx(new), cases (a)–(n), all viaglobal.fetchmock +react-hot-toastspy.- Wire into Stories — cast-row + header + empty-state triggers;
onProfilesChangedprop fromApp.jsx:1074; auto-assign viasetCharacterVoice(row trigger) / refresh-only (header + empty-state). Handle deleted-row + refresh-failure. Docs-sync: update any "create a voice" doc flow in this PR. Tests:StoriesEditor.test.jsx(new) +StoriesEditor.cast.test.ts(pure-store, the import-lightsetCharacterVoicecontract). - Wire into Audiobook — default-voice trigger;
onProfilesChangedprop fromApp.jsx:1080; auto-assign via localsetDefaultVoice. Tests:AudiobookTab.test.jsx(new). Docs-sync note in the same PR if not already covered by slice 3. - (optional, only if #22 ships first) — collapse triggers into
<VoiceSelector>'s "+ Create new voice…" item.
Acceptance criteria
- From Stories cast (including the zero-profiles state, where the dead-end
stories.noProfileshint is replaced by an actionable CTA) a user can open a Create Voice modal, define a voice by audio (mic or upload) or by design (sliders/describe), save it, and the originating cast member is auto-assigned the new voice (setCharacterVoice(castId, profile.id)) without leaving the editor. - From Audiobook, the same modal creates a voice and auto-sets it as the default voice (local
setDefaultVoice(profile.id)). - Saving calls
createProfile()with the correctkind-specificFormData(clone:name+ref_audioBlob; design:name+kind=design+vd_statesJSON-object string+instructstring frombuildDesignInstruct(...).instruct, asserted not"[object Object]"); the success body is{id,name,kind}only andonCreatedreads no other field; the profile list refreshes viaonProfilesChangedand the new voice appears in the relevant<select>. - The main Voice studio (
CloneDesignTab) renders and behaves identically post-extraction (clone + design + describe + personalities + recording/cleaning states all unchanged, including the describe race guard and the documented/design/describeresponse handling); the newCloneDesignTab.test.jsxpasses (no prior suite existed) and its out-of-order describe-response case proves thecancelledguard survived the extraction. - All new UI text is i18n keys under
createVoice.*inen.json; no hardcoded user-facing strings; the 20 non-enlocales fall back viafallbackLng: 'en'; the CI CJK check (tests/test_no_hardcoded_cjk.py, viaci.yml:67) passes with no allowlist edit (functional CJK lives only in already-allowlistedvoiceInstruct.js/constants.js). - Default-feature parity: the modal behaves identically on macOS/Windows/Linux; the only OS-aware code is the per-OS mic-denied hint, whose behavior is uniform; no opt-in toggle is required because no platform-only feature is introduced.
- Local-first: the feature makes only same-origin localhost calls (
POST /profilesmultipart,POST /clean-audiomultipart, optionalPOST /design/describeJSON); no cloud, accounts, keys, or telemetry; app remains functional offline-except-localhost. - No data-format change: no alembic migration, no localStorage/store schema change, no version-file bump (main stays
0.3.6); existing profiles/projects continue to work. - Every failure/empty path is handled: mic-permission-denied, recording-too-short, denoise-unavailable (raw fallback), oversized clip, unprobeable clip, design
.instruct-empty (Create disabled), design-render-503, 422, and network failure each show a clear toast (or disabled control) and never crash, never silently lose user input, and (for save failures) keep the modal open with inputs preserved. Each path has a named unit test (CreateVoiceModal cases b–n). - No close mid-save: while
createProfileis in flight, ESC/backdrop/Cancel/close-button do not dismiss; they re-enable after the request resolves (unit case i). - Duplicate names are accepted without error (no UNIQUE constraint); assignment keys on
id, so a name collision never breaks assignment. - Any "create a voice" doc flow (
docs/**/README) is updated in the same PR (docs-sync). - CI green:
bunx vitest run(ci.yml:107),uv run pytest tests/andbackend/tests/(ci.yml:67,81), and CodeQL (security.yml:65-108— no new ReDoS surface) all pass; PR is not merged before all checks report green (MEMORY merge-discipline).