ca8a2e8eb8475cd8f0be554f4657d105e35fb2de
433
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
ca8a2e8eb8 |
feat(audiobook): PDF ingest for /audiobook/import (ebook-in core value) (#459)
The audiobook importer accepted .txt/.md/.epub but not PDF — the single most common "ebook in" format. Add a pure `pdf_to_chapter_script(data)` that extracts the text layer page-by-page and runs it through the existing chapterizer, so PDFs land in the same `# Heading` + body grammar EPUB and plaintext already produce (one front door onto the unchanged render pipeline). - Dep: `pypdf>=4.0` — pure-Python, MIT, zero native deps, so PDF import behaves identically on macOS/Windows/Linux (default-feature cross-platform rule). EPUB + plaintext stay stdlib-only; only PDF needs a real parser. - Robustness, surfaced as actionable 400s rather than silent empty imports: corrupt file, password-protected (empty-password decrypt attempted first), scanned/image-only (no text layer → clear "scanned PDF" message), and a page-count ceiling. A single unparseable page is skipped, not fatal. - Route: `.pdf` branch in audiobook_import; frontend accept filter + api-client doc updated to `.txt,.md,.epub,.pdf`. tests/test_longform_import.py: 5 PDF cases (extract+chapterize, no-marker single chapter, corrupt, image-only, page-cap) using a hand-built in-memory PDF — no PDF-authoring test dep, mirroring the in-memory-EPUB approach. 16 passed; frontend suite 401; CJK guard green. |
||
|
|
142b4bc25a |
feat(dub): wire second-pass timing QC into the dub editor UI (#458)
The Wave 3.3 QC backend was complete but unreachable from the UI: the
`POST /dub/qc/{job_id}` route (re-recognizes the dubbed audio, scores per-line
drift vs the target text, annotates segments with qc_drift/qc_flagged/
qc_recognized/qc_measured_start-end), the `dubQc()` API client, and the
DubSegmentRow "Verify" badge all existed — but nothing ever called the route,
so the badge never lit and the measured timings were never surfaced.
Add a "Verify dub timing" action to the dub editor header (shown once
dubStep === 'done'):
- Calls `dubQc(jobId, lang)` for the currently-previewed language.
- Merges the returned per-segment scores back onto dubSegments by id, so
flagged lines light their re-listen badge and carry the measured onsets.
- Toast summary: "{flagged} of {total} lines may need a re-listen", or a
clean-pass success when nothing drifted. Loading + error states handled;
non-destructive (generated text untouched).
i18n: dub.qc_btn / qc_running / qc_result / qc_clean / qc_failed in en.json
(fallbackLng=en covers other locales). Frontend suite green (401).
|
||
|
|
4531e999b1 |
feat(capture): opt-in LLM refinement on REST /transcribe (parity with live dictation) (#457)
The live-dictation socket (capture_ws) already runs the final transcript through the configured local LLM (disfluency/self-correction/punctuation cleanup, Wave 2.1). The REST /transcribe endpoint — the MCP / CLI / file-upload surface — only did the always-on hallucination-loop collapse, so agentic and batch callers couldn't get the same cleaned output. Add an opt-in `refine` form flag that runs the identical `maybe_refine` pipeline off-thread: - OFF by default → existing MCP/CLI callers keep raw-only output and pay no LLM latency (backward-compatible). - Honours the user's Settings → Dictation-refinement config and silently passes through when no LLM backend is configured (cross-platform default parity — identical no-op everywhere with no LLM). - Raw `text` is always returned; `refined_text` is added only when the LLM actually changed the text — same contract the socket emits. tests/test_capture_refine.py: 13 cases — flag-off no-call, refined_text on change, no-op/identical omission, and flag parsing. maybe_refine is patched at its source module since the handler imports it lazily. |
||
|
|
875f840d8e |
chore(issues): structured GitHub Issue Forms (bug / install / feature) + config (#456)
Replace the two flat markdown templates with validated YAML Issue Forms and a chooser config, so reports arrive with the diagnostic fields triage actually needs and "how do I…" traffic routes to chat instead. - `bug_report.yml` — dup-search + latest-version checkboxes; required what/repro/expected; OS / install-method / version / compute-device dropdowns (incl. ROCm + XPU); active-engine; logs (render: text) with the diagnostic- bundle + `--diagnose` tip up top. - `install_problem.yml` — NEW, for the "first-run that just works" core value: a failure-stage dropdown (launch / uv-bootstrap / model-download / engine- install / first-synth), required error + OS/install/version, and a network-conditions dropdown (proxy / restricted-region / offline) since restricted networks are a known bootstrap failure mode. - `feature_request.yml` — problem/solution/alternatives + an Area dropdown, with a local-first/cross-platform constraints note so proposals fit. - `config.yml` — `blank_issues_enabled: false`; contact links to Discord, Discussions, and the private security policy. Removes bug_report.md / feature_request.md (superseded). Forms validated (yaml parse); SECURITY.md backs the security link; CJK guard green. |
||
|
|
d20c24e1e1 |
feat(longform): two-pass loudnorm measure orchestrator + wiring (#28 slice 2) (#455)
* feat(longform): two-pass loudnorm measure orchestrator + wiring (#28 slice 2) Completes accurate ACX/podcast mastering end-to-end (builds on the pure builders from #28 slice 1). - `services/loudness.py` — `measure_loudness(ffmpeg, concat, preset, *, job_id)`: runs ffmpeg's measure pass, parses the loudnorm JSON → MeasuredLoudness. **Never raises** — skip / non-zero rc / rc None / asyncio.TimeoutError / spawn OSError / empty or unparseable stderr / silent program all WARN + return None → single-pass fallback (a slow/broken measure degrades the master, never aborts the render). Logs rc + a static message only, never the raw stderr (path-safe / local-first). UTF-8 decode with replacement (Windows-cp safe). - `_render_longform_sse` (audiobook.py): between the concat write and the mux, when `loudness` is a known preset (acx/podcast; same `.lower()`/no-strip gate as the builders) → emit a `mastering` event, measure, and pass `measured` into `build_render_cmd` (two-pass apply; `None` → single-pass). `done` gains a `loudness` block {preset, target_i, target_tp, two_pass, measured_i} ONLY for a requested preset — off/None paths keep the byte-identical legacy `done` shape. Both front doors (/audiobook + /longform/render) get it via the shared generator. Chapter cache key is deliberately untouched (loudness-agnostic → acx/off reuse the same cached WAVs; no re-render, no cache-layout break). Tests: `test_loudness.py` (14 — happy fixture, skip-without-spawn for off/ unknown/whitespace/None, non-zero/None rc, timeout-not-propagated, OSError, empty/unparseable stderr, non-UTF-8 stderr, job_id+argv forwarding) + 2 e2e cases (mastering event + done.loudness present for acx; absent for off). Orch tests run locally (stubbed run_ffmpeg, no torch); e2e on CI. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(loudness): lazy-import run_ffmpeg so the measure stub survives sys.modules purges test_loudness monkeypatched services.loudness.run_ffmpeg, but the route-shape fresh_app fixture purges services.* from sys.modules, so under the full-suite ordering the patch missed the re-imported module → real ffmpeg ran → 3 failures. Lazy-import run_ffmpeg inside measure_loudness and patch it at its source (services.ffmpeg_utils.run_ffmpeg) so the stub is always picked up at call time. Verified by running the purging suite + test_loudness together (31 pass). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
4a75d694e6 |
feat(persona): .ovsvoice manifest + SPDX + consent core (#29 / parity §R3 G1, pure) (#453)
The model-free nucleus of the portable .ovsvoice persona-bundle format: format constants, SPDX normalization, and the manifest/consent builders — all pure (no torch, no I/O), fully locally testable. The audio preview + ZIP pack/unpack + watermark `force=` param + router + frontend are follow-on slices. - Constants: OVSVOICE_FORMAT/SCHEMA_VERSION, MAX_BUNDLE_BYTES (100 MB), DEFAULT_LICENSE (`LicenseRef-OmniVoice-Personal`), the SPDX allowlist. - `normalize_spdx()` — membership + `LicenseRef-` prefix; junk/None/injection → DEFAULT_LICENSE, never raises/400s. No regex over the SPDX string (CodeQL-clean). - `build_manifest()` — mirrors the legacy `_bundle_metadata` persona fields into the manifest + format discriminator + normalized license + tags + engine / preview / members blocks. seed/vd_states pass through (None-safe; vd_states is a JSON string, never re-parsed). `BundleError(status, detail)` for the router. - `build_consent_json()` — designed-synthetic for `kind='design'`, self-recorded for an attested clone, None when nothing to attest; `recorded_at` coerced. Fields are advisory by design — real verification needs the actual consent audio member, so verified-own-voice can't be forged by editing a manifest. Tests: 14 cases (SPDX allowlist/prefix/junk/strip, manifest schema + field mirror + None-passthrough + bad-license-normalized, consent design/clone/none/ coerce). Backend pytest green; CJK guard green. Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
53c6845784 |
fix(ui): app-shell scales via zoom and always fills the viewport — permanent black-band fix (#452)
Root cause: uiScale DEFAULTS to 1.3, so the shell's `width: calc(100vw/--ui-scale)` + `transform: scale(--ui-scale)` path is active for every user. On WebKitGTK (the Linux webview) the transform wasn't magnifying the shrunk shell, so `calc(100vw/1.3)` left ~⅓ of the window black — on EVERY view, by default. (The earlier #445 fix addressed the responsive breakpoints, not this — wrong layer.) Permanent fix: scale via `zoom` and keep the shell at full `100vw × 100vh` (drop the `calc(…/scale)` shrink + the `transform`): - Chromium (mac/win): `zoom` magnifies AND fills (standard browser zoom — same mechanism the bootstrap/wizard wrappers already use). - WebKitGTK (Linux): `zoom` is a no-op → UI renders at 1.0× but the shell is a plain 100vw×100vh element → it FILLS, no band. A missed magnification now degrades to "unscaled but full", never "shrunk + black band". Regression-proofed: `src/test/appShellScale.test.js` fails CI if anyone reintroduces `width: calc(100vw/var(--ui-scale))` or `transform: scale(var(--ui-scale))` on the shell, or drops the zoom/100vw/100vh contract — so a future change can't silently bring the band back. The fix + guard are documented inline in the `.app-container` rule. Full vitest green (398, incl. the 3-case guard); typecheck:ci + vite build clean. |
||
|
|
8e3c1a8bcc |
fix(realtime): probe auth-exempt /health, not gated /model/status (#450) (#451)
The cold-start health probe added in #439 used a raw fetch() to /model/status. Raw fetch does not carry the LAN PIN / remote API-key headers that apiFetch attaches, and /model/status is not in the backend _SHELL_PATHS allowlist, so it is gated by NetworkAccessMiddleware and BearerKeyMiddleware. In LAN-share / remote-API mode the probe gets 401, rejects forever, and the realtime-events WebSocket never opens. Probe /health instead — the auth-exempt liveness endpoint (in _SHELL_PATHS) that returns 200 as soon as Uvicorn is up. Using apiUrl('/health') also avoids a double-slash when the API base has a trailing slash. Default loopback desktop use is unaffected. Adds a regression test asserting the probe targets /health (not a gated path) and only opens the WebSocket after the probe succeeds. Fixes #450 Co-authored-by: mergetest <hashduch@gmail.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
e1c8c3bc0d |
feat(longform): two-pass loudnorm builders + parser (#28 slice 1 — pure) (#449)
Groundwork for accurate ACX mastering: the pure, ffmpeg-free pieces of the
two-pass loudnorm upgrade, layered over the existing single-pass builders
(which stay). The async measure orchestrator + SSE wiring into the render path
is slice 2.
- `MeasuredLoudness` (frozen dataclass: the 5 measure-pass floats).
- `build_loudnorm_measure_filter(preset)` — first pass (+print_format=json);
mirrors build_loudnorm_filter's lookup (no strip) so the same inputs map to
"no filter".
- `parse_loudnorm_measure(stderr)` — extracts the LAST balanced {...} via a
linear brace-depth scan (NO regex → CodeQL-safe), json.loads + coerces the 5
keys to finite floats; returns None on the full failure matrix (absent/empty/
unbalanced/malformed/missing-key/non-numeric/non-finite "-inf"/array/scalar).
Rejecting "-inf" is the silent-clip path → single-pass fallback.
- `build_loudnorm_apply_filter(preset, measured)` — second pass feeding
measured_*/offset back in with linear=true; None for off/unknown OR measured
is None.
- `build_loudnorm_measure_cmd(ffmpeg, concat, filt)` — exact 16-element argv,
input segment byte-identical to build_render_cmd (measured == muxed),
portable `-f null -` sink (no /dev/null or NUL).
- `build_render_cmd` gains `measured: Optional[MeasuredLoudness] = None`: apply
two-pass when present, else single-pass; off-render still emits no -af. The
`measured=None` default keeps every existing caller + argv byte-identical.
Loudness stays opt-in (default None) → default cross-platform behavior unchanged.
Tests: 28 cases — measure-filter goldens + off/unknown/whitespace; parser
success (last-block-wins, ignores extra keys) + full failure matrix +
non-finite rejection; apply-filter golden + None cases; exact measure argv;
build_render_cmd two-pass/single-pass/off branches. Backend pytest green (71).
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
e297cbfee3 |
feat(longform): TranscriptionPicker + shared reader util (#23 slices 1–2) (#448)
Groundwork for "import from a past dictation": the shared store reader + the reusable picker modal, fully unit/RTL-tested. The two-tab wiring (Audiobook Replace/Append prompt + Stories split-panel routing) is slice 3 — deferred for visual verification. Slice 1 — shared reader (`utils/transcriptionsStore.js`): - `loadTranscriptions()` (parse + Array.isArray guard, [] on absent/empty/malformed/non-array/blocked-storage) + `TRANSCRIPTIONS_KEY` / `TRANSCRIPTION_EVENT` consts. Kills the third copy of the localStorage parse. - Refactored `Transcriptions.jsx` + `Projects.jsx` onto it (behavior-preserving; the Array.isArray guard is a superset that only hardens against corrupt blobs). Storage key/shape/200-cap unchanged → no migration. Slice 2 — `components/TranscriptionPicker.jsx`: - Controlled modal wrapping the shared `ui/Dialog` (Radix → focus trap, ESC, backdrop, ARIA inherited). Reads on open, subscribes to the add-event only while open. Per-row display normalization, hides empty-text rows, distinct empty vs empty-search states, case-insensitive `String.includes` search (no RegExp → no ReDoS surface), keyboard-activatable `<button>` rows, Invalid-Date guard. `onPick` gets the original un-normalized entry. Every string via t(). Tests: util edge matrix (3) + picker RTL (7: empty, list+hide-empty, click→onPick+onClose, keyboard rows, search filter + empty-search, bad-timestamp chip omitted, live-refresh on event). Full vitest green; typecheck:ci clean; CJK guard green (new files i18n-only). Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
4817fd1c1f |
fix: poll backend HTTP before WebSocket connect to avoid startup ECONNREFUSED (#439)
The frontend mounts faster than the Python backend (which takes ~14s to import torch/fastapi before Uvicorn starts). useRealtimeEvents was creating a WebSocket immediately, which always failed with code 1006 on the first attempt, triggering an unnecessary exponential-backoff reconnect. Fix: poll /model/status via HTTP fetch before creating the WebSocket. Once the backend responds 200, proceed to open the WS. If the health check fails, schedule a reconnect using the same backoff — but without the noisy 'closed (code=1006)' log. The /model/status endpoint is chosen because it's already polled by the TanStack Query hooks and always returns 200 once Uvicorn is running, even before models are loaded. |
||
|
|
95289b8192 |
fix(ui): scale-aware shell breakpoints — no more cramped/black layout at narrow widths (#445)
The app shell is sized `width: calc(100vw / --ui-scale)` then `transform: scale(--ui-scale)` (the WebKitGTK fix, #407), so its grid lays out against `100vw / scale`. But the responsive collapse used viewport `@media (max-width)` queries, which fire on raw `100vw` — so at any `--ui-scale ≠ 1` they trip at the wrong threshold. In a narrow window the 3-column grid was kept, the sidebar's `min 180px` crushed the main column toward 0, and the content ended up jammed into a left sliver with a black band filling the rest. Fix: drive the breakpoints off the shell's OWN width. A ResizeObserver on the app-container reads `el.clientWidth` (= the pre-transform layout width = 100vw/scale; transforms don't change the layout box) and toggles `shell-narrow` (≤1100) / `shell-mini` (≤600) classes; the `@media` queries become equivalent `.app-container.shell-*` rules. Correct on every engine and at every UI scale. Observer fires on both window resize and scale change (the calc width changes). Needs a visual check in the running app at a couple of window sizes + UI scales. Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
52eeb0b194 |
feat(longform): Story⇄Audiobook convert transforms (#24 slice 1 — pure utils) (#447)
The render-faithful interchange between the two long-form editors, as pure,
unit-tested functions (no UI/store yet — that's slice 2). The store seam for
this (convertMode/projectMode) already shipped in #31a.
- `storyToScript(tracks, cast, {projectName})` → `{script, defaultVoice,
metadata}`. Emits **profile-id** `[voice:]` tags (the backend resolver keys on
id, not display name) so the script renders identically through
/longform/render from either door. Most-used effective voice → defaultVoice
(no tag), deterministic earliest-occurrence tie-break; tags emitted only on
voice change; single-# un-indented headings; inline markup ([pause], SSML-lite,
emotion) passes through verbatim — never re-tokenized (no drift vs the backend
parser). Respects the three client/server divergences (heading depth,
[voice:default] semantics, [pause] dialect): it never synthesizes a pause and
never emits [voice:default].
- `scriptToStory(text, profiles)` → `{tracks, cast}` (persisted StoryTrack shape;
cast always ≥ a narrator clone). One physical line = one track; a leading
[voice:id] becomes the track override + a cast member (named from profiles or
the raw id, which is kept as profileId so it round-trips); mid-line markup +
body text preserved byte-for-byte; CRLF normalized; slug-collision-safe cast
ids; sequential numeric ids.
- No new regex over user input (leading-voice detection is string ops) —
CodeQL-clean; render output stays identical across both doors (the invariant).
Tests: 19 cases incl. the edge matrix + **round-trip equivalence** both
directions (script→story→script reproduces; story→script→story preserves spoken
text + voice mapping). Full vitest green; typecheck:ci clean; CJK guard green.
Deferred (slice 2, needs visual verify): the two UI buttons, store prefill
fields, mount read-clear effects, AppMode 'audiobook' fix, i18n.
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
c1e3031cfa |
feat(audiobook): use shared VoiceSelector for the default-voice picker (#22 migration 1/N) (#446)
First call-site migration onto the shared <VoiceSelector> (#22): the Audiobook default-voice <select> becomes the searchable, grouped picker. Value contract is unchanged ('' = engine default | profileId), already store-bound (#31b), so no behavior or data change — just search + clone/designed grouping. `defaultLabel` preserves the existing "engine default" row label. Stories cast / per-line track / Dub segment pickers are intricate live layouts (custom select CSS, row composition) — deferred to follow-up migrations that can be visually verified, rather than blind-swapped. vitest green (357); typecheck:ci clean; vite build clean. Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
b6f0c73f7a |
feat(audiobook): persist book metadata/script/prefs via LongformProject store (#31b) (#444)
Audiobook's script, default voice, output format, loudness, book metadata (title/author/narrator/genre/year/description) and pronunciation lexicon now bind to the unified store (#31a) instead of component useState — so they **survive a tab switch / reload** (previously all lost). The headline #31 win. - text→script, defaultVoice, format→outputFormat, loudness, meta→setProjectMeta, bound to store selectors. `meta` is default-filled so an empty record never flips a controlled input to uncontrolled. - Lexicon rows stay LOCAL (half-typed rows aren't junk-persisted); the filtered dict flushes to the store on change and hydrates back into rows on mount. - Transient state (plan, generating, progress, output, chapter previews) stays component-local — correctly NOT persisted. Deferred (noted): coverRef persistence (a File/blob can't go to localStorage); the "Save as named project" affordance + Projects-list card + App `onOpenStory` mode-aware routing (criterion 4 — re-open from Projects). This slice lands the working-state persistence (criterion 3); save/reopen is the next slice. Full frontend vitest green (357); typecheck:ci clean; CJK guard green. Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
0a72a75ed2 |
feat(ui): shared VoiceSelector component + SearchableSelect grouping (#22) (#442)
A single searchable, grouped voice picker to replace the per-tab <select>s
across Stories / Audiobook / Dub. This slice ships the COMPONENT + the two
backward-compatible SearchableSelect extensions it needs; the call-site
migrations are a follow-up slice (component lands first, tested in isolation).
- `SearchableSelect` gains two opt-in, back-compat props (the two existing
call sites are untouched, both render-identically):
- `renderGroupHeaders` (default false) — emits a `.ss-group-label` header on
the first MAIN row of each new `option.group` with a non-empty `groupLabel`
(pinned recent/popular rows never trigger one; empty groups never emit a
stray header).
- `isRecentable` (default `() => true`) — gates which committed values get
recorded as recents.
- `VoiceSelector` builds a group-ordered options array (default → fromVideo →
clone → designed → preset) over the EXISTING value contract
('' | id | preset:<id> | auto:<slug>) — byte-identical to what every call
site already sends, so project data stays compatible. Clone-vs-designed
splits on the runtime `.instruct` string (matching VoicePreview), not
`.kind`. Renders optional preview / gallery-jump / create adornments (the
component owns no audio and makes no API call — it only emits the value and
fires the parent's callbacks). A deleted-but-referenced voice renders a
"Voice not found (re-pick)" ghost row WITHOUT auto-clearing the value.
`isRecentable` excludes '' / preset: / auto: so only real voices are recents.
- i18n keys under `voiceSelector.*` (en.json; other locales fall back via
fallbackLng, matching the project's established pattern).
Tests: 9 RTL cases — grouping/headers, value contract for id/preset/auto,
from-video slug parity, ghost row (no auto-clear), recents guard (sentinels
excluded, real ids kept), preview button presence/value/loading. Full frontend
vitest green (362); typecheck:ci clean; CJK guard green.
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
83dcad5878 |
feat(store): unified LongformProject store + v4→v5 migration (#31a) (#443)
Introduces one project concept both long-form editors bind to: Stories (cast+tracks) and Audiobook (raw script + book metadata), discriminated by a `projectMode`. Store-only, no UI behavior change — Audiobook is not yet bound (its inputs still use local state; that's the #31b follow-up). Ships the data model + migration + the `convertMode` seam #24 will consume. - `storiesSlice.ts` → `longformSlice.ts`: `StoryProject` → `LongformProject` (gains mode/script/meta/lexicon/coverRef/outputFormat/loudness/defaultVoice); new working fields + actions (setScript, setProjectMeta [merge], setLexicon [replace], setOutputPrefs [merge], setCoverRef, convertMode). `loadProject` restores the FULL surface default-filled (old records never surface undefined to a controlled input); `newProject(mode?)` clears it. `SLICE_DEFAULTS` + `genProjectId` exported (the migrate fn imports genProjectId). Deprecated aliases (`StoryProject`/`StoriesSlice`/`createStoriesSlice`) re-exported so the rename breaks no import. - **Field names kept** (`storyProjects`/`storyTracks`/`cast`) so all 6 consumers and every existing localStorage blob keep working with zero change — the persisted KEY is unchanged; only the per-project SHAPE is enriched. - The project-mode working field is named **`projectMode`**, NOT `mode` — `mode` is already the app navigation field (uiSlice/AppMode); the spec's `mode` would collide (TS error + duplicate partialize key). The stored `LongformProject.mode` (nested) keeps its name. - persist `version: 4 → 5` + a `version < 5` migrate branch (the localStorage analog of an alembic upgrade): enriches each saved project with defaults (spread `...sp` last so id/name/cast/tracks/updatedAt win), drops malformed entries, never throws. v4 users see the same projects, same names/cast/tracks. Tests: ported the back-compat suite (Stories unchanged) + new coverage — default-fill on a v4-shaped record, no-stale-carryover, merge-vs-replace semantics, convertMode idempotency/guard, snapshot+restore of the new fields. Full frontend vitest green (357); typecheck:ci clean; CJK guard green (new slice scanned). No app version-file change (the persist version is the localStorage schema, not the release). Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
2a1c3eee3d |
feat(routing): synth-time no-silent-fallback gating at all TTS entry points (#21 follow-up) (#440)
Closes the last #21 gap: a per-request engine=/model= override bypasses the /engines/select host-gate, so an engine that can't use this host's GPU could still be triggered at synth time and silently fall back to CPU (or die mid- synth). Now enforced at every TTS synth entry point, reusing the SAME probe + resolver — never re-deriving routing. Shared helpers (services/engine_routing.py): - `routing_notice(result)` → (status, reason) to surface, or None. Fires for cpu_fallback (always) and accelerated-with-caveat (driver/arch); silent for cpu_only / clean-accelerated / n/a. - `header_safe_reason(reason)` → scrubbed + ASCII-sanitized (headers are latin-1; a non-ASCII device name would 500 otherwise) + ≤256 chars. No regex. Entry points: - REST `POST /generate` (generation.py): after engine resolution, resolve routing once; `unavailable` → 400; cpu_fallback / accelerated-caveat → 200 + `X-OmniVoice-Routing` + `X-OmniVoice-Routing-Reason` headers on the WAV StreamingResponse; benign → no headers. Covers OmniVoice + adapter branches. - OpenAI-compat `POST /v1/audio/speech` (openai_compat.py): same gate + same headers; the tts-1/tts-1-hd alias inherits the active engine's routing. - WebSocket `/ws/tts` (tts_stream.py): no headers → frames. `unavailable` → `{"type":"error",...}` + skip stream; cpu_fallback / caveat → one `{"type":"routing","status","reason"}` frame before any audio. - `select_engine` response now echoes routing_status / effective_device / routing_reason (PR #432 added the gate; this adds the fields so the UI can warn on a cpu_fallback pick). New fields on SelectEngineResponse. Frontend: `useTTS` reads the X-OmniVoice-Routing header and shows a one-time, non-blocking toast (in-memory de-dup by status — a 50-clip batch fires once, no localStorage). i18n keys `tts.routingFallback`/`tts.routingCaveat`. Tests: routing_notice + header_safe_reason (ASCII/length/scrub) unit tests; REST synth gate (unavailable→400, cpu_fallback→headers, cpu_only→none) via the fake-engine harness with a mocked host; select response routing fields. Deferred (small follow-up): dub-pipeline ASR routing note on the preflight_error SSE channel — separate path, not a TTS synth entry point. No frontend /ws/tts client exists today (the routing frame serves external API consumers). Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
63a4b897ee |
test(longform): real-ffmpeg + stub-TTS e2e for the chapterized renderer (#34) (#441)
#34 runtime-verify Layer 1 — the cheap regression net over the audiobook / stories convergence. Drives the REAL `_render_longform_sse` generator + REAL ffmpeg with a stub CPU-tone synth (no GPU/model), and ffprobes the muxed output. Covers happy m4b (full SSE sequence + 2 tagged chapters), mp3 container, per-chapter partial failure (chapter_error isolates ch.0, surviving chapter still muxes), total failure (error + NO file), empty plan, and the no-ffmpeg branch. Gated on ffmpeg present (skip otherwise; runs in CI). Like the other endpoint tests it imports the app+torch stack, so it's validated on CI (local pytest segfaults on the pre-existing torch/Triton import). Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
747507ff61 |
feat(ui): Engine Compatibility Matrix routing display (#21 PR 5/5) (#434)
Surfaces the /engines routing data (PR 3) in the matrix so users see the device each engine will actually use on THIS machine. - The chip matching `effective_device` is highlighted (accent ring + bold), with a "Runs on X on this machine" tooltip. - A status-toned routing badge: accelerated→success "GPU active", cpu_fallback→warn "CPU fallback" (reason in tooltip), cpu_only→neutral "CPU". The badge is SUPPRESSED for unavailable rows (the availability badge already says so) and for legacy payloads with no routing_status (renders exactly as before). An unknown/future status falls back to a neutral "Unknown" badge. - LLM rows (routing 'n/a') render a single neutral "Remote" badge instead of device chips — no false GPU claim. - types.ts: EngineBackend gains effective_device / routing_status / routing_reason; GPUTarget gains `xpu`; new EffectiveDevice + RoutingStatus unions. Corrected the stale "only TTS migrated" comment (all 3 families now emit the full shape). - i18n keys in en.json (other locales fall back to en via fallbackLng until translated — no key-parity gate). xpu chip color in the matrix CSS. Tests: 5 new RTL cases (accelerated highlight+badge, cpu_fallback badge, unavailable suppression, legacy no-badge, LLM Remote). Full frontend vitest green (350); typecheck:ci clean. Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
c6a55794da |
feat(routing): active-engine GPU verdict in preflight + diagnose (#21 PR 4/5) (#433)
Surfaces a routing verdict for the CURRENTLY-SELECTED TTS engine in the two system-health surfaces, so a CPU fallback / unavailable-GPU is heard about before a slow or failed synth — the no-silent-fallback contract, read-only. - `tts_backend.active_routing()` + `gpu_routing_verdict()`: the active engine's routing derived from list_backends() (byte-identical to the matrix) plus the host compute summary (family + VRAM from the canonical probe). Never raise. - `/system/diagnose` gains a `gpu_routing` check: accelerated→ok, accelerated-with-caveat / cpu_fallback→warn (+ actionable hint), cpu_only→ok (no-GPU host is the expected normal state — never noise-warns), unavailable→ fail, no-engine→warn. ASCII-safe detail strings (the text dump enforces ASCII). - `/setup/preflight` gains an "Active engine routing" check + an explicit `gpu_routing` object on PreflightResponse (a real field — the response has no extra="allow", so it would otherwise be dropped). `device` gains `gpu_family` (ROCm-vs-CUDA aware) + `vram_gb`. New `GpuRouting` schema. Tests: gpu_routing_verdict (host + active-engine + degraded), diagnose status mapping across all 6 states + never-raises, preflight gpu_routing object + check + device.gpu_family. Existing diagnose/preflight tests stay green (checks are additive; the report's top-level key set is unchanged). Deferred (documented): synth-time routing headers/WS-frames at the 3 synth entry points. Selection is already hard-gated (PR 3 select_engine), and the matrix (PR 5) + this preflight/diagnose verdict surface the situation — the synth-time signal is incremental belt-and-suspenders for the env-var-pinned edge and is best validated interactively. Tracked as a #21 follow-up. Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
8c8d525397 |
feat(routing): wire effective-device into /engines + select gate (#21 PR 3/5) (#432)
* feat(routing): wire effective-device + routing_status into /engines (#21 PR 3/5) Surfaces the PR-1 probe + resolver through the engine registries so the matrix UI (PR 5) and the no-silent-fallback gates can consume it. - `engine_routing.routing_fields()`: shared helper returning the three serialization-ready keys, centralizing the scrub rule — routing_reason is scrubbed via `core.scrub.scrub_text` only when truthy, so a None reason stays JSON `null` (never coerced to ""). - TTS/ASR `list_backends()` each gain `effective_device` / `routing_status` / `routing_reason`, computed from a SINGLE `detect_host_caps()` call per request (host caps are constant per process). ASR is brought to full TTS parity: it now also carries `install_hint` / `last_error` / `isolation_mode` and a SCRUBBED `reason` (closing a pre-existing ASR token-leak gap) — an identical 11-key shape across families. ASR also gains the same is_available()-raises resilience TTS has (degrade to available:false, never 500). - LLM `list_backends()` reaches 11-key parity too but emits literal `effective_device:"network"` / `routing_status:"n/a"` / `routing_reason:null` (NOT via resolve_routing — LLM runs no local GPU model). `LLMBackend.gpu_compat = ()`. "network" is a label, not a probe — nothing here touches the network. - `select_engine` host-routing gate: refuses a pick whose `routing_status` is `unavailable` on this host (400 with an actionable detail), while ALLOWING `cpu_fallback` (it runs, just slower). LLM is never gated. Defensive `.get` so legacy payloads still select. New typed `SelectEngineResponse`. Tests: 11-key shape across all 3 families, well-formed tts/asr routing keys (+ None-not-"" contract), LLM network/n/a labels, select gate (block unavailable / allow cpu_fallback / never-gate LLM). Updated the registry exact-shape test for the 3 new keys. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * test(cjk): allowlist docs/specs/ in the hardcoded-CJK guard PR #429 merged the longform design specs, which legitimately quote functional CJK (test-fixture descriptions, CosyVoice speaker IDs, multilingual sample text). The CJK guard scans every tracked file, so those docs turned main red. Specs are documentation, not shipped UI strings — allowlist the docs/specs/ prefix, matching the individually-allowlisted docs already in the set. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
e0b59f3984 |
docs(longform): implementation specs for the 14 roadmap tasks (#21–#34) (#429)
* docs(longform): implementation specs for the 14 roadmap/integration tasks (#21–#34) Per-task implementation specs under docs/specs/longform/ for the remaining longform + #346-roadmap work: GPU compat matrix, shared VoiceSelector, Transcriptions import, Story⇄Audiobook export, inline Create Voice, gallery handoff, parser unification, two-pass ACX, .ovsvoice format, Dub→Stories, unified LongformProject store, phone calls, cue-sheet, runtime-verify. Authored by a draft + iterative-refinement workflow (codebase-grounded: exact file:line anchors, API/data shapes, test plans, constraints, deps, risk, PR slices). NOTE: the 10-round refinement was cut to ~rounds 4–5 by an account session limit; rounds 5–10 (incl. the final de-bloat/polish pass) are pending — the specs carry per-round revision-note preambles that the polish round trims. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs(longform): strip accreted (this-revision) note preambles from specs Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
c3b2346759 |
fix(engines): MLX platform gate (#390) + ASR gpu_compat + IndexTTS2 (#21 PR 2/5) (#431)
Builds on the device probe from PR 1. Backend-only; the routing keys are wired into /engines in PR 3. - #390 closed: MLXAudioBackend / MLXWhisperBackend now call the shared `core.device_caps.mlx_supported()` gate FIRST, before importing the package. On Linux/Windows/mac-Intel they report unavailable and never advertise a usable `mps` route, even with a stray mlx wheel installed. Replaces the ASR backend's ad-hoc inline MPS check with the one shared rule. (The Wave-4.4 OSError/RuntimeError import-guard is preserved — it now lives behind the platform gate; its test forces the gate open so the guard stays the path under test.) - `ASRBackend` ABC gains `gpu_compat: tuple[str, ...] = ("cpu",)` mirroring TTSBackend, and each subclass declares its real targets: whisperx/faster-whisper → (cuda,cpu); mlx-whisper → (mps,cpu); pytorch-whisper → (cuda,mps,cpu); nemo/funasr → (cuda,cpu); moonshine → (cpu,). Inert until PR 3 serializes them. - IndexTTS2 declares `gpu_compat = ("cuda","cpu")` so it stops advertising the inherited CPU-only default. - ROCm is deliberately NOT claimed for any ASR engine (or for IndexTTS2): CTranslate2 has no upstream HIP build, and an unverified `rocm` claim would route ROCm hosts to a broken GPU path — strictly worse than the honest `cpu_fallback` the resolver already emits ("declares CUDA only; ROCm not in its compat set"). The per-engine TTS ROCm audit is a tracked follow-up that will verify each path before claiming it. Tests: MLX gate regression (both backends, on/off Apple), ASR gpu_compat tuples + no-false-rocm invariant, IndexTTS2 override; existing MLX import-guard test updated for the new gate ordering. Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
b62d1f5073 |
refactor(longform): share the SSE stream consumer across Stories + Audiobook (#436)
Stories and Audiobook are two authoring frontends over one server-side
renderer (_render_longform_sse), emitting the same chapter-progress events.
Both hand-rolled the identical read/decode/splitSSEBuffer/parseSSELine loop.
Extract utils/longformStream.consumeLongformStream(res, onEvent, {isAborted}):
one place owns the SSE protocol; each editor keeps only its own per-event state
handling (Stories: export %; Audiobook: {current,total,title,assembling,done}).
Behaviour unchanged — Audiobook keeps its abort check via isAborted.
The rest of the two editors stay distinct on purpose (cast/dialogue vs
manuscript/EPUB authoring), per docs/specs/2026-06-13-stories-audiobook-maturity.
Tests: frontend/src/test/longformStream.test.js (chunk-boundary parsing, abort,
no-body). Full vitest: 348 passed; typecheck:ci clean.
Co-authored-by: mergetest <test@local>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
000010ebb8 |
feat(dub): dedicated Dub home (projects/history) + project rename (#435)
The dub Projects + History rail (WorkspaceProjects/WorkspaceHistory) used to
sit beside the editor at all times. Now it's a landing: shown only when no
project is being edited (dubStep === 'idle'); opening/creating one switches to
a full-width editor. (The global Sidebar is already hidden in dub mode, so the
studio-right rail is the only surface — no Sidebar change needed.)
Adds project rename:
- backend: PATCH /projects/{id} updates just the name (400 on empty, 404 on
missing) — lighter than PUT which rewrites the whole state blob.
- api: renameProject(id, name); App.jsx renameProject handler (updates the
active-project label + refreshes the list).
- UI: inline rename on each project card (pencil → edit → Enter/Save / Esc).
Verified: PATCH create→rename→list / 400 / 404; frontend typecheck:ci clean.
Co-authored-by: mergetest <test@local>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
e61665fe34 |
feat(routing): host device probe + routing resolver (#21 PR 1/5) (#430)
* feat(routing): canonical host device probe + routing resolver (#21 PR 1/5) Foundational, backend-only slice of the GPU compatibility matrix (#21). No API or UI change — wiring lands in PRs 3–5. - `core/device_caps.py`: single source of truth for host accelerator capability. `detect_host_caps()` distinguishes ROCm from CUDA (unlike the gguf hardware_probe), never raises, makes no network call, stays kernel-free on cold start, and caches per process. Enumerates the full degradation contract (torch-unimportable→probe_ok=False, CUDA-init raises, device_count==0, multi-GPU, mem_get_info failure, arch mismatch, MPS, XPU, DirectML). Plus shared `mlx_supported()` gate (#390 groundwork) — exact-string platform check, no regex. - `services/engine_routing.py`: pure `resolve_routing(gpu_compat, caps)` → `{effective_device, routing_status, routing_reason}`; deterministic and byte-identical across OSes. Rules for accelerated / cpu_fallback (the no-silent-fallback signal) / cpu_only / unavailable, incl. the ROCm-not-in-set, DirectML-neutral, and XPU edges. - `get_best_device()` delegates its family decision to the probe so the loader and probe can never disagree; keeps the ROCm HSA env override and DirectML device-string return (probe reads, loader writes). String contract unchanged. - 39 unit tests (probe / resolver / mlx gate / reason-scrub contract); no new regex (CodeQL-clean), English-only (CJK guard green). The gguf hardware_probe rebase is a deliberate follow-up: it has its own torch-mocked suite and a VRAM-driven quant table unaffected by the family rename, so it stays out of this zero-risk slice. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(routing): address review — full available_families + empty-except comments CodeRabbit / CodeQL review on PR 1: - `available_families` no longer drops secondary accelerators on hybrid hosts (e.g. NVIDIA + Intel-iGPU-via-IPEX). The probe now detects every accelerator independently and picks `family` by priority at the end, instead of short-circuiting after the first hit. Routing is unaffected (it keys off `family`), but the field is now honest. + hybrid-host test. - Annotated every `except: pass` in device_caps with an explanatory comment (CodeQL py/empty-except). - Removed the unused `_MIN_NVIDIA_DRIVER` constant — the driver-version check stays in wizard preflight (no subprocess on the probe path); documented why. - `get_best_device()` now checks MPS before DirectML, mirroring the probe's family-priority order so loader and probe never disagree. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
dc1d36fe5f |
refactor(models): model-management v2 cleanup (mm2, all tiers) (#428)
One coherent lifecycle surface over the in-process model, diarization, and subprocess sidecars; fixes the engine-switch VRAM leak; tightens download robustness. Backend-only, response shapes preserved, no new deps. Tier 1 — correctness: - MM2-01: get_active_tts_backend() caches one instance per backend id and unload()s the outgoing engine on switch (fixes the VRAM leak behind #278); adds reset_active_backend(). - MM2-02: OmniVoiceBackend.unload() releases the shared model_manager singleton + free_vram(); SubprocessBackend.unload() -> unload_sidecar(self.id), inherited by all sidecar engines. Idempotent + preload-safe. - MM2-03: /model/loaded ASR row reports the real device + a note explaining the disabled unload button. Tier 2 — single surface: - MM2-04: new services/model_lifecycle.py owns list_loaded/unload/unload_all/ free_vram; system.py routers are thin delegations (shapes unchanged). - MM2-05: idle timeouts (in-process + sidecar) resolve via prefs.resolve (env wins, no restart); removed the duplicated _IDLE_TIMEOUT_SECONDS. Tier 3 — robustness/observability: - MM2-06: _install_cooldowns swept (1h TTL) + cleared on success — bounded. - MM2-07: per-extension weight floors (onnx 64KB, tensors 5MB) OR the original >=5MB catch — small ONNX no longer false-flagged, #352 still caught. - MM2-08: indextts GPU sidecar self-reports vram_mb in pong; parent surfaces it in list_live_sidecars (0 = CPU/unmeasured). - MM2-09: is_cached scan_cache_dir->disk fallback logs WARNING w/ exc type (#117/#118), was invisible at DEBUG. Tests: tests/test_mm2_lifecycle.py (15). Full suite: 1379 passed. Plan/summary: .planning/quick/260613-mm2-clean-model-management-v2/. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
4cc55ab852 |
Fast model downloads: Xet fast path + accurate progress (FDL W0–W2 + W4) (#424)
* feat(downloads): Xet fast path + accurate progress (FDL W0–W2)
Make model downloads fast and show accurate downloaded/remaining/speed.
Research confirmed hf-xet already implements the IDM/uGet technique
(content-defined chunking, parallel byte-range gets, dedup, resume), and
the spike found all 25 catalog repos are Xet-backed — so the win is
driving Xet well + accurate progress, not a custom downloader.
W1 — maximize + guarantee Xet:
- pin huggingface_hub>=1.7 + hf-xet>=1.1 (was transitive); no hf_transfer
- drive snapshot_download with explicit tqdm_class + max_workers + endpoint
- opt-in HF_XET_HIGH_PERFORMANCE / HDD sequential-write knobs (default off)
- /system/info reports fast_download {xet_enabled, xet_version, high_perf}
W2 — accurate progress:
- dry_run preflight -> install_plan event (exact total/cached/remaining)
- utils/download_aggregator.py: one overall bar; byte bars (by id) vs the
"Fetching N files" count bar; windowed rate; emits one 'aggregate' event
- frontend overall bar (speed/remaining/ETA), cached-skip, ⚡ fast badge
Known limit (verified live): under Xet+hf_hub 1.7.2 per-file byte bars
never advance/close via tqdm, so mid-download the bar is file-granular and
bytes flush to the exact total on completion. Classic-LFS/mirror repos get
true byte progress (W4).
Drive-by: download.py used os.walk without importing os (latent NameError
in _validate_snapshot_has_weights on every install) — fixed.
Tests: tests/backend/setup/test_download_preflight.py (10). Spike + plan
under .planning/quick/260613-fdl-fast-model-downloads/.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(downloads): opt-in mirror + cancel + docs (FDL W4)
- mirror (FDL-10): snapshot_download(endpoint=) honours prefs hf_endpoint /
env HF_ENDPOINT on preflight + download (per-call, no process-wide env).
Documented as the classic-LFS path (no Xet) for restricted networks.
- cancel (FDL-11): POST /models/install/cancel {repo_id} stops further
retries at the next boundary, emits install_cancelled, clears the cooldown
(cancel is intent, not failure). Frontend treats it as a terminator.
- docs (FDL-12): docs/downloading-models.md (Xet fast path, progress
semantics + byte-speed limitation, opt-in tuning, mirror, cancel,
troubleshooting) + README pointer. Docs-sync rule satisfied.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* docs(planning): model-management v2 cleanup plan (mm2)
GSD plan for cleaning the model-management subsystem: registry unload-on-
switch + per-engine unload() (fixes VRAM leak), model_lifecycle facade,
unified idle/timeout config, bounded cooldowns, sidecar VRAM self-report,
cache-fallback logging. Planning artifact only — no code.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(downloads): reconcile with main's HF_HUB_DISABLE_XET; honest status
Rebasing onto main surfaced that main forces HF_HUB_DISABLE_XET=1 (classic
LFS) because Xet progress bypasses the tqdm hook — the same limitation found
here. Reconcile instead of fight:
- /system/info fast_download now reports runtime truth: xet_installed +
xet_active (installed AND not HF_HUB_DISABLE_XET) + xet_enabled alias. The
⚡ badge only shows when Xet actually runs; startup log says
"downloads: Xet disabled → legacy LFS".
- complete(): clear the rate window before the final flush so crediting the
full size in one step can't emit an absurd instantaneous rate.
- docs/downloading-models.md rewritten: default is legacy LFS for accurate
progress; Xet is opt-in via HF_HUB_DISABLE_XET=0. hf-xet pin stays (ready
for a future Xet progress hook).
W2 (preflight total/remaining + aggregate bar + exact completion) is the
value on either path; W1's "maximize Xet" is dormant by main's design.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(downloads): opt-in segmented multi-connection accelerator (FDL W3)
Since main forces Xet off (HF_HUB_DISABLE_XET=1), the default path is
single-stream legacy LFS — so a segmented downloader is the way to get BOTH
parallel speed and live byte progress.
- services/segmented_download.py: async multi-connection Range downloader for
one file — parallel byte-ranges, resume (.part + manifest), per-segment
short-read truncation guard, optional sha256/etag verify, cancel, and a
single-stream fallback when the server won't range. Auth-safe: the HF
Authorization header is sent only to huggingface.co/hf.co and never
forwarded to a CDN host on redirect (unit-tested).
- dispatch (download.py): opt-in via prefs segmented_downloader / env
OMNIVOICE_SEGMENTED_DOWNLOAD (default off). When on and Xet inactive,
fetches each file into the HF cache mirroring hf_hub_download (blobs +
snapshot symlinks + refs/main), feeding real bytes to the aggregator. Any
failure falls back to snapshot_download — never breaks a correct install.
- fix: complete() was adding a full total on top of accumulated segmented
bytes (2x); now replaces byte bars so the sum is exactly total.
Verified live (accelerator on): real byte progress to ~16.6 MB/s, final
bytes==total, /models installed=True, delete frees correctly.
Tests: test_segmented_download.py (7) + aggregator double-count regression.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(downloads): relocate FDL tests to top-level; loop-isolate segmented test
CI runs the full suite, which exposed a pre-existing test-isolation leak:
several tests/backend/** fixtures purge core.*/services.* from sys.modules
under a temp OMNIVOICE_DATA_DIR and never restore, leaving core.config/core.db
bound to a dead temp dir. It only bites when collection order puts a purging
test ahead of a real-DB reader (test_longform_jobs). Adding tests under
tests/backend/setup/ reordered collection and tripped it.
Fix without touching the shared (fragile) fixtures or risking class-identity
breakage from a blanket sys.modules restore:
- move the two FDL test files to top-level tests/ (tests/test_fdl_*.py) so
tests/backend/** collection order is identical to main — longform passes.
- rewrite the segmented test to run each case under asyncio.run() (fresh loop)
instead of asyncio.get_event_loop(), which an earlier async test can leave
closed in the full suite.
Full suite green locally: 1364 passed, 0 failed.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: mergetest <test@local>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
df94af888f |
feat(longform): cohesion quick-wins — Audiobook launchpad card + Stories in Projects (#426)
* fix(ui): align Audiobook + Stories controls to the design tokens The hand-written tabs used a bare `.btn` class (which has NO CSS rule → bright white browser-default buttons) and an unstyled `.field-label`, so the buttons, labels, and selects looked off-theme. (Other tabs use the Button/ui-btn system, which is why only these looked wrong.) Found via a design-token audit workflow. AudiobookTab: - Import / Preview plan / Add word / Add cover / Download → `ui-btn ui-btn--subtle`; Create → `ui-btn ui-btn--primary`; cover-remove / lexicon-remove / chapter-play → `ui-btn ui-btn--icon` (the app's themed button variants from ui/Button.css). - AudiobookTab.css: define `.audiobook-tab .field-label` (chrome mono/uppercase via --chrome-* tokens) + header serif title / muted subtitle (--font-serif, --text-xl, --color-fg/-muted). Selects/inputs already used `.input-base` (the canonical chrome look) — left as-is. StoriesEditor: - Format `<select>` now uses `.input-base` (canonical chrome select + arrow); trimmed the bespoke `.stories-editor__format` rule to just the toolbar sizing. Build clean; 345 frontend tests green. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(longform): cohesion quick-wins — Audiobook launchpad card + Stories in Projects Make Stories/Audiobook feel wired into the app (integration-map plan, quick-win tier): - Launchpad: an Audiobook ActionCard (was NavRail-only; Stories already had one). - Projects/OmniDrive: saved Stories projects now appear as a "Stories" category (line + voice counts) and open via onOpenStory → loadProject + setMode('stories'), mirroring onOpenDub. App.jsx reads storyProjects/loadProject from storiesSlice. - Live profile sync (QW1) confirmed already working: both tabs map the `profiles` prop in render (no mount snapshot), so a voice cloned/designed/imported anywhere shows up live in the cast/default pickers — no code needed. Deferred (no trigger yet): QW4 create-voice handoff to these tabs needs an inline create/gallery "use here" affordance first (QW3/M3). Build clean; 345 frontend tests green; en.json valid. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
8bd2f149a0 |
fix(ui): align Audiobook + Stories controls to the design tokens (#425)
The hand-written tabs used a bare `.btn` class (which has NO CSS rule → bright white browser-default buttons) and an unstyled `.field-label`, so the buttons, labels, and selects looked off-theme. (Other tabs use the Button/ui-btn system, which is why only these looked wrong.) Found via a design-token audit workflow. AudiobookTab: - Import / Preview plan / Add word / Add cover / Download → `ui-btn ui-btn--subtle`; Create → `ui-btn ui-btn--primary`; cover-remove / lexicon-remove / chapter-play → `ui-btn ui-btn--icon` (the app's themed button variants from ui/Button.css). - AudiobookTab.css: define `.audiobook-tab .field-label` (chrome mono/uppercase via --chrome-* tokens) + header serif title / muted subtitle (--font-serif, --text-xl, --color-fg/-muted). Selects/inputs already used `.input-base` (the canonical chrome look) — left as-is. StoriesEditor: - Format `<select>` now uses `.input-base` (canonical chrome select + arrow); trimmed the bespoke `.stories-editor__format` rule to just the toolbar sizing. Build clean; 345 frontend tests green. Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
e196d790cc |
fix(longform): evict oldest chapters from the render cache (review fast-follow) (#423)
The content-addressed longform_cache/ accumulated uncompressed chapter WAVs across every render with no bound (a review finding). Add prune_cache_dir() — LRU-by-mtime eviction down to a 2 GB ceiling (OMNIVOICE_LONGFORM_CACHE_MAX_GB); best-effort, never raises. Called at the start of each render job, before its chapters are written, so the fresh ones are never the eviction target. Tests: under-cap no-op, evicts-oldest-keeps-newest, missing-dir safe. 38 green. Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
0c761ee991 |
feat(audiobook): pronunciation editor + markup reference UI (#422)
Makes the lexicon backend (#419) and SSML-lite markup (#421) usable from the tab. - A "Pronunciation" editor in the full-width side pane: add/remove {word → say it as…} rows, compiled to a lexicon dict sent with both the full render and per-chapter preview (so previews match the final output). - A collapsible "Markup reference" listing the script syntax (# chapter, [voice:], [pause], [slow]/[fast]/[emphasis]/[spell]). - api/audiobook.ts: lexicon field on the generate + preview bodies. Build clean; 345 frontend tests green. Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
62b2a6fab9 |
feat(longform): SSML-lite prosody markup — [slow]/[fast]/[emphasis]/[spell] (PR 8b) (#421)
Inline delivery hints within a narration line, wired into BOTH front doors so
Audiobook and Stories behave identically.
- services/ssml_lite.py (parallel-built, 18 tests): parse_ssml_lite splits a
line into {text, speed, spell, emphasis} segments — nesting (innermost wins),
unclosed-to-EOL, stray-close ignored, adjacent-merge; ReDoS-safe literal
alternation. + spell_out().
- _parse_spans (audiobook script path) now applies SSML-lite as the innermost
layer (precedence: [voice:] → [pause] → SSML); each segment becomes a Span
with its speed (threaded to the renderer) and spelled-out text for [spell].
Trailing pause attaches to the run's last segment.
- frontend/src/utils/ssmlLite.js: client port (kept in sync with the .py) +
storyToSpans applies it per chunk — inline speed OVERRIDES the per-line slider,
falls back to it otherwise.
Tests: parse_ssml_lite (18 py + 10 js), script-level prosody parse, Stories
SSML compile (override + spell). 70 backend + 345 frontend green; build clean.
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
8555c510b8 |
fix(ui): full-width/height Audiobook + Stories layouts (match other tabs) (#420)
Both tabs rendered as narrow centered columns (Audiobook maxWidth:860, Stories max-width:1040 margin-auto) while the rest of the studio is full-bleed. - AudiobookTab: rebuilt into a full-height two-pane layout (new AudiobookTab.css) — header with the action buttons, a left script editor that grows to fill the window height, and a right settings+results pane (voice/format/loudness, cover & metadata, progress/output/plan) that scrolls independently. Collapses to one column under 900px. Removed the inline 860px cap. - StoriesEditor: dropped the `max-width:1040px; margin-inline:auto` cap → fills edge-to-edge like the dub/projects/transcripts tabs. Build clean; 334 frontend tests green. Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
dde43de5a4 |
feat(longform): pronunciation lexicon — per-render word respelling (PR 8a) (#419)
Lets a render correct hard-to-say words (e.g. {"GIF":"jiff","Dr":"Doctor"}).
Backend wiring; the editor UI folds into the full-width Audiobook redesign.
- services/pronunciation.py (parallel-built, 19 tests): apply_lexicon —
whole-word, case-insensitive, longest-first, word-boundary, single ReDoS-safe
re.sub pass; + normalize/load/save_lexicon (JSON).
- synthesize_chapter gains a `lexicon` kwarg, applied to each span's text before
chunk splitting (None/empty = no-op → backward compatible).
- _render_chapter_cached folds the normalized lexicon into the chapter cache key
(a lexicon edit re-renders); threaded through _render_longform_sse + the
/audiobook, /audiobook/preview, /longform/render request models.
Tests: synthesize_chapter respells via lexicon; pronunciation module (19);
75 related backend tests green.
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
18e4c2347a |
fix(longform): correctness + robustness fixes from adversarial review (#418)
* fix(longform): correctness + robustness fixes from adversarial review Fixes the confirmed findings from a multi-agent review of the convergence: HIGH (correctness/output): - MP3 + cover produced a corrupt file (-map 2:v -c:v copy is invalid for mp3). Cover art is now embedded for M4B only; mp3 skips it (m4b is the cover format). - Chapter cache key omitted ref_text — editing only a profile's ref_text served stale audio. ref_text is now part of the voice signature. - Preview wrote audiobook_cache/ but the render reads longform_cache/ (rename missed in PR 5) → cache-warming silently broke. Unified to longform_cache/. Robustness (DoS/OOM guards): - /audiobook/import caps upload at 64 MB; epub_to_chapter_script bounds per-entry (25 MB) and cumulative (300 MB) uncompressed reads (zip-bomb guard). - /longform/render rejects > 10,000 chapters (422). Frontend leaks: - StoriesEditor.removeTrack revokes the line's preview blob URL. - AudiobookTab revokes the cover blob URL on replace/unmount. Deferred fast-follows (also from review): render-cache disk eviction; restoring the standalone chapter cue-sheet export (needs chapter times in the done event). Tests: mp3-drops-cover, epub entry/total caps, import + chapter-count limits; updated the cache-hit test for the 4-field voice sig. 70 backend + 334 frontend green. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * test(longform): pass EPUB caps as params, not monkeypatch (CI import-path fix) The cap tests monkeypatched module constants, but in the full-suite CI context the module loads under a different import path so the patch missed the function (it used the real 300 MB cap → tests failed). epub_to_chapter_script now takes max_entry_bytes/max_total_bytes kwargs (default to the constants); tests pass small values directly — deterministic regardless of import path. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
36e7fb12fc |
feat(longform): job library — finished books/stories in Projects (PR 7/8) (#417)
Surfaces finished Audiobook + Story renders so they're re-downloadable from the Projects view — closing the resume/history loop of the convergence. Backend (new, no migration — reads existing job_store rows): - routers/longform_jobs.py: GET /longform/jobs lists finished audiobook/story jobs newest-first, recovering output/chapters/duration from each job's persisted 'done' SSE event. Pure build_longform_library() over the job_store callables; defensive (skips unparseable jobs, never 500s). Registered in main.py. Frontend: - Projects.jsx: new "Audiobooks" category fed by /longform/jobs; each row opens the rendered file (/audio/<output>) with type/chapters/duration. Offline-safe (empty on fetch failure). en.json keys added. Built via parallel worktree agent; backend tests/test_longform_jobs.py (9) green; 334 frontend tests + build clean. Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
00f400e4c7 |
fix(stories): thread per-line speed through the shared renderer (PR 6/8) (#416)
PR 5 moved Stories' full export to /longform/render but dropped per-line **speed** — the old client export sent each line's speed to /generate; the converged path silently ignored it. This restores it end-to-end. - Span gains an optional `speed`; synthesize_chapter passes it to the injected synth (signature now `synth(text, voice_id, speed)`); both engine paths (OmniVoice model + generic TTSBackend) forward it to generate(speed=…). - chapter_cache_key now includes speed (a speed change re-renders; tuples accept an optional 4th element so existing 3-tuple callers/tests still work). - LongformSpan + /longform/render carry speed; storyToSpans emits each line's speed onto its spans. Emotion note: per-line tone is already model-native via inline tags ([laughter] etc.) inserted into the text, so no separate emotion→instruct plumbing is needed — the dead `emotion` store field stays unused/superseded. Tests: storyToSpans speed passthrough (8); cache-key speed sensitivity; synth stubs updated for the 3-arg signature. 65 backend + 334 frontend green; build clean. Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
0f67895585 |
feat(stories): full export → shared server-side renderer (PR 5/8) (#413)
* feat(stories): full export → shared server-side renderer (PR 5/8)
The convergence core. Stories' full export no longer stitches audio in the
browser (Web Audio, capped by RAM, no resume/loudness/markers) — it compiles
cast + lines into a chapter/span plan and streams through the same chapterized
renderer the Audiobook tab uses.
Backend:
- Extracted the audiobook SSE job into a shared `_render_longform_sse(plan, …)`
generator (resume cache, per-chapter fault isolation, mux). /audiobook is now
a thin caller.
- New POST /longform/render — accepts a pre-built {chapters:[{title,spans:
[{voice_id,text,pause_ms_after}]}]} plan (+ format/loudness/cover/metadata) and
renders it. Pause-only spans (empty text) are kept as silence. job_type=story.
- Shared content-addressed cache renamed longform_cache (one render per unique
chapter across both front doors).
Frontend:
- storyToSpans(tracks, cast) — pure compiler: `# ` lines → chapters; each line
resolves its cast/override voice; inline [voice:]/[pause] split into spans;
pauses fold into the previous span.
- StoriesEditor.generateAll now posts via longformRender and downloads the
server file (chaptered M4B / MP3). Single-line preview stays client-side;
stems export unchanged. Format select WAV→M4B.
Deferred to PR 6 (with the component split): per-line regenerate, emotion→instruct.
Tests: storyToSpans (7) — cast resolution, chapters, per-line + inline voice,
pause folding, empty-drop. 64 backend + 333 frontend green; build clean.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(audiobook): confine cover_path to OUTPUTS_DIR + don't leak exception text (CodeQL)
- _safe_cover_path() restricts the user-supplied cover to OUTPUTS_DIR before it
reaches ffmpeg (py/path-injection).
- SSE error events now emit a generic message and log the detail server-side
(py/stack-trace-exposure); empty best-effort excepts annotated.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(audiobook): cover path via basename+fixed dir (clears CodeQL py/path-injection)
CodeQL didn't recognize realpath+startswith as a barrier; os.path.basename is a
recognized sanitizer. Covers only come from /audiobook/cover (OUTPUTS_DIR/
audiobook_covers), so rebuilding from the basename onto that fixed dir is both
CodeQL-clean and strictly tighter — no caller path can escape it.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(audiobook): regex-allowlist cover filename (clears CodeQL py/path-injection)
basename alone wasn't a barrier CodeQL credits. Restrict the cover name to the
exact pattern /audiobook/cover emits (12 hex + jpg/jpeg/png) before joining onto
the fixed covers dir — an anchored-regex guard CodeQL recognizes as sanitizing,
and strictly tighter than before.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(audiobook): commonpath-confine resolved cover path (CodeQL py/path-injection)
Add an os.path.realpath + os.path.commonpath containment check on the resolved
cover path (the barrier static analysis recognizes), on top of the regex
allowlist + basename. Defense in depth; the path provably cannot escape
OUTPUTS_DIR/audiobook_covers.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
d674084510 |
fix(ci): make Docker Hub description sync non-fatal (#414)
The image build+push succeeds, but the "Update Docker Hub description" step 403s (Forbidden) — DOCKERHUB_TOKEN can push yet lacks description-edit scope, a common limitation of fine-grained Docker Hub tokens. That cosmetic overview sync was failing the whole Docker (GHCR) run on main. Mark the step continue-on-error so a creds-scope mismatch no longer reds-out an otherwise-successful build. To actually sync the overview, the token needs read/write (incl. description) scope, or use the account password. Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
ea6833138b |
feat(audiobook): text + EPUB import → auto-chapter (PR 4/8) (#412)
Spec PR 4. A front door onto the existing chapter parser: import a file, get a
chapter-delimited script in the editor.
Backend (new services/longform_import.py — pure, stdlib only, no new dep):
- chapterize_plaintext(text): inserts `# ` headings ahead of short standalone
chapter-title lines (Chapter/Part/Prologue/…); no-op if the text already has
H1s; long "Chapter …" sentences stay prose. ReDoS-safe (anchored, per-line).
- epub_to_chapter_script(bytes): parses EPUB (zipfile + ElementTree +
html.parser) in spine order → `# Title` + stripped body per document; skips
empty/nav pages; the heading becomes the chapter title (not narrated). Raises
ValueError on a malformed EPUB. ET.fromstring annotated `# nosec B314` (local
user file, no external-entity expansion).
- POST /audiobook/import (UploadFile) → {text, chapters}.
Frontend: an Import button (.txt/.md/.epub) that fills the script editor.
Tests: tests/test_longform_import.py (9) incl. an in-memory synthetic EPUB
(spine order, empty-doc skip, tag stripping, bad-zip). 64 backend + 326 frontend
green; build clean; en.json valid.
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
bd62659b9f |
docs(docker): maintain Docker Hub overview in-repo + auto-sync on main (#410)
The hub.docker.com/r/palashdeb/omnivoice-studio overview was managed by hand and had gone stale (stuck at the sha-f86beb0 era, missing the tag table, audiobook/long-form, Supertonic-3, server-mode networking notes). Add deploy/dockerhub-overview.md as the source of truth and a peter-evans/dockerhub-description step in docker.yml that pushes it to Docker Hub on main pushes. Gated identically to the image push: only when DOCKERHUB_TOKEN is set, so forks / GHCR-only runs are unaffected. Overview adds the :latest=preview / :stable=release tag semantics (matching docs/install/docker.md), the current feature set, server-mode + LAN networking notes, and shields badges. Short description is 98/100 chars. Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
7af5143fac |
feat(audiobook): per-chapter preview + resume + chapter fault-isolation (PR 3/8) (#411)
* feat(audiobook): per-chapter preview + resume + chapter fault-isolation (PR 3/8) Builds on the shared core (#408) and metadata UI (#409). Chapter-level control, the spec's PR 3. Shared core: - chapter_cache_key(spans, sr, engine_id, voice_sig) — deterministic content hash of a chapter's audio inputs. Same inputs → reuse; any change (text, voice, order, pauses, sr, engine, resolved-voice signature) → re-render. Backend (audiobook router): - Chapter WAVs are now content-addressed in OUTPUTS_DIR/audiobook_cache. A re-run after a failure/interruption reuses already-rendered chapters and only synthesizes the missing/changed ones (resume). Job emits `cached` per chapter and `cached_chapters`/`failed_chapters` on done. - Per-chapter fault isolation: a chapter that throws emits `chapter_error` and the job continues; the m4b assembles from the successful chapters. Re-running retries only the failed (un-cached) chapters. - POST /audiobook/preview — render a single chapter to audition it; shares the same cache so a preview warms the full run and a re-preview is instant. - _build_synth now exposes resolve + engine_id; _prepare_synth unifies the omnivoice/generic paths for both the job and preview. Frontend: - Plan view: a ▶ preview button per chapter with inline playback. - Done panel: "reused N chapters" + "N failed — click Create to retry" notes. Tests: chapter_cache_key determinism + sensitivity (8); preview validation + cache-hit-skips-synth (3). 55 backend + 326 frontend green; build clean. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(audiobook): mark cache-key SHA1 usedforsecurity=False (bandit B324) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
086ac08592 |
feat(audiobook): metadata, cover art, format + loudness UI (PR 2/8) (#409)
Surfaces the shared-render-core capabilities (PR 1, #408) in the Audiobook tab. Backend: - POST /audiobook/cover — multipart cover upload (jpg/png, 8 MB cap), returns a server-side path passed back as cover_path. Unit-tested via the handler directly (no main+torch import). Frontend: - api/audiobook.ts: AudiobookGenerateBody (format/loudness/cover_path/metadata) + audiobookUploadCover(file). - AudiobookTab: format select (M4B/MP3), loudness select (off/ACX/podcast, default off), and a "Cover & details" panel — cover picker with preview + title/author/narrator/year/genre/description. On create, the cover uploads first, then the job runs with metadata + format + loudness. - en.json: audiobook.* keys for the new controls. Tests: tests/test_audiobook_cover.py (4) green; frontend vitest 326 green; prod build clean; CJK + i18n-parity gates pass. Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
e9481ef307 |
feat(longform): shared render core — loudness, metadata, cover art (PR 1/8) (#408)
First slice of the Stories+Audiobook convergence (spec: docs/specs/2026-06-13-stories-audiobook-maturity.md). Both features will compile to one server-side chapterized renderer; this lands the shared pure builders and wires them behind Audiobook. New `backend/services/longform_render.py` (all pure, unit-tested without ffmpeg/torch): - build_ffmetadata(chapters, global_meta) — FFMETADATA1 with an optional global tag block (title/author→artist/narrator→composer/year→date/genre/description→ comment) + chapter table. - build_loudnorm_filter(preset) — `-af loudnorm` for ACX (~-19 LUFS, -3 dBTP) or podcast (-16 LUFS); off/unknown → None. Opt-in, so default behavior stays platform-identical. - validate_cover_image — jpg/png + 8 MB cap guard. - build_render_cmd — generalizes the m4b mux: m4b|mp3, optional cover (attached_pic) + loudness, bitrate validated. - build_concat_list — moved here. `services/audiobook.py`: build_chapter_ffmetadata / build_m4b_cmd / build_concat_list are now backward-compatible wrappers over the core (existing imports + tests unchanged). `POST /audiobook`: now accepts optional `format` (m4b|mp3), `loudness`, `cover_path`, and `metadata` and passes them through — backend-complete; the UI for these lands in PR 2. Tests: tests/test_longform_render.py (28) + existing test_audiobook.py (11) green. Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
f86beb041c |
fix(ui): UI scale via transform:scale, not zoom — fixes WebKitGTK black bands (#407)
CSS `zoom` is honoured by Chromium (the macOS/Windows webview) but IGNORED by WebKitGTK (the Linux webview). The shell sized itself to `100vw/scale` × `100vh/scale` expecting `zoom` to magnify it back to full size; on Linux the magnification never happened, so at the default uiScale of 1.3 the whole app rendered at 1/1.3 ≈ 77% of the window, leaving black bands on the right and bottom (a cross-platform default-parity P0 — 1.3 ships out of the box). Switch to `transform: scale(var(--ui-scale))` + `transform-origin: top left`, which scales identically on every engine and doesn't alter how vw/vh resolve, so `declared (100vw/scale) × scale` fills the viewport exactly. Drop the inline `zoom` (keep setting the `--ui-scale` CSS var the transform reads). Verified on the real WebKitGTK webview (Tauri debug build, localStorage uiScale=1.3): shell now fills edge-to-edge — header, content, and logs footer all reach the window edges; no black bands. Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
599f3bcc5c |
feat(engines): on-demand unload of subprocess-engine sidecars (Action 13) (#406)
Completes the dynamic engine load/unload slice. The idle reaper (#401) frees sidecar VRAM after 5 min; this adds a user-initiated "free VRAM now" path so multi-engine users don't have to wait: - subprocess_backend: `list_live_sidecars()`, `unload_sidecar(id)`, `unload_all_sidecars()` via a shared `_force_reap(predicate)` — busy-guarded exactly like the idle reaper (non-blocking lock; a sidecar mid-synth is skipped, never interrupted; next request respawns it). - system.py: `/model/loaded` now surfaces live sidecars as unloadable rows; `/model/unload/{sidecar:<id>|sidecars}` frees one or all. The existing generic flush panel picks these up with zero frontend change. Also refresh CLAUDE.md stale version notes: main is 0.3.6 (latest release v0.3.5 + 1 patch); the v0.3.0-as-unreleased framing in the project/cadence notes is corrected to the v0.3.x continuous-to-main reality. Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
6704d062fc |
fix(persona): preserve design kind + vd_states across share/import (Wave 5 §R3) (#405)
The persona-gallery surface already exists (VoiceGallery Community zone + community.py manifest + marketplace .omnivoice bundles). The blocker for §R3's 'synthetic-only' gate was data integrity: a *designed* persona lost its kind='design' (and vd_states) when imported from the community gallery or round-tripped through a bundle — silently demoting it to a clone. - community.py /use: a 'preset' (rendered from instruct) imports as kind='design'; a 'voice' (real reference clip) as 'clone'. - marketplace.py: extract a pure _bundle_metadata() (dedupes export+publish) that captures kind + vd_states; import restores them. Old bundles without the keys import as 'clone' (backward-compatible). This makes 'accept only designed/synthetic voices' enforceable instead of everything defaulting to clone. No new persona-gallery feature was built — that would duplicate the existing community/marketplace surface. 4 torch-free tests (isolated DB): _bundle_metadata captures design + defaults to clone; import round-trip preserves design kind+vd_states; legacy bundle → clone. docs §R3 status updated. Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
151f73f794 |
feat(audiobook): Audiobook tab — script → plan → m4b (Wave 5 UI) (#404)
Frontend for the audiobook backend (#402/#403): a dedicated Audiobook tab. - pages/AudiobookTab.jsx: script textarea + default-voice picker (reuses the app's profiles), 'Preview plan' (POST /audiobook/plan → chapter list) and 'Create' (POST /audiobook → reads the SSE stream, shows per-chapter progress + assembling, then an <audio> player + m4b download via the /audio mount). - api/audiobook.ts: typed plan() + generate() (returns the raw streaming Response). - utils/sseParse.js: pure splitSSEBuffer/parseSSELine helpers for reading the POST event-stream (EventSource is GET-only) — unit-tested (the buffer/line handling is the easy thing to get subtly wrong). - NavRail + App.jsx wiring (lazy tab, hideSidebar); i18n keys in en.json. All strings via i18n (CJK gate green). 7 new SSE tests; full vitest 326 + vite build green. Runtime-unverifiable here (Tauri webview) — wants an in-app pass. docs §R3 updated. Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |