ca8a2e8eb8475cd8f0be554f4657d105e35fb2de
130
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
ca8a2e8eb8 |
feat(audiobook): PDF ingest for /audiobook/import (ebook-in core value) (#459)
The audiobook importer accepted .txt/.md/.epub but not PDF — the single most common "ebook in" format. Add a pure `pdf_to_chapter_script(data)` that extracts the text layer page-by-page and runs it through the existing chapterizer, so PDFs land in the same `# Heading` + body grammar EPUB and plaintext already produce (one front door onto the unchanged render pipeline). - Dep: `pypdf>=4.0` — pure-Python, MIT, zero native deps, so PDF import behaves identically on macOS/Windows/Linux (default-feature cross-platform rule). EPUB + plaintext stay stdlib-only; only PDF needs a real parser. - Robustness, surfaced as actionable 400s rather than silent empty imports: corrupt file, password-protected (empty-password decrypt attempted first), scanned/image-only (no text layer → clear "scanned PDF" message), and a page-count ceiling. A single unparseable page is skipped, not fatal. - Route: `.pdf` branch in audiobook_import; frontend accept filter + api-client doc updated to `.txt,.md,.epub,.pdf`. tests/test_longform_import.py: 5 PDF cases (extract+chapterize, no-marker single chapter, corrupt, image-only, page-cap) using a hand-built in-memory PDF — no PDF-authoring test dep, mirroring the in-memory-EPUB approach. 16 passed; frontend suite 401; CJK guard green. |
||
|
|
4531e999b1 |
feat(capture): opt-in LLM refinement on REST /transcribe (parity with live dictation) (#457)
The live-dictation socket (capture_ws) already runs the final transcript through the configured local LLM (disfluency/self-correction/punctuation cleanup, Wave 2.1). The REST /transcribe endpoint — the MCP / CLI / file-upload surface — only did the always-on hallucination-loop collapse, so agentic and batch callers couldn't get the same cleaned output. Add an opt-in `refine` form flag that runs the identical `maybe_refine` pipeline off-thread: - OFF by default → existing MCP/CLI callers keep raw-only output and pay no LLM latency (backward-compatible). - Honours the user's Settings → Dictation-refinement config and silently passes through when no LLM backend is configured (cross-platform default parity — identical no-op everywhere with no LLM). - Raw `text` is always returned; `refined_text` is added only when the LLM actually changed the text — same contract the socket emits. tests/test_capture_refine.py: 13 cases — flag-off no-call, refined_text on change, no-op/identical omission, and flag parsing. maybe_refine is patched at its source module since the handler imports it lazily. |
||
|
|
d20c24e1e1 |
feat(longform): two-pass loudnorm measure orchestrator + wiring (#28 slice 2) (#455)
* feat(longform): two-pass loudnorm measure orchestrator + wiring (#28 slice 2) Completes accurate ACX/podcast mastering end-to-end (builds on the pure builders from #28 slice 1). - `services/loudness.py` — `measure_loudness(ffmpeg, concat, preset, *, job_id)`: runs ffmpeg's measure pass, parses the loudnorm JSON → MeasuredLoudness. **Never raises** — skip / non-zero rc / rc None / asyncio.TimeoutError / spawn OSError / empty or unparseable stderr / silent program all WARN + return None → single-pass fallback (a slow/broken measure degrades the master, never aborts the render). Logs rc + a static message only, never the raw stderr (path-safe / local-first). UTF-8 decode with replacement (Windows-cp safe). - `_render_longform_sse` (audiobook.py): between the concat write and the mux, when `loudness` is a known preset (acx/podcast; same `.lower()`/no-strip gate as the builders) → emit a `mastering` event, measure, and pass `measured` into `build_render_cmd` (two-pass apply; `None` → single-pass). `done` gains a `loudness` block {preset, target_i, target_tp, two_pass, measured_i} ONLY for a requested preset — off/None paths keep the byte-identical legacy `done` shape. Both front doors (/audiobook + /longform/render) get it via the shared generator. Chapter cache key is deliberately untouched (loudness-agnostic → acx/off reuse the same cached WAVs; no re-render, no cache-layout break). Tests: `test_loudness.py` (14 — happy fixture, skip-without-spawn for off/ unknown/whitespace/None, non-zero/None rc, timeout-not-propagated, OSError, empty/unparseable stderr, non-UTF-8 stderr, job_id+argv forwarding) + 2 e2e cases (mastering event + done.loudness present for acx; absent for off). Orch tests run locally (stubbed run_ffmpeg, no torch); e2e on CI. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(loudness): lazy-import run_ffmpeg so the measure stub survives sys.modules purges test_loudness monkeypatched services.loudness.run_ffmpeg, but the route-shape fresh_app fixture purges services.* from sys.modules, so under the full-suite ordering the patch missed the re-imported module → real ffmpeg ran → 3 failures. Lazy-import run_ffmpeg inside measure_loudness and patch it at its source (services.ffmpeg_utils.run_ffmpeg) so the stub is always picked up at call time. Verified by running the purging suite + test_loudness together (31 pass). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
2a1c3eee3d |
feat(routing): synth-time no-silent-fallback gating at all TTS entry points (#21 follow-up) (#440)
Closes the last #21 gap: a per-request engine=/model= override bypasses the /engines/select host-gate, so an engine that can't use this host's GPU could still be triggered at synth time and silently fall back to CPU (or die mid- synth). Now enforced at every TTS synth entry point, reusing the SAME probe + resolver — never re-deriving routing. Shared helpers (services/engine_routing.py): - `routing_notice(result)` → (status, reason) to surface, or None. Fires for cpu_fallback (always) and accelerated-with-caveat (driver/arch); silent for cpu_only / clean-accelerated / n/a. - `header_safe_reason(reason)` → scrubbed + ASCII-sanitized (headers are latin-1; a non-ASCII device name would 500 otherwise) + ≤256 chars. No regex. Entry points: - REST `POST /generate` (generation.py): after engine resolution, resolve routing once; `unavailable` → 400; cpu_fallback / accelerated-caveat → 200 + `X-OmniVoice-Routing` + `X-OmniVoice-Routing-Reason` headers on the WAV StreamingResponse; benign → no headers. Covers OmniVoice + adapter branches. - OpenAI-compat `POST /v1/audio/speech` (openai_compat.py): same gate + same headers; the tts-1/tts-1-hd alias inherits the active engine's routing. - WebSocket `/ws/tts` (tts_stream.py): no headers → frames. `unavailable` → `{"type":"error",...}` + skip stream; cpu_fallback / caveat → one `{"type":"routing","status","reason"}` frame before any audio. - `select_engine` response now echoes routing_status / effective_device / routing_reason (PR #432 added the gate; this adds the fields so the UI can warn on a cpu_fallback pick). New fields on SelectEngineResponse. Frontend: `useTTS` reads the X-OmniVoice-Routing header and shows a one-time, non-blocking toast (in-memory de-dup by status — a 50-clip batch fires once, no localStorage). i18n keys `tts.routingFallback`/`tts.routingCaveat`. Tests: routing_notice + header_safe_reason (ASCII/length/scrub) unit tests; REST synth gate (unavailable→400, cpu_fallback→headers, cpu_only→none) via the fake-engine harness with a mocked host; select response routing fields. Deferred (small follow-up): dub-pipeline ASR routing note on the preflight_error SSE channel — separate path, not a TTS synth entry point. No frontend /ws/tts client exists today (the routing frame serves external API consumers). Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
c6a55794da |
feat(routing): active-engine GPU verdict in preflight + diagnose (#21 PR 4/5) (#433)
Surfaces a routing verdict for the CURRENTLY-SELECTED TTS engine in the two system-health surfaces, so a CPU fallback / unavailable-GPU is heard about before a slow or failed synth — the no-silent-fallback contract, read-only. - `tts_backend.active_routing()` + `gpu_routing_verdict()`: the active engine's routing derived from list_backends() (byte-identical to the matrix) plus the host compute summary (family + VRAM from the canonical probe). Never raise. - `/system/diagnose` gains a `gpu_routing` check: accelerated→ok, accelerated-with-caveat / cpu_fallback→warn (+ actionable hint), cpu_only→ok (no-GPU host is the expected normal state — never noise-warns), unavailable→ fail, no-engine→warn. ASCII-safe detail strings (the text dump enforces ASCII). - `/setup/preflight` gains an "Active engine routing" check + an explicit `gpu_routing` object on PreflightResponse (a real field — the response has no extra="allow", so it would otherwise be dropped). `device` gains `gpu_family` (ROCm-vs-CUDA aware) + `vram_gb`. New `GpuRouting` schema. Tests: gpu_routing_verdict (host + active-engine + degraded), diagnose status mapping across all 6 states + never-raises, preflight gpu_routing object + check + device.gpu_family. Existing diagnose/preflight tests stay green (checks are additive; the report's top-level key set is unchanged). Deferred (documented): synth-time routing headers/WS-frames at the 3 synth entry points. Selection is already hard-gated (PR 3 select_engine), and the matrix (PR 5) + this preflight/diagnose verdict surface the situation — the synth-time signal is incremental belt-and-suspenders for the env-var-pinned edge and is best validated interactively. Tracked as a #21 follow-up. Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
8c8d525397 |
feat(routing): wire effective-device into /engines + select gate (#21 PR 3/5) (#432)
* feat(routing): wire effective-device + routing_status into /engines (#21 PR 3/5) Surfaces the PR-1 probe + resolver through the engine registries so the matrix UI (PR 5) and the no-silent-fallback gates can consume it. - `engine_routing.routing_fields()`: shared helper returning the three serialization-ready keys, centralizing the scrub rule — routing_reason is scrubbed via `core.scrub.scrub_text` only when truthy, so a None reason stays JSON `null` (never coerced to ""). - TTS/ASR `list_backends()` each gain `effective_device` / `routing_status` / `routing_reason`, computed from a SINGLE `detect_host_caps()` call per request (host caps are constant per process). ASR is brought to full TTS parity: it now also carries `install_hint` / `last_error` / `isolation_mode` and a SCRUBBED `reason` (closing a pre-existing ASR token-leak gap) — an identical 11-key shape across families. ASR also gains the same is_available()-raises resilience TTS has (degrade to available:false, never 500). - LLM `list_backends()` reaches 11-key parity too but emits literal `effective_device:"network"` / `routing_status:"n/a"` / `routing_reason:null` (NOT via resolve_routing — LLM runs no local GPU model). `LLMBackend.gpu_compat = ()`. "network" is a label, not a probe — nothing here touches the network. - `select_engine` host-routing gate: refuses a pick whose `routing_status` is `unavailable` on this host (400 with an actionable detail), while ALLOWING `cpu_fallback` (it runs, just slower). LLM is never gated. Defensive `.get` so legacy payloads still select. New typed `SelectEngineResponse`. Tests: 11-key shape across all 3 families, well-formed tts/asr routing keys (+ None-not-"" contract), LLM network/n/a labels, select gate (block unavailable / allow cpu_fallback / never-gate LLM). Updated the registry exact-shape test for the 3 new keys. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * test(cjk): allowlist docs/specs/ in the hardcoded-CJK guard PR #429 merged the longform design specs, which legitimately quote functional CJK (test-fixture descriptions, CosyVoice speaker IDs, multilingual sample text). The CJK guard scans every tracked file, so those docs turned main red. Specs are documentation, not shipped UI strings — allowlist the docs/specs/ prefix, matching the individually-allowlisted docs already in the set. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
000010ebb8 |
feat(dub): dedicated Dub home (projects/history) + project rename (#435)
The dub Projects + History rail (WorkspaceProjects/WorkspaceHistory) used to
sit beside the editor at all times. Now it's a landing: shown only when no
project is being edited (dubStep === 'idle'); opening/creating one switches to
a full-width editor. (The global Sidebar is already hidden in dub mode, so the
studio-right rail is the only surface — no Sidebar change needed.)
Adds project rename:
- backend: PATCH /projects/{id} updates just the name (400 on empty, 404 on
missing) — lighter than PUT which rewrites the whole state blob.
- api: renameProject(id, name); App.jsx renameProject handler (updates the
active-project label + refreshes the list).
- UI: inline rename on each project card (pencil → edit → Enter/Save / Esc).
Verified: PATCH create→rename→list / 400 / 404; frontend typecheck:ci clean.
Co-authored-by: mergetest <test@local>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
dc1d36fe5f |
refactor(models): model-management v2 cleanup (mm2, all tiers) (#428)
One coherent lifecycle surface over the in-process model, diarization, and subprocess sidecars; fixes the engine-switch VRAM leak; tightens download robustness. Backend-only, response shapes preserved, no new deps. Tier 1 — correctness: - MM2-01: get_active_tts_backend() caches one instance per backend id and unload()s the outgoing engine on switch (fixes the VRAM leak behind #278); adds reset_active_backend(). - MM2-02: OmniVoiceBackend.unload() releases the shared model_manager singleton + free_vram(); SubprocessBackend.unload() -> unload_sidecar(self.id), inherited by all sidecar engines. Idempotent + preload-safe. - MM2-03: /model/loaded ASR row reports the real device + a note explaining the disabled unload button. Tier 2 — single surface: - MM2-04: new services/model_lifecycle.py owns list_loaded/unload/unload_all/ free_vram; system.py routers are thin delegations (shapes unchanged). - MM2-05: idle timeouts (in-process + sidecar) resolve via prefs.resolve (env wins, no restart); removed the duplicated _IDLE_TIMEOUT_SECONDS. Tier 3 — robustness/observability: - MM2-06: _install_cooldowns swept (1h TTL) + cleared on success — bounded. - MM2-07: per-extension weight floors (onnx 64KB, tensors 5MB) OR the original >=5MB catch — small ONNX no longer false-flagged, #352 still caught. - MM2-08: indextts GPU sidecar self-reports vram_mb in pong; parent surfaces it in list_live_sidecars (0 = CPU/unmeasured). - MM2-09: is_cached scan_cache_dir->disk fallback logs WARNING w/ exc type (#117/#118), was invisible at DEBUG. Tests: tests/test_mm2_lifecycle.py (15). Full suite: 1379 passed. Plan/summary: .planning/quick/260613-mm2-clean-model-management-v2/. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
4cc55ab852 |
Fast model downloads: Xet fast path + accurate progress (FDL W0–W2 + W4) (#424)
* feat(downloads): Xet fast path + accurate progress (FDL W0–W2)
Make model downloads fast and show accurate downloaded/remaining/speed.
Research confirmed hf-xet already implements the IDM/uGet technique
(content-defined chunking, parallel byte-range gets, dedup, resume), and
the spike found all 25 catalog repos are Xet-backed — so the win is
driving Xet well + accurate progress, not a custom downloader.
W1 — maximize + guarantee Xet:
- pin huggingface_hub>=1.7 + hf-xet>=1.1 (was transitive); no hf_transfer
- drive snapshot_download with explicit tqdm_class + max_workers + endpoint
- opt-in HF_XET_HIGH_PERFORMANCE / HDD sequential-write knobs (default off)
- /system/info reports fast_download {xet_enabled, xet_version, high_perf}
W2 — accurate progress:
- dry_run preflight -> install_plan event (exact total/cached/remaining)
- utils/download_aggregator.py: one overall bar; byte bars (by id) vs the
"Fetching N files" count bar; windowed rate; emits one 'aggregate' event
- frontend overall bar (speed/remaining/ETA), cached-skip, ⚡ fast badge
Known limit (verified live): under Xet+hf_hub 1.7.2 per-file byte bars
never advance/close via tqdm, so mid-download the bar is file-granular and
bytes flush to the exact total on completion. Classic-LFS/mirror repos get
true byte progress (W4).
Drive-by: download.py used os.walk without importing os (latent NameError
in _validate_snapshot_has_weights on every install) — fixed.
Tests: tests/backend/setup/test_download_preflight.py (10). Spike + plan
under .planning/quick/260613-fdl-fast-model-downloads/.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(downloads): opt-in mirror + cancel + docs (FDL W4)
- mirror (FDL-10): snapshot_download(endpoint=) honours prefs hf_endpoint /
env HF_ENDPOINT on preflight + download (per-call, no process-wide env).
Documented as the classic-LFS path (no Xet) for restricted networks.
- cancel (FDL-11): POST /models/install/cancel {repo_id} stops further
retries at the next boundary, emits install_cancelled, clears the cooldown
(cancel is intent, not failure). Frontend treats it as a terminator.
- docs (FDL-12): docs/downloading-models.md (Xet fast path, progress
semantics + byte-speed limitation, opt-in tuning, mirror, cancel,
troubleshooting) + README pointer. Docs-sync rule satisfied.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* docs(planning): model-management v2 cleanup plan (mm2)
GSD plan for cleaning the model-management subsystem: registry unload-on-
switch + per-engine unload() (fixes VRAM leak), model_lifecycle facade,
unified idle/timeout config, bounded cooldowns, sidecar VRAM self-report,
cache-fallback logging. Planning artifact only — no code.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(downloads): reconcile with main's HF_HUB_DISABLE_XET; honest status
Rebasing onto main surfaced that main forces HF_HUB_DISABLE_XET=1 (classic
LFS) because Xet progress bypasses the tqdm hook — the same limitation found
here. Reconcile instead of fight:
- /system/info fast_download now reports runtime truth: xet_installed +
xet_active (installed AND not HF_HUB_DISABLE_XET) + xet_enabled alias. The
⚡ badge only shows when Xet actually runs; startup log says
"downloads: Xet disabled → legacy LFS".
- complete(): clear the rate window before the final flush so crediting the
full size in one step can't emit an absurd instantaneous rate.
- docs/downloading-models.md rewritten: default is legacy LFS for accurate
progress; Xet is opt-in via HF_HUB_DISABLE_XET=0. hf-xet pin stays (ready
for a future Xet progress hook).
W2 (preflight total/remaining + aggregate bar + exact completion) is the
value on either path; W1's "maximize Xet" is dormant by main's design.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(downloads): opt-in segmented multi-connection accelerator (FDL W3)
Since main forces Xet off (HF_HUB_DISABLE_XET=1), the default path is
single-stream legacy LFS — so a segmented downloader is the way to get BOTH
parallel speed and live byte progress.
- services/segmented_download.py: async multi-connection Range downloader for
one file — parallel byte-ranges, resume (.part + manifest), per-segment
short-read truncation guard, optional sha256/etag verify, cancel, and a
single-stream fallback when the server won't range. Auth-safe: the HF
Authorization header is sent only to huggingface.co/hf.co and never
forwarded to a CDN host on redirect (unit-tested).
- dispatch (download.py): opt-in via prefs segmented_downloader / env
OMNIVOICE_SEGMENTED_DOWNLOAD (default off). When on and Xet inactive,
fetches each file into the HF cache mirroring hf_hub_download (blobs +
snapshot symlinks + refs/main), feeding real bytes to the aggregator. Any
failure falls back to snapshot_download — never breaks a correct install.
- fix: complete() was adding a full total on top of accumulated segmented
bytes (2x); now replaces byte bars so the sum is exactly total.
Verified live (accelerator on): real byte progress to ~16.6 MB/s, final
bytes==total, /models installed=True, delete frees correctly.
Tests: test_segmented_download.py (7) + aggregator double-count regression.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(downloads): relocate FDL tests to top-level; loop-isolate segmented test
CI runs the full suite, which exposed a pre-existing test-isolation leak:
several tests/backend/** fixtures purge core.*/services.* from sys.modules
under a temp OMNIVOICE_DATA_DIR and never restore, leaving core.config/core.db
bound to a dead temp dir. It only bites when collection order puts a purging
test ahead of a real-DB reader (test_longform_jobs). Adding tests under
tests/backend/setup/ reordered collection and tripped it.
Fix without touching the shared (fragile) fixtures or risking class-identity
breakage from a blanket sys.modules restore:
- move the two FDL test files to top-level tests/ (tests/test_fdl_*.py) so
tests/backend/** collection order is identical to main — longform passes.
- rewrite the segmented test to run each case under asyncio.run() (fresh loop)
instead of asyncio.get_event_loop(), which an earlier async test can leave
closed in the full suite.
Full suite green locally: 1364 passed, 0 failed.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: mergetest <test@local>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
e196d790cc |
fix(longform): evict oldest chapters from the render cache (review fast-follow) (#423)
The content-addressed longform_cache/ accumulated uncompressed chapter WAVs across every render with no bound (a review finding). Add prune_cache_dir() — LRU-by-mtime eviction down to a 2 GB ceiling (OMNIVOICE_LONGFORM_CACHE_MAX_GB); best-effort, never raises. Called at the start of each render job, before its chapters are written, so the fresh ones are never the eviction target. Tests: under-cap no-op, evicts-oldest-keeps-newest, missing-dir safe. 38 green. Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
dde43de5a4 |
feat(longform): pronunciation lexicon — per-render word respelling (PR 8a) (#419)
Lets a render correct hard-to-say words (e.g. {"GIF":"jiff","Dr":"Doctor"}).
Backend wiring; the editor UI folds into the full-width Audiobook redesign.
- services/pronunciation.py (parallel-built, 19 tests): apply_lexicon —
whole-word, case-insensitive, longest-first, word-boundary, single ReDoS-safe
re.sub pass; + normalize/load/save_lexicon (JSON).
- synthesize_chapter gains a `lexicon` kwarg, applied to each span's text before
chunk splitting (None/empty = no-op → backward compatible).
- _render_chapter_cached folds the normalized lexicon into the chapter cache key
(a lexicon edit re-renders); threaded through _render_longform_sse + the
/audiobook, /audiobook/preview, /longform/render request models.
Tests: synthesize_chapter respells via lexicon; pronunciation module (19);
75 related backend tests green.
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
18e4c2347a |
fix(longform): correctness + robustness fixes from adversarial review (#418)
* fix(longform): correctness + robustness fixes from adversarial review Fixes the confirmed findings from a multi-agent review of the convergence: HIGH (correctness/output): - MP3 + cover produced a corrupt file (-map 2:v -c:v copy is invalid for mp3). Cover art is now embedded for M4B only; mp3 skips it (m4b is the cover format). - Chapter cache key omitted ref_text — editing only a profile's ref_text served stale audio. ref_text is now part of the voice signature. - Preview wrote audiobook_cache/ but the render reads longform_cache/ (rename missed in PR 5) → cache-warming silently broke. Unified to longform_cache/. Robustness (DoS/OOM guards): - /audiobook/import caps upload at 64 MB; epub_to_chapter_script bounds per-entry (25 MB) and cumulative (300 MB) uncompressed reads (zip-bomb guard). - /longform/render rejects > 10,000 chapters (422). Frontend leaks: - StoriesEditor.removeTrack revokes the line's preview blob URL. - AudiobookTab revokes the cover blob URL on replace/unmount. Deferred fast-follows (also from review): render-cache disk eviction; restoring the standalone chapter cue-sheet export (needs chapter times in the done event). Tests: mp3-drops-cover, epub entry/total caps, import + chapter-count limits; updated the cache-hit test for the 4-field voice sig. 70 backend + 334 frontend green. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * test(longform): pass EPUB caps as params, not monkeypatch (CI import-path fix) The cap tests monkeypatched module constants, but in the full-suite CI context the module loads under a different import path so the patch missed the function (it used the real 300 MB cap → tests failed). epub_to_chapter_script now takes max_entry_bytes/max_total_bytes kwargs (default to the constants); tests pass small values directly — deterministic regardless of import path. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
36e7fb12fc |
feat(longform): job library — finished books/stories in Projects (PR 7/8) (#417)
Surfaces finished Audiobook + Story renders so they're re-downloadable from the Projects view — closing the resume/history loop of the convergence. Backend (new, no migration — reads existing job_store rows): - routers/longform_jobs.py: GET /longform/jobs lists finished audiobook/story jobs newest-first, recovering output/chapters/duration from each job's persisted 'done' SSE event. Pure build_longform_library() over the job_store callables; defensive (skips unparseable jobs, never 500s). Registered in main.py. Frontend: - Projects.jsx: new "Audiobooks" category fed by /longform/jobs; each row opens the rendered file (/audio/<output>) with type/chapters/duration. Offline-safe (empty on fetch failure). en.json keys added. Built via parallel worktree agent; backend tests/test_longform_jobs.py (9) green; 334 frontend tests + build clean. Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
00f400e4c7 |
fix(stories): thread per-line speed through the shared renderer (PR 6/8) (#416)
PR 5 moved Stories' full export to /longform/render but dropped per-line **speed** — the old client export sent each line's speed to /generate; the converged path silently ignored it. This restores it end-to-end. - Span gains an optional `speed`; synthesize_chapter passes it to the injected synth (signature now `synth(text, voice_id, speed)`); both engine paths (OmniVoice model + generic TTSBackend) forward it to generate(speed=…). - chapter_cache_key now includes speed (a speed change re-renders; tuples accept an optional 4th element so existing 3-tuple callers/tests still work). - LongformSpan + /longform/render carry speed; storyToSpans emits each line's speed onto its spans. Emotion note: per-line tone is already model-native via inline tags ([laughter] etc.) inserted into the text, so no separate emotion→instruct plumbing is needed — the dead `emotion` store field stays unused/superseded. Tests: storyToSpans speed passthrough (8); cache-key speed sensitivity; synth stubs updated for the 3-arg signature. 65 backend + 334 frontend green; build clean. Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
0f67895585 |
feat(stories): full export → shared server-side renderer (PR 5/8) (#413)
* feat(stories): full export → shared server-side renderer (PR 5/8)
The convergence core. Stories' full export no longer stitches audio in the
browser (Web Audio, capped by RAM, no resume/loudness/markers) — it compiles
cast + lines into a chapter/span plan and streams through the same chapterized
renderer the Audiobook tab uses.
Backend:
- Extracted the audiobook SSE job into a shared `_render_longform_sse(plan, …)`
generator (resume cache, per-chapter fault isolation, mux). /audiobook is now
a thin caller.
- New POST /longform/render — accepts a pre-built {chapters:[{title,spans:
[{voice_id,text,pause_ms_after}]}]} plan (+ format/loudness/cover/metadata) and
renders it. Pause-only spans (empty text) are kept as silence. job_type=story.
- Shared content-addressed cache renamed longform_cache (one render per unique
chapter across both front doors).
Frontend:
- storyToSpans(tracks, cast) — pure compiler: `# ` lines → chapters; each line
resolves its cast/override voice; inline [voice:]/[pause] split into spans;
pauses fold into the previous span.
- StoriesEditor.generateAll now posts via longformRender and downloads the
server file (chaptered M4B / MP3). Single-line preview stays client-side;
stems export unchanged. Format select WAV→M4B.
Deferred to PR 6 (with the component split): per-line regenerate, emotion→instruct.
Tests: storyToSpans (7) — cast resolution, chapters, per-line + inline voice,
pause folding, empty-drop. 64 backend + 333 frontend green; build clean.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(audiobook): confine cover_path to OUTPUTS_DIR + don't leak exception text (CodeQL)
- _safe_cover_path() restricts the user-supplied cover to OUTPUTS_DIR before it
reaches ffmpeg (py/path-injection).
- SSE error events now emit a generic message and log the detail server-side
(py/stack-trace-exposure); empty best-effort excepts annotated.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(audiobook): cover path via basename+fixed dir (clears CodeQL py/path-injection)
CodeQL didn't recognize realpath+startswith as a barrier; os.path.basename is a
recognized sanitizer. Covers only come from /audiobook/cover (OUTPUTS_DIR/
audiobook_covers), so rebuilding from the basename onto that fixed dir is both
CodeQL-clean and strictly tighter — no caller path can escape it.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(audiobook): regex-allowlist cover filename (clears CodeQL py/path-injection)
basename alone wasn't a barrier CodeQL credits. Restrict the cover name to the
exact pattern /audiobook/cover emits (12 hex + jpg/jpeg/png) before joining onto
the fixed covers dir — an anchored-regex guard CodeQL recognizes as sanitizing,
and strictly tighter than before.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(audiobook): commonpath-confine resolved cover path (CodeQL py/path-injection)
Add an os.path.realpath + os.path.commonpath containment check on the resolved
cover path (the barrier static analysis recognizes), on top of the regex
allowlist + basename. Defense in depth; the path provably cannot escape
OUTPUTS_DIR/audiobook_covers.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
ea6833138b |
feat(audiobook): text + EPUB import → auto-chapter (PR 4/8) (#412)
Spec PR 4. A front door onto the existing chapter parser: import a file, get a
chapter-delimited script in the editor.
Backend (new services/longform_import.py — pure, stdlib only, no new dep):
- chapterize_plaintext(text): inserts `# ` headings ahead of short standalone
chapter-title lines (Chapter/Part/Prologue/…); no-op if the text already has
H1s; long "Chapter …" sentences stay prose. ReDoS-safe (anchored, per-line).
- epub_to_chapter_script(bytes): parses EPUB (zipfile + ElementTree +
html.parser) in spine order → `# Title` + stripped body per document; skips
empty/nav pages; the heading becomes the chapter title (not narrated). Raises
ValueError on a malformed EPUB. ET.fromstring annotated `# nosec B314` (local
user file, no external-entity expansion).
- POST /audiobook/import (UploadFile) → {text, chapters}.
Frontend: an Import button (.txt/.md/.epub) that fills the script editor.
Tests: tests/test_longform_import.py (9) incl. an in-memory synthetic EPUB
(spine order, empty-doc skip, tag stripping, bad-zip). 64 backend + 326 frontend
green; build clean; en.json valid.
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
7af5143fac |
feat(audiobook): per-chapter preview + resume + chapter fault-isolation (PR 3/8) (#411)
* feat(audiobook): per-chapter preview + resume + chapter fault-isolation (PR 3/8) Builds on the shared core (#408) and metadata UI (#409). Chapter-level control, the spec's PR 3. Shared core: - chapter_cache_key(spans, sr, engine_id, voice_sig) — deterministic content hash of a chapter's audio inputs. Same inputs → reuse; any change (text, voice, order, pauses, sr, engine, resolved-voice signature) → re-render. Backend (audiobook router): - Chapter WAVs are now content-addressed in OUTPUTS_DIR/audiobook_cache. A re-run after a failure/interruption reuses already-rendered chapters and only synthesizes the missing/changed ones (resume). Job emits `cached` per chapter and `cached_chapters`/`failed_chapters` on done. - Per-chapter fault isolation: a chapter that throws emits `chapter_error` and the job continues; the m4b assembles from the successful chapters. Re-running retries only the failed (un-cached) chapters. - POST /audiobook/preview — render a single chapter to audition it; shares the same cache so a preview warms the full run and a re-preview is instant. - _build_synth now exposes resolve + engine_id; _prepare_synth unifies the omnivoice/generic paths for both the job and preview. Frontend: - Plan view: a ▶ preview button per chapter with inline playback. - Done panel: "reused N chapters" + "N failed — click Create to retry" notes. Tests: chapter_cache_key determinism + sensitivity (8); preview validation + cache-hit-skips-synth (3). 55 backend + 326 frontend green; build clean. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(audiobook): mark cache-key SHA1 usedforsecurity=False (bandit B324) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
086ac08592 |
feat(audiobook): metadata, cover art, format + loudness UI (PR 2/8) (#409)
Surfaces the shared-render-core capabilities (PR 1, #408) in the Audiobook tab. Backend: - POST /audiobook/cover — multipart cover upload (jpg/png, 8 MB cap), returns a server-side path passed back as cover_path. Unit-tested via the handler directly (no main+torch import). Frontend: - api/audiobook.ts: AudiobookGenerateBody (format/loudness/cover_path/metadata) + audiobookUploadCover(file). - AudiobookTab: format select (M4B/MP3), loudness select (off/ACX/podcast, default off), and a "Cover & details" panel — cover picker with preview + title/author/narrator/year/genre/description. On create, the cover uploads first, then the job runs with metadata + format + loudness. - en.json: audiobook.* keys for the new controls. Tests: tests/test_audiobook_cover.py (4) green; frontend vitest 326 green; prod build clean; CJK + i18n-parity gates pass. Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
e9481ef307 |
feat(longform): shared render core — loudness, metadata, cover art (PR 1/8) (#408)
First slice of the Stories+Audiobook convergence (spec: docs/specs/2026-06-13-stories-audiobook-maturity.md). Both features will compile to one server-side chapterized renderer; this lands the shared pure builders and wires them behind Audiobook. New `backend/services/longform_render.py` (all pure, unit-tested without ffmpeg/torch): - build_ffmetadata(chapters, global_meta) — FFMETADATA1 with an optional global tag block (title/author→artist/narrator→composer/year→date/genre/description→ comment) + chapter table. - build_loudnorm_filter(preset) — `-af loudnorm` for ACX (~-19 LUFS, -3 dBTP) or podcast (-16 LUFS); off/unknown → None. Opt-in, so default behavior stays platform-identical. - validate_cover_image — jpg/png + 8 MB cap guard. - build_render_cmd — generalizes the m4b mux: m4b|mp3, optional cover (attached_pic) + loudness, bitrate validated. - build_concat_list — moved here. `services/audiobook.py`: build_chapter_ffmetadata / build_m4b_cmd / build_concat_list are now backward-compatible wrappers over the core (existing imports + tests unchanged). `POST /audiobook`: now accepts optional `format` (m4b|mp3), `loudness`, `cover_path`, and `metadata` and passes them through — backend-complete; the UI for these lands in PR 2. Tests: tests/test_longform_render.py (28) + existing test_audiobook.py (11) green. Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
599f3bcc5c |
feat(engines): on-demand unload of subprocess-engine sidecars (Action 13) (#406)
Completes the dynamic engine load/unload slice. The idle reaper (#401) frees sidecar VRAM after 5 min; this adds a user-initiated "free VRAM now" path so multi-engine users don't have to wait: - subprocess_backend: `list_live_sidecars()`, `unload_sidecar(id)`, `unload_all_sidecars()` via a shared `_force_reap(predicate)` — busy-guarded exactly like the idle reaper (non-blocking lock; a sidecar mid-synth is skipped, never interrupted; next request respawns it). - system.py: `/model/loaded` now surfaces live sidecars as unloadable rows; `/model/unload/{sidecar:<id>|sidecars}` frees one or all. The existing generic flush panel picks these up with zero frontend change. Also refresh CLAUDE.md stale version notes: main is 0.3.6 (latest release v0.3.5 + 1 patch); the v0.3.0-as-unreleased framing in the project/cadence notes is corrected to the v0.3.x continuous-to-main reality. Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
6704d062fc |
fix(persona): preserve design kind + vd_states across share/import (Wave 5 §R3) (#405)
The persona-gallery surface already exists (VoiceGallery Community zone + community.py manifest + marketplace .omnivoice bundles). The blocker for §R3's 'synthetic-only' gate was data integrity: a *designed* persona lost its kind='design' (and vd_states) when imported from the community gallery or round-tripped through a bundle — silently demoting it to a clone. - community.py /use: a 'preset' (rendered from instruct) imports as kind='design'; a 'voice' (real reference clip) as 'clone'. - marketplace.py: extract a pure _bundle_metadata() (dedupes export+publish) that captures kind + vd_states; import restores them. Old bundles without the keys import as 'clone' (backward-compatible). This makes 'accept only designed/synthetic voices' enforceable instead of everything defaulting to clone. No new persona-gallery feature was built — that would duplicate the existing community/marketplace surface. 4 torch-free tests (isolated DB): _bundle_metadata captures design + defaults to clone; import round-trip preserves design kind+vd_states; legacy bundle → clone. docs §R3 status updated. Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
9441274ab6 |
feat(audiobook): synth job → chapterized m4b, SSE progress (Wave 5) (#403)
Completes the audiobook backend: POST /audiobook renders each chapter through the active TTS engine (synthesize_chapter + chunked_tts), writes per-chapter WAVs, then muxes a chapterized m4b (FFMETADATA1 chapters via build_m4b_cmd + concat demuxer). Progress streams as SSE (started/chapter/assembling/done/ error), recorded to job_store. ffmpeg-gated — emits an error event and stops when ffmpeg is absent (m4b is the only output). - services/audiobook.build_concat_list: pure ffmpeg concat-list builder with proper single-quote escaping (no arg injection). Unit-tested. - router: voice resolution (compact form of generation.py's locked/design/ clone cases) cached per id; OmniVoice native model path + generic TTSBackend path; chapter synthesis runs on the GPU pool, ffmpeg via run_ffmpeg. Reuses the tested building blocks from #402 (parser, synthesize_chapter, FFMETADATA + m4b argv builders) — the new router glue is thin and import-checked by CI. Deferred: epub/pdf ingest, ACX loudnorm mastering, crash-resume, UI. 15 audiobook tests (added concat-list); docs §R3 updated. Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
34b47282af |
feat(audiobook): chapterized audiobook core + plan preview (Wave 5) (#402)
* feat(audiobook): chapterized audiobook core + plan preview (Wave 5) First cut of the long-form vertical (parity §R3). Engine-agnostic core in services/audiobook.py: - parse_audiobook_script: pure parser. Markdown '# H1' headings → chapters; inline [voice:NAME] switches the narrator; [pause …] is delegated to the shared omnivoice.utils.text.parse_pause_markers so audiobooks and single-shot synthesis keep one pause dialect. Returns a chapter/span plan. - synthesize_chapter: orchestration via an injected synth(text, voice) callable (reuses chunked_tts split + crossfade, stitches inter-span silence) — so it's unit-testable with a stub backend, no model/GPU. - build_chapter_ffmetadata + build_m4b_cmd: pure FFMETADATA1 [CHAPTER] builder and faststart-m4b concat-demux argv (bitrate-validated, no injection). POST /audiobook/plan returns the parsed plan (no TTS/ffmpeg, no side effects). Deferred (follow-ups): the streaming synth job + chapterized-m4b run, epub/pdf ingest (new dep), ACX loudnorm mastering, crash-resume, UI. 14 tests: parser (chapters/voice/pause/intro/empties/to_dict), FFMETADATA offsets+escaping, m4b argv + bitrate guard, and stub-synth orchestration (span+silence stitching, voice threading). docs §R3 status updated. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(audiobook): linear-time regexes (CodeQL ReDoS) CodeQL flagged polynomial backtracking on user-provided input in three regexes reachable from the new POST /audiobook/plan endpoint: - _VOICE_RE: \s*(...)\s* → single [^\]]* class, stripped in code. - _HEADING_RE: trailing [ \t]* removed; title captured greedily + stripped. - _PAUSE_RE (omnivoice/utils/text.py): the numeric spec is now an atomic group (?>…) so its leading \s+ can't backtrack against the trailing \s*. Behavior-preserving (Python >=3.11 already required); 14 pause tests + 14 audiobook tests green. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(audiobook): require non-space heading title start (CodeQL ReDoS) The previous _HEADING_RE '[ \t]+(.+)' still let the leading whitespace class and the title '.+' both match the same tab run (overlap → polynomial). Anchor the title capture with \S so the two can't overlap. 14 audiobook tests green. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(audiobook): exclude '[' from voice-tag content (CodeQL ReDoS) [^\]]* still matched '[', so a run of nested [voice: prefixes produced overlapping finditer match attempts → O(n^2). Excluding both brackets ([^\]\[]) makes matches non-overlapping and linear. A voice name never contains a bracket. 14 audiobook tests green. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
e8705a106d |
feat(dictation): opt-in NLMS AEC for dictate-over-playback (Wave 8b) (#399)
* feat(dictation): opt-in NLMS AEC for dictate-over-playback (Wave 8b) Dictating while OmniVoice plays audio (TTS preview, dub, video) leaks the loudspeaker signal into the mic, and the streaming ASR transcribes that bleed. Browser echoCancellation varies per platform/webview — it can't be a cross-platform default — so this adds a server-side canceller that behaves identically everywhere. services/aec.py ports Patter's NlmsEchoCanceller (MIT): a time-domain NLMS adaptive filter with a Geigel double-talk detector, warm-up step ramp, and far-end staleness pass-through. /ws/transcribe gains an opt-in '?aec=1[&sr=]' mode: frames are raw int16 mono PCM tagged with a 1-byte prefix (0x00 mic, 0x01 playback reference); the mic is cleaned against the reference before buffering, and the cleaned PCM is muxed via stdlib wave (not ffmpeg). Without the param the protocol and behaviour are byte-for-byte unchanged. Backend ships dark (no new deps — numpy already pinned); frontend far-end streaming is a follow-up. Tests cover echo attenuation, double-talk preservation, cold/stale pass-through, param validation, and the framing helpers — all pure-numpy/stdlib so they skip the torch ASR stack. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(capture_ws): stubs accept the new pcm_sr kwarg _transcribe_buffer/_transcribe_buffer_full gained an optional pcm_sr kwarg for the AEC PCM path; the protocol-test stubs had fixed signatures and raised TypeError on it, so the handler sent 'error' instead of 'final'. Accept **kw in the stubs. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
4aa4d983aa |
feat(settings): Hugging Face mirror (HF_ENDPOINT) for restricted networks (Wave 4.3) (#391)
The model manager already lists/deletes cached models; this adds the remaining high-value slice — an in-app HF mirror setting so users behind restricted networks (e.g. the Great Firewall) can route downloads through hf-mirror.com or any HF_ENDPOINT. Persisted to the durable per-user env (survives Tauri/Finder launches); HF reads HF_ENDPOINT at import, so the override applies on restart (surfaced in the UI). - GET/PUT /api/settings/hf-mirror (loopback-gated): presets (official + hf-mirror.com), http(s) validation, empty clears to official. - Models-tab panel with quick-picks + free-text field + restart note. (Skipped 'hf cache verify' — version-fragile across huggingface_hub releases and low value vs the mirror, which the China/Russia network research flagged as the real gap.) 3 endpoint tests (default, set+trim+clear, non-http rejection). Spec §R4(c) / parity program Wave 4.3. Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
3161166328 |
fix: stale-chunk preload recovery (#380) + surface unsupported-GPU-arch in notifications (#284) (#385)
- #380: vite:preloadError (old hashed assets after an update) triggers a one-time reload to pick up the fresh manifest; session flag prevents loops - #284: check_device_compatibility's warning (e.g. Blackwell sm_120 on a pre-cu128 torch) now appears in the notification panel as an error with the pip fix — a log line never reached affected users while synthesis silently produced noise. Cached once per process. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
f5579b40aa |
fix: issue-triage batch — timeline box flicker, truncated-model detection, stale history pruning (#381)
- #373: drop will-change:transform on the segment lane (persistent compositor layer made the semi-transparent boxes vanish during playback/drag on some Windows GPUs) + raise region alpha 0.30→0.45 - #352: validate a finished snapshot actually contains weights (>5 MB file) so interrupted downloads fail at install time with a re-download hint; loader translates the opaque transformers error into the same guidance - GET /history prunes rows whose audio file is gone instead of serving dead 404 players forever Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
66f2ea7e50 |
feat(profiles): unified profile model — kind discriminator + stored design params (spec P3) (#376)
Migration 0005_unified_profiles (0004 taken by mcp bindings):
- voice_profiles.kind TEXT DEFAULT 'clone' ('clone' | 'design'), backfilled
- voice_profiles.vd_states TEXT NULL — JSON of design category picks
- mirrored in _BASE_SCHEMA; idempotent _has_column guards; downgrade drops
POST /profiles:
- ref_audio now optional; kind + vd_states form fields with validation
(clone requires audio; design requires vd_states JSON object + instruct)
- design profiles render a deterministic identity sample (seed 42) through
the shared archetype renderer — one TTS code path
POST /generate:
- profile resolution branches on profile.kind (authoritative) instead of
the brittle is_locked/instruct inference; legacy pre-0005 rows keep the
old inference as fallback; history.mode records profile.kind
Frontend:
- 'Save design as profile' in the Design tab (vd_states + buildDesignInstruct)
- selecting a design profile restores its sliders (vd_states) for re-editing
Also unforks the alembic chain (0004_mcp + my 0004 both revised 0003 →
multiple heads broke alembic upgrade head and the 0003 migration tests).
Tests: tests/test_profile_unification.py — validation, design-create with
mocked renderer, migration up/backfill/downgrade. 18/18 profile tests,
312/312 frontend, related backend suite green.
Note: docs/specs/voice-studio-unification.md (on feat/studio-ux-overhaul)
still says 0004 — renumber to 0005 when branches meet.
Co-authored-by: mergetest <test@local>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
d6562d6f30 |
feat(dub): Smart Fit phase B — per-segment video retime export, drift absorption, fitted subtitles (#350)
* feat(dub): Smart Fit phase B — per-segment video retime export, drift absorption, fitted subtitles Executes the video side of the Smart Fit plans persisted by Phase A (job["fit_plans"], #347) at export and preview time. Backend: - services/video_retime.py (new, clean-room): two-tier retime executor. ≤48 chunks → the proven single-pass split/trim/setpts/concat filter_complex; above → batches of 40 chunks rendered to intermediate slices (identical libx264 medium/crf20 params, keyframe at t=0) joined losslessly with the concat demuxer. Slices are CFR-resampled (fps=) because setpts leaves VFR-ish timestamps that broke tpad and drifted a frame per retimed chunk on ffmpeg 7.x. Temp slices cleaned on success AND failure/abort. - Drift absorption: fitted track longer than retimed video → freeze-frame tail (tpad=stop_mode=clone) predicted into the last slice / single-pass graph, with residual mux-side tpad; video longer → silence-pad the dub audio chain (apad=whole_dur). ±50 ms tolerance. - VFR guard: probe r_frame_rate vs avg_frame_rate; normalise with fps= before trim/setpts; probe failure degrades gracefully. - Plan resolution: _video_retime_plan_for spans legacy video_stretch_plans (byte-identical resolution + command construction) and fit_plans, gated on the track's own timing_strategy so stale plans never retime a track re-generated under another strategy. - Fitted subtitles: /dub/srt + /dub/vtt accept ?lang= and serve cue times from fitted_segments for Smart Fit tracks; _write_burn_srt does the same for burn-in. burn_subs+retime is now allowed for smart_fit (burn runs AFTER the retime graph); still rejected for legacy stretch_video. - /dub/preview-video resolves the same plan so in-app preview matches export. - Fallback ladder: batch encode failure/timeouts → un-retimed export with a structured core.failure warning (X-Dub-Export-Warning header + job["last_export_warning"]); concat join rejection → one single-pass retry while ≤96 chunks; abort → 409 + proc kill via run_ffmpeg job_id registration (/dub/abort reaches export encodes now) + temp cleanup. Frontend: - Export drawer passes ?lang= on subtitle exports and shows an i18n'd re-encode cost note (~0.5–2× video length on CPU) when a retiming strategy is active — translated in all 21 locales. Tests: tests/test_smart_fit_export.py — plan resolution, batch math, graph parity + new stages, fitted-cue SRT/VTT/burn selection, burn policy, VFR detection; ffmpeg-gated integration renders both executor tiers (batch size forced to 2) and the real /dub/download endpoint, ffprobing durations within ±50 ms across both pad branches. All existing dub export/subtitle/preview/timing tests pass unchanged. Refs docs/competitive-analysis.md Action 1 (dub-length fitting v2); completes Smart Fit (Phase A = #347). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(security): sanitize Smart Fit retime work paths at every sink (CodeQL py/path-injection) The job_id-derived retime work path (retimed_*.mp4 / preview_retimed_*.tmp.mp4) flowed unguarded from dub_export into prepare_smart_fit_video / render_retimed_video and their derived slice/concat paths and ffmpeg argv. Apply the repo's proven inline realpath+startswith containment pattern (helpers/commonpath are not recognized — see #309/#328/#329/#348): - dub_export.py: validate work_path against DUB_DIR at both construction sites (export + preview) and pass the validated realpath onward. - video_retime.py: make both entry points self-defending — realpath + DUB_DIR containment on out_path/work_path before any derivation, raising RetimeError(stage="plan") on escape; slices_dir/slice_path/list_path and RetimeDecision.file_path now all derive from the sanitized value. DUB_DIR is read via module attribute so test fixtures reloading core.config work. - ffmpeg_utils.py: document that all caller-assembled argv paths are realpath-validated upstream. - tests: sandbox DUB_DIR in the executor integration tests (tmp_path) so the new guard sees the test workspace. No behavior change for valid (server-built) paths — the guard only fires on traversal. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(smart-fit): patch DUB_DIR on video_retime's own config ref — survives suite-wide reload The retime guard reads video_retime._config.DUB_DIR at call time; the sandbox fixture patched a fresh 'import core.config' instead. Another test reloads core.config in the full suite, so the two module refs diverged — the patch missed and the guard rejected the test's tmp paths (green in isolation, red in CI's full run). Patch the exact ref the guard dereferences. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(dub): resolve DUB_DIR live at call time in retime guards — survive full-suite reload The path-containment guards bound DUB_DIR via a module-level 'from core import config as _config'. Other tests importlib.reload() core.config (sandboxing OMNIVOICE_DATA_DIR), after which the guard checked containment against a stale DUB_DIR while dub_export built the path under the reloaded one — every retime path then 'escaped the dub workspace' (green file-alone, red full-suite: the 5 integration failures CI hit). Re-import DUB_DIR locally in each guard so it always reads the current sys.modules value; simplify the sandbox fixture to patch the canonical module. Verified: full backend suite green on the Smart Fit tests (the 2 remaining settings_store failures are pre-existing on main, unrelated — local data-dir artifact). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(security): clear CodeQL alerts on Smart Fit export — job_id allowlist, proc-registry decouple - py/path-injection (8, video_retime.py): validate job_id with a strict inline regex allowlist (re.fullmatch [A-Za-z0-9_-]{1,64}) at the entry of dub_download and dub_preview_video, before it reaches any filesystem path or ffmpeg argv. The existing realpath containment guards stay as defense-in-depth; the regex barrier is the sanitizer CodeQL recognizes through the service-module call chain. - py/log-injection (4): newline-strip job_id inline at the logger calls in ffmpeg_utils.run_ffmpeg and the two retime-fallback logger.error sites in dub_export. - py/empty-except (3): best-effort cleanup os.remove handlers now log the OSError at debug instead of bare pass (video_retime + both dub_export mux finally blocks; _discard_tmp too for consistency). - py/cyclic-import (2): break the dub_pipeline <-> ffmpeg_utils cycle for real — the subprocess registry (register_proc/unregister_proc/ kill_job_procs/has_active_procs + state) moves to a new stdlib-only leaf module services/proc_registry.py. ffmpeg_utils now imports it at module top (no lazy import); dub_pipeline re-exports every name so dub_core aliases and tests keep working unchanged. No behavior change for valid inputs; invalid job ids now get a clean 400 instead of a 404/containment error. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(dub): address #350 review — cancelled-vs-failed retime, logged best-effort excepts, redacted probe logs, narrowed test assert - rc<0 (killed by user cancel) now raises RetimeError(stage='aborted') instead of reporting an ordinary render failure (CodeRabbit) - best-effort cleanup/QC-event excepts log at debug instead of bare pass (CodeQL empty-except x3) - probe failure logs use basename, not full user paths (CodeRabbit/CodeQL) - test_render_cleans_slices_on_failure asserts RetimeError, not Exception Rebuttals (no change needed, see PR comment): fitted-cue subtitles track the fitted AUDIO timeline which is correct even on retime fallback; the planner only emits stretch ratios >1 so the early-exit guard is a true no-op check; '\'' is ffmpeg's own utility quoting for concat lists; has_active_procs is an intentional re-export (noqa'd). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
825f4f7ac6 |
feat(dub): regenerate subtitle timeline on the fitted timeline (Wave 3.1) (#371)
Smart Fit Phase A (planner) + the export-side video retime + audio stretch already shipped (#347 + dub_export stretch filter). The last piece of Spec 1 was the subtitle timeline: under stretch_video the dubbed audio plays at FITTED positions, but the standalone SRT/VTT export still used the original segment times — so external subtitles drifted against the dubbed video. - services/fitted_subtitles.py (pure, tested): map_time_to_fitted() + fitted_cues() remap original cue times onto the same per-chunk {orig→new, stretch_ratio} plan the video stretch uses, with a monotonicity guard. - dub_export SRT + VTT endpoints: when a job used stretch_video, cues are regenerated from the plan (subtitles track actual dub placement); no plan → original times, unchanged. New optional ?lang= selects the track. 7 pure tests (chunk-bound mapping, linear interpolation, unit-rate tail, fitted cues, monotonicity, empty-plan identity). Spec 1 (remaining) / parity program Wave 3.1. Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
a12492af07 |
feat(dub): second-pass ASR QC — flag lines whose dub drifts from target (Wave 3.3) (#370)
After a dub is generated, re-recognize the synthetic audio and compare what
the ASR heard against what we asked the TTS to say. Lines that drift are
flagged for the user to re-listen / re-dub — turning subtitle timing and
pronunciation from trusted math into measured truth, and doubling as an
automatic dub-quality check.
Design delta from pyvideotrans (which lets recognized text REPLACE the
subtitles wholesale): we keep the generated text authoritative and use the
second pass only for MEASUREMENT — a per-line drift score + measured
start/end that feed the incremental re-dub loop, never silently overwriting
the translation.
- services/dub_qc.py (pure, tested): word_error_rate (normalized token edit
distance, case/punct-insensitive, script-agnostic) + score_dub (matches
recognized segments to dub segments by time overlap, concatenates the
hypothesis, scores drift, derives measured bounds).
- POST /dub/qc/{job_id}: runs the active ASR backend on the dubbed track in
the GPU pool, annotates each segment with qc_drift/qc_flagged/
qc_recognized/qc_measured_start-end (non-destructive — content untouched),
persists, emits a qc_done job event. Opt-in, never fatal.
- Frontend: dubQc() API fn + a red 'Verify' badge on flagged segment rows
(en.json keys; other locales fall back).
12 pure scoring tests (identical/substitution/empty/no-overlap/multi-segment
matching/measured-timing); endpoint validated in CI.
Spec 5 / parity program Wave 3.3.
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
8cce99298a |
feat(dub): per-segment clone references (Wave 3.2) (#369)
Cut each long-enough dub segment's clone reference from the isolated vocals
at that segment's own timestamps, so the dub of each line carries the
prosody/emotion of its source line — finer than one reference per speaker.
Reimplemented from the clean-room spec (pyvideotrans per-line ref idea); our
design delta is a quality floor with fallback.
- services/speaker_clone.py: extract_segment_refs() keyed by segment id;
reference transcript is the SOURCE text (text_original), since the vocals
slice is source-language audio. Floor at MIN_SEGMENT_REF_DURATION_S=3.0
(not the per-speaker 5.0, which most dialogue lines fall under) — shorter
lines are omitted and fall back to the per-speaker clone, so it's a strict
improvement, never a regression.
- dub_core: run extraction at transcribe (per_segment_refs query param,
default on), store job['segment_clones'], default each unassigned
segment's profile_id to 'auto-seg:{id}' when it has its own ref, else the
existing 'auto:{speaker}'. Forcing per-speaker (per_segment_refs=false)
is supported for long-form consistency.
- dub_generate _gen: resolve 'auto-seg:' from segment_clones, ahead of the
per-speaker 'auto:' path. profile_id is already a fingerprint field, so
flipping the mode re-dubs automatically (no _GEN_INPUT_FIELDS change).
7 pure tests over a synthetic vocals wav (own-ref for long lines,
short-line omission/fallback, source-text transcript, bounds clamping,
floor boundary). Pipeline wiring validated in CI.
Spec 4 / parity program Wave 3.2.
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
99357e8c5b |
feat(mcp): MCP server v1 — mount on /mcp, per-agent voice binding, stdio shim (Wave 2.2) (#368)
* feat(mcp): MCP server v1 — mount on /mcp, per-agent voice binding, stdio shim (Wave 2.2) The FastMCP server (previously dead code, never mounted) is now mounted on the main FastAPI app at /mcp via Streamable HTTP, with its session manager composed into the app lifespan through an AsyncExitStack (best-effort: a missing mcp package or OMNIVOICE_MCP_DISABLE=1 never breaks startup). streamable_http_path set to '/' so the sub-mount lands at /mcp, not /mcp/mcp. Adds the 'mcp' dependency (1.27.x). Per-agent voice binding (Spec 2 headline): each MCP client sends an X-OmniVoice-Client-Id header; generate_speech resolves the voice as explicit arg > the client's binding > global default > app default. New mcp_client_bindings table (alembic 0004 + _BASE_SCHEMA, additive/idempotent), services/mcp_bindings.py (CRUD + resolve_voice + best-effort last_seen), and a loopback-gated REST router (/api/mcp/bindings) the Settings panel drives. New transcribe tool (base64 audio in, 200 MB cap). Stdio shim (backend/mcp_shim, httpx-only, ported from voicebox MIT) proxies stdio clients to the mounted endpoint and forwards OMNIVOICE_CLIENT_ID as the binding header. Settings → Sharing gains an MCP bindings panel. Docs: docs/mcp.md (both connection modes + binding REST) and docs/mcp.json updated to the shim form. Tests: bindings service + resolution precedence + migration up/down (pure, run locally); REST CRUD + mount-not-404 + disable-flag (main-importing, validated in CI). MCP build + mount + initialize handshake verified out-of-band (no torch). Spec: docs/competitive-analysis.md Spec 2 / parity program Wave 2.2. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(mcp): assert /mcp mount via app.routes, not a lifespan client The two main-importing mount tests ran the app lifespan, which now starts the FastMCP session manager and binds asyncio queues to the test loop — contaminating later lifespan-running tests ('bound to a different event loop'). The mount happens at import time, so inspecting app.routes for the /mcp Mount is the correct loop-free assertion. Same fix shape as the Wave 0.2 consent tests. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(mcp): stop reload-main poisoning across the MCP test files Root cause of the CI failure: the bindings REST fixture set OMNIVOICE_MCP_DISABLE=1 and reloaded main but never restored it, so a later 'from main import app' in test_mcp_mount saw /mcp un-mounted ({'/audio','/voice_audio'}). Reloading main mutates the shared module for every subsequent test. - REST fixture: drop the disable flag (the mount is harmless without a lifespan), yield the client, and restore main (+ core.config/db) to the default data dir in teardown so the global module is clean again. - test_main_mounts_mcp_route: reload main with the disable flag cleared so the assertion is independent of any earlier reload. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
c8fdcb619a |
fix(settings): remove stray rebase conflict marker in settings.py (#367)
A '>>>>>>>' marker from the #365 rebase was committed at the tail of the LLM-endpoint block, making the module unparseable. Strip it; settings.py parses clean and the endpoint tests pass. Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
d0b46f249e |
feat(settings): remote LLM endpoint UI — Ollama/vLLM/LM Studio (Wave 2.4) (#365)
A focused Settings panel for the OpenAI-compatible LLM that powers cinematic translate, glossary auto-extract, and dictation refinement (Wave 2.1). Persistence reuses the existing TRANSLATE_BASE_URL / TRANSLATE_MODEL / TRANSLATE_API_KEY env vars (already in system.py PERSISTENT_KEYS, restored at startup), so llm_backend/translator resolution is unchanged — vLLM is a verified drop-in, Ollama ignores the key, vLLM/LM Studio require it. - GET/PUT /api/settings/llm-endpoint (loopback-gated): read shape returns base_url, model, masked key, and live availability; PUT treats a null field as unchanged and an empty string as clear (so the key isn't wiped by a base-url-only save). Key is masked to last-4 in the read path, never echoed. - Credentials-tab panel with one-click presets (Ollama/LM Studio/vLLM/ OpenAI), base URL + model + optional key fields, and a reachable/not status badge. 6 endpoint tests (read shape, set+mask, null-unchanged, empty-clears, local-url-no-key, short-key masking); availability assertions guarded on openai being installed. Spec: parity program Wave 2.4 / competitive-analysis §R2 rung 4. Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
22ba348f17 |
feat(remote): backend URL + bearer key + Tailscale docs (Wave 2.3) (#364)
Run inference on a remote GPU box, drive it from the desktop app — opt-in,
off by default (loopback-only is unchanged when no key is set).
Backend:
- BearerKeyMiddleware (main.py): when OMNIVOICE_API_KEY is set, every
non-loopback HTTP + WebSocket request must present it (Authorization:
Bearer, ?api_key=, or the ov_key cookie set on first auth). Pure ASGI
(no response buffering), loopback always bypasses, SPA shell stays
reachable. Constant-time compare, never logged.
- ws_remote_authorized() in dependencies; capture_ws lets a keyed
non-loopback client through its inline loopback guard (the thin-client
dictation case: mic local, GPU remote).
Frontend:
- api/client.ts: ov_backend_url (localStorage) is the top-precedence base
override; new wsUrl() derives ws scheme + host from the API base (not
window.location, which lies in the Tauri webview) and appends ?api_key.
apiFetch attaches the bearer header. Both WS call sites (dictation,
events) routed through wsUrl; the HTTP transcribe fallback through
apiFetch.
- Settings > Sharing > Remote backend panel: URL + key fields, a
test-connection probe against {url}/health, save-and-reload.
Docs: docs/remote-gpu.md — the Tailscale recipe (MagicDNS + Serve, never
Funnel, headscale note, plain-HTTP-is-sniffable warning, PIN-vs-key split).
Tests: 10 bearer-middleware cases (inert without env, loopback bypass,
401 without/pass with key via header+query, wrong key, shell exemption,
plain-ASGI guard, WS handshake reject/accept). Validated in CI.
Spec: parity program Wave 2.3 / competitive-analysis §R2 rungs 1-3.
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
10806fea4f |
feat(dictation): optional local-LLM refinement of finals (Wave 2.1) (#363)
Phase 2 of Spec 3, on top of Wave 1.1's deterministic collapse. Prompt design ported from voicebox (MIT): 'text filter, not an assistant' base instruction + three toggleable sections (smart_cleanup, self_correction, preserve_technical) + 7 few-shot examples passed as STRUCTURED chat turns (small local models echo inline examples). Runs through the user's own Ollama/LM Studio/OpenAI-compat endpoint via llm_backend — new additive chat_messages() on the adapter; chat() now delegates to it. Pass-through is the contract: with no LLM configured (backend 'off'), on any error/timeout, or on an empty reply, the raw transcript stands — identical default behavior on every platform. Refinement runs off-thread on FINALS only; the WS final dict gains optional refined_text and the dictation pill pastes refined_text ?? text (raw kept in history). Settings: GET/PUT /api/settings/dictation-refinement (loopback-gated, persisted in the settings table) + a Capture-tab panel with the master switch + per-flag toggles and a 'no LLM configured' hint. 15 new unit tests: prompt sections per flag, structured few-shot message shape, and the full maybe_refine pass-through matrix (off backend, disabled config, LLM failure, empty reply, empty input). Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
9162f2b9e7 |
feat(stream): sentence-by-sentence /ws/tts via ported chunker (Wave 1.4) (#358)
Ports Patter's SentenceChunker (MIT, attribution header) behavior-identical — all 61 upstream golden parity scenarios ship as fixtures and pass, including documented quirks (current_behavior xfail semantics mirrored from their parity runner). Terminator tables carry functional CJK; file added to the test_no_hardcoded_cjk allowlist per convention. /ws/tts now splits the request into sentences and synthesizes each in turn, streaming the first sentence's PCM while later sentences are still generating — the time-to-first-audio win on multi-sentence input. Single-sentence requests behave exactly like the old single-shot path; 'start' metadata still waits for the first generation so lazy-loading engines report their true sample rate. Italian comma-decimal guard hard-disables aggressive first-clause flush per upstream. Spec 8a (docs/competitive-analysis.md) / parity program Wave 1.4. Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
454affb6e9 |
feat(tts): unlimited-length generation — sentence-boundary chunking + crossfade (Wave 1.2) (#357)
Ports voicebox's chunked TTS (MIT, attribution header) with two deliberate changes: the concat half is reworked for torch tensors (matching what our inference helpers feed the effect chain, incl. multi-channel on the last axis), and the sample rate comes from the engine's declared rate instead of the first chunk (fixes a latent upstream bug). Long text (> max_chunk_chars, default 800) splits at sentence boundaries (abbreviation/decimal-aware, bracket tags atomic, fullwidth enders via unicode escapes for the CJK gate) -> per-chunk generation with deterministic seed variation (seed+i) -> linear crossfade join (default 50 ms, 0 = hard cut) -> effect chain + watermark once on the joined audio. Wired into BOTH inference paths (OmniVoice-native _run_inference and the engine-adapter _run_backend_inference) beside the existing [pause] stitcher; [pause] inputs keep their dedicated path. Short text is byte-for-byte the old single-shot path; max_chunk_chars=0 disables. New /generate form params: max_chunk_chars (>=0, default 800), crossfade_ms (0-1000, default 50). Tests: 15 model-free unit tests (split priorities, abbreviation/decimal/ tag guards, crossfade math incl. multichannel + clamping) + 3 stubbed- engine endpoint tests (long text fans out with no words lost, short text single-shot, 0 disables). Endpoint tests validated in CI — this machine has a pre-existing local torch/Triton segfault on any main-importing test. Spec: voicebox deep dive 1 / parity program Wave 1.2 / #346 unlimited-length item. Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
93723c2789 |
feat(dictation): collapse Whisper hallucination loops in final transcripts (Wave 1.1) (#356)
Deterministic pre-pass ported from voicebox (MIT, attribution header): word-level (token repeated >=6x, punctuation-normalized) + character-level (2-60-char unit repeated >=6x, catches multi-word and no-space-script loops). Rhetorical repeats below 6 survive; no LLM involved; identical on every platform. Applied to the FINAL text in /ws/transcribe and POST /transcribe — segments keep raw recognition so timings stay truthful. Phase 1 of Spec 3 (docs/competitive-analysis.md); the optional local-LLM refinement pass (phase 2) lands with parity program Wave 2.1 in the same module. Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
7422f20a63 |
feat(profiles): consent-locked voice profiles — verified_own_voice + spoken consent flow (Wave 0.2) (#354)
* feat(profiles): consent-locked voice profiles — verified_own_voice + spoken consent flow (Wave 0.2)
A profile becomes 'verified own voice' when its owner records themselves
reading a consent statement (spoken attestation, not a checkbox). Agentic
features and gallery sharing will gate on the flag; plain local synthesis
never does.
- alembic 0003 (additive, PRAGMA-guarded, downgrade supported) +
_BASE_SCHEMA columns: verified_own_voice, consent_text,
consent_audio_path, consent_recorded_at
- POST/DELETE /profiles/{id}/consent — stores the recording as provenance
in VOICES_DIR ({id}_consent.*), replaces on re-record, cleans up on
revoke and on profile delete; 422 on empty statement / too-short audio
- VoiceProfile page: Verified badge + Voice ownership panel (record via
the existing useRecording denoise flow, revoke with confirm); en.json
keys only (other locales fall back per the advisory i18n parity policy)
Spec: docs/competitive-analysis.md Action 22 / parity program Wave 0.2.
Prerequisite for agentic v2/v3 and the persona gallery.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(profiles): harden consent paths against py/path-injection; drop lifespan in tests
- _voices_path(): resolve DB-stored filenames strictly inside VOICES_DIR
(bare-filename check + realpath containment); extension whitelist on the
uploaded consent filename (fallback .wav) so a crafted filename can never
steer the on-disk path. Applied to write, re-record cleanup, revoke, and
profile-delete cleanup. New test: malicious upload filename falls back.
- Test fixture no longer runs the app lifespan: startup/shutdown touched
module-level asyncio primitives bound to another module's event loop,
making the suite order-dependent in full-suite CI. init_db() is called
directly; endpoints under test need only the schema.
Fixes the CodeQL (3x py/path-injection high) and full-suite event-loop
failures on PR #354.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
65fc5245dc |
feat(dub): timeline segment editor — drag, snap-to-onset, keyboard a11y (#280) (#348)
* feat(dub): full-track speech-onset detection + GET /dub/onsets/{job_id} (#280)
detect_speech_onsets() lists every speech rise across the track (frame RMS,
adaptive threshold, 150ms hysteresis) — powers the timeline editor's
snap-to-onset ticks. Route prefers the Demucs vocals stem, falls back to the
mix, and caches onsets.json per job (mtime-invalidated).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat(dub): timeline editor math core — windowing, snap, clamp, fingerprint-safe commit (#280)
Pure helpers for the segment track: binary-search windowing, snapTime with
deterministic ties, neighbour/min-duration clamps with Alt-overlap (<=200ms),
commitMoveResize with fingerprint parity (move touches only start/end; resize
sets speed exactly like the old Regions handler and DELETES the key at 1.0 so
_canon_value's missing-vs-1.0 hashing can't mark untouched segments stale),
and overlap detection.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat(dub): SegmentTrack editing lane replaces the Regions plugin (#280)
Custom DOM segment boxes (6px edge handles, body-drag move, speaker colors,
stale/fresh tint, hatched overlap warning) virtualized by time over a single
{pxPerSec, scrollLeft} alignment source read off WaveSurfer's wrapper.
Snap-to-onset ticks on a viewport-sized canvas light up in snap range;
Ctrl/Cmd-wheel zooms centered on the cursor; double-click plays the slot via
playRange (timeupdate watcher pauses at slot end). Roving-tabindex listbox
keyboard model (arrows / Enter / Shift / Alt / Delete / S) with polite
aria-live announcements. WebKit fallback keeps a self-scrolling lane at a
fixed px/sec. timeline.* strings translated in all 21 locales.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat(dub): wire timeline editor — per-gesture undo, id fix, table selection sync (#280)
segmentMoveResize() pushes undo ONCE per gesture (drag commits on pointerup;
keyboard nudges coalesce per focus session) and matches by String(id) — the
old parseInt('seg-3_a') path edited the wrong segment after a split. Commits
go through commitMoveResize for fingerprint parity, and the existing
recomputeIncremental effect picks up every commit. Clicking a timeline box
scrolls + highlights its row in DubSegmentTable; 'preview dub here' parks
the player at the slot start, then synthesizes the line.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(dub): inline the onsets-cache containment guard — CodeQL can't track helpers
Same lesson as #328/#329: the realpath+startswith sanitizer must sit at
the sink, not behind a function return. Unused helper removed.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: mergetest <test@local>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
4b21f82619 |
feat(dub): Smart Fit timing strategy — planner, fingerprints, generate path (phase A) (#347)
* feat(dub): Smart Fit planner, fit fingerprints, shared ffmpeg stretch helpers - services/fit_planner.py: pure, I/O-free planner for dub-length fitting v2 — slack absorption (gap guard), audio-only band (<=1.2x), geometric 50/50 audio/video split capped at 1.5x / 2.0x, residual overflow accounting, and a stretch_video-compatible video_plan + fitted timeline cursor. Clean-room reimplementation from a published description. - services/incremental.py: fit_fingerprint() over the fit params with the same _canon_value canonicalisation as segment hashes (#281 class). Fit params stay OUT of segment_fingerprint — a fit change re-mixes, never re-TTSes. - services/ffmpeg_utils.py: move _atempo_chain/_pitch_preserving_stretch out of the dub_generate router (lazy torch/numpy imports) so the Phase B export pipeline can reuse them; add probe_duration() ffprobe helper. - schemas/requests.py: timing_strategy gains "smart_fit"; optional fit_options knob overrides default server-side. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(dub): smart_fit branch in the generate path TTS loop unchanged (dur_s=None, natural-rate WAVs on disk). After the loop, plan_fit() decides per segment; the mix loop applies audio_rate via the pitch-preserving atempo pipe (linear-interp fallback), trims residual overflow with the existing fades, and places audio at the planned new_start on a fitted-length canvas. Truthful fit_status entries (audio_rate / video_ratio / overflow_s) feed the row badges. Persists job["fit_plans"][lang] = {plan (exact _build_video_stretch_filter_graph shape), fitted_segments (cue times from ACTUAL stretched sample positions), total/orig duration, params, fit_fp} and mirrors fit_fp on dubbed_tracks[lang]. video_stretch_plans untouched. Strategy-transition guard: job["seg_wav_kind"] records whether on-disk seg WAVs are natural or slot-squeezed; a smart_fit partial regen over slotted (or unknown) WAVs forces one full regen instead of double-compressing. Old strategies and old persisted jobs are byte-identical (all new reads via .get()). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(ui): Smart Fit option in the dub timing picker (all 21 locales) - prefsSlice: TimingStrategy union gains 'smart_fit'; optional FitOptions overrides (null by default — backend defaults apply identically on every platform); persisted alongside timingStrategy. - DubTab: Segmented gains Smart Fit with i18n label + tooltip. - useDubWorkflow: sends fit_options only when set and strategy is smart_fit. Default strategy stays 'concise' — no default behaviour change on any platform. - locales: dub.timing_smart_fit{,_title} translated in all 21 languages. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(dub): fit planner unit + golden suites, smart_fit generate-path integration - test_fit_planner.py: threshold boundaries (0.9/1.0/1.2/1.21/4.0), cap saturation -> overflow, slack absorption incl. gap guard, last-segment tail, cursor monotonicity, allow_video_retime=False, video_plan fed straight into _build_video_stretch_filter_graph, fit_fingerprint canonicalisation (int vs float, omitted vs default — the #281 class) and a pinned stable digest. - tests/fixtures/fit_planner/*.json: 4 golden FitPlans; algorithm drift is a deliberate fixture diff, never a silent change. - test_smart_fit_generate.py: hermetic end-to-end runs (mock TTS, no ffmpeg) covering audio-only stretch, hybrid timeline growth + persisted plan shape, fit_options override, strict_slot->smart_fit forced regen then zero-TTS fit-only re-mix, and concise back-compat. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(competitive): dub-length fitting row reflects Smart Fit Phase A Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(incremental): mark fingerprint hashes usedforsecurity=False — dedup keys, not security (Bandit) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
3c780dced9 |
feat(dub): speech-onset alignment + regional dialect targeting (#280) (#330)
Items 1 and 2 from the improvement list: 1. Synchronization — Whisper-family ASR stretches segment starts back over leading non-speech (intro music, silence), so the dub starts at 0:00 while the speaker starts at 0:02-0:03. New onset_align service snaps each segment start forward to the first audible vocal onset (adaptive RMS threshold over the Demucs-isolated vocals when available). Forward-only and conservative: never moves a start earlier, ignores sub-100ms shifts, preserves minimum duration, leaves silent-window segments untouched. Pure NumPy — identical across platforms. 2. Accent/vocabulary by country — a Dialect picker in the Dub panel (BCP-47 codes per target language) injects a regional instruction into LLM translation prompts (OpenAI/Ollama engines and the Cinematic refine pass): Argentina yields 'Vos sos muy listo', not 'Tú eres muy listo'. Non-LLM engines show a clear hint that the dialect needs an LLM. New i18n keys translated in all 21 locales. Item 3 (segment rectangles: move/crop/stretch on the timeline) is a larger editor feature and stays open on #280. Co-authored-by: Claude Fable 5 <noreply@anthropic.com> Co-authored-by: mergetest <test@local> |
||
|
|
c0924f5eba |
fix(tts): torch.compile failures fall back to eager — generation never fails on unsupported GPUs (#278) (#327)
* fix(tts): torch.compile failures fall back to eager — generation never fails on unsupported GPUs (#278) On GPU architectures the bundled Triton doesn't support (e.g. Blackwell sm_120 / RTX 5060), the compiled model dies mid-generation inside the Dynamo/Inductor/Triton/cudagraph stack — previously surfaced as a fake 'ran out of memory' error and a dead Archetype preview. Now: - up-front arch gate: skip compile when the GPU's compute capability is not in this torch build's arch list (OMNIVOICE_FORCE_TORCH_COMPILE=1 overrides for PTX forward-compat setups) - runtime fallback: model.generate is wrapped once; a compile-stack failure (classified by exception chain: module, message, traceback paths — the cudagraph case is a bare AssertionError) logs a warning, restores the eager module, disables compile for the session, resets dynamo state, and retries eagerly. Non-compile errors propagate unchanged. - the /generate OOM handler no longer mislabels compile crashes as OOM and points users at the actual remedy. Fixes #278 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Potential fix for pull request finding 'CodeQL / Empty except' Co-authored-by: Copilot Autofix powered by AI <62310815+github-advanced-security[bot]@users.noreply.github.com> * Potential fix for pull request finding 'CodeQL / Empty except' Co-authored-by: Copilot Autofix powered by AI <62310815+github-advanced-security[bot]@users.noreply.github.com> * Update backend/api/routers/generation.py Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> Co-authored-by: Copilot Autofix powered by AI <62310815+github-advanced-security[bot]@users.noreply.github.com> Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com> Co-authored-by: mergetest <test@local> |
||
|
|
853b9eefc7 |
fix(dub): burn translated subtitles, fix subtitle save JSON error (#309) (#328)
* fix(dub): burn translated subtitles, fix subtitle save JSON error (#309) Two symptoms, one root: the job kept the original-language ASR transcript while the editor only sent translated/edited text in the generate request. - dub_generate now persists the segments the dub was actually generated from back onto the job (metadata carried over by stable id, fallback index; text_original retained for dual-subtitle layouts) — SRT/VTT export and ffmpeg burn-in now render the dub language, not the source. - The SRT/VTT export endpoints honor the save_path query param the Tauri save dialog appends (like every other export) and return the standard JSON envelope — previously they ignored it and returned the raw body, so the frontend's JSON.parse choked on the SRT cue index ('Unexpected non-whitespace character after JSON'). - Frontend guards the save response content-type so any future raw-body response surfaces as a clear error. Fixes #309 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Potential fix for pull request finding 'CodeQL / Uncontrolled data used in path expression' Co-authored-by: Copilot Autofix powered by AI <62310815+github-advanced-security[bot]@users.noreply.github.com> * fix(dub): use the file's established realpath+startswith containment idiom (CodeQL) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(dub): write subtitle saves from the Tauri process, not the backend (#309) The backend save_path variant on /dub/srt and /dub/vtt routed a user-controlled destination through the loopback HTTP surface — six new CodeQL path-injection flows plus two log-injection flows. Subtitles are small text bodies, so the frontend now fetches them raw and writes the file via a new save_text_file Tauri command: the OS save dialog in the trusted process is the write authorization, and the backend never sees a destination path. Binary exports keep the established save_path flow. Also strips newlines from user-derived values in the two flagged log lines. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(dub): leave _native_save byte-identical to main The newline-strip on the log line moved a path sink onto a changed line, which made CodeQL re-attribute the long-standing binary-export flow to this PR as a new alert. The subtitle endpoints no longer feed this function at all, so restore the exact original line — the baseline alert stays baseline, and hardening pre-existing flows belongs in its own PR. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> Co-authored-by: Copilot Autofix powered by AI <62310815+github-advanced-security[bot]@users.noreply.github.com> |
||
|
|
668d824e86 |
feat(setup): unified first-run journey — install gate, studio-console wizard, platform awareness (#295)
* feat(setup): first-run install gate — nothing installs until the user confirms a plan New `setup` module parks first runs in BootstrapStage::AwaitingSetup instead of auto-installing. complete_setup validates the user's InstallPlan and only then starts the existing bootstrap: - install modes: installed (platform dirs) / portable (one folder next to the exe / AppImage, config.json travels with it) - user-chosen storage: env dir, data dir (OMNIVOICE_DATA_DIR), model cache (OMNIVOICE_CACHE_DIR) — None = legacy default, byte-identical behavior - minimum-space gate: per-volume free-space check (fs4 statvfs), grouped by filesystem so dirs sharing a disk sum their requirements; install refused when short (9 GiB env + 7 GiB models + 1 GiB data, measured + headroom) - custom mirrors (PyPI index, HF endpoint, python-build-standalone) take precedence over region presets in the venv/sync/backend env wiring - ROCm torch variant selectable via config (env var still wins) - existing installs migrate silently: venv present → setup_complete=true, no questions re-asked; dev trees skip the gate entirely 19 unit tests (disk probing, space grouping, mirror validation, legacy config compat). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(setup): first-run setup screen — mode, storage with space gate, mirrors, compute FirstRunSetup renders when the Rust side reports awaiting_setup (lazy-loaded; regular launches pay nothing). One screen, defaults all work: - language picker first (rest re-renders translated), 21 locales shipped - Installed / Portable mode cards (portable disabled with reason when the exe-adjacent folder isn't writable) - storage rows with live per-path free-space probes (debounced check_install_target), 'needs ~X / Y free' readouts, folder pickers - client mirrors the Rust per-volume space gate: Start installation is disabled with an explicit reason until every volume fits - compute (CUDA-auto / ROCm), update channel, region + custom mirror URLs - complete_setup errors surface inline; on success the normal bootstrap progress UI takes over on the next status poll Verified on a wiped machine: gate parks (no spawn, no downloads), screen renders, 450 GB ≥ 17 GB requirement → Start enabled. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(setup): studio-console redesign of the first-run screen The setup screen now reads as powering on studio hardware rather than a web form — true to a voice studio, and self-sufficient offline (every font and asset is bundled; a first run may be on a restricted network): - breathing waveform masthead (CSS-only, deterministic speech-cadence silhouette, staggered per-bar delays) - Source Serif 4 display headline + engraved IBM Plex Mono panel labels + Inter body — the three faces the app already ships - rack-unit panels with corner screws, engraved title rules, serial plate (OVS · vX.Y.Z) - disk space as segmented LED capacity meters: lit = what the install consumes, alarm-blink red on insufficient volumes - mode cards with indicator LEDs; 'armed' Start button — LED lights and a halo pulses only once every volume passes the space gate - atmosphere: corner accent glows + SVG film grain; staggered rise-in choreography on load - all motion transform/opacity only; prefers-reduced-motion holds every frame still; theme-token derived colors; focus-visible rings throughout No logic changes: same IPC calls, same i18n keys, same space-gate math. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(setup): wide desktop deck, hardware-aware Compute + Update channel cards Three pieces of feedback addressed: - width: the console is now a 1240px two-column deck (storage rail left, decision rail right) that uses desktop real estate; collapses to one column under 980px and stacks fully under 620px - no outer chassis box: panels float directly on the atmospheric backdrop, each carrying its own rack-unit treatment - Compute and Update channel split into separate cards with real information: get_setup_state now detects hardware (nvidia-smi → CUDA name, /sys/class/drm vendor 0x1002 → AMD/ROCm, Apple Silicon → MPS, CPU cores + RAM via sysinfo; best-effort, never blocks) — the Compute card shows a live 'Detected: …' readout, badges the option that matches the machine, and pre-selects ROCm on AMD boxes; both cards use LED radio options with full descriptions (6 new i18n keys × 21 locales) Also pins playwright-core as an explicit devDep — bun did not materialize it through @playwright/test, breaking programmatic browser use. 20/20 Rust tests · vite build · CJK guard green. Verified live (gate engaged, responsive single-column) and at 1600×1000 via mocked-IPC browser shot (two-column deck). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(setup): move network (region + mirrors) into the masthead with language Language and download region are the two 'where am I' choices — they now sit together top-right of the masthead, with the custom-mirrors disclosure tucked beneath the subtitle. The Network panel is gone, leaving a balanced deck: Install mode + Storage left, Compute + Update channel right. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(setup): strip the boxes — fills and rules carry the structure One design rule now: borders only where state demands them. Panels lose their boxes entirely (engraved mono title + rule separates sections); option cards, storage rows, selects/inputs, the hw readout, the version plate and the ghost buttons are all flat fills; active options glow with an accent tint + LED; blocked rows and errors use a red tint + 2px inset edge bar instead of a border. The badge chip is fill-only too. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(setup): quiet pass — every element earns its visual weight - waveform becomes a whisper: 22px trace, 2px bars, ~half opacity — an ambient signature instead of a billboard - storage readouts collapse to one mono line ('needs ~9 GB · 449 GB free'); the LED meter now appears only when it carries information (install would consume >35% of free space, or the volume is blocked) — at 449 GB free a bar was a meaningless sliver - Change… buttons go text-quiet (transparent until hover) - custom-mirrors disclosure right-aligns under the region select it extends, instead of floating under the subtitle - version plate moves to the footer next to the disk total — the masthead keeps only title, subtitle, and the two locale/region selects Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(setup): platform-matrix awareness — distro+arch detection, ROCm gated to Linux, no Windows console flash The install matrix is OS family × distro × arch × GPU vendor, and the setup screen now both shows it and only offers choices valid for it: - HardwareInfo gains os_name (distro PRETTY_NAME from /etc/os-release on Linux, macOS/Windows elsewhere) and arch (x86_64/aarch64) — the detected line reads 'CachyOS x86_64 · NVIDIA RTX 4070 · 32×CPU · 31 GB RAM', exactly what bug reports cite - SetupState gains os; the ROCm option renders on Linux only (wheels don't exist elsewhere) and complete_setup clamps rocm→auto on non-Linux as the server-side backstop - nvidia-smi probe gets CREATE_NO_WINDOW on Windows — no cmd flash on the first screen a user ever sees - Apple Silicon → MPS, Intel mac → CPU, ARM Linux → CPU: all matrix cells resolve through the same base constructor Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(setup): unify the whole first-run journey under the studio-console system Setup → Installing → Model wizard now read as one continuous experience: the same atmosphere, whisper waveform masthead, serif/mono type, LED language and quiet fills across all three acts. - Installing (BootstrapSplash): rebuilt in frs-* — segmented LED journey meter (completed steps + live byte progress), LED step rail (done=green, active=pulsing accent, pending=dim), engraved ACTIVITY panel with the quiet mono log (collapse/copy as text-quiet actions), failure act with red-tint error + hints + armed Retry. All logic untouched: stage poll, event subscription + backfill, dedupe, hints, region/language selects. - Model wizard (SetupWizard): same masthead with the step rail as engraved mono LED steps top-right, welcome cards as option-card surfaces, preflight as LED check rows (pass/warn/fail), frs nav buttons with armed primaries, embedded Model Store / Engines / Dictation panels scroll inside the act. Old 556-line stylesheet replaced by ~60 lines of glue; BootstrapSplash.css reduced to a resolving stub. - FirstRunSetup.css is now the journey's shared design system (step rails, log panel, banners, hints, wizard chrome, check rows appended). - 2 new strings (Installing / Activity) translated across all 21 locales. Validated end-to-end on this machine: setup screen → Start installation → real venv bootstrap (~10 min) → backend healthy on 3900 → model wizard. 20/20 Rust tests · vite build · CJK guard green · installing act verified via mocked-IPC screenshot at stage=installing_deps. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(setup): --setup re-entry flag + make the install-plan screen un-stealable The setup stage is first-run-only by design (completed installs skip it), but it must be reachable on demand and must actually win the mount when engaged. Three fixes: - 'omnivoice-studio --setup' parks the bootstrap in AwaitingSetup on any launch — checked before the attach-to-healthy-backend shortcut, so a running backend can't skip past it - App routing: awaiting_setup now outranks everything (a live backend answering /setup/status used to route straight to the model wizard); the wizard additionally requires stage === 'ready' so it can't mount during the initial stage race - useBootstrapStage: a transient IPC miss no longer permanently declares 'ready' (which killed the poll loop and silently skipped the setup / progress screens) — it retries up to 5 ticks before conceding Plus journey-wide titlebar clearance (content never sits under the GTK headerbar / macOS traffic lights / Windows controls) and drag-region mastheads on all three acts. Verified: mocked-IPC harness with stage=awaiting_setup + a LIVE backend answering /setup/status renders the setup screen, not the wizard. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * style(setup): remove backdrop decoration — flat surface, state-only emphasis The corner accent glows and SVG film grain rendered as visible banding / noise artifacts on many panels — both gone; the journey now sits on a clean flat chrome background. Also swept the remaining decorative bloom: the active option card drops its glow shadow (flat accent tint + LED carry the state), and the armed Start button loses its pulsing halo (the lit LED already signals actionable). Remaining shadows are functional micro-detail only: 6px LED glows, meter track inset, red edge bars. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(setup): journey rail + verbosity diet — clean, smooth, elegant The setup page is now visibly stage 1 of the install flow: a quiet breadcrumb rail (SETUP → INSTALLING → MODELS & ENGINES) sits between the waveform and the headline on both the setup and installing acts, LEDs marking done/active/pending — one continuous story across the journey. Verbosity halved without hiding information: - option descriptions unfold (260ms ease) only on the selected card; the page shows exactly one explanation per group, collapsed cards keep the text as a tooltip - storage rows drop their always-on caption (label + path + readout + Change… on one line; caption lives in the row tooltip) The whole page now fits a laptop window without scrolling. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(setup): merge Models + Engines into one wizard act Two tabs weren't necessary: models are the required gate, engines the optional extras — now two stacked panels in a single 'Models & engines' step (label reuses the journey-rail key, translated in 21 locales). Wizard shrinks to 4 steps: Welcome → System check → Models & engines → Dictation. Continue still gates on models_ready only; engines stay optional. Welcome cards updated to the 3 remaining acts; static cards keep their descriptions visible (the active-only fold is for radios). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(setup): wizard was skipped after first-run install — probe /setup/status on bootstrap ready The models-needed probe started at mount with a ~30s retry ceiling. On a first run, mount happens at the setup page — by the time the user reads it and the multi-minute install finishes, the attempts were long burned, so setupChecked landed as 'no wizard needed' and the studio rendered with zero models on disk. The probe is now keyed on bootstrapStage and runs when it hits 'ready' — the first moment a backend exists to answer. Normal launches (backend up quickly) behave exactly as before. Caught by running the full journey three times end-to-end: rounds 2–3 skipped Models & engines after install; with the fix the wizard mounts with models_ready=false (Whisper large-v3 listed missing). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(setup): drop the Welcome step — wizard opens on System check The welcome act had nothing left to say: the journey rail names the stages, the setup page already oriented the user, and the cards repeated both. The wizard is now three steps — System check (auto-runs on mount) → Models & engines → Try dictation — landing the user directly on live preflight results instead of a page about the pages to come. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(setup): true unified library — models + engines as ONE list 'Merge them' meant one list, not two panels stacked — fair criticism. The wizard's Models & engines act is now a purpose-built WizardLibrary: every installable is a row of the same grammar (LED · name · chip · size · action): - required models lead (REQUIRED chip, Download action, live SSE progress bar + percent, green LED when installed) — they gate continue - TTS engines follow (ENGINE chip): active engine glows accent, available ones offer one-click Use (selectEngine), heavy installs defer honestly to Settings ('install later in Settings' + reason tooltip) - the optional-model tail folds behind 'Show N optional models' The full management surface (search, HF token, deletes, sorting) stays in Settings — a first run needs a checklist, not a store. 9 new strings × 21 locales. Verified against the live backend via the browser harness: required/installed/engine/active/Use/defer states all render in one list. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(diagnostics): local-first self-check, error journal, and bug-report pipeline (#296) * feat(diagnostics): local self-check + scrubbed bug-report pipeline Closes the gap between 'something broke' and 'a useful GitHub issue exists' — entirely within the local-first constraint: the only outbound path remains the user's own browser opening a prefilled issues/new URL. Backend: - core/scrub.py: privacy scrubber for anything leaving the machine — env-var secret values (*TOKEN*|*KEY*|*SECRET*|*PASSWORD*), credential shapes (hf_/ghp_/github_pat_/sk-), home dirs on all three OSes - core/diagnose.py: 9-check self-check (device+GPU, ffmpeg, HF token, disk, data-dir writability, RAM, engine registry, hub reachability), pre-scrubbed, ASCII-safe output - GET /system/diagnose + 'python main.py --diagnose' (exit 0/1) - /system/info: hardware inventory (os_version, cpu_model, cpu_count, ram_total_gb, gpu_name, vram_total_gb, disk_free_gb), cached statics Frontend: - utils/bugReport.js: single source for the prefilled-URL builder — scrubText twin, hardware context capture, scrubbed error+stack embed, URL-length cap; ReportBugButton refactored onto it - ErrorBoundary 'Report this bug' action with the error attached - utils/errorToast.jsx toastErrorWithReport(); wired into export toasts - Settings > About 'Run self-check' with per-check status badges Tests: 27 pytest (scrub, diagnose) + 15 vitest (bugReport); existing suites green; verified live (--diagnose, TestClient, vite build). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(diagnostics): error journal, diagnostic bundle, crash notice, global handlers Second slice of the bug-tracking work — still zero outbound paths beyond the user's own browser/file manager. - core/error_journal.py: deduped ring of recent unhandled backend errors (fingerprint counts, error_class triage: GPU_OOM, HF_AUTH_FAILED, PYANNOTE_LICENSE_REQUIRED, DISK_FULL, FFMPEG_MISSING, NETWORK_ERROR), scrubbed, JSONL-persisted so the error that killed the last run survives restart. Wired into the global exception handler; 500 bodies now carry error_class; GET /system/errors/recent. - core/diagnostic_bundle.py + POST /system/diagnostic-bundle + Settings > About 'Save diagnostic bundle': zip of self-check report, error journal, scrubbed log tails — drag onto a GitHub issue; bypasses the ~8k prefill-URL ceiling. - crash-on-next-launch: /system/notifications flags a crash logged before this session started (size vs acked-size in prefs, mtime vs process start); POST /system/crash/ack; LogsFooter acks on action click. - utils/globalErrorHandlers.js: uncaught errors + unhandled rejections get a throttled, noise-filtered 'Report this bug' toast. - sidecar log parity fix: _tauri_log_candidates() now lists the Rust sidecar's backend.log/backend_err.log on Linux (XDG state dir) and Windows (LOCALAPPDATA) — sidecar crashes were only visible on macOS. Tests: +19 pytest (journal, bundle); suite at 102 passed. Vitest 124 passed; vite build green. Live-verified: journal recorded and classified a real HF 401 from the test run (HF_AUTH_FAILED, paths scrubbed). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(diagnostics): breadcrumbs, deep self-check, report sweep, issue search Final slice of the bug-tracking work. - toastErrorWithReport adopted at the high-traffic failure sites: TTS generation, dub upload/ingest/transcribe, engine install, engines-matrix load, voice profile save/delete/test, batch enqueue/cancel/delete. Validation toasts and cancellations stay plain on purpose. - utils/breadcrumbs.js: local-only ring of the last 20 action names (closed-set names only — never content or paths), embedded as a 'Recent actions' section in the prefilled report. Instrumented: view changes, generate, dub pipeline, export, engine switch. - deep self-check: /system/diagnose?deep=true and --diagnose --deep load the active engine and synthesize a short utterance (num_step=4) — catches 'installed but broken'. 180s time-box, skips during model load, scrubbed failure detail. Verified live: cold-loaded omnivoice and produced 2.2s of audio in 43.9s on CUDA. - 'Search similar issues' action on the ErrorBoundary: scrubbed, noise-stripped GitHub issue search URL — dedupe before filing. - bug_report.md template now points at the diagnostic bundle and the --diagnose CLI so manual reports arrive with the same evidence. Tests: pytest 107 passed (4 new deep-check tests, CJK gate green); vitest 218 passed (breadcrumbs + issue-search suites); vite build green. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs(diagnostics): self-diagnosis section in troubleshooting + README pointer Settings > About self-check / --diagnose / --deep / diagnostic bundle are now the documented first step before the per-error entries — and the support team's first ask on every issue. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(setup): flush sticky action bar, global dbl-click maximize, open maximized First-run polish on the studio-console journey: - FirstRunSetup: fixed-footer / scrollable-middle layout — mast + decision grid live in a dedicated .frs__scroll region; the install action bar is the last flex item, so it sits flush at the window's bottom edge and nothing (e.g. an expanded compute-option description) can render beneath it on small windows. - Double-click-to-maximize on the custom borderless titlebar now works on EVERY drag region (splash, first-run, wizard, main header) via one delegated listener in main.jsx, on all platforms; removed App.jsx's redundant inline handler so it doesn't double-toggle. Skips interactive controls in the bar. - Window opens maximized to the available desktop size (tauri.conf.json). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(diagnostics): quiet Bandit on the journal hash and hub probe The journal fingerprint is a dedup key, not a security boundary — usedforsecurity=False. The hub reachability probe gets an explicit https scheme guard on its constant URL so the urlopen sink is audited. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(setup): address PR #295 review findings — security, lifecycle, privacy, i18n Security: - setup.rs valid_mirror: reject plaintext http:// mirror URLs (MITM supply-chain path into UV_PYTHON_INSTALL_MIRROR / UV_INDEX_URL / HF_ENDPOINT); explicit http://localhost / 127.0.0.1 / [::1] exceptions only. Tests extended incl. loopback-lookalike hosts. - setup.rs detect_hardware: AMD vendor ID alone no longer maps to kind="rocm" — a cheap ROCm userspace probe (/opt/rocm or rocminfo on PATH) gates it; bare AMD GPUs report kind="amd" so the UI offers ROCm without pre-selecting it ("matches this machine" only when verified). Functional: - lib.rs/setup.rs --setup re-entry: complete_setup now kills any backend still serving on the port before retry_bootstrap, so changed env/mirror/layout settings actually apply instead of re-attaching. - setup.rs: nvidia-smi probe runs behind a 3 s timeout thread — a wedged driver degrades to CPU instead of hanging the first-run IPC. - setup.rs: is_first_run is now a pure read; the existing-install migration write moved to migrate_existing_install_if_needed, invoked only from the bootstrap thread (get_setup_state no longer writes). - setup.rs complete_setup: config save errors now abort setup and surface in the UI instead of bootstrapping into a stale on-disk layout. - setup.rs complete_setup: logs default-vs-custom flags instead of the user's absolute env/data/models paths (privacy rule). - scrub.py + bugReport.js: also redact forward-slash Windows homes (C:/Users/<name>, file:///C:/Users/...), ordered before the macOS pattern so "C:~" residue can't form. Tests added on both sides. - bugReport.js: context fetches bounded by a 2.5 s AbortController timeout so report assembly degrades to partial context instead of hanging on a stalled backend. - system.py: crash ack is now {size, mtime} (legacy size-only ack still honored) and /system/logs/clear drops the ack — truncation can no longer permanently suppress 'crash-last-session'. - system.py: Linux Tauri-log probe honors XDG_DATA_HOME. - setup.ts/WizardLibrary.jsx: SetupProgressEvent type now documents the full phase taxonomy actually emitted (per-file start/progress/done + install_*/delete_* lifecycle); reducer verified correct against the backend stream and annotated — a file-level 'done' must not clear the repo row. - SetupWizard.jsx: step rail clamps to the highest unlocked step (preflight/models gates) — no more jumping straight to "Enter studio". Polish: - BootstrapSplash.jsx: Waveform heights wrapped in useMemo([bars]) like its siblings. - BootstrapSplash.jsx: detectHints returns i18n keys (bootstrap.hint_*) rendered through t(); translated in all 21 locales. - SetupWizard.jsx: step rail aria-label localized (setup.step_aria / setup.step_completed) in all 21 locales. - FirstRunSetup.css: deprecated word-break: break-word → overflow-wrap: anywhere; reduced-motion override also stops the frs-hw-pulse LEDs (.frs-step.is-active LED + .swiz-lib__led--busy). Deferred (design-level, follow-up PR): --setup re-entry round-tripping of custom dirs/mirrors into the form (setup.rs), and worker-thread leak on timed-out deep checks (diagnose.py). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(i18n): translate back-filled keys in all 20 locales, drop inline fallbacks The reconciliation merge back-filled 16 new keys (about.self_check*, about.*bundle*, dub.num_speakers_*, errors.*) with English text in every non-English locale — CodeRabbit flagged 9 locales; fixed all 20. Interpolation tokens preserved and asserted during the rewrite. Also removed the two inline English fallback strings in App.jsx (firstrun.first_sound_*) so copy lives only in locales/*.json. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Co-authored-by: mergetest <test@local> |
||
|
|
2ef42ee629 |
feat(design): free-text 'describe your voice' field maps to design parameters (#317) (#331)
Parity with the hosted omnivoice.app describe field, implemented fully locally: a deterministic, ordered synonym-table mapper (no model, no network, stdlib only) projects a natural-language description onto the existing six-category design space (Gender/Age/Pitch/Style/EnglishAccent/ ChineseDialect). Every emitted token is validated at import time against the engine taxonomy, so the mapper can never produce an instruct item the engine validator would reject; Chinese token forms are derived from the taxonomy, never hardcoded (the one functional pinyin->dialect mapping is allowlisted in test_no_hardcoded_cjk.py with justification). UI: a describe textarea in the Design tab fills the attribute picker live (hand-tuning still possible afterwards); parts of the description the taxonomy can't express are listed back to the user as 'ignored' instead of failing silently. New i18n keys in all 21 locales. Fixes #317 Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
48ae4dae1d |
fix(dub): re-dub honors transcript edits — fingerprints canonicalised, preview cache-busted, atomic mux (#281) (#329)
* fix(dub): re-dub honors transcript edits — fingerprints canonicalised, preview cache-busted, mux made atomic (#281) Three symptoms, three causes: 1. Edited line, unchanged result: the dubbed preview-video URL was identical across re-dubs, so the WebView kept serving the previous dub. A generation nonce now cache-busts the preview after every completed generation. 2. Preview stuck loading forever: overlapping preview requests ran ffmpeg against the same output path and the mtime cache check saw the half-written file as valid. The mux now runs under a per-path lock, writes to a temp file, and os.replace()s into place. 3. One edit re-dubs all lines: server-side fingerprints were computed from pydantic-parsed segments (defaults filled in) but recomputed client-side from raw dicts (keys omitted), so every segment always looked stale and incremental degraded to a full re-dub. Values are now canonicalised on the backend and the frontend builds generation inputs through one shared helper (utils/segments.js) for both the generate request and the incremental plan. Fixes #281 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Potential fix for pull request finding 'CodeQL / Uncontrolled data used in path expression' Co-authored-by: Copilot Autofix powered by AI <62310815+github-advanced-security[bot]@users.noreply.github.com> * fix(dub): realpath containment for job-derived preview paths (CodeQL) Request-supplied job_id/lang flowed into the preview mux output path. Both now pass a realpath containment guard against DUB_DIR (the file's existing per-segment pattern) and lang is allowlist-validated before it lands in a filename. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(dub): inline the containment guard — CodeQL can't track it through a helper Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> Co-authored-by: Copilot Autofix powered by AI <62310815+github-advanced-security[bot]@users.noreply.github.com> |
||
|
|
433f1ba617 |
fix(tts): /generate honors the selected TTS engine (#312) (#324)
* fix(tts): /generate honors the selected TTS engine (#312) The /generate route always ran the OmniVoice model directly, ignoring both the Settings engine selection and any per-request override. It now resolves the active backend (env var > Settings selection > default), supports an explicit `engine` form field (same pattern as /ws/tts and /v1/audio/speech), reuses the per-process engine instance cache, keeps inline [pause Nms] markers working on every engine, and honors applies_own_mastering so studio engines skip the broadcast mastering chain. The OmniVoice default path is byte-identical to the old behavior — existing API consumers see no change. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(312): resolve modules at run time, drop lifespan client — fixes full-suite isolation tests/backend/** runs before tests/test_*.py and pollutes sys.modules (re-imports the services tree), so module-level imports bound at pytest collection pointed at a stale services.tts_backend — registry patches landed on a dict the routes no longer read ('Unknown TTS engine' in CI). Modules are now resolved through sys.modules inside each test. The client fixture also drops the module-scoped lifespan context manager that bound event_bus queues to this module's loop (teardown 'Queue bound to a different event loop') — plain function-scoped TestClient, the test_api.py pattern. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |