d91beef0fd314250d8d9b94de86dfea019a8bd96
21
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
bda169c900 |
feat(workers): pin work to the chosen GPU, and say when its model is missing
Three phases that only make sense together: a job that names a worker, a worker that reports honestly what it can actually run, and the small defects that made both lie. **Pinning** (Phase 1). `pinned_worker_id` is now honoured in both places that choose a worker — `eligible_workers` and `select_worker` build independent lists, so applying it to one silently leaked work onto whichever machine was least busy. The pin persists across a restart via an additive column, deliberately not alembic (justified in the code, per the precedent already in db.py): quitting mid-render used to drop it without a word. `max_attempts=1` was rejected as the mechanism — it makes the FIRST failure terminal, including the penalty-free ones a stale advisory view produces routinely. Cancel now actually reaches the worker. `WorkerServicer.cancel` had zero callers, so cancelling released the slot while the GPU thread kept running, and a late result could resurrect the task as COMPLETED — `commit_result` assigned that state directly, bypassing the transition table where CANCELLED is terminal by construction. **Honest capabilities** (Phase 4). A worker now probes whether weights are actually present, and a job stops BEFORE dispatch with a typed 409 naming the model and the machine, instead of failing mid-task. The probe fails OPEN: `is_cached`/`cache_is_complete` cannot see a user-managed clone outside the HF layout, so only a positive "absent" refuses. Refusing an engine that works today would break the compatibility promise. `pool.supports` deliberately still ignores `downloaded` — had it not, the scheduler would drop the worker and answer with a terminal NO_CAPABLE_WORKER, which tells the user to check their install when the truth is one download away. The frontend no longer offers "Report this bug" for that state; it offers the download. Catalog tags resolve against the TARGET's OS/arch/backend, not this machine's. From a Mac control plane, a CUDA worker's model list was showing the mlx-community repos it cannot run and hiding the ones it needs. **And the quiet ones** (Phase 0 leftovers): a model's human label rides its own proto field so renaming it cannot orphan breaker history; an empty model_id no longer forks the capacity slot key into two slots for one model; the idle sweep cannot evict an engine out from under a live LOCAL render. Verified on real hardware, which is the only verification that has ever caught anything here: 2025 characters, default settings, routed to an RTX 4090 over the wire — 100% GPU utilisation on the remote box, 119.6 s of 24 kHz audio returned in 16.6 s, 5.7 MB delivered out of band through the artifact path rather than the control stream. Backend 5259 passed, frontend 1808 passed. |
||
|
|
4f4d9c6e3e |
refactor(workers): give Remote workers its own System entry; ignore remote/
Remote workers was nested under Sharing, which reads backwards: everything in Sharing is about letting something else reach THIS machine (a remote backend, an MCP client, a share PIN), while remote workers sends work OUT to machines you own. It is now its own System entry. Docs-sync: every "Settings → Sharing → Remote workers" reference is updated — the guide, the changelog, the two API error messages that tell a user where to generate a token, and the agent's not-enrolled error. Also ignores remote/ (local goal docs, review briefs, council reports) and repoints the code comments that cited remote/goal_v2.md at the shipped docs/remote-workers.md, so no committed file references a path that is not in the repo. |
||
|
|
43de1c794c |
feat(workers): remote GPU workers over a versioned gRPC protocol
Send individual jobs to GPUs on your other machines while everything else stays local. Opt-in, off by default: with the toggle off there is no listening socket, no certificate and no background loop. Design follows remote/goal_v2.md, the council-revised goal doc. The decisions that shaped the code, and why: * A disconnect is an unknown outcome, not a failure. The original design reassigned on disconnect while also describing the case where the worker had already finished — following both guarantees duplicate execution. An attempt now holds a grace window; a worker returning inside it commits its result and no second attempt is ever made. * At-least-once execution, exactly-once result commit. The result is persisted BEFORE it is acknowledged, so a crash between the two cannot silently lose a finished render. * Deadlines are phased (accept -> model load -> execute -> deliver) and liveness is a progress lease. The old fixed 30s execution budget was two orders of magnitude below what this product actually does; silence is the failure signal, not slowness. * Capacity is derived from free VRAM, never configured: a static value corrupts output under torch.compile thread affinity (#315) and aborts the process on small cards (#567). * A circuit breaker replaces the reliability-score/quarantine machinery, which had no recovery path (no probation workload exists in a TTS product) and penalised consumer networks for existing. * Identity is a keypair the worker generates and never sends. A server-assigned id is a name, not an authenticator, so revocation of one would be theatre. Enrollment tokens are single-use and carry the control plane's certificate fingerprint for pin-on-first-use. Adds the domain core, scheduler, durable task store, gRPC transport, worker agent, management API, Settings panel, and docs. Protobuf reserves the tenant/trace/usage fields a hosted control plane would need, since adding them later means upgrading a whole fleet. Includes tests for the failure paths that matter: duplicate delivery, stale-session fencing, reconnect reconciliation, grace expiry, breaker attribution, and a real end-to-end TLS round trip. |
||
|
|
5cab8e0149 |
feat: rename the product to VoiceStudio (previously OmniVoice-Studio)
Renames what users see. The app, the installers, the window title, the
docs and all 21 locales now say VoiceStudio, with "(previously
OmniVoice-Studio)" noted near the title of each doc surface so people
recognise it.
Deliberately NOT renamed, because renaming any of them silently breaks
an existing install — there is no legacy-path fallback anywhere in this
codebase:
- bundle identifier com.debpalash.omnivoice-studio (MSI UpgradeCode,
macOS TCC grants, managed venv, WebView localStorage, the
single-instance lock)
- data directories OmniVoice / .omnivoice and omnivoice.db
- the ~150 OMNIVOICE_* environment variables
- the X-OmniVoice-* HTTP headers (a wire protocol)
- the published Docker image paths
- the OmniVoice ENGINE, which is a model name and not this product
tests/test_identity_paths_survive_the_rename.py pins every one of those
so a future well-meaning sweep cannot orphan a user's library.
Linux .deb users install a new package name and should apt remove
omnivoice-studio; that note is in the changelog.
|
||
|
|
a23e69d014 |
chore: point every repo reference at github.com/debpalash/VoiceStudio (#1394)
The repository was renamed. 724 references across 59 files now point at the new URL — README badges, docs, install guides, the updater's releases API call, CONTRIBUTING, the Colab link and the probe harness. GitHub redirects the old URLs, so nothing was broken in the meantime. Deliberately NOT renamed, because each breaks something on a user's machine: the Tauri bundle identifier (the path to every existing user's data), /usr/lib/omnivoice-studio and the compose container names, and the published Docker image paths. The image path needed a code change to STAY still: docker.yml derived it from github.repository, so the next build would have published to ghcr.io/debpalash/voicestudio while Docker Hub, a hardcoded literal, stayed put — everyone pulling the documented GHCR path would have kept receiving the last pre-rename image forever. It is now pinned, with a test that fails if it ever derives from the repo name again. Also makes the probe's repo-name assertion shape-based: it hardcoded the old name and failed on every PR after the rename while the code it tests worked perfectly. |
||
|
|
933743e336 |
fix: resolve open issue batch — #1172 #1173 #1174 #1185 #1186 #1188
- validate managed binaries before exec; 0-byte GGUF placeholders fail actionably instead of "Exec format error" (#1172) - KittenTTS: tokenizer-measured chunking to the ONNX 512-token cap; clear 400 for unspeakable input (#1173) - clean SIGTERM during weight load: shutdown-aware loader, benign cancelled-load classification, lifespan hardening, scoped log silencers (transformers load + alembic fileConfig) (#1174) - broken ASR deep-imports (lightning_fabric) mark the engine unavailable with a repair hint and fall through (#1185) - uv cache + managed Python follow the chosen install drive on Windows (cherry-picked cross-drive class fix + spaces/D: tests) (#1186) - adaptive silence-removal ladder for quiet clone references; localized actionable error for truly silent clips, all 21 locales (#1188) - CHANGELOG: consolidated Unreleased into the quiet one-liner style Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
efc99be337 |
feat(studio): generation takes — star, replay, and restore past takes; capped history retention (#1052)
Every generate already recorded a generation_history row; now that history is usable: a takes rail in the workspace history lists recent takes with star/ unstar, replay, and one-click restore as the active output. Alembic migration 0009 adds the starred column (the startup schema self-heal covers pre- migration DBs), a retention cap (setting, default 200) prunes the oldest UNstarred rows — starred takes are never pruned — and history WAVs are only deleted when no other row references them. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
d57ecea804 |
fix(test): DB migration-safety tests leak-proof against full-suite order (#909 follow-up) (#917)
The PR #909 data-safe-update tests passed in isolation but failed only in full-suite CI order. Two independent, order-dependent leaks were at play: 1. Module-identity leak (the #878/#894 class). The `isolated_db`/`fresh_app`/ `fresh_resolver` fixtures in tests/backend/** purge `core.*`/`services.*` from `sys.modules` and never restore them, so `sys.modules["core.db"]` afterward is a DIFFERENT object than the one the migration-safety tests imported at collection. `monkeypatch.setattr("core.db.DB_PATH", ...)` re-resolved the dotted string to the re-imported module, while `_run_alembic_upgrade`/`init_db` (bound at collection) kept reading the ORIGINAL module's globals — so the patch missed and the upgrade ran against the ambient session DB. Result: no backup at the asserted path, and the mid-flight-failure injection never hit the expected DB (DID NOT RAISE). The same divergence hit the lazy `from core import db_backup` inside `_run_alembic_upgrade`, so patching `MAX_BACKUP_DB_BYTES` was silently lost. 2. Logger-disable leak. Alembic's env.py called `fileConfig(...)` with the default `disable_existing_loggers=True`, which disabled the already-created `omnivoice.db.backup` logger the first time any earlier test ran a real `alembic upgrade` — so the oversized-DB "Skipping pre-migration DB backup" line was never emitted and the caplog assertion failed. This also silently mutes the live app's logging after a real startup migration. Fixes: - env.py: `fileConfig(..., disable_existing_loggers=False)` so a migration never mutes the app's (or another test's) loggers. - core/db.py: import `db_backup`/`APP_VERSION` at module level so `_run_alembic_upgrade` uses a stable reference immune to a `sys.modules` purge, matching what tests patch at collection. - test_db_migration_safety.py: patch DB_PATH on the imported `core.db` module object rather than the re-resolvable dotted string — the correct, self-contained seam. Verified: the four migration-safety tests + the oversized-backup test pass in full-suite order and in isolation; full `pytest tests/` is green (2206 passed, 0 failed). Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
16294fed44 |
feat(updates): data-safe updates — pre-migration DB backups, guarded venv heal, release notes + changelog reader (#909)
Backend: - core/db_backup.py: WAL-safe SQLite snapshot to omnivoice.db.backup-<version>-<n> before pending alembic migrations run; keep newest 3, prune older; skip >500MB with a log line. Restore is never automatic. - core/db.py: _run_alembic_upgrade now plans the run (up_to_date / pending / unknown_revision), snapshots first when migrations will execute, and raises MigrationError on a mid-flight failure — startup stops with the backup path named instead of continuing on a half-migrated DB. The #552/#547 unknown-revision class stays non-fatal (warn + additive reconcile). - core/changelog.py + GET /api/settings/changelog: parse the shipped CHANGELOG.md (single-line and wrapped bullet styles) into structured releases. - GET /api/settings/db-backup: newest pre-migration backup for the panel. Rust (bootstrap.rs): - #314 heal guard: an exit-signature match alone can no longer delete the venv — venv_rebuild_justified requires a structural problem or a failed direct interpreter probe; a venv that probes healthy is kept and the real error surfaced. Drift/repair remains in-place `uv sync` (non-destructive). - CHANGELOG.md now ships as a bundle resource and is copied/refreshed into the project dir so the changelog endpoint works in packaged installs. Frontend (Settings → Updates): - Available update shows its actual release notes (updater metadata body) through a safe markdown-lite renderer (text nodes only, refs stay plain). - "Your data is backed up before every update" line with the latest backup timestamp from the new endpoint. - "What's new" changelog reader (accordion, newest expanded) over the shipped CHANGELOG.md; GitHub releases list reuses the same renderer. - One-time, non-blocking "What's new" footer pill after an update (persisted last-seen version; fresh installs baseline silently). - All strings via t() with en keys (other locales fall back to English). Tests: db backup/rotation/failure-path units, migration-safety units, changelog parser (both bullet styles + real CHANGELOG.md), endpoint tests, route inventory regenerated, Rust decision-logic + probe tests, vitest suites for renderer/viewer/panel/pill logic. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
ea4d5e7839 |
fix(generate): self-heal schema + don't 500 a generated clip on a history-write fail (#710) (#714)
A synth that already produced and saved its audio could still return a 500: 'no such table: generation_history' — a DB that somehow missed schema init (init_db's executescript never took) made the history INSERT raise after the clip was done, losing the user's generation to a logging side-effect. - Add db.ensure_schema(): idempotent CREATE ... IF NOT EXISTS + additive column reconcile (no _migrate/alembic), safe to call from a write path. - Generation history write now self-heals: on a sqlite OperationalError it runs ensure_schema() and retries once; if it still fails it logs and returns the audio anyway. A history-logging failure can never fail the generation. Regression test: the write raises 'no such table: generation_history' before the heal and succeeds after (fail-before/pass-after), plus ensure_schema idempotency. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
ee638bda6c |
feat(tts): user pronunciation dictionary (expressive-tts slice 1) (#685)
* feat(tts): user pronunciation dictionary (expressive-tts slice 1) Per-term, per-language pronunciation overrides applied to text before synthesis, so names, brands, and acronyms come out right across generate, longform, and dub. Closes part of the #1 perceived-quality gap vs ElevenLabs (pronunciation dictionaries). First slice of docs/specs/01-expressive-tts.md. - Schema: additive `pronunciation_entries` table (alembic 0008, mirrored into _BASE_SCHEMA; tested upgrade — idempotent, downgrade, converge, back-compat). - Service: extend pronunciation.py to load enabled entries (cached) and apply longest-first, word-boundary-aware, per-language (global '*' + lang match, lang overrides global), reusing the existing ReDoS-safe matcher. - Inline one-off `[[term|replacement]]` overrides that don't persist and don't collide with [voice:]/[pause]/[Name]/SSML-lite (resolved pre-chunking). - API: /pronunciation CRUD + /test dry-run + import/export (loopback-guarded). - Apply point: generation.py after language resolves, before chunking — covers native + pluggable engines. - UI: PronunciationPanel in Settings → General; all strings via i18n. - Tests: migration lifecycle, CRUD, per-language, precedence, inline override, apply-at-synth. Route snapshot regenerated (+7). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(security): bound inline-override regex (ReDoS) + annotate parameterized UPDATE CodeQL flagged py/polynomial-redos on the [[...]] inline-override regex: [^\]] also matches [, so an unterminated run of [ allowed O(n) rescans from O(n) positions. Bound the inner class to {0,256} (linear; an inline override is a short respelling). Annotate the dynamic UPDATE (B608) — its column fragments are fixed literals and every value is a bound parameter; not an injection vector. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: mergetest <test@local> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
915256cb8c |
fix(db): self-heal additive schema columns so consent_audio_path 500s stop (#557)
A profile/persona/consent endpoint 500'd with "no such column: consent_audio_path" (#552/#547) — and the same class for kind/vd_states/is_demo. The 0003/0005 migrations exist and are wired, but init_db's CREATE TABLE IF NOT EXISTS never adds columns to a pre-existing table, the legacy _migrate only knows pre-0.3 columns, and _run_alembic_upgrade swallows every failure. So on a DB whose alembic_version is stamped at a revision no longer in versions/ (common after running a preview build) or where alembic isn't importable, the alembic-era columns silently never land. Fix the whole class: add _reconcile_additive_columns(conn), which builds the canonical schema from _BASE_SCHEMA in-memory and ALTER TABLE ADD COLUMN any column an existing table is missing (additive only — never drops/retypes; names/types from _BASE_SCHEMA so injection-safe). Call it in init_db (so the schema converges regardless of alembic) and again in the alembic-failure branch. Also correct the false "_BASE_SCHEMA guarantees the schema regardless" comment. Test (fail-before/pass-after): init_db on a legacy voice_profiles whose alembic_version is a removed revision now lands consent_audio_path/kind/etc. without raising; reconcile converges to the canonical column set; idempotent + additive-only. 21 passed (incl. existing 0003/0005 migration tests). Co-authored-by: mergetest <test@local> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
66f2ea7e50 |
feat(profiles): unified profile model — kind discriminator + stored design params (spec P3) (#376)
Migration 0005_unified_profiles (0004 taken by mcp bindings):
- voice_profiles.kind TEXT DEFAULT 'clone' ('clone' | 'design'), backfilled
- voice_profiles.vd_states TEXT NULL — JSON of design category picks
- mirrored in _BASE_SCHEMA; idempotent _has_column guards; downgrade drops
POST /profiles:
- ref_audio now optional; kind + vd_states form fields with validation
(clone requires audio; design requires vd_states JSON object + instruct)
- design profiles render a deterministic identity sample (seed 42) through
the shared archetype renderer — one TTS code path
POST /generate:
- profile resolution branches on profile.kind (authoritative) instead of
the brittle is_locked/instruct inference; legacy pre-0005 rows keep the
old inference as fallback; history.mode records profile.kind
Frontend:
- 'Save design as profile' in the Design tab (vd_states + buildDesignInstruct)
- selecting a design profile restores its sliders (vd_states) for re-editing
Also unforks the alembic chain (0004_mcp + my 0004 both revised 0003 →
multiple heads broke alembic upgrade head and the 0003 migration tests).
Tests: tests/test_profile_unification.py — validation, design-create with
mocked renderer, migration up/backfill/downgrade. 18/18 profile tests,
312/312 frontend, related backend suite green.
Note: docs/specs/voice-studio-unification.md (on feat/studio-ux-overhaul)
still says 0004 — renumber to 0005 when branches meet.
Co-authored-by: mergetest <test@local>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
99357e8c5b |
feat(mcp): MCP server v1 — mount on /mcp, per-agent voice binding, stdio shim (Wave 2.2) (#368)
* feat(mcp): MCP server v1 — mount on /mcp, per-agent voice binding, stdio shim (Wave 2.2) The FastMCP server (previously dead code, never mounted) is now mounted on the main FastAPI app at /mcp via Streamable HTTP, with its session manager composed into the app lifespan through an AsyncExitStack (best-effort: a missing mcp package or OMNIVOICE_MCP_DISABLE=1 never breaks startup). streamable_http_path set to '/' so the sub-mount lands at /mcp, not /mcp/mcp. Adds the 'mcp' dependency (1.27.x). Per-agent voice binding (Spec 2 headline): each MCP client sends an X-OmniVoice-Client-Id header; generate_speech resolves the voice as explicit arg > the client's binding > global default > app default. New mcp_client_bindings table (alembic 0004 + _BASE_SCHEMA, additive/idempotent), services/mcp_bindings.py (CRUD + resolve_voice + best-effort last_seen), and a loopback-gated REST router (/api/mcp/bindings) the Settings panel drives. New transcribe tool (base64 audio in, 200 MB cap). Stdio shim (backend/mcp_shim, httpx-only, ported from voicebox MIT) proxies stdio clients to the mounted endpoint and forwards OMNIVOICE_CLIENT_ID as the binding header. Settings → Sharing gains an MCP bindings panel. Docs: docs/mcp.md (both connection modes + binding REST) and docs/mcp.json updated to the shim form. Tests: bindings service + resolution precedence + migration up/down (pure, run locally); REST CRUD + mount-not-404 + disable-flag (main-importing, validated in CI). MCP build + mount + initialize handshake verified out-of-band (no torch). Spec: docs/competitive-analysis.md Spec 2 / parity program Wave 2.2. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(mcp): assert /mcp mount via app.routes, not a lifespan client The two main-importing mount tests ran the app lifespan, which now starts the FastMCP session manager and binds asyncio queues to the test loop — contaminating later lifespan-running tests ('bound to a different event loop'). The mount happens at import time, so inspecting app.routes for the /mcp Mount is the correct loop-free assertion. Same fix shape as the Wave 0.2 consent tests. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(mcp): stop reload-main poisoning across the MCP test files Root cause of the CI failure: the bindings REST fixture set OMNIVOICE_MCP_DISABLE=1 and reloaded main but never restored it, so a later 'from main import app' in test_mcp_mount saw /mcp un-mounted ({'/audio','/voice_audio'}). Reloading main mutates the shared module for every subsequent test. - REST fixture: drop the disable flag (the mount is harmless without a lifespan), yield the client, and restore main (+ core.config/db) to the default data dir in teardown so the global module is clean again. - test_main_mounts_mcp_route: reload main with the disable flag cleared so the assertion is independent of any earlier reload. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
7422f20a63 |
feat(profiles): consent-locked voice profiles — verified_own_voice + spoken consent flow (Wave 0.2) (#354)
* feat(profiles): consent-locked voice profiles — verified_own_voice + spoken consent flow (Wave 0.2)
A profile becomes 'verified own voice' when its owner records themselves
reading a consent statement (spoken attestation, not a checkbox). Agentic
features and gallery sharing will gate on the flag; plain local synthesis
never does.
- alembic 0003 (additive, PRAGMA-guarded, downgrade supported) +
_BASE_SCHEMA columns: verified_own_voice, consent_text,
consent_audio_path, consent_recorded_at
- POST/DELETE /profiles/{id}/consent — stores the recording as provenance
in VOICES_DIR ({id}_consent.*), replaces on re-record, cleans up on
revoke and on profile delete; 422 on empty statement / too-short audio
- VoiceProfile page: Verified badge + Voice ownership panel (record via
the existing useRecording denoise flow, revoke with confirm); en.json
keys only (other locales fall back per the advisory i18n parity policy)
Spec: docs/competitive-analysis.md Action 22 / parity program Wave 0.2.
Prerequisite for agentic v2/v3 and the persona gallery.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(profiles): harden consent paths against py/path-injection; drop lifespan in tests
- _voices_path(): resolve DB-stored filenames strictly inside VOICES_DIR
(bare-filename check + realpath containment); extension whitelist on the
uploaded consent filename (fallback .wav) so a crafted filename can never
steer the on-disk path. Applied to write, re-record cleanup, revoke, and
profile-delete cleanup. New test: malicious upload filename falls back.
- Test fixture no longer runs the app lifespan: startup/shutdown touched
module-level asyncio primitives bound to another module's event loop,
making the suite order-dependent in full-suite CI. init_db() is called
directly; endpoints under test need only the schema.
Fixes the CodeQL (3x py/path-injection high) and full-suite event-loop
failures on PR #354.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
8b00dc1f4f |
feat: onboarding demos, opt-in bug reporting, error-docs deeplinks + issue triage (#133)
* feat: onboarding demos, opt-in bug reporting, error-docs deeplinks + issue triage Working-tree snapshot bundling several in-flight workstreams (v0.3.0): - Onboarding/demo system: DemoPresetGrid, DictationDemo, DubbingDemo components + tests, render scripts (render_demos_omnivoice.py, build_demos.sh, build_dub_demo.sh), personalities preview URLs, alembic 0002 voice-profile demo fields. - Opt-in bug reporting: ReportBugButton (prefilled GitHub-issue URL path). - Error transparency UX: errorDocsMap deeplinks + BootstrapSplash/error wiring. - Dub workspace: DubSegmentRow/Table, WaveformTimeline, dubSlice tweaks. - Issue triage: .planning/issue-clusters/ (plan-01..05 root-cause masters, GH #128-#132). - CLAUDE.md: hard rule — everything ships on v0.3.0, no version bumps. KNOWN GAP (why this is a draft): the generated demo audio assets are NOT in this tree, and backend/assets/samples/demo_voice.wav is deleted. onboarding.py guards the missing file (skips seeding the demo profile with a warning), so no crash — but first-run Launchpad will be empty and /demo_audio/ preview URLs 404 until assets are regenerated via scripts/build_demos.sh. Do not merge before regenerating + committing the demo assets. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * feat(dub): timing strategies — kill audio compression, add Concise + Stretch Video Replaces the current audio time-compression default (atempo squeeze to fit slot) that produced chipmunk/alien output on high-density target languages like Bengali. Two new user-selectable modes; legacy behaviour kept behind an explicit "Strict slot" choice. New `DubRequest.timing_strategy` enum (default "concise"): - "concise" Translator trims text to fit at natural rate; if it still overflows, hard-trim at slot with a fade so we never overlap the next speaker. Surface overflow_s per segment so the user can shorten the text. - "stretch_video" Audio plays at natural 1.0× rate. Backend computes a per-segment new timeline; persists a video_stretch_plan on the job. Mux step (dub_export) builds an ffmpeg trim+setpts+concat filter graph that stretches each segment's video portion to match the natural-rate dub audio. Gaps/pre-roll/tail pass through at 1.0×. Sub burn under stretch_video is skipped in one pass (cues would drift). - "strict_slot" Legacy atempo squeeze. Retained for back-compat. Director rate-bias side-effect (seg_speed *= bias) now gated on strict_slot only, so "urgent"/"slow" direction tokens keep their instruct effect in the new modes without chipmunking. Per-segment fit_status emitted in the SSE done event: {status: "fits" | "overflows" | "video_stretched", overflow_s?, stretch_ratio?} DubSegmentRow's "Sync: 100%" badge (which was lying — sync_ratio was always ~1.0 because the TTS loop pre-trimmed to slot) is replaced with a truthful "Fits / Overflows +Ns / Video 1.18×" label. Frontend: - prefsSlice.timingStrategy (persisted, store v3→v4 with safe migrate). - DubTab footer Segmented control: "Concise · Stretch Video · Strict slot". - useDubWorkflow passes timing_strategy on /dub/generate; consumes fit_status. Tests: tests/test_dub_timing_strategy.py — 13 cases covering schema defaults/validation, _build_video_stretch_filter_graph (pre-roll, gap, tail, empty-plan early return, post-subtitle chain-in), and _video_stretch_plan_for guards. 30/30 existing dub tests still pass. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(waveform): surface missing source as "Source media missing" instead of code-4 black box When a project's underlying media file is gone (moved or deleted between save and reload) the <video> element fires MediaError code 4 and the companion audio fetch returns HTTP 404 — both were silently warned to the console while the user stared at an unresponsive black panel and an empty waveform. - WaveformTimeline now flips loadError when the video element rejects code 3 (decode) or 4 (src not supported), and tracks `sourceMissing` separately so the error UI can name the actual problem. - The audio decode fallback chain catches HTTP 404 specifically and treats it as source-missing instead of loading silent empty peaks — an empty waveform on a deleted source is more confusing than a clear "Re-upload the video to continue" message. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(tray): "Show OmniVoice" reloads when the webview is blank When the dev Vite server restarts (or the main window is created before the backend is ready), the webview load fails and the window is left with `<body></body>` plus a "Could not connect to the server" console error. Clicking "Show OmniVoice" from the tray menu just re-showed the broken window — there was no recovery path short of quit+relaunch. Now the show handler runs a tiny eval after `show()`/`set_focus()` that calls `location.reload()` only when `document.body.childElementCount === 0`. A healthy window doesn't blink (body is non-empty); a blank one self-recovers as soon as the user clicks Show. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(#133): bug-report diagnostics field mapping + drop unused imports Address PR #133 review: - ReportBugButton: /system/info exposes `platform` + `device`, not `os`/`torch_device`/`gpu` — those reads silently dropped OS/GPU from every bug report. Map to the real fields (CodeRabbit). Also remove the dead `home` local in stripHome (CodeQL unused-variable). - DictationDemo: drop unused `Loader` import (CodeQL unused-import). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
4a6b978df9 |
Phase 1 Wave 1: HF token persistence + redactor (closes #35) (#91)
* feat(01-01): encrypted settings store + alembic migration (AUTH-02, T-01-01) Adds the SQLite-backed encrypted settings store that Phase 1 token resolver will read from. Closes the at-rest plaintext risk for HF tokens (T-01-01). - backend/services/settings_store.py: get_hf_token / set_hf_token / clear_hf_token using Fernet symmetric AEAD. Stored value column never contains the literal "hf_" substring. - backend/services/_secret_key.py: per-install Fernet key derived via scrypt(machine-id + 16-byte random salt). machine-id resolution covers macOS (ioreg IOPlatformUUID), Linux (/etc/machine-id and dbus fallback), Windows (HKLM Cryptography MachineGuid via winreg). Final fallback to hostname+user with a warn log. - backend/migrations/versions/0001_phase1_settings_table.py: alembic migration adding `settings(key, value, updated_at)`. Idempotent — checks for an existing table so fresh installs (where _BASE_SCHEMA already created it) and v0.2.7 upgrades both succeed. - backend/core/db.py: _BASE_SCHEMA grows the settings table for fresh installs; init_db() now runs `alembic upgrade head` after the CREATE. - backend/migrations/env.py: honours an externally-set sqlalchemy.url so tests can point alembic at a fixture DB; falls back to core.config DB_PATH for production. - pyproject.toml: cryptography>=41 added explicitly (RESEARCH.md Assumption A1 was checked at execute-time and proved false; the dep was not present transitively, so the install would fail without this). Tests (10 cases, all green): - Round-trip encryption + plaintext-leakage check (T-01-01 invariant) - Salt persistence across clear/set cycles - InvalidToken decrypt path returns None (Open Question #5 resolution) - Concurrent reads consistent under sqlite WAL - Alembic upgrade on a hand-built v0.2.7 fixture DB preserves all existing tables + seeded rows (CLAUDE.md backward-compat constraint) - Alembic downgrade -1 drops only the settings table Refs #35. * feat(01-01): 3-source HF token resolver + log redactor + 5 read sites patched Closes the #35 bug class (bare os.environ.get('HF_TOKEN') reads) by routing every backend HF-token consumer through one resolver, and mitigates T-01-02 (info disclosure via logs) by stripping `hf_[A-Za-z0-9]{30,}` substrings from every log record at the root logger. backend/services/token_resolver.py: - resolve(skip) — 3-source cascade (App → Env → HF-CLI), each source validated via huggingface_hub.whoami(); first valid wins. - on_401(active) — invalidate cache and re-resolve skipping the source that just 401'd (AUTH-06). - state() — three SourceState rows for the Settings UI: set, masked preview (hf_…<last 3>), whoami_user, whoami_ok. - save_app_token / clear_app_token — wraps settings_store + calls huggingface_hub.login(add_to_git_credential=False) per Pitfall #2. - 300-second whoami cache so repeated Settings-page renders don't hit the HF API. backend/core/logging_filter.py: - HFTokenRedactor(logging.Filter) — regex `hf_[A-Za-z0-9]{30,}` so real tokens are masked but `hf_hub` / `hf_token` literals survive. - install_redaction_filter() — idempotent attach to root + every handler. backend/main.py: install the redactor at startup, BEFORE the file handler is added. Re-installed after the file handler attaches so the handler-attached filter list includes it too. Read-side call sites patched (per Pitfall #1 — every HF token read must flow through token_resolver.resolve()): - backend/api/routers/dub_core.py:540 (the original #35 site) - backend/api/routers/system.py:38 (_has_hf_token notification) - backend/services/model_manager.py:480 (diarization pipeline auth) - backend/services/sonitranslate.py:143 (Popen env for SoniTranslate child) - backend/services/sonitranslate.py:217 (gradio_client predict call) New endpoint: - GET /system/hf-token/state — returns the 3-source cascade state with masked tokens for the Wave 2 Settings UI panel. Grep gate confirmed clean: zero `os.environ.get("HF_TOKEN")` reads remain outside token_resolver.py. Tests (17 new cases, all green): - tests/backend/services/test_token_resolver.py: priority cascade, 401 skip mid-resolve, on_401 fallback, state() shape, save+login invariant (add_to_git_credential=False), HUGGING_FACE_HUB_TOKEN alias acceptance. - tests/backend/core/test_logging_filter.py: msg + args redaction, multi-token redaction, non-string args pass-through, short-token literals preserved, install_redaction_filter idempotence. Refs #35. * feat(01-01): Settings hf-token API endpoints + subprocess env injection (AUTH-03/04) Backend half of the Wave 2 Settings → API Keys UI plus the AUTH-04 subprocess env-injection invariant. backend/api/routers/settings.py: - POST /api/settings/hf-token — body {token: str} → save_app_token - DELETE /api/settings/hf-token — also_clear_hf_cli query → clear_app_token - GET /api/settings/hf-token/state — same shape as token_resolver.state() All three are gated by `Depends(require_loopback)` at the router level (threat T-01-03 mitigation; non-loopback Host → 403). backend/main.py: router mounted alongside existing API routers. Subprocess env injection (AUTH-04, threat T-01-04 disposition=accept): - backend/services/sonitranslate.py already updated in Task 2 to read via token_resolver.resolve() and inject HF_TOKEN + YOUR_HF_TOKEN into the SoniTranslate child env block. - backend/services/gpu_sandbox.py: NOT patched — the GPU sandbox runs in-process TTS generation that uses the parent's already-loaded HF state. Adding env injection there is a no-op (parent and child share state via multiprocessing.Pipe before any HF API call). - backend/services/model_manager.py:480 (Task 2): resolves in-process, no subprocess crosses here. - backend/api/routers/exports.py: subprocess.Popen calls only spawn `open` / `explorer` / `xdg-open` — file-manager launchers with no HF needs. Skipped per Task 3 conservative-patching rule. So the canonical AUTH-04 site for this milestone is sonitranslate.py. Future SubprocessBackend work in Phase 2 will inherit the same pattern. Tests (8 new cases, all green): - tests/backend/test_engine_spawn_token.py * POST /hf-token loopback → 200 + state.active == "app" * POST /hf-token non-loopback → 403 ("loopback origin required") * DELETE /hf-token clears settings_store + state.active == None * GET /hf-token/state returns 3 source rows in priority order * GET /hf-token/state non-loopback → 403 * env block contains HF_TOKEN + YOUR_HF_TOKEN when resolver returns one * env block does NOT contain an injected empty HF_TOKEN when resolver returns None * source-level check that backend/services/sonitranslate.py still reads via token_resolver.resolve() (regression guard against silent reverts of the AUTH-04 wiring) Full Wave 1 test suite: 35/35 green. Phase 0 smoke tests still green. Refs #35. * docs(01-01): SUMMARY + STATE update for Phase 1 Wave 1 completion Records execution outcome of the 3-task plan: 10 files created, 9 modified, 35 new test cases, 5 read sites patched, grep gate clean. Documents the two Rule-3/Rule-2 deviations applied (cryptography dep, env.py URL override), the subprocess-launcher inventory for Phase 2, and the known stray edit to the main repo's pyproject.toml that needs a one- line user action to revert. Updates STATE.md current-position table, progress bar, and open TODOs to point at Wave 2 (Plan 01-02) and Wave 3 (Plan 01-03) as the next steps. |
||
|
|
c77bf18ac4 |
feat: onboarding demo profile, voice personalities, i18n framework
- Onboarding: seed 'OmniVoice Demo' profile on first run (empty DB) with bundled reference audio so Launchpad isn't empty - Voice Personalities: 6 built-in presets (Narrator, Casual, News Anchor, Storyteller, Corporate, Energetic) with instruct text auto-fill in Voice Design mode - i18n: react-i18next with English locale, browser language detection, Launchpad & CloneDesignTab strings extracted to en.json - DB migration v4: personality TEXT column on voice_profiles - New API: GET /personalities returns preset list - CSS: demo callout banner + personality picker strip |
||
|
|
52d68d05dc | refactor: update backend architecture, expand frontend state management, and synchronize voice-pro research modules. | ||
|
|
2c3e12d8de | feat: implement responsive layout adjustments for small screens and add cached status visualization to dubbing workflow | ||
|
|
67328d04fe |
refactor: split backend into api/core/services/schemas, harden security + fd pressure, add searchable language picker, fix segment fragmentation
Backend:
- Split monolithic main.py into backend/{api/routers,core,schemas,services}
- core/db.py: allowlist-gated migrations, db_conn context manager (kills SQL injection on ALTER)
- core/tasks.py: lock-guarded listener add/remove/push, snapshot-before-iterate
- services/ffmpeg_utils.py: run_ffmpeg helper with concurrency semaphore, EAGAIN retry, guaranteed reap
- services/segmentation.py: Bengali/CJK/Arabic punctuation, ultra-short tier, stitch_adjacent_shorts,
bounded-loop merge; public clean_up_segments API
- services/model_manager.py: robust lock.locked() handling
- api/routers/dub_core.py: job_id traversal guard, thread-safe _active_procs, timeouts on ffmpeg/demucs,
POST /dub/cleanup-segments endpoint
- api/routers/dub_export.py: guarded SSE listener remove, ffmpeg timeouts via run_ffmpeg
- api/routers/exports.py: destination_path validation, safe source resolver, subprocess list-form
- api/routers/generation.py: contextlib.suppress on tempfile cleanup, db_conn usage, safe output-path helper
- api/routers/system.py: try/finally tmp cleanup, subprocess timeouts
- schemas/requests.py: TranslateSegment.id int->str to match hex segment IDs
- main.py: threading.Lock around crash log writes
Frontend:
- components/SearchableSelect.jsx: popover combobox with search, keyboard nav, popular+recent pins, 200-item cap
- App.jsx: wire SearchableSelect for dub language / ISO code / voice-gen language; Clean Up segments button;
fix blob URL leak (object-shaped prev in setter, unmount cleanup via ref)
- components/WaveformTimeline.jsx: explicit <video> detach instead of innerHTML='' to release decoder
- index.css: ss-* combobox styles matching Gruvbox theme
Tests:
- tests/test_segmentation.py (26 cases), test_dub_transcribe.py, test_dub_export_unique.py, conftest.py
Chore:
- .gitignore: exclude omnivoice.zip, /research/ reference clones
- Remove tracked stray root test scripts + crash_log.txt
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|