d91beef0fd314250d8d9b94de86dfea019a8bd96
41
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
08569397d3 |
fix(asr): secure configured endpoints and refresh guidance (#1751)
Refreshes README and linked docs with accurate installation, platform, privacy, API, and model-license guidance; documents local gigastt; and pins OpenAI-compatible ASR traffic to the configured secure origin. Closes #1736. |
||
|
|
48c9a3b1f8 |
feat(settings): compute-device override (auto / CUDA / ROCm / XPU / MPS / CPU) (#1557)
* feat(settings): compute-device override — auto | CUDA | ROCm | XPU | MPS | CPU Auto-detect stays the default; the override kills the 'auto-detect picked wrong' issue class. Applied at the single choke point (_probe()'s family selection) so routing, get_best_device(), and every badge inherit it. Resolution: OMNIVOICE_DEVICE env > Settings pick (prefs.json) > auto (#981 pattern). An override can steer, never invent hardware: a family the host lacks is noted and ignored; cpu is always honorable. Applies at next backend start (host caps are immutable per process — same restart contract as the rest of the Performance tab, RestartBadge shown). GET/PUT /api/settings/compute-device (admin-gated) reports resolved vs applied so the panel shows restart-required truthfully and disables itself under an env pin instead of pretending. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(changelog): entry for the compute-device override (#1557) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(device-override): harvest — the override reaches CT2 ASR, full i18n, honest edge states - _ctranslate2_cuda_ok() and the ASR sidecar now gate on the probe's family, so a cpu pin (or ROCm host) can never hand CTranslate2 a CUDA device — the override reaches every CT2 loader through one shared gate - override_ignored exposed by the API and shown by the panel (env pin naming a device this machine lacks: auto is in effect, restart won't change it) - all 8 panel strings + 5 device-family labels translated into all 21 locales; failed saves keep their error visible through the re-sync - test isolation: cleanup drops OMNIVOICE_DEVICE before re-probing so no overridden caps leak into later tests; panel tests wait for loaded state - xpu/intel search keywords; oxfmt formatting Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(device-override): round 2 — fail-safe probe fallbacks, complete i18n, combined pin state - a broken capability probe now means CPU everywhere (CT2 gate + ASR sidecar) — never a torch-derived guess that would bypass a cpu pin or re-open #1529 on ROCm; regression test added - env-pinned AND not-detected shows both facts in one subtitle - device_load_failed/perf_save_failed translated into all 21 locales; CJK/th/vi/ar strings no longer say literal 'Auto' - test_ctranslate2_never_gets_cuda_on_a_rocm_build pins the probe family (it was order-dependent on the lru_cache before) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(test): pin the probe family in the faster-whisper OOM-fallback test Same class as the rocm-build test: it mocked torch but not the probe the new override gate consults first, so on a cpu-family CI host the CUDA fallback chain under test was unreachable. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
cd54113173 |
fix(security): close server-mode admin bypasses (#1525)
* fix(security): require keys for remote admin actions * fix(frontend): guard unavailable scrollIntoView * docs: link changelog to PR 1525 * fix(security): align PIN-only discovery policy * fix(security): preserve strict sidecar boundary * fix(security): normalize remote API keys * fix(auth): normalize credential fallback order |
||
|
|
3e292ba831 | Merge remote-tracking branch 'origin/main' into fix/ghas-log-safety | ||
|
|
501450e8ec | Merge remote-tracking branch 'origin/main' into codex/pr1442 | ||
|
|
44bfc94c49 |
Merge remote-tracking branch 'origin/main' into fix/ghas-log-safety
# Conflicts: # CHANGELOG.md # backend/api/routers/dub_core.py # backend/api/routers/settings.py # backend/services/speech_rate.py |
||
|
|
bebe462dc6 |
Merge remote-tracking branch 'origin/main' into fix/ghas-disclosure-streaming
# Conflicts: # CHANGELOG.md |
||
|
|
ae46f187d0 |
fix(security): move host-path authority into the desktop shell (#1448)
Close unauthenticated remote mutation paths and keep filesystem destinations behind one-shot Tauri capabilities. CodeRabbit and Greptile findings were fixed on-branch; all review threads are resolved. Full CI, Security, Rust, and cross-platform smoke checks are green. |
||
|
|
96a6c7574a | fix(security): stabilize streamed failure responses | ||
|
|
854a637b83 | fix(security): make untrusted log values single-line | ||
|
|
4c3ac1ad59 |
feat(pockettts): licence-accept gate, gated-weights preflight, license dialog
- Licence-accept gate: add pockettts to _LICENSE_ALLOWED_ENGINES + PocketTTSLicenseDialog (MIT code + CC-BY-4.0 weights + gated-access notice). - Gated-weights preflight: POCKETTTS_GATED_WEIGHTS in core/failure.py (hint + classify rule), so a gated-repo download surfaces as a typed error naming the agreement, not a raw failure. - Frontend: PocketTTSLicenseDialog registered in EngineCompatibilityMatrix; i18n keys in en.json. Remaining deferred items: CI smoke (stub sidecar integration test), four-platform install verification. |
||
|
|
5cab8e0149 |
feat: rename the product to VoiceStudio (previously OmniVoice-Studio)
Renames what users see. The app, the installers, the window title, the
docs and all 21 locales now say VoiceStudio, with "(previously
OmniVoice-Studio)" noted near the title of each doc surface so people
recognise it.
Deliberately NOT renamed, because renaming any of them silently breaks
an existing install — there is no legacy-path fallback anywhere in this
codebase:
- bundle identifier com.debpalash.omnivoice-studio (MSI UpgradeCode,
macOS TCC grants, managed venv, WebView localStorage, the
single-instance lock)
- data directories OmniVoice / .omnivoice and omnivoice.db
- the ~150 OMNIVOICE_* environment variables
- the X-OmniVoice-* HTTP headers (a wire protocol)
- the published Docker image paths
- the OmniVoice ENGINE, which is a model name and not this product
tests/test_identity_paths_survive_the_rename.py pins every one of those
so a future well-meaning sweep cannot orphan a user's library.
Linux .deb users install a new package name and should apt remove
omnivoice-studio; that note is in the changelog.
|
||
|
|
1a4b95890a |
fix(llm): LM Studio translation failed on a placeholder model name (#1332) (#1359)
* fix(llm): LM Studio translation failed on a placeholder model name (#1332) Reported as a clean A/B: translation works through Ollama and fails through LM Studio on the same machine. The difference is one line in the provider table. LM Studio shipped `local-model` as its default_model, which is a placeholder, not a model id — LM Studio serves whatever the user has loaded and 404s a name it does not know. Ollamas default is `llama3.1`, a real name people actually pull, so the identical code path worked there. No name we ship can be right, because the answer depends on what the user loaded. So ask the server: resolve_model now discovers from /v1/models for providers whose default is a placeholder, positioned BELOW any explicit env or stored setting so it can never override a deliberate choice, and above the default so a server that is down leaves the caller where it was. Cached per provider — translation resolves the model per segment and a round-trip each time would trade a broken setup for a slow one — and dropped whenever a base_url or model edit could invalidate it, since a stale id would make the users change look like it did nothing. Also: a 404 from a LOCAL provider is almost never a wrong URL, because the request reached the server. The generic "check the model name and Base URL path" sends the user to audit a URL that works, so a local 404 now names the models that ARE loaded, or says the server has none. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(llm): bound the discovery cache both ways; do not over-claim a 404 Three review findings, all valid, all about the cache being permanent in one direction or absent in the other. - A failed probe was not remembered, so a stopped LM Studio cost a 5s timeout on EVERY translated segment — a 200-segment dub would spend 1000s discovering nothing, worse than the bug being fixed. Remembered for 30s: short enough that starting the server recovers in seconds rather than needing a restart. - A successful discovery was cached forever, so swapping the loaded model inside LM Studio 404d every translation until an app restart. Now a 300s TTL, plus an immediate invalidation when a local 404 proves the cached name is one the server rejects. - _local_models collapsed a FAILED listing into [], which let the error say "reports no loaded models" about a lookup that never happened — a confident wrong diagnosis replacing a vague right one. None vs [] are now distinct, and the generic 404 text stands when nothing was established. The cache therefore cannot be a dict[str, str]: "no entry" and "we looked and there was nothing" have to be distinguishable for the negative case to be cacheable at all. Five tests, three failing before this change. They age the cache entry rather than patching time.monotonic — that name is the stdlib`s, shared with sqlite and logging, and freezing it breaks the settings store underneath the test. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
23f1767e3c |
feat(analytics): in-repo publishable token — source builds get the same consent-gated analytics (#1193)
Owner-sanctioned reversal: the publishable write-only PostHog client key is committed as the in-repo default in backend/core/analytics.py and frontend/src/utils/analytics.ts (env / baked release token still wins), so source builds show the same first-run consent ask as installers — skip = off, nothing is ever sent without an explicit yes. Adds an install_channel property (installer / docker / source) to lifecycle events, stamped by the desktop shell via OMNIVOICE_INSTALL_CHANNEL and by the Docker image's existing OMNIVOICE_SERVER_MODE marker. Guard tests now pin the two-canonical-files allowlist + same-token invariant, and the uninstall-ping info file works on the default token. Fixes #1193 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
7036e101e0 |
feat(analytics): first-run consent prompt — the opt-in becomes visible, never default-on
The opt-in PostHog analytics existed but was buried in Settings → Privacy, so release builds produced no OSS-useful data. Consent is now ASKED, once: - SetupWizard gains a consent step (between models and dictation) with two equal-weight Yes/No buttons — shown only in builds that ship a destination token and only if the user was never asked. Skipping the wizard or jumping past via the rail = not prompted = analytics stays OFF. - Existing installs (the wizard never reruns) get a one-time dismissible banner on app start with the same choice; dismiss counts as No. Any explicit choice marks analytics_prompted, so the ask never repeats. - Backend: set_opted_in() also persists the new analytics_prompted pref; GET /api/settings/analytics reports it. A broken prefs file reads as 'not prompted' (may re-ask) but never as consent (fails closed). - All strings via i18n across the 21 locales. The privacy promise is unchanged: nothing is sent without an explicit yes, silence is not consent, and source builds (no token) never even ask. |
||
|
|
ea7a7b39b4 |
feat(privacy): opt-in analytics — hardened, off by default, enforced by code (#1120)
The rejected PR #1110 had a genuinely careful PII-free event design, but shipped three things a local-first app can't: exception autocapture ON (raw tracebacks — home paths, and in this codebase HF tokens out of exception messages — bypassing core.failure.sanitize() entirely), no user consent or disclosure, and 3,069 lines of PostHog wizard scaffolding. This is the same capability with those fixed. core/analytics.py, three rules, each enforced and tested rather than promised: 1. OFF unless the user says yes. TWO gates must both be true: a build-provided POSTHOG_PROJECT_TOKEN *and* the user's analytics_enabled pref, default False. A default install transmits nothing, so "nothing leaves your machine" stays literally true for everyone who doesn't opt in. A broken prefs file fails CLOSED. OMNIVOICE_ANALYTICS_DISABLED=1 is a hard kill switch above both. Withdrawing consent tears the client down immediately — no restart. 2. NO exception autocapture. Explicitly disabled; a test asserts the constructor arg, because the SDK's default is the leak. 3. Metadata ONLY, by allowlist. Every property passes sanitize_properties(), which DROPS any key not on _ALLOWED_PROPS and refuses long strings — so no future caller can leak a take's text, a path, or a voice name by adding a field. text_length is the LENGTH; the text itself has no way through. The person id is a random per-install UUID — not hardware, hostname, or username. UI: Settings → Privacy → "Help improve OmniVoice" states in the panel exactly what is sent, exactly what never is, and that it can be turned off — rather than burying it in a policy. No destination in the build (any source build) → the toggle isn't shown, because an inert switch would be a lie. Docs: README FAQ answers "does OmniVoice collect any data about me?" honestly. Also fixed a bug I'd introduced in my own wiring: the generation event referenced variables not in scope, and the call site's bare `except: pass` swallowed the NameError — so the event would have silently never fired. The call site now logs. 12 tests (default-off / opt-in without token still can't transmit / both gates / kill switch / consent withdrawal / prefs failure fails closed / allowlist drops text+paths+names / long strings refused / autocapture OFF / never raises / random install id). Backend 2936 passed; frontend 1211 passed. Refs #1110 Co-authored-by: mergetest <nizam4103@gmail.com> |
||
|
|
23367cccaf |
fix: first-run wizard version + mirror-unreachable rescue + lifecycle-aware backend reachability (#1094)
Three fixes from the same first-run session report:
- SetupWizard shows v{APP_VERSION} in its masthead (same identity mark as
the install splash footer), so setup screenshots identify the build.
- A dead configured HF mirror no longer strands the wizard: the
install_error SSE now carries docs_topic (core.failure.classify), and
WizardLibrary renders the MirrorRescue quick-pick (extracted from
SetupWizard, now including the official preset) next to the failed row,
retrying it the moment a new endpoint is applied. PUT /hf-mirror clears
the install cooldowns (no 429 on the immediate retry) and clearing to
official also drops the legacy hf_endpoint pref that silently kept the
dead mirror in effect. The hint's false "applied when the app starts"
claim is corrected: downloads resolve the endpoint per call, retry
first, restart only if it still fails.
- "Can't reach the local OmniVoice backend" stops firing during real
start/restart windows: a respawn takes 10-20+s (venv spawn + torch
import) but the transport cascade gave up at ~2.9s. apiFetch now asks
the shell (bootstrap_status via utils/backendLifecycle) whether a
start/restart is in progress and keeps retrying while it is (capped at
120s, matching the supervisor's respawn budget); the new
BackendRestartBanner finally implements the reconnecting banner the
#567 supervisor has emitted events for all along. Truly dead backends
(or non-Tauri deploys) still error promptly.
Regression tests for all three layers; docs synced
(downloading-models.md, troubleshooting.md §14b); CHANGELOG [Unreleased].
Co-authored-by: mergetest <test@local>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
3d8799d9d2 |
feat(asr): complete activation flow for the OpenAI-compatible remote ASR engine (#1087)
The openai-compat-asr backend (#877) shipped with settings routes but no discoverable activation path: the config panel hid in Settings → Models, its hint text claimed "there's no in-app engine picker for ASR yet" (stale — the matrix has one), and there was no way to check a server actually answers before pointing a dub/dictation run at it. Configure → test → activate now live on one screen, Settings → Engines: - The ASR family tab mounts the config panel (URL / model / optional API key) below the engine matrix; saving refetches the matrix via a new reloadToken prop so the engine row flips unavailable → available and its "Use" button appears without a manual refresh. - New "Test connection" button + loopback-gated POST /api/settings/asr-openai-compat/test: saves first (same stale-config contract as /llm-providers/{id}/test), then probes GET {base_url}/models — no audio leaves the machine. The structured verdict maps to localized, actionable messages: latency + whether the configured model is listed on success; classified auth_failed / http_error / timeout / unreachable / ok_no_models failures. detail is core.scrub-ed; the key is never logged or echoed. - Engine reads persisted config fresh per transcribe (regression test) — config changes need no backend restart. Never default-active: ASR auto-detect only picks local engines. - i18n for every new string (en.json); no hardcoded CJK; identical behavior on macOS/Windows/Linux (pure HTTP + React). - Docs-sync: docs/engines/openai-compatible-asr.md rewritten around the one-screen flow with LM Studio / llama.cpp / Groq / OpenAI examples and the privacy note; README engine table cell updated. Verified end-to-end against a fake OpenAI-compatible server: UI drive (configure → test → row flip → Use) plus a real transcription through the backend's /v1/audio/transcriptions immediately after a config change, no restart. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
9c81e3389d |
feat(network): automatic Hugging Face endpoint selection — probe, pick, remember (#1082)
* feat(network): automatic Hugging Face endpoint selection — probe, pick, remember Restricted-network first-runs (the #984 class: huggingface.co unreachable, user dead-ends before discovering the mirror setting) now self-heal by default, while explicit endpoint choices are never second-guessed. - New backend/services/endpoint_race.py: parallel HTTPS reachability + latency probes of huggingface.co and the hf-mirror.com community mirror (3s timeouts). Probes are the only signal — no geo-IP, no third-party calls. Reachable beats unreachable; with both reachable the official endpoint wins unless the mirror is decisively faster (anti-flap hysteresis). The pick is cached in prefs and re-raced only on first run, a network-classified download failure, staleness (>7 days), or an explicit "Test again". - Manual mode is sacred: HF_ENDPOINT env, an hf_endpoint pref, or any explicit Settings pick disables auto-switching entirely; OMNIVOICE_HF_ENDPOINT_MODE=manual is a hard opt-out. - Wiring: the wizard preflight races endpoints when nothing is configured (honest copy when the mirror wins; warn-not-block when nothing is reachable); Model Store installs and the model-cache auto-repair resolve their per-call endpoint= through the cached decision, and a network-classified failure re-races once per repo per process and retries on the new winner (same guard pattern as the cache-recovery ladder). - Settings → Models → Hugging Face mirror gains "Auto (recommended)": shows the current pick, measured latency, last-checked time, and a "Test again" button (POST /api/settings/hf-mirror/test). Existing explicit configs surface as the matching manual mode. Panel notes that hf_hub checksums every download regardless of endpoint. - Tests: policy/cache/failover matrices in tests/test_endpoint_race.py, preflight + settings + repair-failover integration with mocked probers, HFMirrorPanel mode tests, and a suite-wide conftest guard that pins the probers so no test can hit the real network. - Docs: downloading-models.md and install/troubleshooting.md describe the automatic default and both opt-outs. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(changelog): Unreleased entry for automatic HF endpoint selection Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(tests): endpoint-probe pin uses an isolated MonkeyPatch and clears the decision cache; dtype guard tolerates stubbed torch The autouse probe pin requested the shared monkeypatch fixture, hoisting its setup earlier for every test and reordering teardown against the fp16 guard — which then ran torch.get_default_dtype() on test_torch_compile_gate's SimpleNamespace stub. The pin now uses its own MonkeyPatch context and also clears the prefs-cached endpoint decision per test (one test's auto pick leaked into other tests' preflight labels on CI ordering). The dtype guard additionally skips non-module torch stubs outright. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(tests): endpoint env vars can no longer leak out of the mirror-settings suite set_hf_mirror writes os.environ[HF_ENDPOINT] during the test, and monkeypatch.delenv(raising=False) on an absent var records nothing to undo — so the write leaked process-wide and flipped later suites' preflight network checks into the explicit-endpoint branch (the CI-order failures). Guaranteed save/restore autouse fixture at the source, plus defensive env shedding in the preflight suite. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
e816a2c24e |
fix(settings): make provider/token panels honest — real Test now, gated probes, MCP bindings i18n + confirm (#1064)
HuggingFace token (ApiKeysPanel):
- "Test now" actually re-runs whoami: GET /api/settings/hf-token/state gains
?fresh=1 which drops the resolver's 300s validation cache (the invalidate
hook existed but was never wired to any endpoint), so a fixed network or
rotated token no longer shows a stale verdict for up to 5 minutes. Plain
panel mounts keep the cache.
- Initial load renders a "Checking token sources…" placeholder instead of
flashing a false amber "not set" for all three sources.
- Source rows are now a valid ARIA list (the old role="table" had rows with
no cells, hiding the status from screen readers).
- Enter in the token input respects the in-flight guard the Save button
already had (no duplicate POSTs).
LLM Providers:
- Test / Fetch models abort when the implicit save fails, instead of probing
the previously-stored config and pairing a green "Test ok" badge with a
save error.
- A failed initial load now offers a Retry button instead of dead-ending
until the panel remounts.
LLM Skills: the per-skill provider Select carries an accessible name
("Provider for <skill>") instead of announcing as an unlabeled combobox.
MCP voice bindings:
- All user-facing strings go through i18n (the panel was the only Settings
surface with hardcoded English throughout).
- First-run guidance moved out of per-row hints (which never rendered with
zero bindings and duplicated per row) into the section header + an
InfoHint that links to docs/mcp.md; an empty state invites the first add.
- Delete asks for confirmation via the shared askConfirm, disables the row's
button while in flight, and re-syncs the list even when the DELETE fails
(a 404 row no longer lingers on screen).
- The add row exposes the optional label the API already accepted (the row
title rendered b.label without any way to set it); default_engine stays
MCP-side-only and is documented as such.
- First component test file for the panel (load/empty/add/delete/error/a11y).
Tests: backend fail-before/pass-after for the fresh=1 cache bust; new
frontend coverage for every behavioral change above.
Co-authored-by: mergetest <test@local>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
b603b9f78d |
fix(settings): factory reset covers all prefs, guarded log clearing, temp reclaim, full Performance i18n (#1061)
Settings system-group cleanup — every fix keeps existing behavior contracts and adds a fail-before/pass-after regression test: - Factory reset now does what it promises: clears every locally-persisted preference via a single registry (utils/prefKeys.js) instead of only the zustand blob — nav-rail side, capture live-typing, stories speed, logs footer state, last settings category, dismissed tips, donate prompts, and the legacy omni_ui blob included. User data and connection state (omni_transcriptions, ov_backend_url, ov_api_key) are explicitly preserved, and prefKeys.test.js scans the source tree so any future localStorage key must be categorized or CI fails. The failure toast now carries the actual error message. - Disk-usage "Clear logs" is confirm-gated with the same wording as Settings → Logs — it truncates the crash log (the bug-report artifact), so it can no longer be a single stray click. - Temporary files got a reclaim action: a confirmed "Clear temp files" button backed by POST /api/settings/storage/temp/clear, which deletes only the omnivoice* entries in the OS temp dir (symlinks unlinked, never followed) and invalidates the cached report. - Performance panel goes through i18n end to end (title, row, note, hint, errors, aria-label) — it was the last fully hardcoded panel; the non-Windows subtitle now reads "Windows only — not needed on this platform" instead of "not applicable". - History retention: GET failures now surface an alert and hold Save until a load succeeds (404 from older backends stays silent), Enter saves, the dead !res.ok branch is gone, and the bespoke button is the shared Button. - Logs tab: "Open folder" reveals the log file, "Copy visible log" copies the tail, the viewer autoscrolls to the newest lines, and the scroll box is keyboard-focusable (role=log) with a labelled source switcher. - Storage paths: the app-data row is labelled "App data stored at" (it was borrowing the Privacy tab's "Uploads stored at"), and all three path rows gained Open folder. - i18n stragglers routed through t(): storage load/open/clear fallbacks, the backend-status badge, and the frontend log buffer label. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
efc99be337 |
feat(studio): generation takes — star, replay, and restore past takes; capped history retention (#1052)
Every generate already recorded a generation_history row; now that history is usable: a takes rail in the workspace history lists recent takes with star/ unstar, replay, and one-click restore as the active output. Alembic migration 0009 adds the starred column (the startup schema self-heal covers pre- migration DBs), a retention cap (setting, default 200) prunes the oldest UNstarred rows — starred takes are never pruned — and history WAVs are only deleted when no other row references them. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
5a7d9cc05c |
feat(asr): generic OpenAI-compatible transcription backend (#877) (#1003)
First slice of the community's two-track proposal for #877: a generic OpenAI-compatible ASR backend that works TODAY, without waiting on transformers to ship a direct Qwen3-ASR integration (tracked separately, still blocked upstream). Points OmniVoice's transcription at any server exposing POST /v1/audio/transcriptions — a self-hosted Qwen3-ASR/ FunASR/SenseVoice server, or OpenAI's own API. - New OpenAICompatASRBackend (backend/services/asr_backend.py): a pure network client, no local model, no install. Prefers response_format=verbose_json for real per-segment timestamps, degrades to plain text (matching MoonshineASRBackend's shape) when a minimal server rejects that format. Never leaks a raw SDK/httpx exception to the caller (#977 convention) — wraps network/auth failures in a clean, actionable RuntimeError naming the server. - Settings persist via the same encrypted-secret convention as services/llm_providers.py (settings_store.set_secret for the API key — Fernet-encrypted, never a .env row, never echoed back; get_text/ set_text for base_url/model). New GET/PUT /api/settings/ asr-openai-compat, loopback-gated like every other settings route. - Frontend: a small settings panel (Settings → Models) mirroring HFMirrorPanel's exact structure. No ASR engine picker exists yet for ANY ASR backend (only TTS has one) — activating this engine still needs OMNIVOICE_ASR_BACKEND=openai-compat-asr; documented plainly rather than pretending otherwise. - README's ASR Engines table (9 → 10 engines) and docs/features.yaml's drift-checker inventory updated; the '9 engines, all fully local' claim corrected since this one genuinely isn't. - docs/engines/openai-compatible-asr.md: setup steps + an explicit privacy note (unlike every other ASR engine, audio leaves the machine to whatever server is configured). Regression tests: tests/test_asr_openai_compat_877.py (12 tests) — is_available() gating, verbose_json + plain-text response adaptation, network-failure error hygiene, SDK retry disabling, and the settings endpoints' persist/mask/clear-vs-unchanged semantics. Fixed two real full-suite-only failures found during verification (not brushed aside): the API route inventory snapshot needed regenerating for the two new routes, and this file's own tests had a module- staleness bug — a collection-time settings_store import went stale relative to a test-time-fresh fixture when another test elsewhere in the ~2400-test suite reimports the module — fixed by making settings_store itself a fixture resolved at test-run time, same lesson already applied to tests/test_mm2_lifecycle.py earlier this session. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
c47633a409 |
fix(settings): a saved LLM provider survives restart — explicit save activates, stale TRANSLATE_* prefs stop hijacking (#963) (#965)
* fix(settings): a saved LLM provider survives restart — explicit save activates, stale TRANSLATE_* prefs stop hijacking (#963) "Ollama works until I restart OmniVoice" had three stacked causes: 1. Only "Save & use for translation" persisted the selection. Plain "Save" and "Test" sent make_active:false, and on restart active_provider_id() deliberately excludes local providers (Ollama/LM Studio) from auto-select — so a saved-and-tested Ollama was never resolved active again. The PUT handler now also claims the active slot on an explicit save when the user has never chosen a provider (new llm_providers.stored_active_provider_id(): the stored row only — no env pin, no legacy fallback, no auto-detect). An explicit prior choice is never stolen; an unconfigured provider can't claim the slot; make_active:true still flips. 2. Users of the retired (≤v0.3.7) Translation-LLM panel had env.TRANSLATE_* rows in prefs.json, re-imported into os.environ every launch — and a live TRANSLATE_BASE_URL resolves the active provider to "custom" ahead of auto-select on every restart. New startup migration (llm_providers.migrate_legacy_translate_prefs, run in main.py BEFORE the prefs→env import) moves those values into the custom provider's own settings-store rows (only where the store has no value yet) and deletes the prefs rows. Real process env vars are never touched; a failed store write keeps the prefs row and retries next boot. The legacy endpoint keeps working — via the store, without hijacking the active slot. 3. The panel read as "done" after a green Test even when another provider stayed active. It now shows a notice after save/Test when the edited provider is not the effective active one (suppressed while LLM_DEFAULT_PROVIDER pins the choice — the env banner already covers that). Tests (fail-before): 7 new backend tests fail on the old code (save-activates, never-steals, migration semantics, env untouched, end-to-end ollama-beats-legacy-env) and the new panel test fails without the notice; all pass after. Full LLM/settings suites, frontend vitest (890), typecheck:ci, oxlint, oxfmt and vite build are green. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(changelog): add LLM-provider persistence fix under [Unreleased] (#965) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
8f71c90f20 |
feat(settings): LLM Skills — per-feature enable/route control for every LLM call (#912)
New Settings → System → LLM Skills area: every LLM-powered capability
(Cinematic & Autofit translation, speech-rate slot fitting, glossary
auto-extract, direction parsing, dictation cleanup) becomes a "skill" the
user can toggle or route to a specific provider (local Ollama/LM Studio vs
a remote key) instead of everything riding the one global active provider.
Backend:
- services/llm_skills.py — skill registry + settings_store persistence
(llm_skill.<id>.enabled / .provider), resolution precedence
override > active > none, resolve_skill_client() (OpenAI-compat client
bound to the effective provider; None when disabled/unconfigured) and
skill_backend() (OffBackend when disabled — the exact no-LLM object every
caller already degrades on).
- All five consumption points wired through the registry; a disabled skill
degrades exactly like "no LLM configured" today (Fast translation
fallback, refinement pass-through, heuristic direction parse, no-llm slot
fit, 503 on glossary auto-extract). No new degradation modes; defaults
(enabled + no override) keep existing setups byte-identical.
- OpenAICompatBackend gains an optional bound provider (None = active, the
historical behavior).
- GET /api/settings/llm-skills + PUT /api/settings/llm-skills/{skill_id}
(404 unknown skill/provider); route snapshot updated.
Frontend:
- LLMSkillsPanel (Sparkles, next to LLM Providers): one row per skill —
i18n name/description, enable toggle, provider Select ("Use active
provider" + configured providers, local ones tagged), ready /
needs-setup badge linking to LLM Providers. All strings via t()
(settings.llmskills_*).
Tests: 30 backend (precedence, per-consumption-point disabled semantics,
endpoint round-trips, validation) + 4 panel render/PUT tests. Docs:
translation-engines.md gains an LLM Skills section.
Co-authored-by: mergetest <test@local>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
75864a597f |
fix(dictation): refinement never stalls a final (~51s→≤4s), REST polish parity, real ASR preload reuse (#911)
P0 — Refinement blocked every dictation final with no timeout. With refinement auto:true and a slow/dead LLM endpoint, maybe_refine ran unbounded and blocked the final send in all three capture_ws handlers (~51s measured; the pill hung "Transcribing…" until the widget's 15s fallback fired). Fix the class: a hard, env-tunable budget (OMNIVOICE_REFINE_TIMEOUT_S, default 4s) via a new maybe_refine_async — a slow/dead endpoint now falls back to the unrefined (but polished) text within the budget and can NEVER delay the final beyond it. The LLM HTTP call is bounded to the same budget so the orphaned worker unwinds instead of holding a connection for the client's full 45s. Refinement is now also fully best-effort in the legacy handler (it can't turn a good final into an error frame). P1 — REST /transcribe lacked polish parity. capture.py never applied polish_text, so REST returned raw "…test" while the WS returned "…test." Apply text_polish.polish_text to `text` and `refined_text` (segments stay raw), so the widget POST fallback and MCP/CLI callers match the live socket. P1 — The #888 "instant first dictation" preload was a no-op. The preload called warmup() only `if hasattr`, but SherpaDictationBackend had none, and the WS handlers built a FRESH backend per session so a warm singleton wasn't reused. Add SherpaDictationBackend.warmup() (builds the recognizer) and share one warm recognizer per model id across sessions (get_sherpa_dictation_backend, same invalidation + a shared lock as the capture singleton); each session keeps its own decode stream. First dictation no longer pays the 1.3–2.5s load. P1 — llm_ready is a lie (feeds the P0). It only means "an endpoint is configured", so a placeholder key reads as ready. The P0 timeout makes a dead endpoint harmless; add last_refine_status so RefinementPanel flags a configured-but-failing LLM and links to LLM Providers → Test. Regression tests (fail-before/pass-after): slow-LLM WS final arrives < budget; maybe_refine_async hard timeout + status; REST polish parity + refined_text polish; warmup builds the recognizer and a second session reuses it; the panel honesty note. Backend refinement/capture_ws/capture/sherpa suites, CJK + route inventory gates, full vitest (733), lint (0 errors) and format all green. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
16294fed44 |
feat(updates): data-safe updates — pre-migration DB backups, guarded venv heal, release notes + changelog reader (#909)
Backend: - core/db_backup.py: WAL-safe SQLite snapshot to omnivoice.db.backup-<version>-<n> before pending alembic migrations run; keep newest 3, prune older; skip >500MB with a log line. Restore is never automatic. - core/db.py: _run_alembic_upgrade now plans the run (up_to_date / pending / unknown_revision), snapshots first when migrations will execute, and raises MigrationError on a mid-flight failure — startup stops with the backup path named instead of continuing on a half-migrated DB. The #552/#547 unknown-revision class stays non-fatal (warn + additive reconcile). - core/changelog.py + GET /api/settings/changelog: parse the shipped CHANGELOG.md (single-line and wrapped bullet styles) into structured releases. - GET /api/settings/db-backup: newest pre-migration backup for the panel. Rust (bootstrap.rs): - #314 heal guard: an exit-signature match alone can no longer delete the venv — venv_rebuild_justified requires a structural problem or a failed direct interpreter probe; a venv that probes healthy is kept and the real error surfaced. Drift/repair remains in-place `uv sync` (non-destructive). - CHANGELOG.md now ships as a bundle resource and is copied/refreshed into the project dir so the changelog endpoint works in packaged installs. Frontend (Settings → Updates): - Available update shows its actual release notes (updater metadata body) through a safe markdown-lite renderer (text nodes only, refs stay plain). - "Your data is backed up before every update" line with the latest backup timestamp from the new endpoint. - "What's new" changelog reader (accordion, newest expanded) over the shipped CHANGELOG.md; GitHub releases list reuses the same renderer. - One-time, non-blocking "What's new" footer pill after an update (persisted last-seen version; fresh installs baseline silently). - All strings via t() with en keys (other locales fall back to English). Tests: db backup/rotation/failure-path units, migration-safety units, changelog parser (both bullet styles + real CHANGELOG.md), endpoint tests, route inventory regenerated, Rust decision-logic + probe tests, vitest suites for renderer/viewer/panel/pill logic. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
e2c4ea93b0 |
fix(models-settings): surface async install errors, disk-space guard, cancel wiring, honest restart (#908)
Live-audit fixes for the Models settings surface — the P1s were cases where the feature silently didn't work for the user. P1-A — Async install errors were invisible. The `install_error` SSE event carries excellent mirror-aware text (#890 core/failure.py), but the Model Store auto-purged the errored row ~800ms later (same as a success) and the first-run WizardLibrary DELETED the row without ever reading `ev.error`. The SSE→rowState reduction is now a pure, tested reducer (downloadReducer.js / reduceWizardDownloadEvent); only SUCCESS terminals auto-purge (isAutoPurgeTerminal), an error persists on the row with inline text + Retry + Dismiss (Model Store) / a Retry (wizard). P1-B — No disk-space check on install. `POST /models/install` now compares the FDL-05 plan's exact `to_download_bytes` (+ MIN_FREE_GB headroom) against `shutil.disk_usage(cache).free` BEFORE downloading and emits an actionable install_error naming the sizes (needs X, headroom Y, have Z) instead of failing mid-download. `/models` also surfaces `disk_free_gb` in the header. MIN_FREE_GB + disk_free_bytes are single-sourced in setup/models.py (wizard delegates). P2-A — Wired the orphaned cancel. `POST /models/install/cancel` (FDL-11) had zero frontend refs; the in-progress row now shows a Cancel button that calls it and transitions the row to install_cancelled. P2-B — Honest restart_required. The HF-mirror PUT returned restart_required:true unconditionally; it now returns true only when the persisted value actually changed, with accurate copy (Model Store downloads use the new mirror immediately — resolved per-call; only transformers model loads need a restart). P3 — i18n the un-localized panels (HFMirrorPanel, ApiKeysPanel source labels/help/status, MODEL_ROLE_LABEL) via new en.json keys; other locales fall back to en. Tests: new tests/test_install_disk_space.py (reject-when-over-budget incl. the worker wiring; allow-when-fits; degrade on unknown size/unprobeable volume), updated tests/test_hf_mirror_settings.py (change-only restart_required), and new frontend reducer + column-render tests for install_error persistence, Retry, Dismiss, and Cancel. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
b3c18db33f |
feat(settings): Storage panel — real disk usage, category breakdown, and low-space warnings (#906)
Settings → Storage now opens with a Disk usage panel backed by a new loopback-gated GET /api/settings/storage endpoint: - Per-volume totals (grouped by st_dev) + du-style sizes for everything the app owns: the HF model cache (with its ~10 largest models), the app data dir broken into voices/outputs/dub_jobs/batch/preview/ database/logs/other subtotals, engine venvs (backend/engines/*/.venv + the app venv), and omnivoice* entries in the OS temp dir. - Bounded scanning: per-category 10 s deadline → partial totals with an "unreadable" warning instead of a hung request; results cached in-process for 5 minutes, ?refresh=1 forces a rescan; the walk runs in a worker thread so the event loop never blocks. - Server-side warnings reuse the setup wizard's MIN_FREE_GB: free < min → critical, free < 2×min → low, volume holding the cache/data >90% full → volume_pressure, unreadable/timed-out paths → unreadable. The panel renders severity-colored banners, a data-volume gauge, proportion bars per category, Open-folder buttons (existing /export/reveal pattern), a Model Store jump for reclaiming model space, and the existing clear-logs action on the logs row. A critical warning is also surfaced outside Settings via the app-wide toast — once per session. All strings via i18n (en fallback). Tests: tests/test_storage_report.py (sizes, thresholds, cache/refresh, timeout partials, endpoint wiring) + StorageUsagePanel.test.jsx (categories, banners, once-per-session toast, refresh=1, error state); route added to tests/fixtures/api_routes.txt via the dump script. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
e7fc37d438 |
fix(settings): retire legacy LLM endpoint panel, surface env overrides, fix Cloudflare account + fast-fail probes (#907)
Live-audit fixes for Settings → LLM Providers / Translation.
Retire the legacy LLMEndpointPanel from the UI (backend endpoint kept).
TranslationTab no longer embeds the inline endpoint panel — it now points to
Settings → LLM Providers (openSettingsTab('llm-providers')), which fully covers
it via the `custom` provider (a lone TRANSLATE_BASE_URL still resolves to
`custom`). Kills the panel's lying "reachable" badge, its hardcoded-English
strings, and one of three duplicate TRANSLATE_* surfaces. The third duplicate —
TranslationTab's "Provider keys" collapsible — drops the TRANSLATE_* trio
(now owned by LLM Providers) and keeps only the DeepL/Microsoft translator
keys; its toast no longer claims "saved for session" (these are in
PERSISTENT_KEYS, restored at startup). GET/PUT /api/settings/llm-endpoint is
untouched (DubTab gates Cinematic off it; tests + route inventory cover it).
Surface env overrides. describe() now reports base_url_from_env / model_from_env
/ active_from_env (mirroring key_from_env). The panel disables env-pinned
base_url/model/account fields with an explainer, and — when
LLM_DEFAULT_PROVIDER pins the active provider — disables make-active and shows a
banner, instead of silently reverting the user's edit / no-oping the button.
Fix the Cloudflare account-id flow (broken two ways): describe() now returns the
stored account_id (the field no longer resets to empty) and shows the RAW
base_url template ({account_id} kept literal) instead of the substituted value;
save_overrides drops a base_url override equal to the built-in default, so the
UI posting the shown value back can't freeze the URL — later account-id changes
take effect again (also self-heals if a default URL changes in a release).
Fast-fail the Test / Fetch-models probes. Pass max_retries=0 to the probe
OpenAI clients so a 429/timeout returns in seconds instead of ~34s on the SDK's
default retry ladder. /models now returns truncated:true when capped at 200 and
the UI hint reads "first 200 shown".
Tests: registry env-flag + Cloudflare round-trip/no-freeze regressions; router
truncation + max_retries=0 assertions; panel disabled+explained + banner;
new TranslationTab test (pointer wired, legacy panel gone, TRANSLATE_* dropped).
Co-authored-by: mergetest <test@local>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
da9315815d |
feat(settings): LLM provider testing pass — latency + classified errors, model discovery, full i18n, router tests (#887)
Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
29269b9cf0 |
feat: LLM Providers page + Autofit translation quality (fit-to-segment-time) (#838) (#854)
* feat(llm): multi-provider LLM registry + encrypted key storage + settings API (v0.3.8, phase 1)
Foundation for the LLM Providers settings page and timing-aware (Autofit)
translation. Every provider in the shipped .env is OpenAI-compatible, so one
client drives all of them via a registry instead of a class-per-provider.
- llm_providers.py: registry of 16 providers (OpenAI, OpenRouter, Groq,
Cerebras, Google AI, Mistral, Cohere, NVIDIA, GitHub Models, Cloudflare,
HuggingFace, SambaNova, SiliconFlow, + local Ollama/LM Studio + Custom).
Field resolution precedence env → encrypted store → default; active-provider
selection (LLM_DEFAULT_PROVIDER → stored → first keyed remote; local requires
explicit pick so we never assume a local server is up). Legacy TRANSLATE_*
maps to the Custom provider (keyless-with-base_url preserved).
- settings_store.py: generic ENCRYPTED secrets (get/set/clear_secret,
list_secret_names) reusing the HF-token Fernet path; get_text/set_text now
refuse the secret namespace (no ciphertext leak).
- llm_backend.py: OpenAICompatBackend resolves the active provider's
base_url/key/model from the registry. Backward-compatible.
- settings API: GET /llm-providers, PUT /llm-providers/{id} (encrypted key +
overrides), POST /llm-providers/active, POST /llm-providers/{id}/test.
Loopback-gated; never returns key material.
- 11 registry tests; existing llm-endpoint/openai-available tests still green.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(settings): LLM Providers page — configure any provider's key/URL/model + Test + set active (v0.3.8, phase 2)
New Settings → System → LLM Providers pane (Brain icon, searchable). Lists all
16 registry providers; pick one to configure its encrypted API key, base URL,
model (and Cloudflare account id), Test the connection with one round-trip, and
'Save & use for translation' to make it the active provider for Cinematic/
Autofit. Keys are write-only from the UI (masked placeholder, never echoed);
env-set keys show as read-only. Local providers (Ollama/LM Studio) need no key.
- LLMProvidersPanel.jsx: provider selector + per-provider config + Test/activate,
following the LLMEndpointPanel pattern (apiJson/apiFetch/apiPost, SettingsSection
primitives).
- settingsCategories.jsx: new 'llm-providers' category under System + Brain icon.
- Settings.jsx: route the category to the panel.
- en.json: settings.llm_providers label.
- Frontend build passes.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(translate): Autofit quality style + one-click LLM setup from the dub menu (v0.3.8, phases 3-4)
Autofit = Cinematic + a strict 'never exceed the segment time' fit. The LLM
rewrites each translated line so its target-language reading time fits within
the slot, preserving the video timing without harsh audio time-stretch.
Backend:
- speech_rate.adjust_for_slot(strict=): strict caps the accepted upper ratio at
1.0 (fit within slot) vs Cinematic's 1.08; best-effort, degrades gracefully
with no LLM.
- dub_translate: quality='autofit' takes the LLM refine path and runs the fit
pass with strict=True; reports quality_used accurately.
- TranslateRequest.quality doc note.
Frontend:
- 'autofit' added to the quality control (Settings Translation + dub menu) and
the TranslateQuality type.
- Dub menu: picking Cinematic/Autofit with no LLM no longer dead-ends on a toast
— it offers a one-click 'Set up' that routes to Settings → LLM Providers,
with copy about fitting translations to segment time (#838).
- 4 strict-fit tests; frontend build green; i18n keys added.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* docs(translate): document Autofit quality + the LLM Providers page (v0.3.8, phase 5)
- CHANGELOG [0.3.8] Added: Autofit style + LLM Providers page.
- docs/dubbing/translation-engines.md: Fast/Autofit/Cinematic quality section
and an LLM Providers setup section (16 providers, encrypted keys, offline
Ollama/LM Studio, env overrides).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(api): add /api/settings/llm-providers routes to the route-inventory snapshot
Regenerated tests/fixtures/api_routes.txt for the 4 new LLM-providers endpoints
so test_route_inventory_matches_snapshot passes (keep-main-green).
* style(frontend): oxfmt the LLM Providers panel + dub quality control (format:check green)
---------
Co-authored-by: mergetest <test@local>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
4aa4d983aa |
feat(settings): Hugging Face mirror (HF_ENDPOINT) for restricted networks (Wave 4.3) (#391)
The model manager already lists/deletes cached models; this adds the remaining high-value slice — an in-app HF mirror setting so users behind restricted networks (e.g. the Great Firewall) can route downloads through hf-mirror.com or any HF_ENDPOINT. Persisted to the durable per-user env (survives Tauri/Finder launches); HF reads HF_ENDPOINT at import, so the override applies on restart (surfaced in the UI). - GET/PUT /api/settings/hf-mirror (loopback-gated): presets (official + hf-mirror.com), http(s) validation, empty clears to official. - Models-tab panel with quick-picks + free-text field + restart note. (Skipped 'hf cache verify' — version-fragile across huggingface_hub releases and low value vs the mirror, which the China/Russia network research flagged as the real gap.) 3 endpoint tests (default, set+trim+clear, non-http rejection). Spec §R4(c) / parity program Wave 4.3. Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
c8fdcb619a |
fix(settings): remove stray rebase conflict marker in settings.py (#367)
A '>>>>>>>' marker from the #365 rebase was committed at the tail of the LLM-endpoint block, making the module unparseable. Strip it; settings.py parses clean and the endpoint tests pass. Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
d0b46f249e |
feat(settings): remote LLM endpoint UI — Ollama/vLLM/LM Studio (Wave 2.4) (#365)
A focused Settings panel for the OpenAI-compatible LLM that powers cinematic translate, glossary auto-extract, and dictation refinement (Wave 2.1). Persistence reuses the existing TRANSLATE_BASE_URL / TRANSLATE_MODEL / TRANSLATE_API_KEY env vars (already in system.py PERSISTENT_KEYS, restored at startup), so llm_backend/translator resolution is unchanged — vLLM is a verified drop-in, Ollama ignores the key, vLLM/LM Studio require it. - GET/PUT /api/settings/llm-endpoint (loopback-gated): read shape returns base_url, model, masked key, and live availability; PUT treats a null field as unchanged and an empty string as clear (so the key isn't wiped by a base-url-only save). Key is masked to last-4 in the read path, never echoed. - Credentials-tab panel with one-click presets (Ollama/LM Studio/vLLM/ OpenAI), base URL + model + optional key fields, and a reachable/not status badge. 6 endpoint tests (read shape, set+mask, null-unchanged, empty-clears, local-url-no-key, short-key masking); availability assertions guarded on openai being installed. Spec: parity program Wave 2.4 / competitive-analysis §R2 rung 4. Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
10806fea4f |
feat(dictation): optional local-LLM refinement of finals (Wave 2.1) (#363)
Phase 2 of Spec 3, on top of Wave 1.1's deterministic collapse. Prompt design ported from voicebox (MIT): 'text filter, not an assistant' base instruction + three toggleable sections (smart_cleanup, self_correction, preserve_technical) + 7 few-shot examples passed as STRUCTURED chat turns (small local models echo inline examples). Runs through the user's own Ollama/LM Studio/OpenAI-compat endpoint via llm_backend — new additive chat_messages() on the adapter; chat() now delegates to it. Pass-through is the contract: with no LLM configured (backend 'off'), on any error/timeout, or on an empty reply, the raw transcript stands — identical default behavior on every platform. Refinement runs off-thread on FINALS only; the WS final dict gains optional refined_text and the dictation pill pastes refined_text ?? text (raw kept in history). Settings: GET/PUT /api/settings/dictation-refinement (loopback-gated, persisted in the settings table) + a Capture-tab panel with the master switch + per-flag toggles and a 'no LLM configured' hint. 15 new unit tests: prompt sections per flag, structured few-shot message shape, and the full maybe_refine pass-through matrix (off backend, disabled config, LLM failure, empty reply, empty input). Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
47729057bd |
chore(lint): remove unused imports + variables (ruff F401/F841) (#210)
Autofixes the genuine lint behind the CodeQL py/unused-import and py/unused-local-variable note-level alerts — actually removing the dead code rather than dismissing it. 68 safe fixes via 'ruff check --select F401,F841 --fix' across 29 backend files (dead stdlib/symbol imports like io/sys/json/torch/typing.Optional and unused locals). Only ruff's safe fixes applied — the 9 'unsafe' fixes and the audio_dsp numpy availability import were left untouched. Not touched: empty-except (needs per-site judgement, not autofixable); frontend js/unused-local-variable (eslint no-unused-vars has no autofix); the loopback-low-risk path/log/stack-trace alerts (real, left visible). Verified: full tests/ suite unchanged at 601 passed (the 2 test_supertonic3 failures are pre-existing on main, local .venv state, green in CI). Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
1cfda2f44e |
feat(settings): configurable models directory (#64) (#149)
* feat(settings): configurable models directory (#64) Let users pick where model weights download (the HuggingFace / Torch cache) instead of being pinned to ~/.cache/huggingface — useful when the system drive is small or slow. Backend: - core/user_env.py: durable per-user env file (~/.config/omnivoice/env) helper with upsert/unset that preserves other keys and writes 0600. main.py already loads this at startup before importing torch/HF, so the value takes effect on the next launch. Path resolves at call time via an OMNIVOICE_ENV_FILE override so it's robust to module re-import in tests. - settings.py: GET/PUT /api/settings/storage/models-dir — validates the dir is writable (mkdir + write-probe → 400 if not), persists the choice, and writes OMNIVOICE_CACHE_DIR to the durable env. Empty path clears → reverts to default. Returns restart_required since an in-use cache can't be safely moved mid-process. Loopback-gated like the other settings. Frontend: - StoragePanel: Models tab panel to view/set/reset the directory, shows effective vs configured vs default + a restart note. Cross-platform default parity preserved (default cache path is the HF default on every OS); local-first (no network); backward-compatible (absent setting → existing behavior). No version bump. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(#64): harden models-dir input + clear CodeQL hygiene flags - settings.py: reject control/NUL chars in the path with a 400 before any filesystem call (an embedded NUL otherwise raised ValueError → 500). Also serves as the explicit input-validation barrier for the user-chosen path (loopback-gated same-user local file picker — no cross-privilege boundary). - test_user_env.py: use `with open(...)` so the file is closed and the assert has no side effects. - user_env.py: comment the best-effort chmod except clause. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * refactor(#64): single source of truth for models dir + review fixes Address CodeRabbit + Greptile review on PR #149: - P1 (both bots): the settings_store copy of the models dir was only ever read by this GET endpoint, so it was a redundant cache that could diverge from the durable env file (the value main.py actually reads). Drop it — the per-user env file (OMNIVOICE_CACHE_DIR) is now the single source of truth: PUT writes it, GET reads it back. No divergence possible. - XDG-aware default (CodeRabbit): _default_models_dir now honors XDG_CACHE_HOME, matching huggingface_hub's real default on Linux. - Atomic 0600 write (Greptile, security): user_env writes via an os.open opener that creates the file 0600 from the start — no world-readable window before chmod for a file that can hold HF_TOKEN. - _read_lines only swallows FileNotFoundError; other OSErrors propagate so an upsert can't silently drop existing keys on a transient read failure. - Guard makedirs("") when the env path is a bare filename (no parent). - Best-effort write-probe cleanup in a finally; raise ... from e. - a11y: label the models-dir input via aria-labelledby/aria-describedby. - OS-neutral unwritable-dir test (mock makedirs) instead of Unix-only /dev/null path semantics. 12 tests green. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
93aa66ab0a |
Phase 3 Plan 03-01: Supertonic-3 engine on SubprocessBackend (#101)
* Phase 3 Plan 03-01: Supertonic-3 engine on SubprocessBackend
Adds Supertonic-3 as a 7th opt-in TTS engine on the Phase 2
SubprocessBackend primitive. Closes TTS-01..06 (REQUIREMENTS.md):
* TTS-01 — _REGISTRY["supertonic3"] resolves to Supertonic3Backend,
a SubprocessBackend subclass.
* TTS-02 — `supertonic==1.3.1` lives under [project.optional-dependencies];
default `uv sync --no-dev` does NOT install it. Exactly one
`onnxruntime` row in `uv pip list` after `--extra supertonic`.
* TTS-03 — Model revision pinned by 40-char commit SHA
(724fb5abbf5502583fb520898d45929e62f02c0b — the "Initial
Supertonic 3 release" SHA, same as the SDK's own pin).
Resolver script for intentional bumps:
scripts/resolve_supertonic3_sha.py.
* TTS-04 — Honest CPU-only reporting. `is_available()` message says
"ready (CPU-only via onnxruntime)" and never mentions
"cuda" or "mps". `gpu_compat = ("cpu",)`.
* TTS-05 — License gate via settings_store helpers
(get/set_license_accepted) + Loopback-only
/api/settings/license endpoint + SupertonicLicenseDialog
frontend modal showing MIT (code) and OpenRAIL-M (model).
Wired into EngineCompatibilityMatrix as an "Accept license"
button on rows whose `reason` mentions "license not
accepted".
* TTS-06 — 3 langs (en/ja/ru) × 3 sec smoke test in
tests/test_supertonic3.py::test_smoke_3langs_3sec
(OMNIVOICE_SMOKE-gated; asserts no onnxruntime-gpu row
post-synthesize).
Package legitimacy gate (Task 1 in plan): supertonic on PyPI verified
to be published by Supertone Inc. (ato@supertone.ai), repo
github.com/supertone-inc/supertonic, wheel is pure-Python with no
postinstall scripts. Same publisher ships supertonic-js on npm under
the same maintainer email.
Test results:
* tests/test_supertonic3.py — 10 passed, 3 skipped (network-gated).
* tests/smoke/ — 4 passed.
* tests/ (full, --ignore=tests/manual) — 412 passed, 0 failed.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* ci(tests): uv sync --all-extras so optional-engine tests can import their package
Phase 3 added `supertonic` as an optional dependency. The CI Tests job
runs `uv sync` (no extras), so `test_cpu_only_honest` and `test_license_gate`
in tests/test_supertonic3.py hit the "supertonic package not installed"
fallback instead of the real import path, and fail.
Bare `uv sync` is the right default for users (engines are opt-in), but
the test environment should exercise the full surface. `--all-extras`
keeps the smoke job lean (still bare `uv sync`) while letting Tests
verify the integrated behavior of every optional engine.
Future-proofs against the same failure mode in Phase 4 (GGUF) and any
later optional engines.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
715766cb04 |
Phase 1 Wave 2: per-OS install docs + Settings UI + error→docs deeplinks (#94)
* docs(install): per-OS install pages + drift validator + CI gate
Splits the 600-line README install section into self-contained per-OS docs
under docs/install/{macos,windows,linux,docker}.md plus a Top-10
troubleshooting index. Each OS doc is end-to-end: a user opens it and
reaches a working app following only commands inside that file.
Adds:
- docs/install/{macos,windows,linux,docker}.md (OS-specific install paths)
- docs/install/troubleshooting.md (top 10 install errors)
- docs/engines/cosyvoice.md (closes #55 docs half)
- docs/features/diarization.md (pyannote license flow)
- docs/setup/huggingface-token.md (3-source cascade guide)
- scripts/validate-install-docs.py (INST-06 docs-drift gate)
- tests/scripts/test_validate_install_docs.py (B-5: validator self-tests)
- .github/workflows/ci.yml step running the validator on every PR
Implements INST-02 (README routing), INST-03 (macOS Gatekeeper anchor),
INST-12 docs half (Windows torch-compile-oom anchor), DOCS-01..05.
The validator is a one-way diff: every `<!-- validate -->`-tagged line
in docs must appear in scripts/desktop-prod.sh after normalisation
(prompt-prefix strip, CRLF, trailing whitespace, blank-and-comment skip).
A `<!-- validate: skip -->` marker opts out for human-readability blocks.
Its own 10 unit tests catch regressions in the gate itself.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* feat(deeplinks): links.py + error_docs_map (Python + TS mirror)
Adds the single source of truth for the project repo URL and the 4-class
error → docs taxonomy that both the in-app ErrorBoundary deeplink button
(Wave 2 Task 3) and the Phase 5 bug reporter will consume.
New:
- backend/core/links.py — PROJECT_REPO_URL + BLOB_MAIN resolver
(Tauri config first, pyproject fallback)
- backend/core/error_docs_map.py — lookup(error_class) → docs URL
- frontend/src/utils/errorDocsMap.ts (TS mirror with classifyError helper)
- tests/backend/core/test_links.py + test_error_docs_map.py
- frontend/src/utils/errorDocsMap.test.ts
Resolves checker B-6 (links.py ownership) and Open Question #3 (which fork
the deeplinks resolve to — the Tauri updater endpoint wins, which points
at the desktop app fork debpalash/OmniVoice-Studio).
The TS BASE constant is documented as the second hardcoded URL drift site;
the keys-sync test (`test_keys_match_python_map` equivalent) guards the
4-class taxonomy contract between Python + TS halves.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* feat(ui): Settings → API Keys panel + ErrorBoundary docs deeplink
Wave 2 AUTH-03 UI half + ErrorBoundary deeplink wiring.
ErrorBoundary fallback now renders an "Open docs for this error" button
that classifies the thrown Error message (heuristic: pkg_resources → 401 /
HfHubHTTP → WebKit / white screen → quarantine / Gatekeeper) and opens the
matching docs anchor via Tauri shell.open (with a window.open fallback
in browser dev mode).
ApiKeysPanel consumes the Wave 1 resolver state endpoint:
- 3 source rows (App / Env var / HF CLI) with set/unset indicator,
masked token preview, whoami username + green check
- "Active" badge on whichever source is currently serving the cascade
- App-row only: Save (POST /api/settings/hf-token) +
Clear (DELETE with optional "also clear HF CLI" confirm dialog)
- "Test now" button refetches state (invalidates the resolver's
validation cache via the same endpoint hit)
Panel mounted in the existing Settings → Credentials tab; the legacy
HF_TOKEN row from CREDENTIAL_FIELDS is filtered out so the two paths
don't fight over the same key.
Threat T-02-02: the panel never displays the full token. The masked
value comes from the resolver state endpoint; the full token only
crosses the IPC boundary on Save (POST) and is cleared from local
state on success.
Closes AUTH-03 fully (Wave 1 backend + this Wave 2 UI).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* feat(perf): INST-12 Disable torch.compile (Windows) toggle (backend + UI)
Wave 2 Task 4 — full INST-12 delivery per checker B-2/B-7 v0.3.0 fat-release
decision. Both the docs half (windows.md anchor, shipped in earlier commit)
and the runtime toggle are now in Phase 1.
Backend:
- backend/services/settings_store.py: adds get_text/set_text helpers for
non-secret config (refuses to write to the encrypted hf_token key).
- backend/api/routers/settings.py: GET + PUT
/api/settings/perf/torch-compile-disabled, both under the existing
loopback guard (threat T-02-04).
- backend/services/engine_env.py: new `build_engine_env()` helper that
centralises HF_TOKEN/YOUR_HF_TOKEN injection from the 3-source resolver
AND injects TORCH_COMPILE_DISABLE=1 when the flag is set on win32.
Phase 2 SubprocessBackend launchers should adopt the same helper.
- backend/services/sonitranslate.py: migrated to engine_env.build_engine_env()
while preserving the source-level `env["HF_TOKEN"]` sentinel that
test_sonitranslate_module_uses_resolver checks.
Frontend:
- frontend/src/components/settings/PerformancePanel.{jsx,css,test.jsx}:
toggle UI with the explainer for #65; renders disabled with a "not
applicable" badge on macOS/Linux.
- frontend/src/pages/Settings.jsx: mounts the panel into the Credentials
tab alongside the API Keys panel.
Tests:
- tests/backend/test_perf_settings.py: 7 backend tests (default state,
PUT persistence, T-02-04 non-loopback rejection, settings_store round-
trip, env injection on win32, NO injection on macOS/Linux, NO injection
when disabled).
- frontend PerformancePanel.test.jsx: 5 tests (renders from GET state,
PUT on toggle, disabled on non-Windows platforms, pre-enabled state).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* docs(planning): Wave 2 SUMMARY + REQUIREMENTS status updates
- .planning/phases/01.../01-02-SUMMARY.md: full implementation report
per template (truths, commits, tests, deviations, drift-site
acknowledgments per W-3, launcher seam name for Phase 2,
taxonomy keys for Phase 5).
- .planning/REQUIREMENTS.md: flips Wave 2 closures to Done:
AUTH-03, INST-02, INST-03 (docs half), INST-06, INST-12,
DOCS-01..05.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
4a6b978df9 |
Phase 1 Wave 1: HF token persistence + redactor (closes #35) (#91)
* feat(01-01): encrypted settings store + alembic migration (AUTH-02, T-01-01) Adds the SQLite-backed encrypted settings store that Phase 1 token resolver will read from. Closes the at-rest plaintext risk for HF tokens (T-01-01). - backend/services/settings_store.py: get_hf_token / set_hf_token / clear_hf_token using Fernet symmetric AEAD. Stored value column never contains the literal "hf_" substring. - backend/services/_secret_key.py: per-install Fernet key derived via scrypt(machine-id + 16-byte random salt). machine-id resolution covers macOS (ioreg IOPlatformUUID), Linux (/etc/machine-id and dbus fallback), Windows (HKLM Cryptography MachineGuid via winreg). Final fallback to hostname+user with a warn log. - backend/migrations/versions/0001_phase1_settings_table.py: alembic migration adding `settings(key, value, updated_at)`. Idempotent — checks for an existing table so fresh installs (where _BASE_SCHEMA already created it) and v0.2.7 upgrades both succeed. - backend/core/db.py: _BASE_SCHEMA grows the settings table for fresh installs; init_db() now runs `alembic upgrade head` after the CREATE. - backend/migrations/env.py: honours an externally-set sqlalchemy.url so tests can point alembic at a fixture DB; falls back to core.config DB_PATH for production. - pyproject.toml: cryptography>=41 added explicitly (RESEARCH.md Assumption A1 was checked at execute-time and proved false; the dep was not present transitively, so the install would fail without this). Tests (10 cases, all green): - Round-trip encryption + plaintext-leakage check (T-01-01 invariant) - Salt persistence across clear/set cycles - InvalidToken decrypt path returns None (Open Question #5 resolution) - Concurrent reads consistent under sqlite WAL - Alembic upgrade on a hand-built v0.2.7 fixture DB preserves all existing tables + seeded rows (CLAUDE.md backward-compat constraint) - Alembic downgrade -1 drops only the settings table Refs #35. * feat(01-01): 3-source HF token resolver + log redactor + 5 read sites patched Closes the #35 bug class (bare os.environ.get('HF_TOKEN') reads) by routing every backend HF-token consumer through one resolver, and mitigates T-01-02 (info disclosure via logs) by stripping `hf_[A-Za-z0-9]{30,}` substrings from every log record at the root logger. backend/services/token_resolver.py: - resolve(skip) — 3-source cascade (App → Env → HF-CLI), each source validated via huggingface_hub.whoami(); first valid wins. - on_401(active) — invalidate cache and re-resolve skipping the source that just 401'd (AUTH-06). - state() — three SourceState rows for the Settings UI: set, masked preview (hf_…<last 3>), whoami_user, whoami_ok. - save_app_token / clear_app_token — wraps settings_store + calls huggingface_hub.login(add_to_git_credential=False) per Pitfall #2. - 300-second whoami cache so repeated Settings-page renders don't hit the HF API. backend/core/logging_filter.py: - HFTokenRedactor(logging.Filter) — regex `hf_[A-Za-z0-9]{30,}` so real tokens are masked but `hf_hub` / `hf_token` literals survive. - install_redaction_filter() — idempotent attach to root + every handler. backend/main.py: install the redactor at startup, BEFORE the file handler is added. Re-installed after the file handler attaches so the handler-attached filter list includes it too. Read-side call sites patched (per Pitfall #1 — every HF token read must flow through token_resolver.resolve()): - backend/api/routers/dub_core.py:540 (the original #35 site) - backend/api/routers/system.py:38 (_has_hf_token notification) - backend/services/model_manager.py:480 (diarization pipeline auth) - backend/services/sonitranslate.py:143 (Popen env for SoniTranslate child) - backend/services/sonitranslate.py:217 (gradio_client predict call) New endpoint: - GET /system/hf-token/state — returns the 3-source cascade state with masked tokens for the Wave 2 Settings UI panel. Grep gate confirmed clean: zero `os.environ.get("HF_TOKEN")` reads remain outside token_resolver.py. Tests (17 new cases, all green): - tests/backend/services/test_token_resolver.py: priority cascade, 401 skip mid-resolve, on_401 fallback, state() shape, save+login invariant (add_to_git_credential=False), HUGGING_FACE_HUB_TOKEN alias acceptance. - tests/backend/core/test_logging_filter.py: msg + args redaction, multi-token redaction, non-string args pass-through, short-token literals preserved, install_redaction_filter idempotence. Refs #35. * feat(01-01): Settings hf-token API endpoints + subprocess env injection (AUTH-03/04) Backend half of the Wave 2 Settings → API Keys UI plus the AUTH-04 subprocess env-injection invariant. backend/api/routers/settings.py: - POST /api/settings/hf-token — body {token: str} → save_app_token - DELETE /api/settings/hf-token — also_clear_hf_cli query → clear_app_token - GET /api/settings/hf-token/state — same shape as token_resolver.state() All three are gated by `Depends(require_loopback)` at the router level (threat T-01-03 mitigation; non-loopback Host → 403). backend/main.py: router mounted alongside existing API routers. Subprocess env injection (AUTH-04, threat T-01-04 disposition=accept): - backend/services/sonitranslate.py already updated in Task 2 to read via token_resolver.resolve() and inject HF_TOKEN + YOUR_HF_TOKEN into the SoniTranslate child env block. - backend/services/gpu_sandbox.py: NOT patched — the GPU sandbox runs in-process TTS generation that uses the parent's already-loaded HF state. Adding env injection there is a no-op (parent and child share state via multiprocessing.Pipe before any HF API call). - backend/services/model_manager.py:480 (Task 2): resolves in-process, no subprocess crosses here. - backend/api/routers/exports.py: subprocess.Popen calls only spawn `open` / `explorer` / `xdg-open` — file-manager launchers with no HF needs. Skipped per Task 3 conservative-patching rule. So the canonical AUTH-04 site for this milestone is sonitranslate.py. Future SubprocessBackend work in Phase 2 will inherit the same pattern. Tests (8 new cases, all green): - tests/backend/test_engine_spawn_token.py * POST /hf-token loopback → 200 + state.active == "app" * POST /hf-token non-loopback → 403 ("loopback origin required") * DELETE /hf-token clears settings_store + state.active == None * GET /hf-token/state returns 3 source rows in priority order * GET /hf-token/state non-loopback → 403 * env block contains HF_TOKEN + YOUR_HF_TOKEN when resolver returns one * env block does NOT contain an injected empty HF_TOKEN when resolver returns None * source-level check that backend/services/sonitranslate.py still reads via token_resolver.resolve() (regression guard against silent reverts of the AUTH-04 wiring) Full Wave 1 test suite: 35/35 green. Phase 0 smoke tests still green. Refs #35. * docs(01-01): SUMMARY + STATE update for Phase 1 Wave 1 completion Records execution outcome of the 3-task plan: 10 files created, 9 modified, 35 new test cases, 5 read sites patched, grep gate clean. Documents the two Rule-3/Rule-2 deviations applied (cryptography dep, env.py URL override), the subprocess-launcher inventory for Phase 2, and the known stray edit to the main repo's pyproject.toml that needs a one- line user action to revert. Updates STATE.md current-position table, progress bar, and open TODOs to point at Wave 2 (Plan 01-02) and Wave 3 (Plan 01-03) as the next steps. |