d91beef0fd314250d8d9b94de86dfea019a8bd96
46
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
80aafa3a53 |
fix: preserve startup and media failure evidence (#1722)
* fix: preserve startup and media failure evidence * fix(dictation): bind readiness to installed revision * test(dictation): align cache resolution contract * fix: keep retry failures actionable |
||
|
|
de51120d6a |
fix(models): handle load OOMs safely (#1696)
* fix(models): handle load OOMs safely (#1695) * fix(dub): sanitize streamed generation failures * style(ui): format readiness checklist |
||
|
|
41722afe3b |
refactor(launchpad): quieter, borderless design refresh (#1515)
* refactor(launchpad): quieter, borderless design refresh The launchpad carried decoration from an earlier direction: icon chips, corner-hung count badges, a permanently visible filled arrow, uppercase mono card titles, and a dotted stipple divider — plus a frame that had been invisible since the app-wide border tokens were zeroed. Rework it around what the borderless direction actually implies: - Feature tiles get a whisper-faint surface instead of a dead frame, and read as three bands (bare glyph + count / title + arrow / description). `--card-hue` is spent sparingly — the glyph at rest, the surface, count and arrow only once raised. Titles move to sans sentence case; counts are plain tabular numerals. Lift softened 4px -> 2px, coloured glow -> neutral shadow, plus an explicit focus ring and a staggered entrance. - Hero drops the boxed "646" pill and the filled A/B-Compare button for quiet type, with a hairline standing in for the separation. - Section labels trade the dotted stipple for a single fading hairline; rows are transparent until hover and reveal "Open" on hover/focus (it stays in the DOM, so AT and keyboard always reach it). - Hero, tiles, recent files, callout and project lists now share one 1180px column — previously only the top half was capped, so lists ran edge-to-edge on a wide display while the deck stayed centred. Two bugs found and fixed while doing it: - Buttons that had `border border-solid border-transparent` removed fell back to the UA default border and rendered a visible 1px outline. They now carry `border-0` explicitly. - `.lp-animate` used `animation-fill-mode: both`, so after the entrance it kept owning `transform` — and animation-origin declarations outrank normal ones, which silently killed the card hover lift. Now `backwards`, which still holds the from-state through the stagger delay. Also drops CSS the page has not rendered since #904: the cursor-spotlight layer, the breath ring, and the per-card waveform strip. Verified with headless renders at 1600/1280/940 and the empty state. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(dictation): decode Wayland portal signals and show the capture pill The GlobalShortcuts portal declares Activated/Deactivated as (o session, s shortcut_id, t timestamp, a{sv} options). We decoded the timestamp as u32, so zbus rejected every signal with Signature mismatch: got `(osta{sv})`, expected `(osua{sv})` and the press was dropped as an invalid signal. Registration succeeded and the desktop even reported the bound chord back, so the hotkey looked wired up while doing nothing at all — on every Wayland compositor, for the whole life of the feature (#1490). Decode the 64-bit timestamp, and keep the 32-bit spelling as a fallback so a non-conforming portal degrades to working rather than to silence. With presses arriving, the second half of the failure showed: nothing had shown the widget window since it became a hidden recorder host, so a capture ran with no pill on screen — and a mic or Accessibility failure rendered into a window nobody could see. Add show_dictation_pill, which bottom-centres the capsule on the monitor under the pointer and shows it without taking focus (Windows keeps SW_SHOWNOACTIVATE so paste still lands in the user's document), and call it from the widget for every state but idle. Wayland denies clients their own placement, so the compositor picks the spot there; the pill still appears. dispatch_dictation_capture now logs whether a press was emitted or queued — a press that reaches Rust and produces nothing was otherwise indistinguishable from one the compositor never delivered. Tests: portal signals decode at both timestamp widths (the 64-bit case fails before this change with the exact production error); pill placement centres, respects a second monitor's origin, and clamps rather than going off-screen; the widget shows for a state needing the user, stays hidden while idle, and never shows for a press that arrives while dictation is disabled. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * chore: sync in-progress workspace changes Uncommitted work already in the tree, checkpointed so the branch matches the local machine: - Remote GPU workers: join-from-the-app flow, one-time secrets, QR join codes, a Compute control in the status bar, and the device-list Workers panel (#1516) - Model Catalogue workspace, with Settings pointing at it - Settings sidebar search and keyboard navigation - Demo assets for dubbing, dictation and voice design, plus the scripts that render them - Backend: validation-error handling, ASR request-path degradation, and the accompanying tests - CHANGELOG entries for the above and for the Wayland dictation fix Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(tests): follow Engines to the Model Catalogue, and green the sweep - test_supertonic3 asserted the license gate points at "Settings" while the engine now names Model Catalogue → Engines, which is where the accept button actually lives. The assertion follows the move; what it pins is unchanged — the hint must name a place the user can reach it. - Carries the CJK allowlist entries for the rendered dub bundle (#1517) and the regenerated route snapshot for /workers/agent (#1516), both of which this branch inherits from the workspace sync. - docs/install/linux.md: the dictation capsule is bottom-anchored everywhere except Wayland, where the protocol gives applications no say in their placement. Documented rather than left as a surprise (CodeRabbit). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * ci: stop a flaky dependency fetch from failing green runs en-core-web-sm resolves to a direct GitHub release URL, and github.com intermittently answers `http2 error: refused stream before processing any application logic`. uv's own three retries all land within the same few seconds and fail together, so the whole job dies on a dependency that has nothing to do with the change under test — it cost #1518 and #1517 an otherwise-green run tonight. Two changes: back off between whole `uv sync` attempts, which is what actually clears it, and pass --no-sync to the pytest steps. `uv run` re-resolves the environment before running, so every test step was a fresh chance to hit the same fetch even though the install step had already synced — that is exactly how #1518 failed, in the isolated backend/tests step, with all 5467 tests already passed. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * ci: one retry seam for every uv sync, not just the job that failed last en-core-web-sm resolves to a direct GitHub *release* URL rather than a package index, and github.com intermittently answers `http2 error: refused stream before processing any application logic`. uv's own retries all land inside the same ~10 seconds and fail together, so a job dies on a dependency unrelated to the change under test. Tonight that cost four otherwise-green runs across #1515, #1517 and #1518 — and the first fix only covered the Tests job, so the next failure simply moved to Smoke (Linux), which syncs separately. The fetch is per-job, so the fix has to be per-job: scripts/uv-sync-retry.sh backs off between whole attempts (15s, 45s, 90s) and every workflow that syncs now goes through it — ci.yml (tests + the platform matrix), release.yml, security.yml, evals.yml. It still fails loudly after four attempts, so a genuinely broken lockfile is not disguised as a flake. The Tests job also lacked the UV_HTTP_TIMEOUT / UV_HTTP_RETRIES the smoke matrix has always set, which is part of why it was the one that kept dying; it has them now. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test(ci): pin the Intel-Mac contract by intent, not by command spelling test_ci_verifies_intel_mac_as_the_documented_remote_only_host asserted the literal line `run: uv sync --extra pockettts`, so routing every sync through scripts/uv-sync-retry.sh read as a broken Intel-Mac contract. The contract it exists to protect is that the pockettts extra installs ONLY on backend_supported legs — which the regex now pins, while leaving how the sync is invoked free to change. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * ci: keep every uv run out of the resolver, and bound the retry budget CodeRabbit, #1517: - `uv run` re-resolves before running, so the smoke suite, the worker-artifact tests, the release test run and the eval run were each a fresh chance to hit the flaky direct-URL fetch outside the retry loop. All of them pass --no-sync now; the environment is already synced by the step that owns the retries. security.yml's `uv run --with pip-audit` is deliberately left alone — it layers an ephemeral package rather than running the project's own tests. - The retry count multiplied uv's own budget (UV_HTTP_RETRIES=5 with a 120 s timeout on the smoke matrix). Three attempts and 60 s of total backoff outlast the refusals actually observed while staying well inside the jobs' timeout-minutes. - The Intel-Mac contract test pinned the smoke command literally too, so --no-sync tripped it exactly like the sync line did. Same fix: assert the contract (smoke runs only on backend_supported legs), not its spelling. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
aa1d739843 |
feat(workers): dubbing goes remote, and the protocol stops lying to old workers
The remote-GPU line, verified on hardware rather than asserted. **Dubbing renders on the worker.** dub_generate.py dispatches the coarse `dub_segments` operation through the gateway, following the audiobook pattern: per-unit local fallback after consecutive remote failures, one aggregated notice rather than one per segment. A 40-minute dub that loses its worker at segment 200 degrades instead of producing 200 error rows. **An out-of-date worker is now refused by name.** This was the worst defect in the plan and it was silent: an un-upgraded worker registered cleanly, then ignored `inputs` and rendered a clone with NO reference audio — returned as success. A plausible wrong result with nothing anywhere to surface it. Workers now declare features, and one missing them is turned away with the features named and `no task was run`. Verified live: a worker one commit behind was correctly refused. **"Offline" and "cannot run this" are different facts.** Asking a live worker for an engine it lacks answered "is offline or cannot be reached. Wake the selected worker" — while that worker reported ready, one free slot and 3.6 ms latency. The user was sent to wake a machine that was already awake. The scheduler now distinguishes absent from present-but- incapable, and names the engine rather than the operation, because the engine is the thing a user can install. **An engine with no catalog entry is no longer hidden.** A `repo_ids` non-emptiness check had been implemented as a runtime filter, so a worker silently refused to advertise any engine lacking a models.yaml entry — which is four registered engines, including CosyVoice. Users with those already installed would have lost remote support with only a log line. Empty `repo_ids` now means "not downloadable here", never "not runnable". **And a script so this stops being done by hand.** scripts/verify-remote-worker.sh runs the per-phase acceptance checks against a live worker, non-destructively. Its preconditions are the mistakes that cost the most time: exactly one listener on the control port (two instances silently shared it), and never detecting the worker with a pgrep pattern that matches the ssh shell running it. Its first real run found the dubbing picker claiming remote placement. That turned out to be the CHECK being stale, not the picker — the port had landed since it was written. It now asserts self-consistency instead: the picker may claim remote only for an operation the control plane actually advertises as remotely producible, which cannot rot the next time an op is ported. Backend 5291 passed, frontend 1812 passed. Acceptance script: no automated failures across Phases 4-8 on an RTX 4090. Four checks remain MANUAL by design — true airplane mode, concurrent downloads, killing a worker mid-audiobook, and the model-list UI — and are reported as unverified rather than passed. |
||
|
|
603402afce | Merge remote-tracking branch 'origin/main' into codex/pr1442 | ||
|
|
27a8f477b7 |
fix(security): keep private diagnostics out of API responses (#1454)
* fix(security): keep private diagnostics out of API responses * docs: reference response-safety PR * fix(security): preserve constant recovery guidance * fix(security): keep recovery and logs data-independent * fix(security): close remaining response sinks * test(security): keep SOCKS diagnostics private * fix: keep Tailscale exceptions local * fix: keep Tailscale CLI output private |
||
|
|
501450e8ec | Merge remote-tracking branch 'origin/main' into codex/pr1442 | ||
|
|
a2f587bbc1 |
Merge remote-tracking branch 'origin/main' into fix/1406-corrupt-weights
# Conflicts: # CHANGELOG.md |
||
|
|
9a770b15d1 | test(pockettts): verify pinned install on four platforms | ||
|
|
6079689500 | fix(models): repair corrupt configs and isolate ASR preload (#1437) | ||
|
|
4c3ac1ad59 |
feat(pockettts): licence-accept gate, gated-weights preflight, license dialog
- Licence-accept gate: add pockettts to _LICENSE_ALLOWED_ENGINES + PocketTTSLicenseDialog (MIT code + CC-BY-4.0 weights + gated-access notice). - Gated-weights preflight: POCKETTTS_GATED_WEIGHTS in core/failure.py (hint + classify rule), so a gated-repo download surfaces as a typed error naming the agreement, not a raw failure. - Frontend: PocketTTSLicenseDialog registered in EngineCompatibilityMatrix; i18n keys in en.json. Remaining deferred items: CI smoke (stub sidecar integration test), four-platform install verification. |
||
|
|
b3b2bc3e44 |
fix(errors): reject identifier and dotted-numeric 401 matches
Digit-only boundaries rejected 4012 and 1401 but still accepted x401y, pytest-401 and 401.0. With 401 on the symptom side, an error that merely mentions Hugging Face and carries one of those elsewhere still resolved to HF_AUTH_FAILED — the residual CodeRabbit flagged as Critical. Three guards, one per family the others let through: identifier/path/ dotted-version prefixes, identifier and hyphen suffixes, and dotted numerics. A trailing sentence full stop still reads as punctuation. Also drops a duplicated paragraph in the branch comment. |
||
|
|
21d42aca45 |
fix(errors): three digits in a path are not an authentication failure
classify() matched a bare '401' substring and used it to satisfy BOTH halves of the HF-auth condition, so any message containing those digits anywhere classified as HF_AUTH_FAILED on its own — paths, byte counts, job ids, durations. CI hit it when pytest's numbered temp directory reached pytest-401: an audio-save failure came back telling the user to set a valid HF_TOKEN. That is worse than an unclassified error — a confident wrong instruction with a docs deeplink, in an auto-filed bug report — and because it rides a counter that changes between runs, it passes locally forever. Two independent guards: the digits must be a standalone number (not 4012, 1401, pytest-401's neighbours), and they are no longer sufficient evidence by themselves. A real 401 always arrives with 'Unauthorized' or an HF URL beside it, so requiring that costs nothing. |
||
|
|
2a32ab8de2 |
fix(models): require a weight marker before 'unexpected end of file' counts
zipfile, tarfile, gzip and several parsers share the wording, and any of them can surface inside a model-load chain — where a false positive forces a multi-GB re-download of an undamaged cache. |
||
|
|
05b1d096b4 |
fix(models): repair a weight file that arrived damaged, not just a missing one (#1406)
An interrupted or mangled model download leaves one of two states, and only
one had any handling. A MISSING shard raises transformers' "does not appear
to have a file named …" and gets a whole recovery ladder. A shard that is
PRESENT with wrong bytes — a download stopped mid-file, a shard truncated by
antivirus, an HTML error page saved under its name — opens fine and then
fails inside safetensors:
Error while deserializing header: header too large
That reached the user as a raw 500 on every generation, from voice design and
gallery previews alike, and could not enter the ladder for two independent
reasons: the wording is not the missing-shard wording, and SafetensorError is
a Rust-extension exception rather than an OSError.
It also needs the opposite repair. The ladder RESUMES a download, and a resume
trusts a blob that is already the expected size — so it would never re-fetch
the one file that is actually wrong. The new path forces a full re-download,
then retries the load once.
Both halves now classify as MODEL_CACHE_CORRUPT: one class to the user, one
remedy, two repairs underneath. The cause is matched through the whole
exception chain, since transformers wraps the tensor library's error in its
own before it reaches us.
|
||
|
|
5cab8e0149 |
feat: rename the product to VoiceStudio (previously OmniVoice-Studio)
Renames what users see. The app, the installers, the window title, the
docs and all 21 locales now say VoiceStudio, with "(previously
OmniVoice-Studio)" noted near the title of each doc surface so people
recognise it.
Deliberately NOT renamed, because renaming any of them silently breaks
an existing install — there is no legacy-path fallback anywhere in this
codebase:
- bundle identifier com.debpalash.omnivoice-studio (MSI UpgradeCode,
macOS TCC grants, managed venv, WebView localStorage, the
single-instance lock)
- data directories OmniVoice / .omnivoice and omnivoice.db
- the ~150 OMNIVOICE_* environment variables
- the X-OmniVoice-* HTTP headers (a wire protocol)
- the published Docker image paths
- the OmniVoice ENGINE, which is a model name and not this product
tests/test_identity_paths_survive_the_rename.py pins every one of those
so a future well-meaning sweep cannot orphan a user's library.
Linux .deb users install a new package name and should apt remove
omnivoice-studio; that note is in the changelog.
|
||
|
|
d23e56a5ec |
fix(errors): the transformers-import advice names torchvision now (#1376) (#1377)
The TRANSFORMERS_IMPORT hint and the ASR pipeline error told users to reinstall torch + torchaudio + transformers. torchvision — the package whose ABI mismatch actually produces this exact lazy-import wording (#1357's torchvision::nms, wrapped into "Could not import module 'AutoFeatureExtractor'") — was the one package the advice omitted. Following it to the letter left the broken package untouched (#1376). Both surfaces now name the mismatch as a cause and prescribe the pinned reinstall with literal versions (desktop installs ship no deploy/, so the constraint-file form fails there) targeting the venv explicitly. A lockstep test asserts the exact command on every advice surface against deploy/torch-constraints.txt, so a pin bump stays red until the advice matches. Docs gain the same-wording-different-cause section (1a-bis). |
||
|
|
5d6e05ef1b |
fix(errors): stop giving advice that cannot work (#1347, #1335, #1334) (#1374)
* fix(errors): a failed download is not a broken install (#1347, #1335) Two reports, one shape: the error text carried both a network cause and a downstream symptom, the taxonomy matched the symptom first, and the user was sent to fix something that was never broken. #1347 -- transcription failed with "transformers ASR pipeline failed to import (AutoFeatureExtractor) -- your transformers install is incomplete; reinstall with `uv pip install --reinstall transformers` ... Underlying: Cannot send a request, as the client has been closed." The install is fine. The pipeline was DOWNLOADING the feature extractor when the shared HTTP client closed underneath it (#880). Reinstalling transformers cannot fix a dropped connection, so the advice was not merely unhelpful -- it was work the user could repeat forever without succeeding. New MODEL_DOWNLOAD_INTERRUPTED class, checked before the import rules, requiring the httpx closed-client wording AND an import/transformers term so a bare closed-client error elsewhere is left alone. Its hint says the partial download resumes, since otherwise someone on a slow link assumes retrying restarts a multi-GB fetch. #1335 -- a cut TLS connection reached /generate as a bare 500 carrying `_ssl.c:1016`. core/failure.py has classified that since #1301, but /generate keeps its own taxonomy and never learned it, so it fell to the unrecognized-error catch-all. Added to the network signatures there: it is a dropped download, and the remedy is retry, not Flush. Both changes are orderings rather than new detections -- the cause now beats the symptom -- and both keep the case the original rule existed for: a genuinely broken transformers install still classifies as TRANSFORMERS_IMPORT, and a failed handshake is still distinguished from a cut connection. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(errors): a Windows paging-file limit is not out-of-memory (#1334) Same class as the two fixes already on this branch: advice that cannot work. The reporter asked, reasonably, whether OmniVoice needs an internet connection -- generation failed only when they disconnected, with a bare 500 carrying "The paging file is too small for this operation to complete (os error 1455)". Two separate defects made that unanswerable: 1. /generate matched it in _is_oom_failure and said "Try the Flush button to reload the model". Flush cannot help. The hint we had already written for this exact class says so outright -- "closing other apps usually won't fix it" -- but the generate path never consulted it. Now branched before the OOM check, naming the virtual-memory setting and stating plainly that it is not a network problem. 2. WINDOWS_PAGING_FILE_TOO_SMALL was absent from _CONTEXT_FREE_HINT_CLASSES, so on the raw-500 surface classify() identified it correctly and then attached nothing. The user got the OS sentence and no next step, despite the detailed remedy sitting in _HINTS. Its trigger (1455 with winerror/os error, or the literal phrase) is unmistakable, which is the bar that set requires. Both Python (`WinError 1455`) and Rust (`os error 1455`, from the safetensors mmap) spellings are covered, and a genuine CUDA OOM still gets the Flush hint. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(errors): tighten both new matches, and move the 1455 expectation CI caught a real one, and it was my process error: I ran the full sweep before adding the paging-file change, not after. tests/test_generation_audio_guard.py listed WinError 1455 among the OOM signatures and asserted it yields "ran out of memory / try Flush". The new branch routes it to the paging-file advice instead. That test's INTENT -- a genuine memory failure must never fall through to the unknown catch-all -- is preserved and still asserted; 1455 simply gets a more specific memory message now. Expectation moved, guard kept. Two over-broad matches tightened (CodeRabbit), both in the same direction: a rule that fires too widely replaces correct advice with advice that cannot work, which is the exact defect this branch exists to fix. * The TLS EOF wording is OpenSSL's, but nothing stops an unrelated component saying something similar, and calling a local fault a network problem sends the user to check a connection that was never involved. Now gated on an `ssl` marker; the real message always carries it. * MODEL_DOWNLOAD_INTERRUPTED required "client has been closed" OR "cannot send a request". The latter alone is generic enough to appear beside an unrelated import failure, where overriding TRANSFORMERS_IMPORT would swap correct reinstall advice for a "just retry" that never succeeds. Now requires the closed-client wording itself. Negative regression tests for both, plus the positive cases they must not cost us. Also resolves core.failure through a fixture at call time rather than importing it at module level, per the suite convention -- sibling tests reload these modules and a stale binding makes the file order-dependent. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: debpalash <nizam4103@gmail.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
e60c5364ea |
fix(errors): strip terminal colour codes before a failure reaches the user (#1344) (#1353)
* fix(errors): strip terminal colour codes before a failure reaches the user (#1344) yt-dlp colourizes stderr whenever it thinks a terminal is attached, and the frozen backend's pipes are enough for it to think so. A restricted-video failure surfaced as `download: ^[[0;31mERROR:^[[0m [youtube] …`, which reads as an OmniVoice bug rather than a message from YouTube. Fixed at build_failure, the choke point every surfaced failure passes through, so the whole class is covered — ffmpeg, uv, pip, cargo and anything else that colours its output, not just the reported command. Order matters and is pinned by a test: strip_ffmpeg_banner anchors on "ffmpeg version " at the start of a line, so a leading colour code would hide it and quietly reinstate #1309 for any colour-emitting ffmpeg build. The pattern covers CSI and OSC (window title / hyperlink) sequences, not just SGR colour, and never empties a non-empty message. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(errors): escape-only text falls back to the error class, not to bytes CodeRabbit, both valid: - strip_ansi kept the original when stripping left nothing visible, so build_failure copied raw escape bytes into reason/error/detail — the very thing this PR exists to stop. build_failure already falls back to the exception class name for an empty reason, and a class name is a real answer where a run of escapes is not, so let it do that. - the test file bound `from core import failure` at import time. Sibling suites reload and purge core.* between tests, so that alias can outlive the module the app uses and the file would assert against a stale copy while looking green — #1269 exactly. Resolved via a fixture at call time. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
42d017ef9b |
fix(errors): strip ffmpeg's banner so the failure is the message (#1311)
* fix(errors): strip ffmpeg's banner so the failure is the message ffmpeg and ffprobe print a version + configuration banner to stderr on every invocation, before doing any work. When a command fails we capture that stderr and it becomes the error, so #1309's reporter was shown several hundred characters of build flags — "ffmpeg version N-125781-gacf6b520c1-20260727 … --pkg-config-flags=--static --enable-gpl …" — and not one word about why the extract failed. The diagnosis is always AFTER the banner. Stripped centrally in build_failure() rather than at the extract call site: every stage that shells out to ffmpeg (dub prep, export, retime, media probe) captures the same stderr and had the same problem. Done before classify() runs, too — matching docs topics against a build configuration string is how a real topic gets missed. A message that is ONLY a banner keeps the banner: unhelpful beats empty. Closes #1309 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(errors): assert the banner is GONE, not just that the error is present CodeRabbit: the classification test only checked that the post-banner text appeared in `reason` — which was true before the fix too, since the banner was simply prepended to it. It passed against the code it was written to catch. Now asserts the banner markers are absent and that `reason` STARTS with the real error, since burying the diagnosis after 300 characters of build flags is the actual user complaint. Fails without strip_ffmpeg_banner(). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
7dd7e89e7b |
fix(errors): explain a cut TLS connection instead of printing _ssl.c:1016 (#1312)
* fix(errors): explain a cut TLS connection instead of printing _ssl.c:1016 #1301 surfaced as "500 Internal Server Error: [SSL: UNEXPECTED_EOF_WHILE_ READING] EOF occurred in violation of protocol (_ssl.c:1016)" — meaningless to a user, and unclassified: the existing SSL branch requires "handshake" or "certificate verify failed", so this fell through with no hint at all. Deliberately a SEPARATE class from SSL_HANDSHAKE_FAILURE rather than widening it. That class means a proxy re-signed the certificate with a CA certifi does not trust, and its advice is to set SSL_CERT_FILE or add an antivirus exclusion. Here the handshake never failed on trust — the socket was cut mid-exchange, usually flaky Wi-Fi, a reconnecting VPN, a captive portal, or a server dropping a long transfer. Sending that user to fix their certificate store is sending them to fix something that is not broken. Classified before the handshake branch, because the raw text contains "ssl" and the broader branch would otherwise claim it. Added to _CONTEXT_FREE_HINT_CLASSES since its trigger is an exact OpenSSL string — that matters here, because the raw 500 handler is precisely where the reporter met it. Closes #1301 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(errors): scope the resume guarantee, and require a TLS marker Greptile P1 — the hint promised "OmniVoice resumes partial downloads rather than starting over", unqualified. That is verified for HF model downloads (snapshot_download) and segmented_download, but the hint is static and also reaches media fetches where nothing guarantees it. Shipping an instruction that is not true is the exact class of bug this session has been removing, so the guarantee is now scoped to models. CodeRabbit — the matcher accepted either EOF phrase with no ssl marker. "unexpected EOF" is a phrase a parser or another transport can produce, and those would have been handed VPN/proxy advice. The OpenSSL text always carries the marker, so requiring it costs nothing. Also fixes a tautology: test_not_mistaken_for_a_cert_trust_problem asserted != SSL_HANDSHAKE_FAILURE, which the OLD classifier satisfied by returning "". It now pins the exact class. 2 of the 9 tests fail against the previous commit. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
c3d0d6b123 |
feat(update): announce updates as a toast with actions, not a wall of text (#1272)
* fix(dub): purge order was hash-dependent, and the cap could evict a live marker main went red on my own test. Two distinct defects, both mine. 1. `targets = set(job_ids)` made iteration order depend on PYTHONHASHSEED, so which markers a cap-forced trim discarded was luck. That is why the test passed locally and failed in CI — verified: the old code passes at seeds 0/7/42 and fails at 12345. Now a de-duplicated list in caller order. 2. The size cap could evict markers the CURRENT purge had just recorded. Those are the newest and the likeliest to still be held by a running job, so dropping one is precisely the resurrection this mechanism exists to prevent. A 'clear history' larger than the cap forced exactly that. The cap now never touches the current purge, making the real bound cap + one purge — stated plainly rather than implied. Tests now pin both: identical survivors across runs, and an oversized purge keeping all of its own markers. Verified across six hash seeds; full suite green under 12345, the seed that reddened main. * feat(update): announce updates as a toast with actions, not a wall of text A version's release notes ARE the whole changelog section — v0.4.1's was 42 bullets. Any surface that renders them inline becomes unusable: older builds put them in a blocking OS dialog that filled the screen and had to be dismissed before the app could be touched. Removing that dialog left the opposite failure. The only remaining signal was a 6-pixel dot beside the version number in the footer, which is easy to never notice — so users either got shouted at or told nothing. A toast is the middle: it names the version, offers Install and restart / What's new / Later, and leaves. The notes stay one click away in Settings → Updates, where there is room. Keyed by version so the 6-hourly re-check replaces rather than stacks, and it never auto-dismisses — an update the user hasn't answered is still true. Install declines while a generation is running, since the relaunch would lose it. Strings added to all 21 locales, translated rather than English-filled. Tests pin the shape that matters: the toast takes no notes prop at all, renders under 200 characters, and cannot stack duplicates. * fix(update): don't relaunch while work is in flight; share the busy check Review of #1272 found the restart guard was `dubStep === 'generating'` and nothing else. Installing an update relaunches the process, so that permitted throwing away a dub upload, a transcription, a translation, an export or a standalone TTS synth. The same narrow check was written twice — in the new toast and in UpdatesPanel — so the two could also drift apart. Replaced with a single `isAppBusy(state)` in utils/appBusy, unioning the signals the store actually has: the dub state machine, the floating status pill (which every long background operation already pushes to), and `ttsGenerating` — a new transient store field mirroring useTTS's local `isGenerating`, which lived in a hook where no global check could see it. `dubStep === 'editing'` is deliberately not busy: it waits on the user, and counting it would block updates for as long as a transcript stays open. Also from review: - A failed lazy import of the toast fell into the outer catch, whose setUpdateIdle() erased the update that had just been found. The announcement is optional; the available state is not. - "Dismiss" was machine-translated into the employment sense — terminate an employee — in de/ja/ru/zh-CN/zh-TW, on close buttons and one aria-label. Swept every locale rather than the four lines that were flagged. - Hoisted the test's mock state with vi.hoisted. CodeRabbit's changelog finding is declined: Highlights bullets carry no issue refs by design, which tests/test_changelog_style.py enforces. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(shutdown): a quit mid-generate is not a 500, and not a bug report (#1276) #1174 made a model load interrupted by shutdown benign for the background preload, but a *request* that triggered a load took the generic unhandled- exception path: crash log, ERROR traceback, and an error-journal entry that feeds the bug-report pipeline. Quitting the app with a generate queued surfaced "500 Internal Server Error: model load skipped: backend shutting down" and offered to file a GitHub issue for a normal teardown. Nothing failed — the process is exiting. The handler now answers 503 with Retry-After and an actionable detail, ahead of the crash-log/journal writes. Frontend half of the same bug: toastErrorWithReport offered "Report" for any error. A 503 means "not now, try again" by definition, so it now shows the backend's message without the report action — covering a still-warming backend too, not just this shutdown path. Fail-before/pass-after tests on both sides, including that the shutdown case leaves no crash-log entry and no journal record. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(models): repair a half-downloaded model whichever way it reports (#1273) transformers has two unrelated wordings for "this snapshot has no weight shard", sharing no words: hub load "<repo> does not appear to have a file named …" local dir "Error no file named model.safetensors, … found in directory …" The self-heal and the failure classifier both matched only the first. The second is what a load of a *subfolder* inside a cached snapshot raises — exactly where an interrupted download leaves a half-written repo — so the reporter got neither the automatic repair (delete broken entries → re-download → retry) nor an actionable hint, just a raw 500. Their disk had 10.6 GB free, i.e. a download that ran out of room. The phrase list now lives once in core.failure, so the healer and the error text cannot disagree about what an interrupted download looks like. Both fragments of the second wording must match — "no file named" alone is ordinary English and must not claim the class. Also hardens the #1276 handler found by running both suites in one session: services.model_manager can be imported under two module names, making two distinct ModelLoadInterruptedByShutdown classes and breaking a bare isinstance — which silently restored the 500 this fix exists to remove. Now matched by isinstance OR class name, with a test that raises a same-named class from a different module. Verified against the full tests/ + backend/tests/ session: the only remaining failures are the four pre-existing #1269 isolation leaks, unchanged. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(update): count synths in flight at the chokepoint, not in one caller Review round 2. Both Greptile P1s were the same weakness in my first pass: tracking the synth in useTTS meant only the Generate tab counted (voice previews, the compare modal, the stories editor and profile previews call generateSpeech directly and were invisible), and a boolean meant two overlapping syntheses cleared each other — whichever settled first reported "idle" while the other was still running, so Install and restart discarded it. Moved to `api/generate.ts`, around the one `/generate` call all seven paths share, as a count with a `finally` release (so an abort or a network error frees it too). A future synth caller is covered without opting in. Fail-before verified: the overlap test fails against the boolean version. Also from review: - UpdatesPanel re-reads the busy state at click time; the render-time snapshot only exists to disable the button, and work can start after the last render. - The 503 now carries the allowed-origin CORS headers. Without them the browser reports a bare CORS failure and the actionable detail — the whole point of the fix — never reaches the user. Both error responses build them through one helper now. - Log `request.url.path`, not the full URL: a query string can carry tokens and newlines. - `update.busy` said "finish your dub first" in all 21 locales, but the guard now covers uploads, transcription, translation, export and synth. Rewritten as work-in-progress wording, translated per locale. - ru `common.dismissStatus` was "reject status"; missed in the earlier sweep. Declined: CodeRabbit's request for a `(#NNN)` ref on the Highlights bullet — Highlights carry no refs by design (CLAUDE.md), and tests/test_changelog_style.py enforces it. Also declined gating the cache re-download behind a confirmation: the self-heal already existed and already ran for the sibling wording, this only stops it missing half the class, and it stays behind the existing _hf_offline() check. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(update): grey out Install while work is running CI lint caught `busy` as unused after the click-time check moved into the handler — and that exposed a real gap: `busy` had never been wired to anything. The comment claimed it disabled the button; it didn't. Clicking Install during a synth just bounced a toast back. Now it does what it said: the Install and Restart buttons are disabled while work is in flight, with `update.busy` as the tooltip. The click-time read stays the authority, since work can start between the last render and the click — the disabled state is the explanation, not the safety. Adds the first UpdatesPanel test. The case it pins hardest is the inverse of the bug: a busy predicate stuck at true would make the app permanently un-updatable, which is worse than what this fixes. So it asserts enabled when idle and during dub 'editing' (which waits on the user), disabled during an upload, a translation, and one or more synths. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
b66b09ceaa |
fix(dub): deleting a dub no longer resurrects it (#1252, #1253) (#1270)
* fix(dub): deleting a dub no longer resurrects it (#1252, #1253) Split out of #1264. The other six fixes there are independent error-message changes that needed no corrections; this one is a concurrency change that needed five rounds, each finding something real in work that looked finished and tested: 1. review: merge and save were split, so a delete between them had the row written straight back; 2. review: dict membership cannot express 'withdrawn' — an absent key means 'not written yet' for a new job and 'deleted' for an established one; 3. fail-before check: the race test was not testing the race (it asserted WHAT happened, never WHEN, and 'save after delete' IS the resurrection); 4. review: the gate sat in two ingest helpers while eight direct save_job call sites bypassed it; 5. direct check of the reported scenario: the tombstone covered in-flight ingests only, so a delete during a RENDER — the common case — still resurrected the job. Riding a release on that record is not a good trade, so it ships on its own. What it does now: the withdrawal is recorded when a job is DELETED, held in a bounded LRU (there is no moment at which a delete stops mattering), and checked inside save_job so every caller inherits it. Re-importing an id is the only thing that revives it. The lock is re-entrant because the atomic helpers call save_job while holding it. Carries #1252's message half too: str(KeyError) is the repr of the key, which is how the user's own job id became the entire error text. * fix(dub): expire withdrawal markers by age, not by count Greptile P1 — the sixth real finding on this fix, and reachable by ordinary use. DELETE /dub/history selects every row with no limit, so a user with a large history clearing it mid-render pushed that very job's marker out of a count-bounded LRU, and the render then wrote it straight back. Age is the honest policy: what matters is how long ago the delete happened, not how many others followed it. Six hours outlives any realistic render or transcribe. The count cap stays only as a memory backstop, raised far above any real history and documented as such. Verified fail-before: restoring count-based eviction fails both new tests. * docs(changelog): the dub delete-resurrection fix (#1252, #1253) |
||
|
|
73ebd6518d |
fix(errors): six failures that reached users as raw text (#1262, #1256, #1251, #1247, #1257, #1254) (#1264)
* fix(errors): four failures that reached users as raw OS text (#1262, #1256, #1251, #1252) #1262 — a voice profile named in any non-latin-1 script 500'd every download endpoint with "'latin-1' codec can't encode characters in position 22-25". `attachment; filename="` is exactly 22 characters, so those were the first four characters of the user's own name. The sanitisers in front of the header filtered with str.isalnum(), which is True for every alphabetic script — they stripped punctuation and passed exactly what breaks the header. Ten sites, one RFC 6266 builder, plus a guard so an eleventh can't be hand-written. #1256 — a synth died on FileNotFoundError: 'ffprobe' and was reported as "an error OmniVoice doesn't recognize", on a Mac where the app's own ffprobe was resolvable the whole time. Our call sites pass explicit paths; a dependency shelling out by bare name does not. The resolved directories are now published on PATH, and the failure is classified either way. #1251 — "The paging file is too small" reached the user as a bare 500. It was already counted as an OOM, but that remedy (close apps, lighter engine) is wrong on a 32 GB machine — the fix is a Windows setting, and the hint now says which. Matched on the code in both the Python and Rust spellings. #1252/#1253 — deleting a dub mid-import crashed it with `ingest: 'mgw39lx3'`: str(KeyError) is the repr of the key. The pipeline blind-subscripted a job that DELETE /dub/history/{id} had popped minutes earlier. It now stops quietly, and no exception whose str() is a bare value can present itself that way again. * fix(engines): Unload 400'd, a wrong language said nothing, DRM was retried by hand (#1247, #1257, #1254) #1247 — list_loaded() advertises in-process engines as `engine:<id>` with "unloadable": true, but unload() only ever handled tts/diarization/sidecars. The panel was rendering a button for ids the dispatcher rejected. The engines already implement unload(); only the routing was missing. The contract test written for it immediately found a second instance — `capture-asr`, listed the same way with no branch either — which is why it enumerates the listing rather than hard-coding ids. #1257 — MLXAudioBackend.supported_languages() returns ["multi"] on the stated assumption that "each engine silently ignores languages it doesn't know". It doesn't; the library raises. So the picker offers all 646 languages and the rejection arrived as a bare list of 23 codes, naming neither the engine nor the way out. Enumerating each model's real language set would be a brittle map that goes stale every engine update — name the engine and the fix instead. #1254 — reported as intermittent: the same URL failed as DRM-protected, then succeeded on retry. Real DRM doesn't lapse; the player client varies. That is the same shape as the 403 case which already escalates through _YT_PLAYER_CLIENTS, so DRM now routes into it. If every client still refuses, the failure is classified instead of arriving as a raw yt-dlp line. * fix(review): close the delete race, narrow the tool match, sanitize the fallback Greptile P1 + CodeRabbit Major — verified real, and mine: splitting merge from save left a window where a delete lands between them, so the pending save UPSERTs the row straight back and a dub the user deleted reappears. Now one atomic step under _dub_jobs_lock, with both delete endpoints purging rows and memory under that same lock. That also fixed DELETE /dub/history, which deleted every row but evicted nothing — an in-flight job survived 'clear history' outright and re-saved itself on completion. CodeRabbit Minor (#1256): the media-tool match accepted any message ending in 'ffmpeg'/'ffprobe', so a missing FILE at /tmp/ffmpeg got the 'repair your media engine' remedy. Now requires the name unquoted-and-unqualified. CodeRabbit Major (#1262): `fallback` reached the header verbatim whenever the real name folded away entirely, walking past every guard the name goes through. Folded like the name. CodeRabbit Major (#1256): the PATH log printed resolved directories, and a user-set FFMPEG_PATH sits under their home. Logs a count now. CodeRabbit Minor (#1262): the subtitle-route assertion also passed against the pre-fix header; it now asserts filename*= too. Skipped: 'Highlights bullets must end with (#N)'. CLAUDE.md scopes that to the ### subsections; none of the seven pre-existing highlights carry refs, and tests/test_changelog_style.py encodes the rule already. * fix(review): the remaining unlocked save paths, an over-broad signature, two weak tests Greptile P1 — the mid-pipeline put_job + save_job pairs were still unlocked, so a clear-history landing between them left a ghost row behind the purge. Both go through put_and_save_job now; only the final completion gate decides whether a withdrawn job's work is kept. CodeRabbit Major (#1257) — 'unsupported language' as a bare prefix also matches 'Unsupported language model configuration', handing a model/config failure engine-switch advice it has no use for. The loose wordings now require the rejected thing to be a code or to end there. CodeRabbit Major (#1257) — the OOM test asserted on SOURCE TEXT, which passes even if the call is unreachable or its result discarded; #1224 taught this same lesson on this codebase. Both it and the language rewrite now drive the real _run_backend_inference with a raising backend. CodeRabbit Minor (#1257) — 'or "engine" in message' always passed, since the production template contains the word. Asserts the resolved class name now. CodeRabbit Major (#1252) — the delete-race test deleted the job BEFORE the merge, which only re-tested the absent case and would pass with the two steps still split. It now interleaves a real second thread against a slow save. CodeRabbit Major (#1256) — a hardcoded /tmp literal trips Ruff S108; built from tmp_path instead. * fix(review): a withdrawal must survive the job's first write CodeRabbit Major — the concern is real, though its suggested fix (gate the checkpoint on the job already existing) would break creation: an ingest's FIRST persistence is what creates the entry, so that gate would never pass. The actual defect is that dict membership cannot express 'withdrawn'. An absent key means 'not written yet' for a new job and 'deleted' for an established one — two opposite instructions from one signal. So a clear-history arriving before the first checkpoint was silently undone by that checkpoint recreating the row, and the run then persisted its result into history the user had just cleared. Tombstone it explicitly: the ingest declares itself in flight, a purge marks any in-flight id withdrawn, and both write paths refuse a withdrawn id. Released in , so it's bounded by concurrent ingests and can't poison a later run that reuses the id. That also fixed clear-history properly: a job with no row yet appears in no id list, so only an in-flight sweep can catch it. CodeRabbit Minor — my race test waited on an event that could not be set while the save held the lock, so it burned its full 2s timeout every run and synchronised nothing. It now waits for the purge thread to REACH the purge. * test(dub): the race test was not testing the race Caught by verifying fail-before rather than trusting the test: splitting merge from save — the exact resurrection bug — passed all 22 tests. The assertions checked WHAT happened (the save ran, the row was deleted, the job left memory) but never WHEN. A save landing after the delete is indistinguishable from one landing before if you only assert that both occurred — and 'after' is precisely the resurrection. Now recorded and asserted as an order. With merge+save atomic the purge cannot start until the save finishes, so the sequence is always save-then-delete; split them and it fails with ['delete', 'save']. That is the second time this test needed rewriting: v1 deleted the job before the merge and only re-checked the absent case, v2 interleaved a real thread but asserted the wrong thing. Both looked like tests. Also documents why the DB write sits inside the lock (atomicity beats a rare 5 s sqlite busy-timeout stall) and that no locked region calls another, so the non-reentrant lock cannot deadlock — verified by walking every locked region. * fix(dub): gate the withdrawal at save_job, not at its callers Greptile P1 — and the same class I'd already fixed, unfixed elsewhere. The withdrawal check sat in the two ingest helpers, but eight direct save_job call sites across dub generate / translate / export / core bypass those entirely. Deleting a dub mid-RENDER therefore still resurrected it, which is at least as likely as deleting mid-import. Moved the gate into save_job itself: one choke point, every caller inherits it, and the ninth cannot forget. That needs a re-entrant lock, since the atomic helpers call save_job while already holding it — a plain Lock would deadlock the backend, so a test pins the lock type and another exercises the nested path. Verified fail-before: removing the gate fails the new test. * fix(dub): the withdrawal only covered ingests, so it covered almost nothing Caught by testing the reported scenario directly instead of trusting a green suite: CI passed, 26 tests passed, and a dub deleted during a RENDER was still resurrected. The tombstone was scoped to in-flight ingests. But a dub is imported once and rendered many times, so the realistic delete lands during a render — long after its ingest ended — and end_ingest was CLEARING the tombstone at exactly that point. The rare case was protected and the common one left open. Now scoped to deletions, not ingests. Kept in a bounded LRU rather than cleared on completion, because there is no moment at which a delete stops mattering: any operation still holding that job can persist it. Re-importing an id is the only thing that legitimately revives it. Verified fail-before: the previous scoping fails three of the new tests. * refactor: move the dub delete-resurrection fix to its own PR (#1270) The six fixes left here are independent error-message changes that needed no corrections. The dub concurrency change needed five rounds, each finding something real in work that was already reviewed, tested and CI-green — the last of them being that the fix did not fix the reported case at all. Riding a release on that record is a bad trade, so it ships separately as #1270. This branch keeps #1262, #1256, #1251, #1247, #1257 and #1254; the KeyError message half goes with the dub PR, since it is that issue's other half. |
||
|
|
95fae6df0d |
Merge main into fix/download-resume-oom-1224
# Conflicts: # CHANGELOG.md |
||
|
|
09b60901e4 |
Merge main into fix/lazy-model-import-1229
# Conflicts: # CHANGELOG.md # backend/core/failure.py |
||
|
|
2009ef329c |
Merge main into fix/dub-download-errno22-1225
# Conflicts: # CHANGELOG.md # backend/core/failure.py |
||
|
|
027c08f1ff |
fix(dub): #1225 review — the preflight error had no class, and ENOENT no facts
Two P1s, both correct, and both the same shape as the bug being fixed:
- The preflight OSError I added ("Can't save the download: …") matched neither
an errno nor a download marker, so classify() returned "" and the user got
NO hint — the exact dead end this PR exists to remove. Reworded to carry
both signals; a test now asserts the class and that the hint names the data
directory.
- classify() covered ENOENT but _with_target_facts' own signature list did
not, so a job folder that vanished after preflight produced a
disk-classified error that never named the folder.
That second one is a drift class, not a one-off: two lists answering "is this
a disk problem?" will diverge again. They now share
failure.is_os_write_refusal(), with a test asserting both consumers agree
across all four errnos.
CHANGELOG entries shortened with refs last (CodeRabbit).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
63d40c33ce |
fix(synth): #1227/#1221 review — classify on a marker, not on shared wording
Two P1s, both correct: - The enriched write failure lost its class. _describe_write_failure rewrites the message, so the word "libsndfile" no longer appears — classify() returned "" and the auto bug report and docs deeplink had nothing to name. Worse, my own end-to-end test allowed "" as a pass, which is exactly why it went unnoticed. audio_io now emits a stable AUDIO_WRITE_FAILED_MARKER, failure.py matches that, and the assertion is exact. - "error opening" was far too broad. It appears whenever a model, archive or config file fails to open, so any such failure was handed the audio-file remedy (check your disk, add an antivirus exclusion). Dropped in favour of the marker; a regression test pins that a corrupt-model-archive error keeps its own guidance. A third test pins the marker across the two modules — core/ cannot import services/, so the string is duplicated by necessity, and a reword on either side would silently un-classify every enriched write failure. CHANGELOG entries shortened with refs last (CodeRabbit). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
ed321c56e6 |
fix(download): retry a truncated model download instead of aborting (#1224)
The reporter's captured log tail, just before the backend was SIGKILLed: httpx.RemoteProtocolError: peer closed connection without sending complete message body (received 4084175097 bytes, expected 4580080592) A 4.6 GB model died at 4.0 GB — the single most retry-worthy failure in the download path, and it was retried nowhere: - the installer's loop caught (HfHubHTTPError, LocalEntryNotFoundError, OSError). httpx.RemoteProtocolError inherits Exception, NOT OSError, so it escaped all five attempts. The loop now decides by CLASSIFICATION rather than exception type, so the next transport error with a novel type doesn't reopen the same hole. - is_hf_connectivity_error — the single source of truth for "transient download failure" — had no truncation signature, so a widened catch alone would still have called it permanent. It now knows the httpx wording plus the urllib3/http.client equivalents (IncompleteRead, "connection broken", "response ended prematurely"). - the engine load path had no retry at all. VoxCPM2 and MOSS-TTS-Nano now go through the existing _retry_once_with_fresh_hf_client hook, widened to retry transient download failures with a bounded, backed-off budget. The HF cache is resumable (correctly-sized blobs are skipped by hash), so a retry continues rather than restarting. The closed-client path (#880) keeps its single-shot budget deliberately: it's a client-state bug, not a network condition, so a fresh session hitting it again means repeating won't help. Both budgets are now pinned by tests. Also logs the low-memory advisory in the STREAMING synth path. /generate has done this since the earlier 16 GB-Mac reports, but the streaming path — which the desktop UI tries first — did not, so the load most likely to tip a machine into an OS OOM kill was the one load leaving no trail in the captured stderr tail a SIGKILL report has to go on. Regression test: tests/test_truncated_download_retry.py (9 of 13 fail before). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
8f4914272e |
fix(dub): name the folder when a URL ingest fails on a disk error (#1225)
`download: Unable to download video: [Errno 22] Invalid argument`, hit three times in a row on the same URL. Two things made it a dead end. `classify()` matched the generic "errno 22" rule (#763, written for the ASR temp-WAV path) before any download rule, so the attached hint told the user to check their system TEMP folder — while the failing directory is the job folder under the OmniVoice data dir. The one actionable instruction pointed at the wrong place. There is now a VIDEO_DOWNLOAD_OS_ERROR class, checked first, covering the OS-refusal errnos (22/13/28/2) when the message carries download context; #763's class is untouched for everything else. The message also named neither the target nor the reason, so nothing in it distinguished a full drive from a read-only folder from an antivirus lock. `failure.describe_path_target()` (shared with the #1221 audio-write path) attaches what we can observe — exists / writable / free space — and the download site now: - preflights the job dir and fails immediately when it can already see the write can't succeed, instead of starting a download that can only fail; - enriches an OS-refusal failure with the destination facts, leaving network/format failures untouched so they keep their own (retryable) class. Regression test: tests/test_dub_download_os_error.py, including that a transient network failure still classifies as retryable and that the ASR path keeps OS_INVALID_ARGUMENT. Its yt-dlp calls are blocked outright — the first draft passed standalone and silently made a real network call under full-suite ordering. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
8cef017ee5 |
fix(synth): name the App Control block and the libsndfile failure (#1227, #1221)
`_oom_friendly_reraise` classifies every known way a generate can die; two real reports fell through to its "an error OmniVoice doesn't recognize" catch-all. #1227 — `OSError: [WinError 4551] An Application Control policy has blocked this file`. Windows Smart App Control refused to load a file the engine needs. Now named, with the setting to change (and the caveat that Windows only lets you turn Smart App Control off once). WinError 1260 — the same class from AppLocker / Software Restriction Policies — is matched too, on the numeric codes, since Windows translates the message text. #1221 — `LibsndfileError: System error.`, libsndfile's bare wording for an OS-level audio read/write failure: no path, no errno, no next step. Two parts: - `audio_io._safe_torchaudio_save` now re-raises write failures naming the target, whether its folder exists and is writable, and the drive's free space — the facts that identify a full disk, a removed drive, or an antivirus/OneDrive lock. (It cannot re-use the original type: `LibsndfileError.__init__` takes an integer libsndfile code, so `type(e)(message)` builds an exception whose `str()` raises.) - `_oom_friendly_reraise` covers every other libsndfile surface (reading a reference clip, a decode) with the same causes. Both get a `core.failure` class + hint (WINDOWS_APP_CONTROL_BLOCKED, AUDIO_IO_FAILED) so the auto bug report and docs deeplink name them. Regression test: tests/test_synth_error_classes.py — 11 of 12 fail before. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
b6d92637e7 |
fix(startup): don't let one optional transformers symbol kill the backend (#1229)
`backend/api/routers/profiles.py` imports two pure-stdlib regex helpers from `omnivoice.utils.voice_design`. That import pulled in `omnivoice/__init__`, which eagerly imported `omnivoice.models.omnivoice` — torch, torchaudio, transformers, flex_attention, the whole model definition — including a top-level `from transformers import HiggsAudioV2TokenizerModel`. transformers exposes that class through its lazy module and gates it on the torchaudio backend, so the *attribute access* raises when torchaudio is missing, ABI-mismatched, or installed without discoverable distribution metadata — Colab's system Python, an interrupted `uv pip install`. It raised during `backend/main.py`'s module import, before FastAPI existed: TTS, dubbing, ASR and Settings all dead, the user left with a uvicorn traceback and "Backend did not become healthy within 5 minutes". Two changes, both structural rather than Colab-specific: - `omnivoice/__init__` resolves its model exports lazily (PEP 562). Importing `omnivoice.utils.*` no longer costs — or risks — the model stack. `from omnivoice import OmniVoice` is unchanged; only the timing moves. `backend.spec` already lists `omnivoice.models.omnivoice` as a hidden import, so the frozen build is unaffected. - `HiggsAudioV2TokenizerModel` resolves at its single use site in `from_pretrained`, and a failure there raises an ImportError naming torchaudio and the reinstall. Deferred into a request, `core.failure .classify()` maps it to TRANSFORMERS_IMPORT and attaches a repair hint — whose text now names torchaudio too, instead of only transformers + an ASR workaround irrelevant to this path. Colab notebook: cell 2's sanity check imports the model stack (and prints torchaudio/transformers versions), so a broken env fails there with the real error instead of as a health timeout two cells later. Regression test: tests/test_omnivoice_lazy_model_import.py pins that the utils import loads no heavy module, the lazy exports still resolve, and the deferred failure is actionable and classified. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
23367cccaf |
fix: first-run wizard version + mirror-unreachable rescue + lifecycle-aware backend reachability (#1094)
Three fixes from the same first-run session report:
- SetupWizard shows v{APP_VERSION} in its masthead (same identity mark as
the install splash footer), so setup screenshots identify the build.
- A dead configured HF mirror no longer strands the wizard: the
install_error SSE now carries docs_topic (core.failure.classify), and
WizardLibrary renders the MirrorRescue quick-pick (extracted from
SetupWizard, now including the official preset) next to the failed row,
retrying it the moment a new endpoint is applied. PUT /hf-mirror clears
the install cooldowns (no 429 on the immediate retry) and clearing to
official also drops the legacy hf_endpoint pref that silently kept the
dead mirror in effect. The hint's false "applied when the app starts"
claim is corrected: downloads resolve the endpoint per call, retry
first, restart only if it still fails.
- "Can't reach the local OmniVoice backend" stops firing during real
start/restart windows: a respawn takes 10-20+s (venv spawn + torch
import) but the transport cascade gave up at ~2.9s. apiFetch now asks
the shell (bootstrap_status via utils/backendLifecycle) whether a
start/restart is in progress and keeps retrying while it is (capped at
120s, matching the supervisor's respawn budget); the new
BackendRestartBanner finally implements the reconnecting banner the
#567 supervisor has emitted events for all along. Truly dead backends
(or non-Tauri deploys) still error promptly.
Regression tests for all three layers; docs synced
(downloading-models.md, troubleshooting.md §14b); CHANGELOG [Unreleased].
Co-authored-by: mergetest <test@local>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
9c81e3389d |
feat(network): automatic Hugging Face endpoint selection — probe, pick, remember (#1082)
* feat(network): automatic Hugging Face endpoint selection — probe, pick, remember Restricted-network first-runs (the #984 class: huggingface.co unreachable, user dead-ends before discovering the mirror setting) now self-heal by default, while explicit endpoint choices are never second-guessed. - New backend/services/endpoint_race.py: parallel HTTPS reachability + latency probes of huggingface.co and the hf-mirror.com community mirror (3s timeouts). Probes are the only signal — no geo-IP, no third-party calls. Reachable beats unreachable; with both reachable the official endpoint wins unless the mirror is decisively faster (anti-flap hysteresis). The pick is cached in prefs and re-raced only on first run, a network-classified download failure, staleness (>7 days), or an explicit "Test again". - Manual mode is sacred: HF_ENDPOINT env, an hf_endpoint pref, or any explicit Settings pick disables auto-switching entirely; OMNIVOICE_HF_ENDPOINT_MODE=manual is a hard opt-out. - Wiring: the wizard preflight races endpoints when nothing is configured (honest copy when the mirror wins; warn-not-block when nothing is reachable); Model Store installs and the model-cache auto-repair resolve their per-call endpoint= through the cached decision, and a network-classified failure re-races once per repo per process and retries on the new winner (same guard pattern as the cache-recovery ladder). - Settings → Models → Hugging Face mirror gains "Auto (recommended)": shows the current pick, measured latency, last-checked time, and a "Test again" button (POST /api/settings/hf-mirror/test). Existing explicit configs surface as the matching manual mode. Panel notes that hf_hub checksums every download regardless of endpoint. - Tests: policy/cache/failover matrices in tests/test_endpoint_race.py, preflight + settings + repair-failover integration with mocked probers, HFMirrorPanel mode tests, and a suite-wide conftest guard that pins the probers so no test can hit the real network. - Docs: downloading-models.md and install/troubleshooting.md describe the automatic default and both opt-outs. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(changelog): Unreleased entry for automatic HF endpoint selection Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(tests): endpoint-probe pin uses an isolated MonkeyPatch and clears the decision cache; dtype guard tolerates stubbed torch The autouse probe pin requested the shared monkeypatch fixture, hoisting its setup earlier for every test and reordering teardown against the fp16 guard — which then ran torch.get_default_dtype() on test_torch_compile_gate's SimpleNamespace stub. The pin now uses its own MonkeyPatch context and also clears the prefs-cached endpoint decision per test (one test's auto pick leaked into other tests' preflight labels on CI ordering). The dtype guard additionally skips non-module torch stubs outright. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(tests): endpoint env vars can no longer leak out of the mirror-settings suite set_hf_mirror writes os.environ[HF_ENDPOINT] during the test, and monkeypatch.delenv(raising=False) on an absent var records nothing to undo — so the write leaked process-wide and flipped later suites' preflight network checks into the explicit-endpoint branch (the CI-order failures). Guaranteed save/restore autouse fixture at the source, plus defensive env shedding in the preflight suite. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
7f59d5f8fe |
fix(models): self-heal HF-cache snapshots with broken file links, retry load once (#1056)
A first-run breaker: all blobs download fine, but the snapshots/<rev>/ entries are dangling symlinks (0 KB) — os.path.isfile() is False on a dangling link, so transformers reports the weights missing even though the bytes are on disk, and the existing resume repair can't fix it. New services/hf_cache_repair.py deletes exactly the broken snapshot entries (dangling symlinks + zero-byte weight/config stand-ins; never blobs, never resolving entries) and restores them via snapshot_download, verifying afterwards — if the restore recreates broken links (hub's memoized symlink probe passing while real links come out broken), it forces hub into copy-mode and repairs once more with real files. model_manager retries the load exactly once per repo per process (rung 0 of the cache-recovery ladder); dead-end errors now name the exact models--<org>--<name> folder to delete. failure.py classifies the class as MODEL_CACHE_CORRUPT so the user-facing error and auto bug report explain the automatic repair. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
dfe2bd1bc3 |
fix(install): classify SSL handshake failures + trust the OS cert store (#976) (#992)
Windows users behind a corporate/antivirus TLS-inspecting proxy got a raw `[SSL: SSLV3_ALERT_HANDSHAKE_FAILURE]` on every model install — the TCP connection reaches the server fine, but the handshake fails because the OS trusts the proxy's re-signed root CA and Python's bundled certifi CA list doesn't. A genuinely different failure mode from #984 (that was TCP-level unreachability to a blocked host, before any TLS negotiation). - backend/core/failure.py: new SSL_HANDSHAKE_FAILURE classification (handshake/cert-verify-failed/sslv3_alert/sslcertverificationerror substring markers) with an actionable hint, added to _CONTEXT_FREE_HINT_CLASSES so append_hint() (already called by setup/download.py's install worker) surfaces it without further wiring. - backend/main.py: truststore.inject_into_ssl() at module level, before any huggingface_hub/requests/httpx network I/O — patches ssl.SSLContext to verify against the OS trust store instead of only certifi's bundled CA list. Not platform-gated (correctness improvement everywhere); wrapped in try/except so it never blocks startup. - pyproject.toml/uv.lock: truststore>=0.9 — pure Python, MIT, PyPA- maintained, zero transitive deps, same class of fix as socksio. Verified: uv lock --check + uv sync --frozen clean (lockfile diff is just the one new package); main.py imports cleanly; full backend suite passes; no hiddenimports entry needed (main.py is PyInstaller's direct entry script per backend.spec, so a top-level import traces normally — unlike socksio's case, which was httpx's internal lazy import). Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
de99bc3bd5 |
fix(net): SOCKS-proxy users can synthesize again — ship socksio, cache-first model resolution (#959) (#966)
* fix(net): SOCKS-proxy users can synthesize again — ship socksio, cache-first model resolution, degrade LLM clients (#959) Under ALL_PROXY/HTTPS_PROXY=socks5:// without socksio installed, httpx raises ImportError AT CLIENT CONSTRUCTION ("Using SOCKS proxy, but the 'socksio' package is not installed"). huggingface_hub's get_session() builds exactly that client inside snapshot_download, so POST /generate 500'd with the bare message even for a fully installed model, and preload_model's model_info probe hit the same error and silently skipped warm-up. Latent since v0.3.5 — #947's fresh-process engine spawning unmasked it in v0.3.10 by handing the user's proxy env directly to a clean backend process. Three layers, so the class (any session-construction failure) is dead, not just the reported instance: * Ship SOCKS support: socksio>=1.0 in [project] dependencies (pure Python, MIT, zero transitive deps) AND in backend.spec hiddenimports — httpx imports it lazily in try/except, so PyInstaller's tracer misses it and the frozen installers would stay broken without the explicit entry. uv.lock regenerated; `uv lock --check` and `uv sync --frozen` (the Docker/release bootstrap semantics) verified. * Cache-first model resolution: from_pretrained's snapshot resolution extracted into _resolve_snapshot_dir() — local dir, else snapshot_download(local_files_only=True) (a complete cache resolves with NO HTTP session constructed), else the original network path. preload_model's failed network probe now falls back to a cache-only check and warms up anyway instead of silently skipping (honest log either way). * Class guards: resolve_skill_client wraps OpenAI() construction — env-shaped construction failures degrade to the existing "LLM unavailable" contract instead of 500ing the calling feature; and core.failure learns SOCKS_PROXY_SUPPORT_MISSING with an actionable hint, appended on the raw-string surfaces (global 500 handler, model-install SSE) via the new append_hint(). Fail-before/pass-after verified by reverting the fix: 11 of the 12 new tests fail pre-fix (the remaining one is the unchanged network-fallback contract). 165 tests green across the touched suites. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(changelog): add SOCKS-proxy resilience under [Unreleased] (#966) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
eb188931b5 |
fix(dub): classify EINVAL transcribe failures so they stop dead-ending (#763) (#936)
A per-chunk temp-WAV write that fails with OSError EINVAL ("[Errno 22]
Invalid argument") — a missing/read-only/full temp dir, a removed drive,
or antivirus — collapsed into "Transcription produced no segments.
[Errno 22] Invalid argument" with no next step. classify() now names the
class (OS_INVALID_ARGUMENT) so build_failure attaches an actionable
temp-dir/disk/AV hint at the exact surface the streaming dub path already
feeds it (dub_core.py:672) — same treatment the ffmpeg and compute-type
classes get. Fail-before/pass-after regression added; the errno-22 token
keeps it from colliding with the errno-2 transformers-import class.
Also stamps the [0.3.9] CHANGELOG section with today's release date
(2026-07-04) ahead of tagging.
Co-authored-by: mergetest <test@local>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
||
|
|
14f1257d1f |
fix(errors): name the configured HF mirror when a model download fails (#874) (#890)
When a non-default HF_ENDPOINT (Settings → Models → Hugging Face mirror,
e.g. hf-mirror.com) is configured and a model load/download fails with a
connectivity error, the raw transformers message ("We couldn't connect to
'https://hf-mirror.com' to load the files…") leaked to the UI as a bare 500
with no next step.
Class fix — one shared classifier in core/failure.py covers every surface:
- classify()/build_failure(): new HF_MIRROR_UNREACHABLE class with a dynamic
hint that names the configured mirror, says it may be down, points at
Settings → Models → Hugging Face mirror, suggests the official endpoint
when the model isn't cached, and notes the restart requirement (HF reads
HF_ENDPOINT at backend start). Checked before the video-download network
class so a model download's "timed out" no longer gets the "video server"
hint. Feeds /model/status and every build_failure event (dub, tasks).
- main.py global 500 handler: appends the hint to the surfaced detail, so
ALL routes that can leak a model-load error benefit (generate, dub,
archetypes, …), not just TTS generate.
- setup/download.py install SSE: the install_error event gets the same hint.
- error_journal: "couldn't connect to" / "max retries exceeded" now classify
as NETWORK_ERROR (was UNKNOWN) for auto-attached bug reports.
- model_manager (#886 family): the "cache incomplete and could not be
auto-repaired" message now names WHY the auto-repair failed (mirror
outage, offline mode, full disk no longer read identically), which also
lets the mirror hint fire on that surface when applicable.
Fail-before/pass-after regression tests in tests/test_hf_mirror_error_class.py
(12 of 13 fail on main).
Fixes #874
Co-authored-by: mergetest <test@local>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
2339cc8e85 |
fix(model): actionable "reinstall transformers" hint on a corrupted-install model-load error (#676)
A model load failed with `[Errno 2] No such file or directory:
'…/site-packages/transformers/models/qwen3/modeling_qwen3.py'` — the user's
transformers install was incomplete (the file is missing while a correct 5.3.0
install has it; an interrupted `uv sync` / antivirus / partial update drops it).
The System Check showed the raw path + "Check logs and try restarting", which is
useless — restarting can't restore a missing file.
Two fixes:
1. core.failure.classify(): recognize this corrupted-install variant. It's a
FileNotFoundError, not an ImportError, so the existing TRANSFORMERS_IMPORT
match ("could not import module"/"AutoFeatureExtractor") missed it. Now also
matches a "no such file"/"errno 2" + "transformers" + "site-packages" signal
(substrings checked separately so it works on POSIX `/` and Windows `\`
paths). An unrelated package's missing file is NOT mislabelled.
2. model_manager._load(): build the /model/status error via build_failure so it
carries the classified hint AND strips the home dir, instead of storing the
raw str(exc). The System Check now shows "Your transformers install is
incomplete. Reinstall it (uv pip install --reinstall transformers) or switch
ASR to faster-whisper" — the existing TRANSFORMERS_IMPORT hint.
Docs: troubleshooting §1a documents the error + the reinstall fix.
Test: test_failure_classify.py pins the POSIX + Windows path forms classify as
TRANSFORMERS_IMPORT with a "reinstall" hint, and that an unrelated package's
missing file does not.
Co-authored-by: mergetest <test@local>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
ff168048be |
fix(backend): import omnivoice from source when the editable install is missing (#564) (#573)
#564 ("No module named 'omnivoice'") is the backend failing to import its OWN package at the first model call (the dub SSE error on dub:upload). `omnivoice` is an editable install, so an interrupted/offline `uv sync` that installed deps but never laid the editable record, an antivirus-quarantined `_editable_impl_omnivoice.pth`, or an upgrade where only the lock-gated drift sync ran leaves the venv able to start uvicorn yet unable to import omnivoice — it boots fine and only fails at runtime, so the bootstrap health gate and the exit-based broken-venv self-heal (which only see a process that won't start) never catch it. Fix the whole class at the import layer: main.py now also appends the project root (the parent of backend/, where the desktop layout always copies omnivoice/) to sys.path, guarded on omnivoice/__init__.py existing. The backend then resolves omnivoice from source regardless of the editable-install state — covering every variant above. Appended (not inserted) so a real site-packages/editable install keeps precedence and it can't shadow a different omnivoice; a no-op in Docker (no sibling omnivoice/) and a harmless duplicate in a dev checkout. Also routes "No module named 'omnivoice'" through failure.classify() → BROKEN_VENV so, if it ever still surfaces, the toast points at Clean & Retry instead of a bare import error. Regression test covers the classify mapping and its negative guard (a legitimately-named omnivoice_* helper must not match). Co-authored-by: mergetest <test@local> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
b8cc0de44a |
fix: actionable download errors + self-heal a relocated venv (#554/#536/encodings) (#562)
Two robustness fixes that share the failure taxonomy: - Video downloads (#554 douyin "Unsupported URL", #536 "Broken pipe"): yt-dlp's raw error surfaced with no next step. classify() now names UNSUPPORTED_VIDEO_URL (non-downloadable link shape — paste a direct video page or drop a file) and VIDEO_DOWNLOAD_NETWORK (transient CDN/network drop — just retry; the partial download is already cleaned up), each with an actionable hint. yt-dlp is on a current pin, so this is graceful classification, not a dependency bump. - "No module named 'encodings'" (relocated/copied/restored venv whose interpreter can't bootstrap its stdlib — exit 1, not 106): slipped BOTH #314 self-heal matchers, so the user saw the error forever. Widen backend_exit_indicates_broken_venv to also match the full quoted phrase, routing it into the existing rebuild-once self-heal. Kept narrow so an app-level import of an 'encodings'-prefixed package can't trigger a rebuild. Plus a BROKEN_VENV hint for the case where the rebuild itself can't run. Tests: classify() maps the 3 new classes with hints (a generic reason still ""), and the 'encodings'-prefixed-package negative guard holds; the Rust matcher test gains the encodings positive + negative cases. 5 pytest passed; the matcher is compiled by CI's Tauri shell check. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
a7ab148483 |
fix(asr): float16-unsupported GPUs fall back to int8 instead of "no segments" (#561)
#551: both CTranslate2 ASR backends request compute_type="float16" on CUDA with NO fallback. On GPUs without efficient fp16 (older Maxwell/Pascal, GTX 16xx) or a CTranslate2/cuDNN binary mismatch, WhisperModel/whisperx.load_model raise a ValueError at construction — which escaped the existing OOM-only `except RuntimeError`, so every chunk failed and the user got "Transcription produced no segments". Add a per-device compute_type fallback chain (cuda: float16 → int8_float16 → int8; cpu: int8 → float32) to both backends + the ASR sidecar, alongside (not replacing) the existing OOM→CPU path, with an ASR_COMPUTE_TYPE override for exotic hardware (documented in README). Also in the same ASR-robustness pass: - #549: PyTorchWhisperBackend._ensure_pipe wraps the transformers pipeline load and re-raises an actionable error (reinstall transformers / use faster-whisper) instead of a bare "Could not import module 'AutoFeatureExtractor'". - #516: the /dub/transcribe SSE generator is wrapped so it can NEVER close without a terminal event — any unanticipated exception now yields a structured `error` (with build_failure's hint) + `done`, turning "stream dropped, likely ASR failed" into the real cause + Retry. - failure.py: COMPUTE_TYPE_UNSUPPORTED + TRANSFORMERS_IMPORT classes so the no-segments toast is actionable. Tests (fail-before/pass-after): float16-unsupported → int8 for both WhisperX + FasterWhisper; a generic non-OOM RuntimeError still raises; classify() maps the two new classes; the SSE stream always terminates with error→done. 7 + 1 passed, 17 in the failure suite (no regression). Co-authored-by: mergetest <test@local> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
af3da58584 |
fix(bootstrap): force-reinstall setuptools so pkg_resources repair actually works (#248) (#495)
The auto-repair ran `uv pip install setuptools>=75,<80`, which `uv` treats as "already satisfied" (no-op, "Checked 1 package in 5ms") whenever setuptools' *metadata* is present but its `pkg_resources` files are gone — the common cause being Windows Defender quarantining `pkg_resources/`, or a partial extract on a restricted network. So the repair never restored the files, the post-check failed, and users hit the #248 dead-end. The error message *also* told them to run the same no-op command, so the suggested manual fix didn't work either (reported on Discord, Win11 + RTX 5070 Ti). Fix: both repair sites in bootstrap.rs now use `--reinstall` (the flag already used for the ROCm torch repair), which force re-extracts pkg_resources even when uv thinks setuptools is satisfied. The fail() message and the failure.py hint now suggest `uv pip install --reinstall 'setuptools>=75,<80'` + an antivirus-exclusion note, and docs/install/troubleshooting.md (#pkg_resources-missing) is updated with the real cause (metadata-present/files-missing) + AV guidance. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
b64f53b0af |
feat: pipeline error transparency — no more silent "unknown error" (plan-04, closes #131) (#136)
* docs(plan-04): spec + plan for pipeline error transparency (#131) speckit spec/plan/research/data-model/contract/quickstart for plan-04. Grounds the fix in the real code map: shared failure-event builder (backend/core/failure.py) feeding tasks.py + dub_pipeline.py + dub_core.py, non-empty reason guarantee, sanitized diagnostic block, frontend renderer with docs deeplink. Closes-target: #131 (children #122, #63). Design only — no code changes yet. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(pipeline): structured, non-empty failure events + logged tracebacks (#131) plan-04 backend: no more silent "unknown error". A shared failure helper guarantees a non-empty reason at every emit site and a sanitized, copyable diagnostic block. - backend/core/failure.py: build_failure()/build_failure_event() (reason falls back to the exception class name), sanitize() (reuses the logging_filter HF-token regex + redacts *TOKEN*/*KEY*/*SECRET* env values + home→~), diagnostic() (reuses the env capture), classify() reusing the error_docs_map 5-class taxonomy for the docs deeplink + hint. - core/tasks.py worker: structured event instead of bare str(e); keeps the logged traceback. - services/dub_pipeline.py: enrich download/extract error yields; ADD the missing outer `except Exception` (the #122 path — unhandled ingest errors were never surfaced with stage context); surface the previously-silent demucs/scene/thumbnail degradations as non-fatal `warning` events. - api/routers/batch.py: guaranteed non-empty batch failure reason. SSE payload is additive (legacy `error`/`stage`/`detail` keys preserved), so existing frontends keep working and already show the specific reason. Tests (TDD, fail-before/pass-after): 14 cases — non-empty-reason guarantee, redaction, diagnostic sanitization, and the 3 Test-matrix triggers (worker / extract / url). 483 passed, 0 regressions. Closes #131. Refs #122, #63. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(dub-ui): show specific cause + docs deeplink + copyable diagnostic (#131) plan-04 frontend. The backend now sends a structured, non-empty failure; surface it to the user instead of "extract: unknown error". - dubSlice: DubFailure type + dubFailure state/setter. - useDubWorkflow: capture the structured failure on the SSE error event (reason/error_class/stage/hint/docs_topic/diagnostic); clear on new runs. - DubTab: DubFailureNotice renders the actionable hint, an "Open docs" deeplink (via the existing errorDocsMap classifier), and a "Copy diagnostic" button — shown beneath the error badge in both failure banners. typecheck + build clean; 66 frontend tests pass. Refs #131, #122, #63. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(failure): annotate intentional best-effort excepts (CodeQL) The new security workflow's CodeQL flagged 5 bare `except: pass` blocks. All are deliberate best-effort guards (sanitize/diagnostic must never throw on the failure path; the test cancels the worker to tear it down). Added explanatory comments per CodeQL's py/empty-except rule. No behavior change. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |