Commit Graph
8 Commits
Author SHA1 Message Date
Palash Debnath 27a8f477b7 fix(security): keep private diagnostics out of API responses (#1454)
* fix(security): keep private diagnostics out of API responses

* docs: reference response-safety PR

* fix(security): preserve constant recovery guidance

* fix(security): keep recovery and logs data-independent

* fix(security): close remaining response sinks

* test(security): keep SOCKS diagnostics private

* fix: keep Tailscale exceptions local

* fix: keep Tailscale CLI output private
2026-08-10 12:01:12 +00:00
debpalash 7cc173d8be fix: repair the resolved Hugging Face cache 2026-08-10 07:04:42 +00:00
debpalash fe29fe0966 Merge remote-tracking branch 'origin/main' into fix/ghas-hf-revisions
# Conflicts:
#	CHANGELOG.md
2026-08-10 07:01:42 +00:00
debpalash f6f2bc5dcd fix(security): pin curated Hugging Face revisions 2026-08-09 21:07:01 +00:00
debpalash 05b1d096b4 fix(models): repair a weight file that arrived damaged, not just a missing one (#1406)
An interrupted or mangled model download leaves one of two states, and only
one had any handling. A MISSING shard raises transformers' "does not appear
to have a file named …" and gets a whole recovery ladder. A shard that is
PRESENT with wrong bytes — a download stopped mid-file, a shard truncated by
antivirus, an HTML error page saved under its name — opens fine and then
fails inside safetensors:

    Error while deserializing header: header too large

That reached the user as a raw 500 on every generation, from voice design and
gallery previews alike, and could not enter the ladder for two independent
reasons: the wording is not the missing-shard wording, and SafetensorError is
a Rust-extension exception rather than an OSError.

It also needs the opposite repair. The ladder RESUMES a download, and a resume
trusts a blob that is already the expected size — so it would never re-fetch
the one file that is actually wrong. The new path forces a full re-download,
then retries the load once.

Both halves now classify as MODEL_CACHE_CORRUPT: one class to the user, one
remedy, two repairs underneath. The cause is matched through the whole
exception chain, since transformers wraps the tensor library's error in its
own before it reaches us.
2026-08-08 02:56:08 +05:30
Palash DebnathandClaude Opus 5 c3d0d6b123 feat(update): announce updates as a toast with actions, not a wall of text (#1272)
* fix(dub): purge order was hash-dependent, and the cap could evict a live marker

main went red on my own test. Two distinct defects, both mine.

1. `targets = set(job_ids)` made iteration order depend on PYTHONHASHSEED, so
   which markers a cap-forced trim discarded was luck. That is why the test
   passed locally and failed in CI — verified: the old code passes at seeds
   0/7/42 and fails at 12345. Now a de-duplicated list in caller order.

2. The size cap could evict markers the CURRENT purge had just recorded. Those
   are the newest and the likeliest to still be held by a running job, so
   dropping one is precisely the resurrection this mechanism exists to prevent.
   A 'clear history' larger than the cap forced exactly that. The cap now never
   touches the current purge, making the real bound cap + one purge — stated
   plainly rather than implied.

Tests now pin both: identical survivors across runs, and an oversized purge
keeping all of its own markers. Verified across six hash seeds; full suite green
under 12345, the seed that reddened main.

* feat(update): announce updates as a toast with actions, not a wall of text

A version's release notes ARE the whole changelog section — v0.4.1's was 42
bullets. Any surface that renders them inline becomes unusable: older builds put
them in a blocking OS dialog that filled the screen and had to be dismissed
before the app could be touched.

Removing that dialog left the opposite failure. The only remaining signal was a
6-pixel dot beside the version number in the footer, which is easy to never
notice — so users either got shouted at or told nothing.

A toast is the middle: it names the version, offers Install and restart /
What's new / Later, and leaves. The notes stay one click away in Settings →
Updates, where there is room. Keyed by version so the 6-hourly re-check
replaces rather than stacks, and it never auto-dismisses — an update the user
hasn't answered is still true. Install declines while a generation is running,
since the relaunch would lose it.

Strings added to all 21 locales, translated rather than English-filled.

Tests pin the shape that matters: the toast takes no notes prop at all, renders
under 200 characters, and cannot stack duplicates.

* fix(update): don't relaunch while work is in flight; share the busy check

Review of #1272 found the restart guard was `dubStep === 'generating'` and
nothing else. Installing an update relaunches the process, so that permitted
throwing away a dub upload, a transcription, a translation, an export or a
standalone TTS synth. The same narrow check was written twice — in the new
toast and in UpdatesPanel — so the two could also drift apart.

Replaced with a single `isAppBusy(state)` in utils/appBusy, unioning the
signals the store actually has: the dub state machine, the floating status
pill (which every long background operation already pushes to), and
`ttsGenerating` — a new transient store field mirroring useTTS's local
`isGenerating`, which lived in a hook where no global check could see it.
`dubStep === 'editing'` is deliberately not busy: it waits on the user, and
counting it would block updates for as long as a transcript stays open.

Also from review:

- A failed lazy import of the toast fell into the outer catch, whose
  setUpdateIdle() erased the update that had just been found. The
  announcement is optional; the available state is not.
- "Dismiss" was machine-translated into the employment sense — terminate an
  employee — in de/ja/ru/zh-CN/zh-TW, on close buttons and one aria-label.
  Swept every locale rather than the four lines that were flagged.
- Hoisted the test's mock state with vi.hoisted.

CodeRabbit's changelog finding is declined: Highlights bullets carry no issue
refs by design, which tests/test_changelog_style.py enforces.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(shutdown): a quit mid-generate is not a 500, and not a bug report (#1276)

#1174 made a model load interrupted by shutdown benign for the background
preload, but a *request* that triggered a load took the generic unhandled-
exception path: crash log, ERROR traceback, and an error-journal entry that
feeds the bug-report pipeline. Quitting the app with a generate queued
surfaced "500 Internal Server Error: model load skipped: backend shutting
down" and offered to file a GitHub issue for a normal teardown.

Nothing failed — the process is exiting. The handler now answers 503 with
Retry-After and an actionable detail, ahead of the crash-log/journal writes.

Frontend half of the same bug: toastErrorWithReport offered "Report" for any
error. A 503 means "not now, try again" by definition, so it now shows the
backend's message without the report action — covering a still-warming
backend too, not just this shutdown path.

Fail-before/pass-after tests on both sides, including that the shutdown case
leaves no crash-log entry and no journal record.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(models): repair a half-downloaded model whichever way it reports (#1273)

transformers has two unrelated wordings for "this snapshot has no weight
shard", sharing no words:

  hub load   "<repo> does not appear to have a file named …"
  local dir  "Error no file named model.safetensors, … found in directory …"

The self-heal and the failure classifier both matched only the first. The
second is what a load of a *subfolder* inside a cached snapshot raises —
exactly where an interrupted download leaves a half-written repo — so the
reporter got neither the automatic repair (delete broken entries →
re-download → retry) nor an actionable hint, just a raw 500. Their disk had
10.6 GB free, i.e. a download that ran out of room.

The phrase list now lives once in core.failure, so the healer and the error
text cannot disagree about what an interrupted download looks like. Both
fragments of the second wording must match — "no file named" alone is
ordinary English and must not claim the class.

Also hardens the #1276 handler found by running both suites in one session:
services.model_manager can be imported under two module names, making two
distinct ModelLoadInterruptedByShutdown classes and breaking a bare
isinstance — which silently restored the 500 this fix exists to remove. Now
matched by isinstance OR class name, with a test that raises a same-named
class from a different module.

Verified against the full tests/ + backend/tests/ session: the only remaining
failures are the four pre-existing #1269 isolation leaks, unchanged.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(update): count synths in flight at the chokepoint, not in one caller

Review round 2. Both Greptile P1s were the same weakness in my first pass:
tracking the synth in useTTS meant only the Generate tab counted (voice
previews, the compare modal, the stories editor and profile previews call
generateSpeech directly and were invisible), and a boolean meant two
overlapping syntheses cleared each other — whichever settled first reported
"idle" while the other was still running, so Install and restart discarded it.

Moved to `api/generate.ts`, around the one `/generate` call all seven paths
share, as a count with a `finally` release (so an abort or a network error
frees it too). A future synth caller is covered without opting in.
Fail-before verified: the overlap test fails against the boolean version.

Also from review:

- UpdatesPanel re-reads the busy state at click time; the render-time snapshot
  only exists to disable the button, and work can start after the last render.
- The 503 now carries the allowed-origin CORS headers. Without them the
  browser reports a bare CORS failure and the actionable detail — the whole
  point of the fix — never reaches the user. Both error responses build them
  through one helper now.
- Log `request.url.path`, not the full URL: a query string can carry tokens
  and newlines.
- `update.busy` said "finish your dub first" in all 21 locales, but the guard
  now covers uploads, transcription, translation, export and synth. Rewritten
  as work-in-progress wording, translated per locale.
- ru `common.dismissStatus` was "reject status"; missed in the earlier sweep.

Declined: CodeRabbit's request for a `(#NNN)` ref on the Highlights bullet —
Highlights carry no refs by design (CLAUDE.md), and tests/test_changelog_style.py
enforces it. Also declined gating the cache re-download behind a confirmation:
the self-heal already existed and already ran for the sibling wording, this
only stops it missing half the class, and it stays behind the existing
_hf_offline() check.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(update): grey out Install while work is running

CI lint caught `busy` as unused after the click-time check moved into the
handler — and that exposed a real gap: `busy` had never been wired to
anything. The comment claimed it disabled the button; it didn't. Clicking
Install during a synth just bounced a toast back.

Now it does what it said: the Install and Restart buttons are disabled while
work is in flight, with `update.busy` as the tooltip. The click-time read
stays the authority, since work can start between the last render and the
click — the disabled state is the explanation, not the safety.

Adds the first UpdatesPanel test. The case it pins hardest is the inverse of
the bug: a busy predicate stuck at true would make the app permanently
un-updatable, which is worse than what this fixes. So it asserts enabled when
idle and during dub 'editing' (which waits on the user), disabled during an
upload, a translation, and one or more synths.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-27 14:32:32 -07:00
debpalashandClaude Fable 5 63fd497caf feat: TTS-only first run, platform-curated ASR, guided OS permissions, parakeet-mlx
Only the TTS model (~2.4 GB) is required on first run; ASR models are
per-platform curated picks (curated_on in models.yaml) installed on demand.
Every transcription surface returns a typed asr_model_missing error with a
one-click download CTA instead of silently pulling multi-GB Whisper weights.
Settings -> Models is a grouped, platform-aware catalog. New guided
permissions UX (wizard System Check + Settings -> Permissions + mic
pre-flight) with native mic-state checks and OS settings deep-links. New
parakeet-mlx engine brings Parakeet TDT v3 to Apple Silicon (language-gated
capture preference so multilingual dictation never regresses). Docs:
expressive-speech page, Flush/Unload + CPU-fallback triage, clone-length FAQ.
Hardening: preflight fails open for custom model pins, ROCm curation no
longer inherits NVIDIA picks, Windows mic probe reads the NonPackaged
consent key, CaptureWidget setup race fixed, offline-cache CI simulation
fixes so empty-cache runners stay green.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-17 15:20:54 +05:30
7f59d5f8fe fix(models): self-heal HF-cache snapshots with broken file links, retry load once (#1056)
A first-run breaker: all blobs download fine, but the snapshots/<rev>/
entries are dangling symlinks (0 KB) — os.path.isfile() is False on a
dangling link, so transformers reports the weights missing even though
the bytes are on disk, and the existing resume repair can't fix it.

New services/hf_cache_repair.py deletes exactly the broken snapshot
entries (dangling symlinks + zero-byte weight/config stand-ins; never
blobs, never resolving entries) and restores them via snapshot_download,
verifying afterwards — if the restore recreates broken links (hub's
memoized symlink probe passing while real links come out broken), it
forces hub into copy-mode and repairs once more with real files.
model_manager retries the load exactly once per repo per process
(rung 0 of the cache-recovery ladder); dead-end errors now name the
exact models--<org>--<name> folder to delete. failure.py classifies the
class as MODEL_CACHE_CORRUPT so the user-facing error and auto bug
report explain the automatic repair.

Co-authored-by: mergetest <test@local>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-11 00:35:10 +05:30