d91beef0fd314250d8d9b94de86dfea019a8bd96
6
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
6715551c2a |
fix(routing): #1226 review — scope the VRAM floor to where it was measured
Two P1s, both correct: - MPS mislabel. HostCaps.vram_gb on MPS is a heuristic (system RAM / 2) for a UNIFIED memory pool, so an 8 GB Mac reports 4.0 "VRAM" — comparing that to a floor measured on discrete CUDA hardware would warn every small Mac about an engine that runs fine there. The caveat is now dedicated-VRAM families only (cuda/rocm); MPS has a different memory model and no measured floor. - Engine-agnostic timeout. _timeout_guidance serves EVERY job on the GPU pool (reference transcribe, stream assemble, watermarking, dub steps, CPU-only engines on a GPU host), and a hardcoded 6 GB threshold applied without knowing whose job it is would confidently misdiagnose most of them. The floor is now passed in via run_on_gpu_pool_guarded, defaulting to 0 — so the under-provisioned wording is opt-in and only the TTS generate dispatches opt in. A test asserts every "TTS generate" dispatch passes it, so the branch can't become unreachable in production. CHANGELOG entries reworded to end with their refs (CodeRabbit). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
6cfef5c0cf |
fix(routing): warn about an under-provisioned GPU before the job, not after (#1226, #1222)
Two users on 4 GB cards (GTX 1650 Ti, Quadro P2000) ran the `omnivoice` engine, waited out the full compute budget, and were told the job "was too heavy for the available compute … most often the GPU is VRAM-starved". The 300s-vs-372s spread between the two reports is purely text length (`300 + (len-1200)/40`, so 372s ⇒ ~4080 chars) — one bug, not two. Nothing about the budget is device-aware, and nothing needs to be: the real defect is that until the moment it failed, routing showed a clean green "accelerated". `resolve_routing` matched on GPU *family* only, so a 4 GB card and a 24 GB card were indistinguishable, and no engine declared a VRAM requirement anywhere in the repo. - `TTSBackend.min_vram_gb` — advisory metadata alongside `gpu_compat`. Only `omnivoice` declares one (6 GB), derived from the pool's own measured per-job budget (`_GPU_VRAM_PER_JOB_GB = 5.0`) plus resident weights. Inventing floors for engines with no measured figure would put confident numbers in the UI that nothing backs. - `resolve_routing` takes the floor and emits an accelerated-with-caveat reason when the host is below it. Reuses the existing caveat channel, so the Settings matrix and the synth-time routing notice surface it with no UI change. Advisory, never blocking: drivers page to system RAM, and short inputs fit where long ones don't. Kernel-risk still outranks it, and a failed VRAM probe (0.0) never guesses. - `_timeout_guidance` names the actual card and its VRAM, and leads with "pick a lighter engine" instead of wording that reads as transient contention the user can flush their way out of. Regression test: tests/test_low_vram_advisory.py (8 of 12 fail before), including that the 300/372 spread really is just text length. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
b825d99337 |
fix(engines): revive dead license dialog, refresh matrix on select, surface routing verdict (#905)
Six live-audit fixes for the Engines settings surface: - P1-A: the Supertonic license dialog was dead since #101 — `useState` threw away the state value (`const [, setLicenseDialogFor]`) and the imported dialog was never mounted, so "Accept license" did nothing. Keep the value and render LICENSE_DIALOGS[selected] with open/onClose/ onAccepted (accept → matrix reload). - P1-B: the matrix went stale after "Use" — active badge, Use buttons and family-tab captions stayed old until a manual Refresh. Await onSelect, then reload() so the picked engine reflects immediately. - P2-A: consume the /engines/select routing echo. A `cpu_fallback` pick now shows a warn-tone toast naming the reason ("running on CPU — …"); the plain success toast stays for accelerated/cpu_only. Shared helper used by both Settings→Engines and the first-run WizardLibrary. - P2-B: a CPU-native engine (gpu_compat == ("cpu",)) has nothing to fall back FROM, yet on a GPU/MPS host it was mis-classed cpu_fallback (warn). New routing rule classifies ("cpu",) as cpu_only (neutral) on any accelerator host; multi-target engines that could accelerate elsewhere are untouched. - P3-A: the routing reason was only a badge `title` (unreachable on keyboard/touch) — surface it as small visible text under the badge. - P3-B: an in-process "Test engine" pass is an import/liveness check, not a synthesis test — label it "deps OK" instead of a misleading "0 ms" latency; subprocess rows keep their real ping latency. Adds RTL + unit regression tests for all six and updates the routing unit tests to the corrected cpu-native intent. i18n keys added to en.json. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
2a1c3eee3d |
feat(routing): synth-time no-silent-fallback gating at all TTS entry points (#21 follow-up) (#440)
Closes the last #21 gap: a per-request engine=/model= override bypasses the /engines/select host-gate, so an engine that can't use this host's GPU could still be triggered at synth time and silently fall back to CPU (or die mid- synth). Now enforced at every TTS synth entry point, reusing the SAME probe + resolver — never re-deriving routing. Shared helpers (services/engine_routing.py): - `routing_notice(result)` → (status, reason) to surface, or None. Fires for cpu_fallback (always) and accelerated-with-caveat (driver/arch); silent for cpu_only / clean-accelerated / n/a. - `header_safe_reason(reason)` → scrubbed + ASCII-sanitized (headers are latin-1; a non-ASCII device name would 500 otherwise) + ≤256 chars. No regex. Entry points: - REST `POST /generate` (generation.py): after engine resolution, resolve routing once; `unavailable` → 400; cpu_fallback / accelerated-caveat → 200 + `X-OmniVoice-Routing` + `X-OmniVoice-Routing-Reason` headers on the WAV StreamingResponse; benign → no headers. Covers OmniVoice + adapter branches. - OpenAI-compat `POST /v1/audio/speech` (openai_compat.py): same gate + same headers; the tts-1/tts-1-hd alias inherits the active engine's routing. - WebSocket `/ws/tts` (tts_stream.py): no headers → frames. `unavailable` → `{"type":"error",...}` + skip stream; cpu_fallback / caveat → one `{"type":"routing","status","reason"}` frame before any audio. - `select_engine` response now echoes routing_status / effective_device / routing_reason (PR #432 added the gate; this adds the fields so the UI can warn on a cpu_fallback pick). New fields on SelectEngineResponse. Frontend: `useTTS` reads the X-OmniVoice-Routing header and shows a one-time, non-blocking toast (in-memory de-dup by status — a 50-clip batch fires once, no localStorage). i18n keys `tts.routingFallback`/`tts.routingCaveat`. Tests: routing_notice + header_safe_reason (ASCII/length/scrub) unit tests; REST synth gate (unavailable→400, cpu_fallback→headers, cpu_only→none) via the fake-engine harness with a mocked host; select response routing fields. Deferred (small follow-up): dub-pipeline ASR routing note on the preflight_error SSE channel — separate path, not a TTS synth entry point. No frontend /ws/tts client exists today (the routing frame serves external API consumers). Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
8c8d525397 |
feat(routing): wire effective-device into /engines + select gate (#21 PR 3/5) (#432)
* feat(routing): wire effective-device + routing_status into /engines (#21 PR 3/5) Surfaces the PR-1 probe + resolver through the engine registries so the matrix UI (PR 5) and the no-silent-fallback gates can consume it. - `engine_routing.routing_fields()`: shared helper returning the three serialization-ready keys, centralizing the scrub rule — routing_reason is scrubbed via `core.scrub.scrub_text` only when truthy, so a None reason stays JSON `null` (never coerced to ""). - TTS/ASR `list_backends()` each gain `effective_device` / `routing_status` / `routing_reason`, computed from a SINGLE `detect_host_caps()` call per request (host caps are constant per process). ASR is brought to full TTS parity: it now also carries `install_hint` / `last_error` / `isolation_mode` and a SCRUBBED `reason` (closing a pre-existing ASR token-leak gap) — an identical 11-key shape across families. ASR also gains the same is_available()-raises resilience TTS has (degrade to available:false, never 500). - LLM `list_backends()` reaches 11-key parity too but emits literal `effective_device:"network"` / `routing_status:"n/a"` / `routing_reason:null` (NOT via resolve_routing — LLM runs no local GPU model). `LLMBackend.gpu_compat = ()`. "network" is a label, not a probe — nothing here touches the network. - `select_engine` host-routing gate: refuses a pick whose `routing_status` is `unavailable` on this host (400 with an actionable detail), while ALLOWING `cpu_fallback` (it runs, just slower). LLM is never gated. Defensive `.get` so legacy payloads still select. New typed `SelectEngineResponse`. Tests: 11-key shape across all 3 families, well-formed tts/asr routing keys (+ None-not-"" contract), LLM network/n/a labels, select gate (block unavailable / allow cpu_fallback / never-gate LLM). Updated the registry exact-shape test for the 3 new keys. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * test(cjk): allowlist docs/specs/ in the hardcoded-CJK guard PR #429 merged the longform design specs, which legitimately quote functional CJK (test-fixture descriptions, CosyVoice speaker IDs, multilingual sample text). The CJK guard scans every tracked file, so those docs turned main red. Specs are documentation, not shipped UI strings — allowlist the docs/specs/ prefix, matching the individually-allowlisted docs already in the set. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
e61665fe34 |
feat(routing): host device probe + routing resolver (#21 PR 1/5) (#430)
* feat(routing): canonical host device probe + routing resolver (#21 PR 1/5) Foundational, backend-only slice of the GPU compatibility matrix (#21). No API or UI change — wiring lands in PRs 3–5. - `core/device_caps.py`: single source of truth for host accelerator capability. `detect_host_caps()` distinguishes ROCm from CUDA (unlike the gguf hardware_probe), never raises, makes no network call, stays kernel-free on cold start, and caches per process. Enumerates the full degradation contract (torch-unimportable→probe_ok=False, CUDA-init raises, device_count==0, multi-GPU, mem_get_info failure, arch mismatch, MPS, XPU, DirectML). Plus shared `mlx_supported()` gate (#390 groundwork) — exact-string platform check, no regex. - `services/engine_routing.py`: pure `resolve_routing(gpu_compat, caps)` → `{effective_device, routing_status, routing_reason}`; deterministic and byte-identical across OSes. Rules for accelerated / cpu_fallback (the no-silent-fallback signal) / cpu_only / unavailable, incl. the ROCm-not-in-set, DirectML-neutral, and XPU edges. - `get_best_device()` delegates its family decision to the probe so the loader and probe can never disagree; keeps the ROCm HSA env override and DirectML device-string return (probe reads, loader writes). String contract unchanged. - 39 unit tests (probe / resolver / mlx gate / reason-scrub contract); no new regex (CodeQL-clean), English-only (CJK guard green). The gguf hardware_probe rebase is a deliberate follow-up: it has its own torch-mocked suite and a VRAM-driven quant table unaffected by the family rename, so it stays out of this zero-risk slice. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(routing): address review — full available_families + empty-except comments CodeRabbit / CodeQL review on PR 1: - `available_families` no longer drops secondary accelerators on hybrid hosts (e.g. NVIDIA + Intel-iGPU-via-IPEX). The probe now detects every accelerator independently and picks `family` by priority at the end, instead of short-circuiting after the first hit. Routing is unaffected (it keys off `family`), but the field is now honest. + hybrid-host test. - Annotated every `except: pass` in device_caps with an explanatory comment (CodeQL py/empty-except). - Removed the unused `_MIN_NVIDIA_DRIVER` constant — the driver-version check stays in wizard preflight (no subprocess on the probe path); documented why. - `get_best_device()` now checks MPS before DirectML, mirroring the probe's family-priority order so loader and probe never disagree. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |