main
1
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
36e3397613 |
fix(cuda): stop sending every RTX 40-series card to the CPU (#1289)
* fix(engines): warn about under-provisioned hardware before the synth, not after Six reports are the same story: #1240, #1246, #1248, #1277, #1283, #1284 — 4 GB and 6 GB cards running an engine that wants 6 GB, each one waiting out the full 300s compute budget to be told the job "was too heavy". The routing layer knew the whole time. The error text even names the card and the figure. The caveat only ever surfaced on the engine-PICK toast, so it reached people who changed engines and nobody whose engine was already selected — the default, or one persisted from a previous session. That is most users. /generate does return X-OmniVoice-Routing, but a response header arrives when the job ends, five minutes too late to be a warning. So the check moves to the chokepoint every synth path shares (api/generate.ts, same argument as the in-flight count). Fire-and-forget: never awaited, so it cannot add latency to the request it warns about; never throws, so an unreachable backend costs a warning rather than a generate; once per engine+reason per session, so it informs instead of nagging. Advisory, not blocking — the driver can page to system RAM and short inputs fit where long ones don't. Extracts routingNotice() as the single frontend mirror of the backend's routing_notice(). Two callers now need "is this verdict worth interrupting for", and two inline copies would drift — invisibly, until someone on DirectML or an unavailable engine gets a hardware warning for a normal pick. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(cuda): stop sending every RTX 40-series card to the CPU The SM-arch gate required the device's exact tag in get_arch_list(). NVIDIA's rules are not exact, and PyTorch depends on that: SASS is binary-compatible UPWARD within a major version, so the official wheels ship sm_80/sm_86 and deliberately no sm_89 — the 8.6 kernels already cover Ada. Exact matching therefore declared sm_89 unsupported, check_device_compatibility() returned False, and get_best_device() silently returned "cpu". That is every RTX 4060/4070/4080/4090, not just the reporter's card (#1285) — each one running TTS on the CPU on hardware that works fine, with a message telling them their GPU was unsupported. cuda_build_covers() now applies the real rules: sm_XY covers same-major devices with minor >= Y; compute_XY PTX JITs forward to anything newer; an a/f suffix (sm_90a) is architecture-specific and matches exactly. Unparseable entries are skipped, and an empty arch list still degrades to "compatible" — the pre-existing fail-open contract. The remediation text also pointed at a NIGHTLY index for what is a stable supported card; it now names the stable cu128 index. 12 tests: the Ada regression, Jetson Orin (8.7), downward-within-major and cross-major rejection, PTX forward-JIT, arch-specific suffixes, and a genuine sm_120-on-old-wheel mismatch so the gate is proven to still work. Closes #1285 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(cuda): resolve app modules at call time, not import time The tests/** review contract forbids module-level imports of app modules — they go stale under sys.modules pollution from other suites, which is the live cause of #1269's cross-suite failures. Binds core.device_caps per call. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix: restore generate.ts and generatePreflight.test.js from main Conflict markers were committed in the previous merge — `git add` on the directory staged both files as resolved while the markers were still in them. Both belong to #1288 and are unchanged by this PR, so they take main's version verbatim. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |