* fix(tts): torch.compile failures fall back to eager — generation never fails on unsupported GPUs (#278)
On GPU architectures the bundled Triton doesn't support (e.g. Blackwell
sm_120 / RTX 5060), the compiled model dies mid-generation inside the
Dynamo/Inductor/Triton/cudagraph stack — previously surfaced as a fake
'ran out of memory' error and a dead Archetype preview. Now:
- up-front arch gate: skip compile when the GPU's compute capability is
not in this torch build's arch list (OMNIVOICE_FORCE_TORCH_COMPILE=1
overrides for PTX forward-compat setups)
- runtime fallback: model.generate is wrapped once; a compile-stack
failure (classified by exception chain: module, message, traceback
paths — the cudagraph case is a bare AssertionError) logs a warning,
restores the eager module, disables compile for the session, resets
dynamo state, and retries eagerly. Non-compile errors propagate
unchanged.
- the /generate OOM handler no longer mislabels compile crashes as OOM
and points users at the actual remedy.
Fixes#278
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* Potential fix for pull request finding 'CodeQL / Empty except'
Co-authored-by: Copilot Autofix powered by AI <62310815+github-advanced-security[bot]@users.noreply.github.com>
* Potential fix for pull request finding 'CodeQL / Empty except'
Co-authored-by: Copilot Autofix powered by AI <62310815+github-advanced-security[bot]@users.noreply.github.com>
* Update backend/api/routers/generation.py
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Copilot Autofix powered by AI <62310815+github-advanced-security[bot]@users.noreply.github.com>
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
Co-authored-by: mergetest <test@local>