Two users on 4 GB cards (GTX 1650 Ti, Quadro P2000) ran the `omnivoice`
engine, waited out the full compute budget, and were told the job "was too
heavy for the available compute … most often the GPU is VRAM-starved".
The 300s-vs-372s spread between the two reports is purely text length
(`300 + (len-1200)/40`, so 372s ⇒ ~4080 chars) — one bug, not two. Nothing
about the budget is device-aware, and nothing needs to be: the real defect is
that until the moment it failed, routing showed a clean green "accelerated".
`resolve_routing` matched on GPU *family* only, so a 4 GB card and a 24 GB
card were indistinguishable, and no engine declared a VRAM requirement
anywhere in the repo.
- `TTSBackend.min_vram_gb` — advisory metadata alongside `gpu_compat`. Only
`omnivoice` declares one (6 GB), derived from the pool's own measured
per-job budget (`_GPU_VRAM_PER_JOB_GB = 5.0`) plus resident weights.
Inventing floors for engines with no measured figure would put confident
numbers in the UI that nothing backs.
- `resolve_routing` takes the floor and emits an accelerated-with-caveat
reason when the host is below it. Reuses the existing caveat channel, so
the Settings matrix and the synth-time routing notice surface it with no UI
change. Advisory, never blocking: drivers page to system RAM, and short
inputs fit where long ones don't. Kernel-risk still outranks it, and a
failed VRAM probe (0.0) never guesses.
- `_timeout_guidance` names the actual card and its VRAM, and leads with
"pick a lighter engine" instead of wording that reads as transient
contention the user can flush their way out of.
Regression test: tests/test_low_vram_advisory.py (8 of 12 fail before),
including that the 300/372 spread really is just text length.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>