On a Tesla T4 the backend exited during the first /generate with no traceback and no HTTP response, leaving the client with RemoteDisconnected and every later call with ConnectionRefused. Three separate defects combined, which is why none of the reporter's workarounds helped. 1. torch.compile(mode="reduce-overhead") captures CUDA graphs. T4 (sm_75) passed the existing arch gate, so capture was attempted and aborted the process from inside the native CUDA library — below the interpreter, where neither the #278 eager-fallback wrapper nor any except clause can see it. The compile mode is now resolved per GPU: Ampere (sm_80) and newer keep the cudagraph mode, older cards drop to the non-cudagraph "default" mode and keep their compiled Inductor kernels. Fails open on any probe error, so no GPU that works today loses the optimization. OMNIVOICE_FORCE_CUDAGRAPH=1 restores it. 2. should_torch_compile() never read TORCH_COMPILE_DISABLE. main.py sets it on win32, build_engine_env injected it into subprocesses, and docs/install/windows.md tells users to export it — but the in-process gate ignored it, so the reporter exported the documented variable and still got "torch.compile applied". The gate now honours TORCH_COMPILE_DISABLE / TORCHDYNAMO_DISABLE / TORCHINDUCTOR_DISABLE on every platform, and an env opt-out on the parent propagates to engine subprocesses. The settings DB path is logged alongside the toggle: the reporter had three omnivoice.db files and edited one the backend never opened. 3. Settings -> Performance -> "Disable torch.compile" was rendered disabled outside Windows in both the Tauri and Electron UIs, so the one control that would have stopped this was unreachable for the affected Linux user. The toggle is now live on every platform, and build_engine_env honours it everywhere rather than only on win32. Also arms faulthandler before torch is imported, so a fatal native signal writes the faulting thread's Python stack to backend_err.log instead of the process vanishing silently. This does not prevent a crash; it makes one diagnosable. OMNIVOICE_DISABLE_FAULTHANDLER=1 skips it. Tests fail before / pass after, verified by stashing the source and running the new tests against unfixed code. The crash test kills a real child interpreter with a real SIGSEGV and requires a named Python frame in the output. test_torch_compile_path_gate's fixture now clears the compile-disable env vars: main.py setdefaults them on win32, so on a Windows runner they leaked into os.environ and decided those tests. Not verified on real hardware — no Turing GPU available. The sm_80 floor is inferred from the crash report and from docs/hardware-notes-tesla-t4.md, which already flagged cudagraphs on T4 as attempted by default and never evaluated. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
63 lines
4.2 KiB
Markdown
63 lines
4.2 KiB
Markdown
# Verified Tesla T4 (16GB) inference notes
|
|
|
|
Measured on a real NVIDIA Tesla T4 (16GB, Turing/sm_75), driver 550.163.01 (CUDA 12.8), torch
|
|
2.8.0+cu128, transformers 5.3.0, Python 3.11.15 (uv-managed). Engine under test: the default
|
|
`omnivoice` TTS backend (`OMNIVOICE_TTS_BACKEND=omnivoice`).
|
|
|
|
## Cold-cache first call can time out at 300s
|
|
|
|
The first `generate()` call lazily downloads the ~2.3GB `k2-fsa/OmniVoice` checkpoint, and that
|
|
download happens *inside* the `OMNIVOICE_GENERATE_TIMEOUT_S` budget (default 300s). On a fresh
|
|
install, the very first `POST /v1/audio/speech` can fail like this even though the GPU isn't
|
|
actually short on memory:
|
|
|
|
```
|
|
ERROR [omnivoice.openai_compat] OpenAI TTS failed: OpenAI TTS generate exceeded 300s and was
|
|
abandoned — the backend is running, but the job was too heavy for the available compute.
|
|
... most often the GPU is VRAM-starved ...
|
|
```
|
|
|
|
VRAM sampling during the failure showed a flat ~2GB with 0% GPU utilization for the whole 300s —
|
|
consistent with waiting on a download, not compute. Once the checkpoint is cached, the identical
|
|
request succeeds in ~1s (reproduced 5x: 1.574s / 1.034s / 1.065s / 0.995s / 0.911s).
|
|
|
|
**Workaround (no code change needed, both already exist):**
|
|
- For headless/API-only setups, pre-fetch the checkpoint before your first real TTS request:
|
|
```bash
|
|
curl -X POST http://localhost:3900/models/install \
|
|
-H "Content-Type: application/json" \
|
|
-d '{"repo_id": "k2-fsa/OmniVoice"}'
|
|
```
|
|
(`repo_id` is required — `InstallModelRequest` in `backend/api/schemas.py` rejects a bare/empty
|
|
body — and must match one of the entries in `KNOWN_MODELS`, e.g. the default engine's
|
|
`k2-fsa/OmniVoice`.) Progress streams over the existing `/setup/download-stream` SSE feed.
|
|
- Or raise the compute-time budget in **Settings → Performance & Device** for the first
|
|
request (`OMNIVOICE_GENERATE_TIMEOUT_S` does the same thing from the environment, and
|
|
takes precedence over the setting when both are present).
|
|
|
|
## OpenAI-compatible endpoint doesn't expose `num_step` / `guidance_scale`
|
|
|
|
`POST /v1/audio/speech`'s request schema doesn't declare `num_step` or `guidance_scale` fields —
|
|
sending them in the JSON body returns `200 OK` but they're silently discarded (pydantic's default
|
|
`extra=ignore` behavior). The native multipart `POST /generate` endpoint *does* expose both as
|
|
explicit form fields, so use that endpoint if you need to control them.
|
|
|
|
Separately: the app's own default for `num_step` is 16 — half of the model's documented default of
|
|
32 (see `docs/generation-parameters.md`, "Use 16 for faster inference"). Not a bug, just not stated
|
|
that the app already runs the "fast" preset unless you override it via `/generate`.
|
|
|
|
## T4 acceleration checklist
|
|
|
|
| Option | Status |
|
|
|---|---|
|
|
| dtype | `torch.float16` hardcoded for the `omnivoice` engine (`model_manager.py`) — correct for Turing (no bf16 tensor cores this generation). No env var override for this engine specifically (ASR engines have `ASR_COMPUTE_TYPE`; `dots_tts`/`indextts` have their own precision vars; `omnivoice` doesn't). |
|
|
| Attention | `sdpa`, selected automatically since `flash_attn` isn't installed (`_supports_flash_attn_2=True` is declared but the package itself is absent) — safe on T4. |
|
|
| int8 | No int8 path for this engine (ASR's CTranslate2 `int8` and `sherpa-onnx`'s int8 ONNX models are separate/unrelated). |
|
|
| CUDA Graphs | **Not used on T4 any more (#2135).** Reachable only indirectly via `torch.compile(mode="reduce-overhead")`, which the app used to attempt by default here — and which killed the backend process outright on the first `/generate` (no traceback, no HTTP response). The app now picks the compile mode per GPU and drops to the non-cudagraph `default` mode below sm_80. `OMNIVOICE_FORCE_CUDAGRAPH=1` restores the old behaviour for benchmarking. |
|
|
| torch.compile | Still attempted on T4, in `default` mode — compiled Inductor kernels, no graph capture. Disable entirely with Settings → Performance → "Disable torch.compile" or `TORCH_COMPILE_DISABLE=1`. |
|
|
|
|
## VRAM
|
|
|
|
Peak measured: 2487 MiB (`nvidia-smi`) / 2.050 GB (`torch.cuda.max_memory_allocated()`) for the
|
|
default `omnivoice` engine — comfortably fits even the README's stated "minimum" (4GB) tier.
|