Explain why dedicated-venv engines (IndexTTS2) add disk (a second torch + CUDA libs: Linux cu128 ~0.83 GiB, Windows ~3.2 GiB), and how uv's link-mode dedup (clone on macOS/Linux, hardlink on Windows) shares identical wheels for free — provided UV_CACHE_DIR and the venvs are on the same filesystem. Key policy: pin the same torch build as the parent whenever the engine allows, since only identical wheels dedupe; UV_LINK_MODE=hardlink on Linux ext4. Linked from the IndexTTS engine doc. Spec §R4(a) / parity program Wave 4.5. Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
3.2 KiB
Engine venvs & disk usage
Most engines run in-process in OmniVoice's main environment. A few
(IndexTTS2, and any engine whose dependencies conflict with the parent's
torch/transformers pins) run in a dedicated sidecar venv so their pins
can't break the rest of the app. Those sidecars are where disk adds up — this
page explains why, and how the on-disk cost is kept down.
Why a sidecar needs its own venv
IndexTTS2 pins transformers<5, but OmniVoice requires transformers>=5.3.
You can't have both in one environment, so IndexTTS2 gets its own venv created
on first use (uv venv + uv pip install, see
backend/engines/indextts/bootstrap.py). The cost is a second copy of the
heavy ML stack — most of which is torch + the bundled CUDA libraries.
How big is "a second torch"?
Measured (2026-06):
| Platform | torch CUDA wheel | + bundled CUDA libs |
|---|---|---|
| Linux (cu128) | ~0.83 GiB | several GiB of nvidia-* packages on top |
| Windows (cu128) | ~3.2 GiB (DLLs bundled in the wheel) | — |
So a sidecar that pins a different torch version than the parent is a multi-GB add. A sidecar that pins the same torch + CUDA build shares almost all of it (see below).
uv dedupes identical wheels — for free, with one condition
uv installs packages by linking from a global wheel cache into each venv's
site-packages. The link mode:
- macOS + Linux —
clone(copy-on-write reflink). N venvs that install the same wheel share the bytes until one is modified — effectively one copy on disk. - Windows —
hardlink. Same effect on a single volume.
The one condition: the cache and the venv must be on the same
filesystem. If UV_CACHE_DIR lives on a different drive than the engine
venvs, uv falls back to a full copy (no dedup, slower). Keep them together.
Dedup is per identical wheel. torch==2.6.0+cu124 and torch==2.8.0+cu128
are different wheels → zero sharing → a full extra multi-GB copy. The single
biggest disk decision for a sidecar is therefore: pin the same torch build as
the parent whenever the engine allows it. When it doesn't (IndexTTS2's
transformers<5 forces an older torch line), the second copy is the
unavoidable price of isolation — not a bug.
On Linux, the
nvidia-*CUDA packages are separate wheels, so even across different torch versions anynvidia-*whose pinned version happens to match is still shared. On Windows the CUDA DLLs live inside the one torch wheel, so nothing is shared across torch versions.
Practical guidance
- Keep
UV_CACHE_DIRand the engine venvs on one filesystem (the default — both under your home dir — already satisfies this). - On Linux ext4 (no reflink),
export UV_LINK_MODE=hardlinkguarantees dedup on any single filesystem; the defaultcloneonly dedupes on reflink-capable filesystems (XFS-with-reflink, btrfs, APFS). - Reclaiming space: deleting a sidecar venv (
backend/engines/<id>/.venv/) frees its unique files; shared cache bytes stay untiluv cache prune. - Forward-looking: PyTorch's experimental wheel variants (shipped in 2.8)
will eventually let
uv install torchauto-pick the right CUDA build, and uv already exposes--torch-backend=auto.