* Phase 4 Plan 04-01: SPIKE-01 GGUF — GO + Wave 1 integration Integrates Serveurperso/OmniVoice-GGUF as a hardware-adaptive default voice-cloning engine, with overridable fallback to the in-process OmniVoiceBackend. Spike confirmed GO: the model is a clean quantization of k2-fsa/OmniVoice (Apache-2.0 + MIT runtime, `omnivoice-lm` custom architecture so it does NOT load in vanilla llama.cpp). Pinned SHAs: * Serveurperso/OmniVoice-GGUF revision: 361609388ae572a820d085185bbbe2a2aac4b30e * ServeurpersoCom/omnivoice.cpp master: 886fc079838ca7400cb2b42b36e2a65aa1daabe8 Implements GGUF-01 (hardware probe) through GGUF-05 (default-engine resolver with graceful fallback). The four `bin/omnivoice-tts-*` artifacts are committed as zero-byte placeholders; the new CI matrix job builds the real binaries per platform from the pinned commit SHA and appends a SHA-256 manifest used by `is_available()` for tampering detection (T-04-01). The macos-14 (Apple Silicon) slot is marked `continue-on-error: true` because omnivoice.cpp publishes no `buildmetal.sh` (Pitfall 1 / Assumption A1) — failure feeds into Task 3's GO/NO-GO call. Quant override is allow-listed against quant_map.json entries only (T-04-05). Argv is composed from typed Path objects rooted in HF_HUB_CACHE; never uses `shell=True`. HF token redaction applies to captured stderr before logging (AUTH-05 / T-04-04). Tests: 36 new (8 hardware-probe + 13 GGUF engine + 6 settings_store quant override + grep gate); 428 passed in full suite vs 402+ baseline. ADR Status stays "Proposed (research-supported)" — Task 3 (human checkpoint) flips to Accepted after CI produces real binaries and a reviewer signs off on the GGUF-06 cross-hardware smoke. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * ci: install libopenblas-dev on linux-x86_64 omnivoice-tts build The pinned omnivoice.cpp commit (886fc079...) ships a `buildcpu.sh` that passes `-DGGML_BLAS=ON`. ubuntu-latest has no BLAS implementation preinstalled, so the cmake configure step fails with `Could NOT find BLAS (missing: BLAS_LIBRARIES)` and the job exits in 13 s before producing the linux-x86_64 binary. macOS (Accelerate, built in) and Windows (BLAS off by default in the ggml CMakeLists for non-APPLE platforms — the build script doesn't invoke buildcpu.sh on those slots) are unaffected and stay green. Adds a Linux-gated apt step to install libopenblas-dev + pkg-config before the build, restoring cross-platform parity per the CLAUDE.md "default features must work on every platform" rule. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(gguf): constrain ref_audio to project roots — block /etc/shadow on Linux The GGUF engine's `_build_argv` previously validated ref_audio only via `ref_path.is_file()` — i.e. "does this path exist?" That check is platform-dependent: `/etc/shadow` doesn't exist on macOS (rejected naturally), but it IS a real system file on Linux, so the validation silently accepted it. CI's ubuntu-22.04 runner exposed the gap via `test_generate_blocks_freeform_ref_audio`, which exists precisely to guard the "freeform ref_audio path" attack surface. Fix: confine ref_audio to one of three allowed roots before existence checks: - VOICES_DIR (user-saved voice profiles) - DUB_DIR (per-job auto-clones extracted from source video) - tempfile.gettempdir() (browser-upload temp files; existing `cleanup_ref` flow in generation.py) Anything outside those roots → FileNotFoundError, matching the existing failure-mode contract callers handle. Existence check still runs after, so the test's mocked subprocess.run is never reached and the test passes deterministically on all three platforms. Cross-platform parity (per CLAUDE.md 2026-05-20 rule): identical behaviour on macOS / Windows / Linux — the allow-list is computed from core.config which uses platform-specific path resolution but yields the same logical "project tree" on every OS. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * ci(gguf): mark darwin-x86_64 binary build as experimental GitHub's macos-13 (Intel) runner pool is heavily contended — PR #100 queued for 30+ minutes waiting on darwin-x86_64 while every other platform finished in ~1m. Intel Macs are also fading hardware (Apple's platform momentum is entirely on Apple Silicon), and the GGUF engine's runtime already handles a missing binary gracefully (`is_available()` returns False on Intel Mac with a "binary not bundled for this platform" message, same path used for first-launch before any binaries build). `experimental: true` mirrors what darwin-arm64 (Metal) already has — slot still runs and uploads its binary when successful, but a failure or runner backlog no longer blocks merges. Keeps the GGUF engine shippable across the dominant arm64 / Linux / Windows surface without holding the inbox on a slow-runner queue. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
186 lines
6.0 KiB
Python
186 lines
6.0 KiB
Python
"""
|
|
GPU crash sandbox — subprocess isolation for GPU-intensive operations.
|
|
|
|
Wraps TTS generation in a subprocess so a GPU crash (CUDA OOM, MPS fault,
|
|
driver segfault) kills the worker process but NOT the main backend server.
|
|
The parent process catches the crash and returns a 503 with a clear error
|
|
instead of the entire application dying.
|
|
|
|
Usage:
|
|
from services.gpu_sandbox import sandboxed_generate
|
|
|
|
result = await sandboxed_generate(
|
|
text="Hello world",
|
|
profile_id="voice_123",
|
|
timeout=60,
|
|
)
|
|
# result is a dict with either {"audio_path": ...} or {"error": ...}
|
|
|
|
Architecture:
|
|
Main Process ──fork──► Worker Process (GPU ops)
|
|
◄─pipe── {"audio_path": "/tmp/xxx.wav"} or {"error": "..."}
|
|
|
|
If the worker dies (segfault, OOM), the pipe closes and the main
|
|
process returns a clean error response.
|
|
"""
|
|
from __future__ import annotations
|
|
|
|
import asyncio
|
|
import json
|
|
import logging
|
|
import multiprocessing
|
|
import os
|
|
import sys
|
|
import tempfile
|
|
import time
|
|
|
|
logger = logging.getLogger("omnivoice.sandbox")
|
|
|
|
|
|
def _worker(conn, request: dict):
|
|
"""Run in a subprocess — does the actual GPU work."""
|
|
try:
|
|
# Prevent CUDA from inheriting contexts from parent
|
|
os.environ.setdefault("CUDA_DEVICE_ORDER", "PCI_BUS_ID")
|
|
|
|
import torch
|
|
import torchaudio
|
|
|
|
# Add backend to path
|
|
backend_dir = os.path.join(os.path.dirname(__file__), "..")
|
|
if backend_dir not in sys.path:
|
|
sys.path.insert(0, backend_dir)
|
|
|
|
from services.model_manager import _load_model_sync
|
|
from services.audio_dsp import apply_mastering, normalize_audio
|
|
|
|
model = _load_model_sync()
|
|
|
|
# Build generation kwargs
|
|
gen_kw = {
|
|
"text": request["text"],
|
|
"language": request.get("language"),
|
|
"ref_audio": request.get("ref_audio"),
|
|
"ref_text": request.get("ref_text"),
|
|
"instruct": request.get("instruct"),
|
|
"num_step": request.get("num_step", 16),
|
|
"speed": request.get("speed", 1.0),
|
|
"guidance_scale": request.get("guidance_scale", 2.0),
|
|
}
|
|
|
|
audios = model.generate(**gen_kw)
|
|
audio_out = audios[0]
|
|
|
|
sr = getattr(model, "sampling_rate", 24000)
|
|
mastered = apply_mastering(audio_out, sample_rate=sr)
|
|
final = normalize_audio(mastered, target_dBFS=-2.0)
|
|
|
|
# Write to temp file and return path
|
|
tmp = tempfile.NamedTemporaryFile(delete=False, suffix=".wav")
|
|
torchaudio.save(tmp.name, final, sr, format="wav")
|
|
tmp.close()
|
|
|
|
conn.send({"audio_path": tmp.name, "sample_rate": sr})
|
|
|
|
except Exception as e:
|
|
import traceback
|
|
conn.send({
|
|
"error": f"{type(e).__name__}: {e}",
|
|
"traceback": traceback.format_exc(),
|
|
})
|
|
finally:
|
|
conn.close()
|
|
|
|
|
|
async def sandboxed_generate(
|
|
text: str,
|
|
timeout: float = 120,
|
|
**gen_kwargs,
|
|
) -> dict:
|
|
"""Run TTS generation in a sandboxed subprocess.
|
|
|
|
Returns:
|
|
{"audio_path": str, "sample_rate": int} on success
|
|
{"error": str} on failure (GPU crash, timeout, etc.)
|
|
"""
|
|
parent_conn, child_conn = multiprocessing.Pipe()
|
|
|
|
request = {"text": text, **gen_kwargs}
|
|
|
|
proc = multiprocessing.Process(
|
|
target=_worker,
|
|
args=(child_conn, request),
|
|
daemon=True,
|
|
)
|
|
proc.start()
|
|
|
|
loop = asyncio.get_running_loop()
|
|
|
|
def _wait():
|
|
proc.join(timeout=timeout)
|
|
if proc.is_alive():
|
|
logger.warning("Sandbox worker timed out after %.0fs — killing", timeout)
|
|
proc.kill()
|
|
proc.join(timeout=5)
|
|
return {"error": f"GPU operation timed out after {timeout}s"}
|
|
|
|
if proc.exitcode != 0:
|
|
# Worker crashed (segfault, CUDA OOM, etc.)
|
|
return {
|
|
"error": f"GPU worker crashed (exit code {proc.exitcode}). "
|
|
f"This usually means a CUDA OOM or driver fault. "
|
|
f"Try reducing num_step or restarting the server."
|
|
}
|
|
|
|
if parent_conn.poll(timeout=1):
|
|
return parent_conn.recv()
|
|
|
|
return {"error": "Worker completed but returned no data"}
|
|
|
|
result = await loop.run_in_executor(None, _wait)
|
|
|
|
# Clean up
|
|
parent_conn.close()
|
|
|
|
if result.get("error"):
|
|
logger.error("Sandbox error: %s", result["error"])
|
|
else:
|
|
logger.info("Sandbox success: %s", result.get("audio_path", "?"))
|
|
|
|
return result
|
|
|
|
|
|
def is_sandbox_available() -> tuple[bool, str]:
|
|
"""Check if sandboxing is feasible on this platform."""
|
|
try:
|
|
method = multiprocessing.get_start_method()
|
|
if method == "fork":
|
|
return True, "fork-based sandbox available"
|
|
elif method == "spawn":
|
|
return True, "spawn-based sandbox available (slower cold start)"
|
|
return True, f"sandbox available (start method: {method})"
|
|
except Exception as e:
|
|
return False, f"multiprocessing not available: {e}"
|
|
|
|
|
|
# ── Phase 4 GGUF-01: hardware-capability probe ─────────────────────────────
|
|
#
|
|
# The probe itself lives in ``engines.omnivoice_gguf.hardware_probe`` (it's
|
|
# tightly coupled to the GGUF engine's quant_map.json) but it is a
|
|
# *backend-tier* responsibility per RESEARCH.md "Architectural
|
|
# Responsibility Map" — anything that needs to know "is this machine
|
|
# CUDA / MPS / CPU and how much VRAM does it have?" should import it from
|
|
# this module so we have a single entry point.
|
|
#
|
|
# Re-exported lazily via ``__getattr__`` so importing ``gpu_sandbox`` for
|
|
# the existing CUDA-crash sandbox doesn't drag in the GGUF engine package
|
|
# (which transitively pulls ``huggingface_hub`` + ``soundfile``).
|
|
|
|
|
|
def __getattr__(name: str): # pragma: no cover - exercised via tests
|
|
if name in ("detect_capabilities", "HardwareCapabilities", "ComputeClass"):
|
|
from engines.omnivoice_gguf import hardware_probe
|
|
|
|
return getattr(hardware_probe, name)
|
|
raise AttributeError(f"module {__name__!r} has no attribute {name!r}")
|