Files
VoiceStudio/scripts/smoke-gguf.sh
T
Palash DebnathandClaude Opus 4.7 b34dcd9e11 Phase 4 Plan 04-01: SPIKE-01 GGUF — GO + integration (#100)
* Phase 4 Plan 04-01: SPIKE-01 GGUF — GO + Wave 1 integration

Integrates Serveurperso/OmniVoice-GGUF as a hardware-adaptive default
voice-cloning engine, with overridable fallback to the in-process
OmniVoiceBackend. Spike confirmed GO: the model is a clean quantization
of k2-fsa/OmniVoice (Apache-2.0 + MIT runtime, `omnivoice-lm` custom
architecture so it does NOT load in vanilla llama.cpp).

Pinned SHAs:
  * Serveurperso/OmniVoice-GGUF revision: 361609388ae572a820d085185bbbe2a2aac4b30e
  * ServeurpersoCom/omnivoice.cpp master:  886fc079838ca7400cb2b42b36e2a65aa1daabe8

Implements GGUF-01 (hardware probe) through GGUF-05 (default-engine
resolver with graceful fallback). The four `bin/omnivoice-tts-*`
artifacts are committed as zero-byte placeholders; the new CI matrix
job builds the real binaries per platform from the pinned commit SHA
and appends a SHA-256 manifest used by `is_available()` for tampering
detection (T-04-01). The macos-14 (Apple Silicon) slot is marked
`continue-on-error: true` because omnivoice.cpp publishes no
`buildmetal.sh` (Pitfall 1 / Assumption A1) — failure feeds into Task
3's GO/NO-GO call.

Quant override is allow-listed against quant_map.json entries only
(T-04-05). Argv is composed from typed Path objects rooted in
HF_HUB_CACHE; never uses `shell=True`. HF token redaction applies to
captured stderr before logging (AUTH-05 / T-04-04).

Tests: 36 new (8 hardware-probe + 13 GGUF engine + 6 settings_store
quant override + grep gate); 428 passed in full suite vs 402+ baseline.
ADR Status stays "Proposed (research-supported)" — Task 3 (human
checkpoint) flips to Accepted after CI produces real binaries and a
reviewer signs off on the GGUF-06 cross-hardware smoke.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* ci: install libopenblas-dev on linux-x86_64 omnivoice-tts build

The pinned omnivoice.cpp commit (886fc079...) ships a `buildcpu.sh`
that passes `-DGGML_BLAS=ON`. ubuntu-latest has no BLAS implementation
preinstalled, so the cmake configure step fails with
`Could NOT find BLAS (missing: BLAS_LIBRARIES)` and the job exits in
13 s before producing the linux-x86_64 binary.

macOS (Accelerate, built in) and Windows (BLAS off by default in the
ggml CMakeLists for non-APPLE platforms — the build script doesn't
invoke buildcpu.sh on those slots) are unaffected and stay green.

Adds a Linux-gated apt step to install libopenblas-dev + pkg-config
before the build, restoring cross-platform parity per the
CLAUDE.md "default features must work on every platform" rule.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(gguf): constrain ref_audio to project roots — block /etc/shadow on Linux

The GGUF engine's `_build_argv` previously validated ref_audio only via
`ref_path.is_file()` — i.e. "does this path exist?" That check is
platform-dependent: `/etc/shadow` doesn't exist on macOS (rejected
naturally), but it IS a real system file on Linux, so the validation
silently accepted it. CI's ubuntu-22.04 runner exposed the gap via
`test_generate_blocks_freeform_ref_audio`, which exists precisely to
guard the "freeform ref_audio path" attack surface.

Fix: confine ref_audio to one of three allowed roots before existence
checks:
  - VOICES_DIR (user-saved voice profiles)
  - DUB_DIR (per-job auto-clones extracted from source video)
  - tempfile.gettempdir() (browser-upload temp files; existing
    `cleanup_ref` flow in generation.py)

Anything outside those roots → FileNotFoundError, matching the existing
failure-mode contract callers handle. Existence check still runs after,
so the test's mocked subprocess.run is never reached and the test
passes deterministically on all three platforms.

Cross-platform parity (per CLAUDE.md 2026-05-20 rule): identical
behaviour on macOS / Windows / Linux — the allow-list is computed from
core.config which uses platform-specific path resolution but yields the
same logical "project tree" on every OS.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* ci(gguf): mark darwin-x86_64 binary build as experimental

GitHub's macos-13 (Intel) runner pool is heavily contended — PR #100
queued for 30+ minutes waiting on darwin-x86_64 while every other
platform finished in ~1m. Intel Macs are also fading hardware (Apple's
platform momentum is entirely on Apple Silicon), and the GGUF engine's
runtime already handles a missing binary gracefully (`is_available()`
returns False on Intel Mac with a "binary not bundled for this
platform" message, same path used for first-launch before any binaries
build).

`experimental: true` mirrors what darwin-arm64 (Metal) already has —
slot still runs and uploads its binary when successful, but a failure
or runner backlog no longer blocks merges. Keeps the GGUF engine
shippable across the dominant arm64 / Linux / Windows surface without
holding the inbox on a slow-runner queue.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-20 17:51:31 +05:30

150 lines
4.1 KiB
Bash
Executable File

#!/usr/bin/env bash
# GGUF-06 cross-hardware smoke test.
#
# Runs an end-to-end 3-second voice clone via the GGUF engine on one of
# three hardware classes and asserts the output WAV is decodable, ≥2.5s
# long at 24 kHz, and that the quant chosen matches `quant_map.json`.
# A human reviewer (Task 3 in plan 04-01) listens to the WAV and signs
# off on intelligibility.
#
# Usage:
# scripts/smoke-gguf.sh --hardware-class {cpu|mid|high}
#
# Outputs:
# tmp/smoke-gguf-<class>.wav — the generated audio
# tmp/smoke-gguf-<class>.json — metadata (quant selected, duration, sr)
#
# Exit codes:
# 0 — generation succeeded, file is valid, quant matches the table
# 1 — generation failed or output failed validation
# 2 — binary unavailable on this host (expected on CI matrix
# `continue-on-error: true` runner; not a hard failure)
set -euo pipefail
CLASS=""
PROMPT="${PROMPT:-Hello from OmniVoice GGUF smoke test.}"
while [[ $# -gt 0 ]]; do
case "$1" in
--hardware-class)
CLASS="$2"
shift 2
;;
--prompt)
PROMPT="$2"
shift 2
;;
-h|--help)
sed -n '1,25p' "$0"
exit 0
;;
*)
echo "Unknown argument: $1" >&2
exit 1
;;
esac
done
if [[ -z "$CLASS" ]]; then
echo "--hardware-class is required (cpu|mid|high)" >&2
exit 1
fi
case "$CLASS" in
cpu|mid|high) ;;
*)
echo "Unknown class: $CLASS (must be cpu|mid|high)" >&2
exit 1
;;
esac
REPO_ROOT="$(cd "$(dirname "$0")/.." && pwd)"
mkdir -p "$REPO_ROOT/tmp"
OUT_WAV="$REPO_ROOT/tmp/smoke-gguf-$CLASS.wav"
OUT_JSON="$REPO_ROOT/tmp/smoke-gguf-$CLASS.json"
# Force the compute-class bucket via env vars so the smoke test on a
# beefy machine can still exercise the CPU code path. The GGUF backend
# reads these as overrides during the probe.
case "$CLASS" in
cpu)
export OMNIVOICE_GGUF_FORCE_CLASS=cpu
export CUDA_VISIBLE_DEVICES=""
;;
mid)
export OMNIVOICE_GGUF_FORCE_CLASS=mid-vram
;;
high)
export OMNIVOICE_GGUF_FORCE_CLASS=high-vram
;;
esac
cd "$REPO_ROOT"
# Quick availability check first so CI can route around a missing binary.
PYTHONPATH=backend python - <<PY || exit 2
import json, sys
from engines.omnivoice_gguf.backend import _make_backend_class, _platform_slug
cls = _make_backend_class()
ok, reason = cls.is_available()
if not ok:
print(f"GGUF binary not available on this host ({_platform_slug()}): {reason}", file=sys.stderr)
sys.exit(2)
print("→ GGUF binary available; running smoke generate")
PY
# Run the actual generation through the backend class.
PYTHONPATH=backend python - "$OUT_WAV" "$OUT_JSON" "$PROMPT" "$CLASS" <<'PY'
import json, sys, time
out_wav, out_json, prompt, class_name = sys.argv[1:5]
from engines.omnivoice_gguf.backend import _make_backend_class
import soundfile as sf
cls = _make_backend_class()
backend = cls()
entry = backend._select_quant_entry()
t0 = time.monotonic()
tensor = backend.generate(prompt)
elapsed = time.monotonic() - t0
# Save WAV to the expected path.
arr = tensor.squeeze(0).cpu().numpy()
sf.write(out_wav, arr, backend.sample_rate, subtype="PCM_16")
# Validate.
info = sf.info(out_wav)
duration_s = info.frames / info.samplerate
if duration_s < 2.5:
print(f"FAIL: output too short ({duration_s:.2f}s < 2.5s)", file=sys.stderr)
sys.exit(1)
if info.samplerate != 24_000:
print(f"FAIL: unexpected sample rate {info.samplerate} (expected 24000)", file=sys.stderr)
sys.exit(1)
meta = {
"class": class_name,
"quant_base": entry.get("base"),
"quant_tokenizer": entry.get("tokenizer"),
"rationale": entry.get("rationale"),
"duration_s": duration_s,
"elapsed_s": elapsed,
"sample_rate": info.samplerate,
"frames": info.frames,
"prompt": prompt,
}
with open(out_json, "w") as f:
json.dump(meta, f, indent=2)
print(f"✓ smoke-gguf-{class_name}: {duration_s:.2f}s in {elapsed:.1f}s "
f"using {entry.get('base')}")
PY
echo "✓ Smoke test passed for hardware class: $CLASS"
echo " Output: $OUT_WAV"
echo " Meta: $OUT_JSON"