Files
VoiceStudio/backend/services/settings_store.py
T
Palash DebnathandClaude Opus 4.7 b34dcd9e11 Phase 4 Plan 04-01: SPIKE-01 GGUF — GO + integration (#100)
* Phase 4 Plan 04-01: SPIKE-01 GGUF — GO + Wave 1 integration

Integrates Serveurperso/OmniVoice-GGUF as a hardware-adaptive default
voice-cloning engine, with overridable fallback to the in-process
OmniVoiceBackend. Spike confirmed GO: the model is a clean quantization
of k2-fsa/OmniVoice (Apache-2.0 + MIT runtime, `omnivoice-lm` custom
architecture so it does NOT load in vanilla llama.cpp).

Pinned SHAs:
  * Serveurperso/OmniVoice-GGUF revision: 361609388ae572a820d085185bbbe2a2aac4b30e
  * ServeurpersoCom/omnivoice.cpp master:  886fc079838ca7400cb2b42b36e2a65aa1daabe8

Implements GGUF-01 (hardware probe) through GGUF-05 (default-engine
resolver with graceful fallback). The four `bin/omnivoice-tts-*`
artifacts are committed as zero-byte placeholders; the new CI matrix
job builds the real binaries per platform from the pinned commit SHA
and appends a SHA-256 manifest used by `is_available()` for tampering
detection (T-04-01). The macos-14 (Apple Silicon) slot is marked
`continue-on-error: true` because omnivoice.cpp publishes no
`buildmetal.sh` (Pitfall 1 / Assumption A1) — failure feeds into Task
3's GO/NO-GO call.

Quant override is allow-listed against quant_map.json entries only
(T-04-05). Argv is composed from typed Path objects rooted in
HF_HUB_CACHE; never uses `shell=True`. HF token redaction applies to
captured stderr before logging (AUTH-05 / T-04-04).

Tests: 36 new (8 hardware-probe + 13 GGUF engine + 6 settings_store
quant override + grep gate); 428 passed in full suite vs 402+ baseline.
ADR Status stays "Proposed (research-supported)" — Task 3 (human
checkpoint) flips to Accepted after CI produces real binaries and a
reviewer signs off on the GGUF-06 cross-hardware smoke.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* ci: install libopenblas-dev on linux-x86_64 omnivoice-tts build

The pinned omnivoice.cpp commit (886fc079...) ships a `buildcpu.sh`
that passes `-DGGML_BLAS=ON`. ubuntu-latest has no BLAS implementation
preinstalled, so the cmake configure step fails with
`Could NOT find BLAS (missing: BLAS_LIBRARIES)` and the job exits in
13 s before producing the linux-x86_64 binary.

macOS (Accelerate, built in) and Windows (BLAS off by default in the
ggml CMakeLists for non-APPLE platforms — the build script doesn't
invoke buildcpu.sh on those slots) are unaffected and stay green.

Adds a Linux-gated apt step to install libopenblas-dev + pkg-config
before the build, restoring cross-platform parity per the
CLAUDE.md "default features must work on every platform" rule.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(gguf): constrain ref_audio to project roots — block /etc/shadow on Linux

The GGUF engine's `_build_argv` previously validated ref_audio only via
`ref_path.is_file()` — i.e. "does this path exist?" That check is
platform-dependent: `/etc/shadow` doesn't exist on macOS (rejected
naturally), but it IS a real system file on Linux, so the validation
silently accepted it. CI's ubuntu-22.04 runner exposed the gap via
`test_generate_blocks_freeform_ref_audio`, which exists precisely to
guard the "freeform ref_audio path" attack surface.

Fix: confine ref_audio to one of three allowed roots before existence
checks:
  - VOICES_DIR (user-saved voice profiles)
  - DUB_DIR (per-job auto-clones extracted from source video)
  - tempfile.gettempdir() (browser-upload temp files; existing
    `cleanup_ref` flow in generation.py)

Anything outside those roots → FileNotFoundError, matching the existing
failure-mode contract callers handle. Existence check still runs after,
so the test's mocked subprocess.run is never reached and the test
passes deterministically on all three platforms.

Cross-platform parity (per CLAUDE.md 2026-05-20 rule): identical
behaviour on macOS / Windows / Linux — the allow-list is computed from
core.config which uses platform-specific path resolution but yields the
same logical "project tree" on every OS.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* ci(gguf): mark darwin-x86_64 binary build as experimental

GitHub's macos-13 (Intel) runner pool is heavily contended — PR #100
queued for 30+ minutes waiting on darwin-x86_64 while every other
platform finished in ~1m. Intel Macs are also fading hardware (Apple's
platform momentum is entirely on Apple Silicon), and the GGUF engine's
runtime already handles a missing binary gracefully (`is_available()`
returns False on Intel Mac with a "binary not bundled for this
platform" message, same path used for first-launch before any binaries
build).

`experimental: true` mirrors what darwin-arm64 (Metal) already has —
slot still runs and uploads its binary when successful, but a failure
or runner backlog no longer blocks merges. Keeps the GGUF engine
shippable across the dominant arm64 / Linux / Windows surface without
holding the inbox on a slow-runner queue.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-20 17:51:31 +05:30

325 lines
12 KiB
Python

"""Encrypted settings store for OmniVoice — AUTH-02 + threat T-01-01.
Persists small key/value rows in the SQLite `settings` table. The `hf_token`
row is encrypted at rest via Fernet (symmetric AEAD). The Fernet key itself
is derived per-install in `_secret_key.derive_fernet_key()` so a copied
omnivoice.db is not a smoking gun on another machine.
Public API (consumed by `token_resolver.py`):
get_hf_token() -> Optional[str]
set_hf_token(token: str) -> None
clear_hf_token() -> None
"""
from __future__ import annotations
import logging
import time
from typing import Optional
logger = logging.getLogger("omnivoice.settings_store")
_TOKEN_KEY = "hf_token"
def _fernet():
"""Build a Fernet instance from the derived per-install key.
Lazy import of cryptography keeps the module importable even if the
dep ever falls out of the install (graceful warn instead of ImportError
at module load).
"""
from cryptography.fernet import Fernet
from services._secret_key import derive_fernet_key
return Fernet(derive_fernet_key())
def get_hf_token() -> Optional[str]:
"""Return the decrypted HF token from the settings table, or None.
On `InvalidToken` (e.g. user migrated omnivoice_data/ across machines so
machine-id no longer matches the salt used to encrypt), log a warning
and return None — the caller's cascade (env / HF-CLI) takes over
naturally. Per Open Question #5 in 01-RESEARCH.md.
"""
from core.db import db_conn
try:
with db_conn() as conn:
row = conn.execute(
"SELECT value FROM settings WHERE key = ?", (_TOKEN_KEY,)
).fetchone()
if row is None or not row[0]:
return None
blob = row[0]
try:
from cryptography.fernet import InvalidToken
except ImportError: # pragma: no cover — dep should always be present
logger.error("cryptography unavailable; cannot decrypt HF token")
return None
try:
return _fernet().decrypt(blob.encode("ascii")).decode("utf-8")
except InvalidToken:
logger.warning(
"Stored HF token failed to decrypt — most likely the install "
"moved across machines or the salt row was tampered with. "
"Falling back to env/HF-CLI sources."
)
return None
except Exception:
# SQLite errors must not break the resolver — return None and let
# the resolver cascade.
logger.exception("settings_store.get_hf_token: SQLite read failed")
return None
def set_hf_token(token: str) -> None:
"""Persist an encrypted HF token. The first call also writes the per-install
salt row inside the same transaction (atomic, no torn state)."""
if not token:
# Defense in depth — callers should pass a non-empty string. An empty
# token in the cascade is the same as "no token" and we should not
# write a row that round-trips to the empty string.
clear_hf_token()
return
from core.db import db_conn
# Derive the key first — this lazily generates the salt row on first
# write, inside its own transaction. Subsequent ops use the cached key.
cipher = _fernet()
blob = cipher.encrypt(token.encode("utf-8")).decode("ascii")
with db_conn() as conn:
conn.execute(
"INSERT OR REPLACE INTO settings(key, value, updated_at) "
"VALUES (?, ?, ?)",
(_TOKEN_KEY, blob, time.time()),
)
def clear_hf_token() -> None:
"""Remove the HF token row. The salt row is preserved so a re-save by
the same user on the same machine produces ciphertext the resolver can
still decrypt."""
from core.db import db_conn
with db_conn() as conn:
conn.execute("DELETE FROM settings WHERE key = ?", (_TOKEN_KEY,))
# ── Non-secret text settings ──────────────────────────────────────────────
# Plan 01-02 Task 4 (INST-12): the Performance panel needs to persist a
# boolean toggle (`perf.torch_compile_disabled`). It is NOT a secret — no
# user-recoverable harm comes from a leaked "user disabled torch.compile"
# bit — so we store the raw text directly in the same `settings` table
# without Fernet wrap.
#
# Use these helpers (not `set_hf_token`) for non-secret config:
# set_text("perf.torch_compile_disabled", "1")
# get_text("perf.torch_compile_disabled", default="0")
def get_text(key: str, default: Optional[str] = None) -> Optional[str]:
"""Read a non-encrypted text value from the settings table.
Returns `default` if the row is missing OR if reading the row fails.
The HF-token row is encrypted ciphertext and will round-trip here
looking like opaque bytes — callers MUST use `get_hf_token()` for
secrets and only ever pass non-secret keys to `get_text()`.
"""
if key == _TOKEN_KEY: # defence in depth — never let a misrouted call leak ciphertext
return default
from core.db import db_conn
try:
with db_conn() as conn:
row = conn.execute(
"SELECT value FROM settings WHERE key = ?", (key,)
).fetchone()
if row is None or row[0] is None:
return default
return str(row[0])
except Exception:
logger.exception("settings_store.get_text(%s): SQLite read failed", key)
return default
def set_text(key: str, value: str) -> None:
"""Persist a non-encrypted text value into the settings table.
Use for non-secret config only. For tokens, use `set_hf_token()`.
"""
if key == _TOKEN_KEY:
raise ValueError(
"set_text refuses to write to the encrypted hf_token row; "
"use set_hf_token() for secrets"
)
from core.db import db_conn
with db_conn() as conn:
conn.execute(
"INSERT OR REPLACE INTO settings(key, value, updated_at) "
"VALUES (?, ?, ?)",
(key, value, time.time()),
)
# ── Phase 4 Plan 04-01 (GGUF-04): per-engine quant override ────────────────
#
# Settings row "gguf_quant_override" holds either:
# * "auto" — explicit auto-select (default, equivalent to row absent)
# * a quant filename listed in quant_map.json (e.g. "omnivoice-base-F32.gguf")
# Any other value is rejected at write time (T-04-05 — the quant override UI
# cannot be used to load an attacker-controlled GGUF path).
_QUANT_OVERRIDE_KEY = "gguf_quant_override"
def get_quant_override() -> Optional[str]:
"""Return the user's GGUF quant override, or None when auto-select.
Returns:
``None`` — no row present, or row holds the sentinel ``"auto"``.
``str`` — a quant filename from ``quant_map.json``'s allow-list.
Reads through the same disk-backed SQLite path as ``get_text``, so
the override survives process restart (per GGUF-04 acceptance).
"""
raw = get_text(_QUANT_OVERRIDE_KEY)
if raw is None or raw == "auto":
return None
return raw
def set_quant_override(value: Optional[str]) -> None:
"""Persist a GGUF quant override (or clear it on ``None`` / ``"auto"``).
Args:
value: ``None`` or ``"auto"`` to clear; otherwise a quant filename
from ``quant_map.json``'s allow-list
(e.g. ``"omnivoice-base-F32.gguf"``).
Raises:
ValueError: ``value`` is a string but is not in the allow-list,
or it contains a path separator / ``..``. This protects
against UI input being used to load a quant from an
attacker-controlled path (T-04-05 in the Plan 04-01 threat
model).
"""
from core.db import db_conn
if value is None or value == "auto":
# Clear the row entirely — get_quant_override returns None for
# both "row absent" and the explicit "auto" sentinel; we
# canonicalize on "absent" so the table stays clean.
with db_conn() as conn:
conn.execute(
"DELETE FROM settings WHERE key = ?",
(_QUANT_OVERRIDE_KEY,),
)
return
if not isinstance(value, str):
raise ValueError(
f"set_quant_override expects a str or None, got {type(value).__name__}"
)
# Defence-in-depth: never accept anything that could escape the HF
# cache. The allow-list check below would catch this too, but
# failing fast on path separators surfaces the bug closer to the
# caller.
if "/" in value or "\\" in value or ".." in value:
raise ValueError(
f"set_quant_override refuses path components in {value!r}"
)
# Allow-list against quant_map.json. Importing here keeps the engine
# package off settings_store's hot import path for non-GGUF callers.
try:
from engines.omnivoice_gguf.backend import _allowed_quant_filenames
except Exception as exc: # pragma: no cover - engine package always present
raise ValueError(
f"GGUF engine package unavailable; cannot validate quant override: {exc}"
) from exc
allowed = _allowed_quant_filenames()
if value not in allowed:
raise ValueError(
f"set_quant_override rejects {value!r}: not in quant_map.json "
f"allow-list. Allowed values: {sorted(allowed)}"
)
with db_conn() as conn:
conn.execute(
"INSERT OR REPLACE INTO settings(key, value, updated_at) "
"VALUES (?, ?, ?)",
(_QUANT_OVERRIDE_KEY, value, time.time()),
)
# ── License acceptance helpers (Phase 3 Plan 03-01 / TTS-05) ──────────────
# Tiny wrappers around the plaintext ``set_text``/``get_text`` helpers so
# every engine that needs an acceptance gate (Supertonic-3 today; future
# OpenRAIL-M / non-commercial engines) reads + writes the same key shape:
# ``"<engine_id>_license_accepted"`` -> ``"1"`` | ``"0"``.
#
# Pitfall 6 in 03-RESEARCH.md: ``set_license_accepted`` MUST block until
# the SQLite commit returns. ``db_conn()`` is a context manager that
# commits on ``__exit__`` (Phase 1 contract verified at 01-RESEARCH.md
# settings_store section), so ``set_text`` already satisfies the
# "readable on next call" invariant. We re-read inside this helper as a
# belt-and-braces verification path that tests can rely on.
_LICENSE_KEY_SUFFIX = "_license_accepted"
def _license_key(engine_id: str) -> str:
"""Map an engine id to its license-flag settings key.
Public so tests can assert on the exact stored row. The HF-token
row's name ``hf_token`` will never collide because we always append
the suffix.
"""
if not engine_id or not isinstance(engine_id, str):
raise ValueError(f"engine_id must be a non-empty string, got {engine_id!r}")
# Defence in depth: disallow the literal hf_token key so a misrouted
# call can never overwrite the encrypted-token row.
if engine_id == _TOKEN_KEY:
raise ValueError("engine_id cannot be the reserved HF-token key")
return f"{engine_id}{_LICENSE_KEY_SUFFIX}"
def get_license_accepted(engine_id: str) -> bool:
"""Return True iff the user has accepted the engine's license terms.
Reads from the plaintext ``settings`` table. Missing row, SQLite
read failure, or any non-``"1"`` value all return False so the
callsite (``Supertonic3Backend.is_available()``) defaults to safe.
"""
key = _license_key(engine_id)
raw = get_text(key, default="0")
return raw == "1"
def set_license_accepted(engine_id: str, accepted: bool) -> None:
"""Persist the acceptance flag. Blocks until SQLite commits.
``accepted=False`` writes ``"0"`` (not a delete) so a once-accepted
user who later revokes acceptance still has an explicit row in the
audit-trail-friendly settings table.
"""
key = _license_key(engine_id)
set_text(key, "1" if accepted else "0")
# Re-read invariant ‑‑ Pitfall 6 defence. A failure here means the
# commit silently dropped, which would be a SQLite/db_conn bug; we
# surface it as a hard error rather than hand back a stale state.
actual = get_text(key, default="0")
expected = "1" if accepted else "0"
if actual != expected:
raise RuntimeError(
f"set_license_accepted({engine_id!r}, {accepted!r}) did not "
f"persist (read back {actual!r}, expected {expected!r})"
)