Files
VoiceStudio/backend/services/token_resolver.py
T
Palash Debnath 4a6b978df9 Phase 1 Wave 1: HF token persistence + redactor (closes #35) (#91)
* feat(01-01): encrypted settings store + alembic migration (AUTH-02, T-01-01)

Adds the SQLite-backed encrypted settings store that Phase 1 token resolver
will read from. Closes the at-rest plaintext risk for HF tokens (T-01-01).

- backend/services/settings_store.py: get_hf_token / set_hf_token /
  clear_hf_token using Fernet symmetric AEAD. Stored value column never
  contains the literal "hf_" substring.
- backend/services/_secret_key.py: per-install Fernet key derived via
  scrypt(machine-id + 16-byte random salt). machine-id resolution covers
  macOS (ioreg IOPlatformUUID), Linux (/etc/machine-id and dbus fallback),
  Windows (HKLM Cryptography MachineGuid via winreg). Final fallback to
  hostname+user with a warn log.
- backend/migrations/versions/0001_phase1_settings_table.py: alembic
  migration adding `settings(key, value, updated_at)`. Idempotent — checks
  for an existing table so fresh installs (where _BASE_SCHEMA already
  created it) and v0.2.7 upgrades both succeed.
- backend/core/db.py: _BASE_SCHEMA grows the settings table for fresh
  installs; init_db() now runs `alembic upgrade head` after the CREATE.
- backend/migrations/env.py: honours an externally-set sqlalchemy.url so
  tests can point alembic at a fixture DB; falls back to core.config
  DB_PATH for production.
- pyproject.toml: cryptography>=41 added explicitly (RESEARCH.md
  Assumption A1 was checked at execute-time and proved false; the dep was
  not present transitively, so the install would fail without this).

Tests (10 cases, all green):
- Round-trip encryption + plaintext-leakage check (T-01-01 invariant)
- Salt persistence across clear/set cycles
- InvalidToken decrypt path returns None (Open Question #5 resolution)
- Concurrent reads consistent under sqlite WAL
- Alembic upgrade on a hand-built v0.2.7 fixture DB preserves all
  existing tables + seeded rows (CLAUDE.md backward-compat constraint)
- Alembic downgrade -1 drops only the settings table

Refs #35.

* feat(01-01): 3-source HF token resolver + log redactor + 5 read sites patched

Closes the #35 bug class (bare os.environ.get('HF_TOKEN') reads) by routing
every backend HF-token consumer through one resolver, and mitigates
T-01-02 (info disclosure via logs) by stripping `hf_[A-Za-z0-9]{30,}`
substrings from every log record at the root logger.

backend/services/token_resolver.py:
  - resolve(skip)   — 3-source cascade (App → Env → HF-CLI), each source
    validated via huggingface_hub.whoami(); first valid wins.
  - on_401(active) — invalidate cache and re-resolve skipping the source
    that just 401'd (AUTH-06).
  - state()        — three SourceState rows for the Settings UI: set,
    masked preview (hf_…<last 3>), whoami_user, whoami_ok.
  - save_app_token / clear_app_token — wraps settings_store + calls
    huggingface_hub.login(add_to_git_credential=False) per Pitfall #2.
  - 300-second whoami cache so repeated Settings-page renders don't hit
    the HF API.

backend/core/logging_filter.py:
  - HFTokenRedactor(logging.Filter) — regex `hf_[A-Za-z0-9]{30,}` so real
    tokens are masked but `hf_hub` / `hf_token` literals survive.
  - install_redaction_filter() — idempotent attach to root + every handler.

backend/main.py: install the redactor at startup, BEFORE the file
handler is added. Re-installed after the file handler attaches so the
handler-attached filter list includes it too.

Read-side call sites patched (per Pitfall #1 — every HF token read must
flow through token_resolver.resolve()):
  - backend/api/routers/dub_core.py:540  (the original #35 site)
  - backend/api/routers/system.py:38     (_has_hf_token notification)
  - backend/services/model_manager.py:480 (diarization pipeline auth)
  - backend/services/sonitranslate.py:143 (Popen env for SoniTranslate child)
  - backend/services/sonitranslate.py:217 (gradio_client predict call)

New endpoint:
  - GET /system/hf-token/state — returns the 3-source cascade state with
    masked tokens for the Wave 2 Settings UI panel.

Grep gate confirmed clean: zero `os.environ.get("HF_TOKEN")` reads remain
outside token_resolver.py.

Tests (17 new cases, all green):
  - tests/backend/services/test_token_resolver.py: priority cascade, 401
    skip mid-resolve, on_401 fallback, state() shape, save+login
    invariant (add_to_git_credential=False), HUGGING_FACE_HUB_TOKEN
    alias acceptance.
  - tests/backend/core/test_logging_filter.py: msg + args redaction,
    multi-token redaction, non-string args pass-through, short-token
    literals preserved, install_redaction_filter idempotence.

Refs #35.

* feat(01-01): Settings hf-token API endpoints + subprocess env injection (AUTH-03/04)

Backend half of the Wave 2 Settings → API Keys UI plus the AUTH-04
subprocess env-injection invariant.

backend/api/routers/settings.py:
  - POST /api/settings/hf-token       — body {token: str} → save_app_token
  - DELETE /api/settings/hf-token     — also_clear_hf_cli query → clear_app_token
  - GET /api/settings/hf-token/state  — same shape as token_resolver.state()
  All three are gated by `Depends(require_loopback)` at the router level
  (threat T-01-03 mitigation; non-loopback Host → 403).

backend/main.py: router mounted alongside existing API routers.

Subprocess env injection (AUTH-04, threat T-01-04 disposition=accept):
  - backend/services/sonitranslate.py already updated in Task 2 to read
    via token_resolver.resolve() and inject HF_TOKEN + YOUR_HF_TOKEN into
    the SoniTranslate child env block.
  - backend/services/gpu_sandbox.py: NOT patched — the GPU sandbox runs
    in-process TTS generation that uses the parent's already-loaded HF
    state. Adding env injection there is a no-op (parent and child share
    state via multiprocessing.Pipe before any HF API call).
  - backend/services/model_manager.py:480 (Task 2): resolves in-process,
    no subprocess crosses here.
  - backend/api/routers/exports.py: subprocess.Popen calls only spawn
    `open` / `explorer` / `xdg-open` — file-manager launchers with no
    HF needs. Skipped per Task 3 conservative-patching rule.

So the canonical AUTH-04 site for this milestone is sonitranslate.py.
Future SubprocessBackend work in Phase 2 will inherit the same pattern.

Tests (8 new cases, all green):
  - tests/backend/test_engine_spawn_token.py
    * POST /hf-token loopback → 200 + state.active == "app"
    * POST /hf-token non-loopback → 403 ("loopback origin required")
    * DELETE /hf-token clears settings_store + state.active == None
    * GET /hf-token/state returns 3 source rows in priority order
    * GET /hf-token/state non-loopback → 403
    * env block contains HF_TOKEN + YOUR_HF_TOKEN when resolver returns one
    * env block does NOT contain an injected empty HF_TOKEN when resolver
      returns None
    * source-level check that backend/services/sonitranslate.py still
      reads via token_resolver.resolve() (regression guard against
      silent reverts of the AUTH-04 wiring)

Full Wave 1 test suite: 35/35 green. Phase 0 smoke tests still green.

Refs #35.

* docs(01-01): SUMMARY + STATE update for Phase 1 Wave 1 completion

Records execution outcome of the 3-task plan: 10 files created, 9 modified,
35 new test cases, 5 read sites patched, grep gate clean. Documents the
two Rule-3/Rule-2 deviations applied (cryptography dep, env.py URL
override), the subprocess-launcher inventory for Phase 2, and the
known stray edit to the main repo's pyproject.toml that needs a one-
line user action to revert.

Updates STATE.md current-position table, progress bar, and open TODOs to
point at Wave 2 (Plan 01-02) and Wave 3 (Plan 01-03) as the next steps.
2026-05-20 05:10:37 +05:30

263 lines
9.2 KiB
Python

"""3-source HF token resolver — AUTH-01, AUTH-03, AUTH-06.
Resolution priority (highest → lowest):
1. app — `settings_store.get_hf_token()` (encrypted in SQLite)
2. env — `HF_TOKEN` or the legacy `HUGGING_FACE_HUB_TOKEN` env var
3. hf-cli — `huggingface_hub.get_token()` (canonical ~/.cache/huggingface/token)
For each candidate, the resolver calls `huggingface_hub.whoami(token=...)`
to verify the token is live; any HTTP error (401, 403, network) skips to
the next source. Results are cached per (source, token-sha256) for 300
seconds so repeat reads from the UI/dub_core don't hammer the HF API.
Replaces every bare `os.environ.get("HF_TOKEN")` call site in the backend
(per Pitfall #1 in 01-RESEARCH.md and the grep gate in 01-01-PLAN.md
Task 2 verification).
"""
from __future__ import annotations
import hashlib
import logging
import os
import threading
import time
from dataclasses import dataclass
from typing import Literal, Optional
logger = logging.getLogger("omnivoice.token_resolver")
Source = Literal["app", "env", "hf-cli"]
_PRIORITY: tuple[Source, ...] = ("app", "env", "hf-cli")
_CACHE_TTL_SECONDS = 300.0 # See Open Question #4 — UI "Test now" calls invalidate.
@dataclass(frozen=True)
class ResolvedToken:
token: str
source: Source
username: Optional[str]
@dataclass(frozen=True)
class SourceState:
source: Source
set: bool
masked: Optional[str]
whoami_user: Optional[str]
whoami_ok: bool
# ── module-level cache ────────────────────────────────────────────────────
_VALIDATION_CACHE: dict[tuple[Source, str], tuple[float, Optional[str]]] = {}
_CACHE_LOCK = threading.Lock()
def invalidate_cache() -> None:
"""Drop the whoami validation cache. Called by the Settings UI "Test now"
button (Plan 01-02) and by save/clear API endpoints (Task 3)."""
with _CACHE_LOCK:
_VALIDATION_CACHE.clear()
def _hash(token: str) -> str:
return hashlib.sha256(token.encode("utf-8")).hexdigest()
# ── source readers ────────────────────────────────────────────────────────
def _read_app() -> Optional[str]:
try:
from services import settings_store
return settings_store.get_hf_token()
except Exception:
logger.exception("settings_store read failed")
return None
def _read_env() -> Optional[str]:
# HF docs explicitly accept either name; user may have either exported.
val = os.environ.get("HF_TOKEN") or os.environ.get("HUGGING_FACE_HUB_TOKEN")
return val or None
def _read_hf_cli() -> Optional[str]:
try:
import huggingface_hub
tok = huggingface_hub.get_token()
return tok or None
except Exception:
logger.exception("huggingface_hub.get_token failed")
return None
_READERS: dict[Source, callable] = { # type: ignore[type-arg]
"app": _read_app,
"env": _read_env,
"hf-cli": _read_hf_cli,
}
# ── whoami validation ─────────────────────────────────────────────────────
def _validate(source: Source, token: str) -> Optional[str]:
"""Returns the validated whoami username, or None if the token is invalid.
Caches results for `_CACHE_TTL_SECONDS` per (source, token-hash) so the
Settings panel's repeated state() calls don't hit the HF API every load.
"""
key = (source, _hash(token))
now = time.monotonic()
with _CACHE_LOCK:
cached = _VALIDATION_CACHE.get(key)
if cached is not None:
ts, username = cached
if now - ts < _CACHE_TTL_SECONDS:
return username
import huggingface_hub
try:
info = huggingface_hub.whoami(token=token)
name = (info or {}).get("name") if isinstance(info, dict) else None
with _CACHE_LOCK:
_VALIDATION_CACHE[key] = (now, name)
return name
except Exception as exc:
# Any failure — HfHubHTTPError 401/403, network — disqualifies this source.
# Cache the negative result so we don't slam the API in tight loops;
# the cache TTL is bounded so transient failures still recover.
with _CACHE_LOCK:
_VALIDATION_CACHE[key] = (now, None)
logger.debug("whoami failed for source=%s: %s", source, exc)
return None
def _mask(token: str) -> str:
"""`hf_…<last 3>` — what the Settings UI shows in the "currently set"
field. We never reveal the full token in any read API."""
if not token:
return ""
tail = token[-3:] if len(token) >= 3 else token
return f"hf_…{tail}"
# ── public API ────────────────────────────────────────────────────────────
def resolve(skip: frozenset[Source] = frozenset()) -> Optional[ResolvedToken]:
"""Return the highest-priority valid token, or None if all sources are
empty/invalid. `skip` excludes specific sources — used by `on_401()`
when a previously-resolved token started returning 401 mid-job."""
for source in _PRIORITY:
if source in skip:
continue
token = _READERS[source]()
if not token:
continue
username = _validate(source, token)
if username is None and not _all_validation_skipped():
# Token present but whoami failed — log once at debug and try
# the next source. We do NOT log the token (the redactor would
# mask it anyway, but no need to even emit it).
continue
return ResolvedToken(token=token, source=source, username=username)
return None
def _all_validation_skipped() -> bool:
"""Hook left here as a no-op for now. Originally intended to allow
network-disabled environments to bypass whoami; left in for future
extension and explicit so reviewers see the choice."""
return False
def on_401(active_source: Source) -> Optional[ResolvedToken]:
"""AUTH-06: when the active source started returning 401 mid-job (e.g.
the user rotated the token externally), invalidate the cache and try
resolving again while skipping the offending source."""
invalidate_cache()
return resolve(skip=frozenset({active_source}))
def state() -> dict:
"""Return one SourceState per priority position so the Settings UI can
render the cascade table. Includes a masked token + whoami result;
never includes the raw token."""
rows: list[SourceState] = []
active: Optional[Source] = None
for source in _PRIORITY:
token = _READERS[source]()
if token:
username = _validate(source, token)
ok = username is not None
rows.append(SourceState(
source=source,
set=True,
masked=_mask(token),
whoami_user=username,
whoami_ok=ok,
))
if active is None and ok:
active = source
else:
rows.append(SourceState(
source=source,
set=False,
masked=None,
whoami_user=None,
whoami_ok=False,
))
return {"sources": rows, "active": active}
def save_app_token(token: str) -> None:
"""Persist token to the encrypted settings store AND populate the HF
canonical file via `huggingface_hub.login()`. Per Pitfall #2:
`add_to_git_credential=False` is non-negotiable — the alternative
silently writes the token to the user's global git credential helper,
which is leaks-galore for a desktop app."""
if not token:
clear_app_token()
return
from services import settings_store
settings_store.set_hf_token(token)
try:
import huggingface_hub
huggingface_hub.login(
token=token,
add_to_git_credential=False,
new_session=False,
)
except TypeError:
# Older huggingface_hub may not have new_session kwarg — retry
# without it. The add_to_git_credential=False kwarg is the
# invariant that matters; new_session is just a perf tweak.
try:
import huggingface_hub
huggingface_hub.login(token=token, add_to_git_credential=False)
except Exception:
logger.exception("huggingface_hub.login failed (non-fatal)")
except Exception:
# Hub login failure must not strand the user — the token is still
# in the encrypted store and the resolver will pick it up.
logger.exception("huggingface_hub.login failed (non-fatal)")
invalidate_cache()
def clear_app_token(also_clear_hf_cli: bool = False) -> None:
"""Remove from the encrypted settings store; optionally also call
`huggingface_hub.logout()` to clear the canonical HF file."""
from services import settings_store
settings_store.clear_hf_token()
if also_clear_hf_cli:
try:
import huggingface_hub
huggingface_hub.logout()
except Exception:
logger.exception("huggingface_hub.logout failed (non-fatal)")
invalidate_cache()