Files
VoiceStudio/tests/probe/judges/desktop.py
T
Palash DebnathandClaude Opus 4.8 24a00bea64 test(probe): spec-driven AI-agent test harness (L1–L5 + HTML report + triage) (#245)
* test(probe): add spec-driven AI-agent test harness (L1/L2/L4/L5 + report + triage)

Introduces `tests/probe/`, a portable, mostly-deterministic test harness built
on the Actor/Judge split: AI agents may drive and self-heal, but verdicts are
always deterministic code + metrics — no LLM on the verdict path.

Layers:
- L1 API: Schemathesis property-fuzz over in-process ASGI (enable-on-demand).
- L2 web: Playwright Driver + deterministic self-heal (id→test-id→text, loosened
  CSS) → pluggable Healer; LLMHealer/anthropic_healer for genuine agentic heal.
  Judges + self-heal logic unit-tested offline via FakePage; live browser skips.
- L4 media: audio correctness — exists/decode/duration/not-silent/clipping/NaN,
  round-trip ASR WER (pure-python, faster-whisper backend), speaker similarity.
  No golden-WAV (device-stable metrics only); naturalness is advisory-only.
- L5 env/first-run: fresh-data-dir backend boot in a SUBPROCESS (no session
  contamination), asserts health + DB init + endpoint reachability. Docker gated.

Plus: hybrid YAML spec engine + JudgeResult/registry; self-contained HTML report
that auto-opens (suppressed in CI/headless/PROBE_NO_OPEN); Triager that clusters
failures and drafts a prefilled GitHub issue URL (sanitized, no auto-submit) with
a one-click button in the report.

Dependency-light: runs in the base venv; schemathesis/resemblyzer/playwright/
anthropic are enable-on-demand and skip cleanly. Generated reports gitignored.
Full suite green (657 passed); no contamination of existing tests.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(probe): add L3 desktop layer (Tauri config-integrity + guarded launch)

Per the architecture decision, desktop E2E is substituted by backend-over-HTTP
(L5) + browser (L2) since Tauri has no official macOS WebDriver. L3 guards the
packaging/shell contract a browser test can't see, against the real
tauri.conf.json (with platform-override merge), running on any platform with no
Tauri toolchain:

- version parity between tauri.conf.json and pyproject (release integrity)
- dev/build wiring (devUrl matches the Vite frontend, frontendDist, before* cmds)
- bundled binaries first-run depends on (uv / ffmpeg / ffprobe in externalBin)
- CSP actually permits the local backend origins (desktop-only failure mode:
  packaged app can't reach :3900 while the browser build works)

Adds desktop.py (config load + platform deep-merge + bundle discovery + launch
guard), judges/desktop.py (config_present/config_eq/config_contains/csp_allows),
desktop_smoke.probe.yaml, and tests covering integrity, platform-merge replace
semantics, and a live bundle launch that skips without a built bundle/display.

Full suite green (662 passed).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-02 12:00:19 +05:30

88 lines
3.1 KiB
Python

"""L3 desktop judges — deterministic checks on the Tauri configuration.
Per the architecture decision, desktop E2E is substituted by backend-over-HTTP +
browser testing (Tauri has no official macOS WebDriver). What remains genuinely
desktop-specific — and what these judges guard — is the *packaging/shell
contract*: things a browser test can't see but that break the shipped app:
- version parity between tauri.conf.json and pyproject (release integrity)
- dev/build wiring (devUrl, frontendDist, before* commands)
- the bundled binaries first-run depends on (uv / ffmpeg / ffprobe)
- the CSP actually permitting the local backend origin (a desktop-only failure
mode — get it wrong and the packaged app can't reach :3900, while the browser
build works fine)
All operate on a parsed config dict from the run context, so they run for real
on every platform with no Tauri toolchain.
"""
from __future__ import annotations
from typing import Any
from ..spec import JudgeResult
_MISSING = object()
def _dig(config: Any, path: str) -> Any:
"""Resolve a dotted key path (``build.devUrl``) into a nested dict."""
cur = config
for key in path.split("."):
if not isinstance(cur, dict) or key not in cur:
return _MISSING
cur = cur[key]
return cur
def config_present(config: dict, path: str) -> JudgeResult:
val = _dig(config, path)
ok = val is not _MISSING and val not in (None, "", [], {})
return JudgeResult(
name="config_present",
passed=ok,
measured=path,
detail=f"{path} present" if ok else f"{path} missing/empty",
)
def config_eq(config: dict, path: str, value: Any) -> JudgeResult:
got = _dig(config, path)
shown = None if got is _MISSING else got
ok = got == value
return JudgeResult(
name="config_eq",
passed=ok,
measured=shown,
detail=f"{path}={shown!r} (expected {value!r})",
)
def config_contains(config: dict, path: str, items: list) -> JudgeResult:
got = _dig(config, path)
if not isinstance(got, (list, tuple)):
return JudgeResult(name="config_contains", passed=False, measured=None,
detail=f"{path} is not a list ({got!r})")
missing = [i for i in items if i not in got]
return JudgeResult(
name="config_contains",
passed=not missing,
measured=list(got),
detail=f"{path} contains {items}" if not missing else f"{path} missing {missing}",
)
def csp_allows(config: dict, origins: list, csp_path: str = "app.security.csp") -> JudgeResult:
"""The packaged app's CSP must permit the local backend origins, or invoke()/
fetch to :3900 is blocked in the bundle (but not in a plain browser)."""
csp = _dig(config, csp_path)
if not isinstance(csp, str):
return JudgeResult(name="csp_allows", passed=False, detail=f"no CSP at {csp_path}")
missing = [o for o in origins if o not in csp]
return JudgeResult(
name="csp_allows",
passed=not missing,
measured=None if not missing else missing,
detail="CSP permits local backend origins" if not missing else f"CSP missing {missing}",
)