Files
VoiceStudio/tests/probe/test_report.py
T
Palash DebnathandClaude Opus 4.8 24a00bea64 test(probe): spec-driven AI-agent test harness (L1–L5 + HTML report + triage) (#245)
* test(probe): add spec-driven AI-agent test harness (L1/L2/L4/L5 + report + triage)

Introduces `tests/probe/`, a portable, mostly-deterministic test harness built
on the Actor/Judge split: AI agents may drive and self-heal, but verdicts are
always deterministic code + metrics — no LLM on the verdict path.

Layers:
- L1 API: Schemathesis property-fuzz over in-process ASGI (enable-on-demand).
- L2 web: Playwright Driver + deterministic self-heal (id→test-id→text, loosened
  CSS) → pluggable Healer; LLMHealer/anthropic_healer for genuine agentic heal.
  Judges + self-heal logic unit-tested offline via FakePage; live browser skips.
- L4 media: audio correctness — exists/decode/duration/not-silent/clipping/NaN,
  round-trip ASR WER (pure-python, faster-whisper backend), speaker similarity.
  No golden-WAV (device-stable metrics only); naturalness is advisory-only.
- L5 env/first-run: fresh-data-dir backend boot in a SUBPROCESS (no session
  contamination), asserts health + DB init + endpoint reachability. Docker gated.

Plus: hybrid YAML spec engine + JudgeResult/registry; self-contained HTML report
that auto-opens (suppressed in CI/headless/PROBE_NO_OPEN); Triager that clusters
failures and drafts a prefilled GitHub issue URL (sanitized, no auto-submit) with
a one-click button in the report.

Dependency-light: runs in the base venv; schemathesis/resemblyzer/playwright/
anthropic are enable-on-demand and skip cleanly. Generated reports gitignored.
Full suite green (657 passed); no contamination of existing tests.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(probe): add L3 desktop layer (Tauri config-integrity + guarded launch)

Per the architecture decision, desktop E2E is substituted by backend-over-HTTP
(L5) + browser (L2) since Tauri has no official macOS WebDriver. L3 guards the
packaging/shell contract a browser test can't see, against the real
tauri.conf.json (with platform-override merge), running on any platform with no
Tauri toolchain:

- version parity between tauri.conf.json and pyproject (release integrity)
- dev/build wiring (devUrl matches the Vite frontend, frontendDist, before* cmds)
- bundled binaries first-run depends on (uv / ffmpeg / ffprobe in externalBin)
- CSP actually permits the local backend origins (desktop-only failure mode:
  packaged app can't reach :3900 while the browser build works)

Adds desktop.py (config load + platform deep-merge + bundle discovery + launch
guard), judges/desktop.py (config_present/config_eq/config_contains/csp_allows),
desktop_smoke.probe.yaml, and tests covering integrity, platform-merge replace
semantics, and a live bundle launch that skips without a built bundle/display.

Full suite green (662 passed).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-02 12:00:19 +05:30

66 lines
2.5 KiB
Python

"""Offline tests for the HTML report system itself (no browser is opened)."""
from __future__ import annotations
from . import report as R
from .spec import JudgeResult
def _sample_report() -> R.Report:
outcome = R.SpecOutcome(
name="tts-synthesis",
feature="tts-synthesis",
layer="media",
duration_s=0.42,
results=[
JudgeResult("artifact_exists", True, "exists (1024 bytes)", measured=1024),
JudgeResult("asr_wer_below", False, "WER=0.40 (max 0.15)", measured=0.40),
JudgeResult("speaker_similarity_above", None, "skipped: no embedder"),
JudgeResult("not_clipping", True, "peak=0.7", measured=0.7, advisory=True),
],
)
return R.Report(outcomes=[outcome])
def test_tallies_and_verdict():
rep = _sample_report()
assert (rep.total, rep.passed, rep.failed, rep.skipped, rep.advisory) == (4, 1, 1, 1, 1)
assert rep.ok is False # one blocking failure
# advisory failures never flip the verdict
rep2 = R.Report(outcomes=[R.SpecOutcome(name="x", results=[
JudgeResult("a", True), JudgeResult("b", False, advisory=True)])])
assert rep2.ok is True
def test_render_html_is_self_contained():
html = R.render_html(_sample_report())
assert html.startswith("<!doctype html>")
assert "<style>" in html and "<script>" in html # inline, no external assets
assert "http://" not in html and "https://" not in html # no remote deps
assert "tts-synthesis" in html
assert ">FAIL<" in html and ">PASS<" in html and ">SKIP<" in html
assert "advisory" in html
def test_html_escapes_detail():
rep = R.Report(outcomes=[R.SpecOutcome(name="x", results=[
JudgeResult("inj", False, "heard '<script>alert(1)</script>'")])])
html = R.render_html(rep)
assert "<script>alert(1)</script>" not in html
assert "&lt;script&gt;" in html
def test_write_creates_files_without_opening(tmp_path):
path = R.save_and_open(_sample_report(), out_dir=tmp_path, open_browser=False)
assert path.exists() and path.read_text(encoding="utf-8").startswith("<!doctype html>")
assert (tmp_path / "report-latest.html").exists()
def test_should_open_respects_env(monkeypatch):
monkeypatch.setenv("PROBE_NO_OPEN", "1")
assert R._should_open(None) is False
monkeypatch.delenv("PROBE_NO_OPEN", raising=False)
monkeypatch.setenv("CI", "true")
assert R._should_open(None) is False
assert R._should_open(True) is True # explicit override wins