* test(probe): add spec-driven AI-agent test harness (L1/L2/L4/L5 + report + triage) Introduces `tests/probe/`, a portable, mostly-deterministic test harness built on the Actor/Judge split: AI agents may drive and self-heal, but verdicts are always deterministic code + metrics — no LLM on the verdict path. Layers: - L1 API: Schemathesis property-fuzz over in-process ASGI (enable-on-demand). - L2 web: Playwright Driver + deterministic self-heal (id→test-id→text, loosened CSS) → pluggable Healer; LLMHealer/anthropic_healer for genuine agentic heal. Judges + self-heal logic unit-tested offline via FakePage; live browser skips. - L4 media: audio correctness — exists/decode/duration/not-silent/clipping/NaN, round-trip ASR WER (pure-python, faster-whisper backend), speaker similarity. No golden-WAV (device-stable metrics only); naturalness is advisory-only. - L5 env/first-run: fresh-data-dir backend boot in a SUBPROCESS (no session contamination), asserts health + DB init + endpoint reachability. Docker gated. Plus: hybrid YAML spec engine + JudgeResult/registry; self-contained HTML report that auto-opens (suppressed in CI/headless/PROBE_NO_OPEN); Triager that clusters failures and drafts a prefilled GitHub issue URL (sanitized, no auto-submit) with a one-click button in the report. Dependency-light: runs in the base venv; schemathesis/resemblyzer/playwright/ anthropic are enable-on-demand and skip cleanly. Generated reports gitignored. Full suite green (657 passed); no contamination of existing tests. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * test(probe): add L3 desktop layer (Tauri config-integrity + guarded launch) Per the architecture decision, desktop E2E is substituted by backend-over-HTTP (L5) + browser (L2) since Tauri has no official macOS WebDriver. L3 guards the packaging/shell contract a browser test can't see, against the real tauri.conf.json (with platform-override merge), running on any platform with no Tauri toolchain: - version parity between tauri.conf.json and pyproject (release integrity) - dev/build wiring (devUrl matches the Vite frontend, frontendDist, before* cmds) - bundled binaries first-run depends on (uv / ffmpeg / ffprobe in externalBin) - CSP actually permits the local backend origins (desktop-only failure mode: packaged app can't reach :3900 while the browser build works) Adds desktop.py (config load + platform deep-merge + bundle discovery + launch guard), judges/desktop.py (config_present/config_eq/config_contains/csp_allows), desktop_smoke.probe.yaml, and tests covering integrity, platform-merge replace semantics, and a live bundle launch that skips without a built bundle/display. Full suite green (662 passed). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
69 lines
2.0 KiB
Python
69 lines
2.0 KiB
Python
"""L1/L5 HTTP + filesystem judges — deterministic verdicts on responses and
|
|
on-disk state. No audio, no LLM; just status codes, JSON shape, latency, and
|
|
file existence.
|
|
|
|
These judges take their inputs explicitly from the run context (resolved via
|
|
``$.``) rather than from an audio ``subject``, so they compose into env / API
|
|
specs without colliding with the L4 ``subject`` injection.
|
|
"""
|
|
|
|
from __future__ import annotations
|
|
|
|
import os
|
|
from typing import Any
|
|
|
|
from ..spec import JudgeResult
|
|
|
|
|
|
def status_eq(actual: int, expected: int = 200) -> JudgeResult:
|
|
ok = int(actual) == int(expected)
|
|
return JudgeResult(
|
|
name="status_eq",
|
|
passed=ok,
|
|
measured=actual,
|
|
detail=f"HTTP {actual} (expected {expected})",
|
|
)
|
|
|
|
|
|
def json_has(obj: Any, key: str) -> JudgeResult:
|
|
present = isinstance(obj, dict) and key in obj
|
|
return JudgeResult(
|
|
name="json_has",
|
|
passed=present,
|
|
measured=key,
|
|
detail=f"key {key!r} present" if present else f"key {key!r} missing from response body",
|
|
)
|
|
|
|
|
|
def json_field_eq(obj: Any, key: str, value: Any) -> JudgeResult:
|
|
got = obj.get(key) if isinstance(obj, dict) else None
|
|
ok = got == value
|
|
return JudgeResult(
|
|
name="json_field_eq",
|
|
passed=ok,
|
|
measured=got,
|
|
detail=f"{key}={got!r} (expected {value!r})",
|
|
)
|
|
|
|
|
|
def responds_within_ms(elapsed_ms: float, max: float) -> JudgeResult:
|
|
ok = float(elapsed_ms) <= float(max)
|
|
return JudgeResult(
|
|
name="responds_within_ms",
|
|
passed=ok,
|
|
measured=round(float(elapsed_ms), 1),
|
|
detail=f"{elapsed_ms:.1f} ms (budget {max} ms)",
|
|
)
|
|
|
|
|
|
def path_exists(target: str) -> JudgeResult:
|
|
"""Filesystem existence (file or directory). Named ``target`` (not ``path``)
|
|
so the L4 audio-subject auto-injection never binds to it."""
|
|
ok = bool(target) and os.path.exists(target)
|
|
return JudgeResult(
|
|
name="path_exists",
|
|
passed=ok,
|
|
measured=target,
|
|
detail=f"{target!r} exists" if ok else f"{target!r} does not exist",
|
|
)
|