* test(probe): add spec-driven AI-agent test harness (L1/L2/L4/L5 + report + triage) Introduces `tests/probe/`, a portable, mostly-deterministic test harness built on the Actor/Judge split: AI agents may drive and self-heal, but verdicts are always deterministic code + metrics — no LLM on the verdict path. Layers: - L1 API: Schemathesis property-fuzz over in-process ASGI (enable-on-demand). - L2 web: Playwright Driver + deterministic self-heal (id→test-id→text, loosened CSS) → pluggable Healer; LLMHealer/anthropic_healer for genuine agentic heal. Judges + self-heal logic unit-tested offline via FakePage; live browser skips. - L4 media: audio correctness — exists/decode/duration/not-silent/clipping/NaN, round-trip ASR WER (pure-python, faster-whisper backend), speaker similarity. No golden-WAV (device-stable metrics only); naturalness is advisory-only. - L5 env/first-run: fresh-data-dir backend boot in a SUBPROCESS (no session contamination), asserts health + DB init + endpoint reachability. Docker gated. Plus: hybrid YAML spec engine + JudgeResult/registry; self-contained HTML report that auto-opens (suppressed in CI/headless/PROBE_NO_OPEN); Triager that clusters failures and drafts a prefilled GitHub issue URL (sanitized, no auto-submit) with a one-click button in the report. Dependency-light: runs in the base venv; schemathesis/resemblyzer/playwright/ anthropic are enable-on-demand and skip cleanly. Generated reports gitignored. Full suite green (657 passed); no contamination of existing tests. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * test(probe): add L3 desktop layer (Tauri config-integrity + guarded launch) Per the architecture decision, desktop E2E is substituted by backend-over-HTTP (L5) + browser (L2) since Tauri has no official macOS WebDriver. L3 guards the packaging/shell contract a browser test can't see, against the real tauri.conf.json (with platform-override merge), running on any platform with no Tauri toolchain: - version parity between tauri.conf.json and pyproject (release integrity) - dev/build wiring (devUrl matches the Vite frontend, frontendDist, before* cmds) - bundled binaries first-run depends on (uv / ffmpeg / ffprobe in externalBin) - CSP actually permits the local backend origins (desktop-only failure mode: packaged app can't reach :3900 while the browser build works) Adds desktop.py (config load + platform deep-merge + bundle discovery + launch guard), judges/desktop.py (config_present/config_eq/config_contains/csp_allows), desktop_smoke.probe.yaml, and tests covering integrity, platform-merge replace semantics, and a live bundle launch that skips without a built bundle/display. Full suite green (662 passed). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
64 lines
2.5 KiB
Python
64 lines
2.5 KiB
Python
"""L3 desktop — config-integrity against the REAL tauri.conf.json (runs for real
|
|
on any platform, no Tauri toolchain), platform-merge behaviour, and a guarded
|
|
live bundle launch that skips without a built bundle / display.
|
|
"""
|
|
|
|
from __future__ import annotations
|
|
|
|
import os
|
|
|
|
import pytest
|
|
|
|
from . import desktop
|
|
from . import spec as probe_spec
|
|
from .judges import desktop as dj
|
|
|
|
_SPEC = os.path.join(os.path.dirname(__file__), "specs", "desktop_smoke.probe.yaml")
|
|
|
|
|
|
def test_desktop_config_integrity(probe_report):
|
|
spec = probe_spec.load_spec(_SPEC)
|
|
context = desktop.desktop_context()
|
|
results = probe_spec.run_judges(spec, context)
|
|
probe_report.record(spec, results)
|
|
assert probe_spec.blocking_failures(results) == [], "\n".join(str(r) for r in results)
|
|
|
|
|
|
def test_version_parity_is_actually_checked():
|
|
# Guard against a vacuous pass: the config version really equals pyproject's.
|
|
cfg = desktop.load_tauri_config()
|
|
assert cfg["version"] == desktop.pyproject_version() != ""
|
|
|
|
|
|
def test_platform_override_replaces_targets():
|
|
# Tauri replaces (not merges) array fields from the platform override.
|
|
linux = desktop.load_tauri_config("linux")
|
|
assert "rpm" in linux["bundle"]["targets"] # only present in the linux override
|
|
windows = desktop.load_tauri_config("windows")
|
|
assert set(windows["bundle"]["targets"]) == {"nsis", "msi"}
|
|
# Non-overridden fields survive the merge.
|
|
assert linux["build"]["devUrl"] == "http://localhost:3901"
|
|
|
|
|
|
def test_csp_judge_catches_blocked_backend():
|
|
good = {"app": {"security": {"csp": "connect-src 'self' http://localhost:* http://127.0.0.1:*;"}}}
|
|
assert dj.csp_allows(good, ["http://localhost:*", "http://127.0.0.1:*"]).passed is True
|
|
bad = {"app": {"security": {"csp": "connect-src 'self';"}}}
|
|
assert dj.csp_allows(bad, ["http://localhost:*"]).passed is False
|
|
|
|
|
|
def test_config_contains_judge():
|
|
cfg = {"bundle": {"externalBin": ["binaries/uv", "binaries/ffmpeg"]}}
|
|
assert dj.config_contains(cfg, "bundle.externalBin", ["binaries/uv"]).passed is True
|
|
assert dj.config_contains(cfg, "bundle.externalBin", ["binaries/ffprobe"]).passed is False
|
|
assert dj.config_contains(cfg, "bundle.missing", ["x"]).passed is False # not a list → FAIL
|
|
|
|
|
|
def test_live_bundle_launch_skips_cleanly():
|
|
ok, reason = desktop.can_launch()
|
|
if not ok:
|
|
pytest.skip(f"L3 live launch unavailable: {reason}")
|
|
# A bundle exists and we have a display — smoke that it launches.
|
|
# (Reached only on a machine with a built desktop bundle.)
|
|
assert os.path.exists(reason)
|