Files
ComfyUI/tests-unit/assets_test/services/test_split_policy.py
T
19e1058f4c feat(assets): split asset records from content (#16295)
* review-stack 1/4: code (37 files, +3217/-3958)

Review-and-land stack for synap5e/feat/asset-record-content-split, generated by review-stack.py. Once approved,
merges DOWN into the layer below (a fast-forward); only the bottom layer
squash-merges into the real base. See ~/adocs/review-stack.md.
Rule: path not under tests-unit/ or tests/
Question: Is the logic change right?
Source tip: 7007d18582
Merge-base: 783545f689

* review-stack 2/4: tests-removed (24 files, +274/-8220)

Review-and-land stack for synap5e/feat/asset-record-content-split, generated by review-stack.py. Once approved,
merges DOWN into the layer below (a fast-forward); only the bottom layer
squash-merges into the real base. See ~/adocs/review-stack.md.
Rule: test file deleted, or modified with deleted/(added+deleted) >= 0.9
Question: For each dropped assertion: obsolete by a ruling, or covered by a tests-new test?
Source tip: 7007d18582
Merge-base: 783545f689

* review-stack 3/4: tests-changed (13 files, +1043/-1218)

Review-and-land stack for synap5e/feat/asset-record-content-split, generated by review-stack.py. Once approved,
merges DOWN into the layer below (a fast-forward); only the bottom layer
squash-merges into the real base. See ~/adocs/review-stack.md.
Rule: remaining modified test files (incl. conftest.py / helpers)
Question: Did the edits weaken an existing check?
Source tip: 7007d18582
Merge-base: 783545f689

* review-stack 4/4: tests-new (46 files, +8601/-0)

Review-and-land stack for synap5e/feat/asset-record-content-split, generated by review-stack.py. Once approved,
merges DOWN into the layer below (a fast-forward); only the bottom layer
squash-merges into the real base. See ~/adocs/review-stack.md.
Rule: test file added
Question: Is the code layer well covered?
Source tip: 7007d18582
Merge-base: 783545f689

* review-stack 5/6: code (13 files, +351/-104)

Review-and-land stack for synap5e/feat/assets-di, generated by review-stack.py. Once approved,
merges DOWN into the layer below (a fast-forward); only the bottom layer
squash-merges into the real base. See ~/adocs/review-stack.md.
Rule: path not under tests-unit/ or tests/
Question: Is the logic change right?
Source tip: eca2c74bff
Merge-base: 20d59d2a5f

* review-stack 6/6: tests (8 files, +753/-238)

Review-and-land stack for synap5e/feat/assets-di, generated by review-stack.py. Once approved,
merges DOWN into the layer below (a fast-forward); only the bottom layer
squash-merges into the real base. See ~/adocs/review-stack.md.
Rule: every changed file under tests-unit/ or tests/ (added, modified, or deleted)
Question: Is the code layer well covered, and did any edit weaken an existing check?
Source tip: eca2c74bff
Merge-base: 20d59d2a5f

* review-stack 7/8: ported-fixes (42 files, +1361/-180)

Review-and-land stack for synap5e/feat/assets-di-v2, generated by
review-stack.py conventions (hand-built continuation layer; see the PR body).
Once approved, merges DOWN into the layer below (a fast-forward); only the
bottom layer squash-merges into the real base. See ~/adocs/review-stack.md.
Rule: the 11 base-branch fix/docs commits 595cd6e4..94d7185b cherry-picked across the DI refactor (7efdd1d7 excluded, superseded by layer 8)
Question: was each base fix ported faithfully across the DI refactor?
Source tip: 6841881069284803b902b4a9e33bdcda13126771
Merge-base: 7fdfb40f4b

* review-stack 8/8: defensive-parity (4 files, +36/-3)

Review-and-land stack for synap5e/feat/assets-di-v2, generated by
review-stack.py conventions (hand-built continuation layer; see the PR body).
Once approved, merges DOWN into the layer below (a fast-forward); only the
bottom layer squash-merges into the real base. See ~/adocs/review-stack.md.
Rule: match-or-improve master's dependency defenses — NoAssets selection when DB deps unavailable (7efdd1d7's outcome via the DI seam), requirements warning before assets imports, blake3 in the guarded dependency set
Question: does each degradation path now match or improve master's behavior?
Source tip: ebc2cfeebc
Merge-base: 7fdfb40f4b

* fix(assets): only discard content rows this operation actually inserted

CR-9: Enumerated all six create_content call sites. Only scanner seeding and the three ingest registration paths track IDs for failure cleanup.

* fix(assets): reject hash-only uploads with FEATURE_DISABLED when hashing is off

CodeRabbit finding CR-2: reject hash-only multipart uploads before create_from_hash when hashing is disabled.

* fix(assets): seed persists the stat it verified

CR-7: persist the fresh seed-time restat instead of walk-time spec values.

* fix(assets): route database lock failures to the lock guidance

CR-16: route file-lock startup failures through the existing lock guidance and exit path.

* fix(assets): drop the inaccurate temp-cleanup claim from the shutdown warning

References CR-10.

* fix(assets): walk the output root after execution so undeclared outputs register promptly

Custom nodes that write files into the output directory without declaring
them in output_ui only became assets when the next full walk happened - a
frontend GET /object_info or a restart. Headless and API-only sessions never
trigger either, so those files never converged into the asset database.

The post-execution hook now requests a FULL scan of the output root instead of
an enrich-only pass. The seeder's pending-request queue was generalised from
enrich-specific to carrying a scan phase, so the request starts immediately
when the seeder is idle and coalesces (escalating to FULL on a phase mismatch)
when a scan is already running. queue_output_enrichment is renamed to
queue_output_scan across the protocol, the NoAssets no-op and the call site.

References FIX-6.

* chore(assets): remove seeder paths orphaned by the output-scan change

45c2f96e rerouted both former enrich call sites to start()/enqueue_scan(),
leaving two seeder methods that look live but are not. Review round F2
raised this along with four smaller items; the user's disposition was to
fix all six here.

- Delete start_enrich: zero callers repo-wide after 45c2f96e.
- Delete enqueue_enrich: no production callers; its ~18 call sites in
  tests/test_asset_seeder.py move to enqueue_scan(phase=ScanPhase.ENRICH)
  with their semantics unchanged. The deletion forces the half-done class
  renames (TestEnqueueEnrich* -> TestEnqueueScan*, consistent with the
  already-renamed TestPendingScanDrain) and restores the module docstring
  that was dropped rather than reworded.
- Document at manager.queue_output_scan that ScanPhase.FULL per debounce
  window is the deliberate, user-ratified trade, so it is not optimised
  back to ENRICH without revisiting the decision.
- Document that SeedAssetSpec.size_bytes/mtime_ns are walk-time
  diagnostics only - production persists the seed-time restat since CR-7.
- Export create_content_reporting_insert from the queries facade and fold
  scanner.py's direct-module import into the existing facade block.
- Harden test_queue_output_scan_does_not_duplicate_declared_output against
  a vacuous pass: it now asserts the seeder finished without errors and
  that an undeclared sibling written into the same directory WAS
  registered by the same scan, proving the walk actually ran.

No production behaviour changes beyond the two deletions.

References F2-cleanup.

* chore: comment cleanup

Comment-Gate: 18 quarantined

* fix(assets): preserve pause across the seeder's pending-scan drain

pause() runs before every prompt, while pending-scan enqueue and resume only run inside the debounced gc-interval gate. If the active scan finishes just after the next prompt's pause, its finally block resets the seeder to idle and the pending drain starts a replacement with the run gate open, so resume becomes a no-op.

Capture pausedness under the lock before resetting to idle, then start the drained scan already paused. Setting the state and gate before launching the thread avoids the start-then-reclear window and lets resume release the existing scan checkpoints.

Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>

* test(assets): pin job_id absence for scan-discovered assets

Owner ruling, recorded 2026-09-03 in the stack-9-hardening planning notepad: scan-discovered assets — including undeclared outputs found by the post-execution walk — carry job_id = None, always; only emission-time registration (output_ui declaration) attributes a job; attributing walk finds to the most recent prompt would be a temporal-correlation guess that is wrong exactly when prompts interleave; None is honest provenance. Do NOT add proximity-based attribution heuristics to the scanner. Ratified against Jacob Segal's cross-job-attribution concern (2026-09-08 review meeting) — a wrongly-attributed asset could mean one user's cloud job sees another user's asset.

Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)
Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>

* [review-stack 10/10] assets-tests (#16218)

* test(execution): run the battery with assets enabled and assert asset-system health at teardown

* test(execution): cover list-shaped outputs registering assets

Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>

---------

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>

* [review-stack 11/11] review-fixes (#16261)

* fix(assets): only exit on database file-lock timeout when assets are enabled

* test(assets): pin live_contents_under_prefixes path-filtering semantics

* perf(assets): push live-content prefix filtering into SQL

* test(assets): declare per-entry intent in the path-prefix corpus

* test(assets): normalize POSIX-literal path expectations for Windows

* test(assets): force observable stat changes and close-before-mutate on Windows-sensitive rewrites

* test(assets): force an observable mtime change in the hash-mode split test

---------

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
Co-authored-by: guill <jacob.e.segal@gmail.com>
2026-09-13 12:04:51 -07:00

394 lines
13 KiB
Python

from __future__ import annotations
import os
from collections.abc import Iterator
from contextlib import contextmanager
from dataclasses import dataclass
from pathlib import Path
from unittest.mock import patch
import sqlalchemy as sa
from sqlalchemy import select
from sqlalchemy.orm import Session
import pytest
from app.assets.database.models import Asset, AssetContent
from app.assets.database.queries import create_content, create_record
from app.assets.database.queries.records import (
RecordPageSpec,
fetch_record_tags,
list_records_page,
)
from app.assets.helpers import to_stored_hash
from app.assets.scanner import get_unenriched_assets_for_roots
from app.assets.scanner_changes import (
clear_pending_verifications,
detect_content_change,
drain_pending_verifications,
)
from app.assets.services.hash_mode_state import (
clear_transition_queue,
drain_transition_queue,
enqueue_transition_work,
)
from app.assets.services.lookup import lookup_for_view
from app.assets.services.snapshot_hash import snapshot_hash
@dataclass(frozen=True, slots=True)
class _FakeStat:
st_size: int
st_mtime_ns: int
@contextmanager
def _reuse_session(session: Session) -> Iterator[Session]:
yield session
def _raw_system_metadata(session: Session, record_id: str) -> object:
return session.execute(
sa.text("SELECT system_metadata FROM assets WHERE id = :id"),
{"id": record_id},
).scalar()
def _candidates_under(session: Session, temp_dir: Path, *, compute_hashes: bool) -> set[str]:
with (
patch("app.assets.scanner.create_session", lambda: _reuse_session(session)),
patch(
"app.assets.scanner.get_scan_prefixes_for_root",
return_value=[str(temp_dir)],
),
):
rows = get_unenriched_assets_for_roots(("models",), compute_hashes=compute_hashes)
return {row.record_id for row in rows}
@pytest.fixture(autouse=True)
def _transition_queue_isolation() -> Iterator[None]:
clear_transition_queue()
yield
clear_transition_queue()
@pytest.fixture(autouse=True)
def _input_base_is_temp_dir(temp_dir: Path, monkeypatch: pytest.MonkeyPatch) -> None:
monkeypatch.setattr("folder_paths.get_input_directory", lambda: str(temp_dir))
@pytest.fixture(autouse=True)
def _pending_verification_isolation() -> Iterator[None]:
clear_pending_verifications()
yield
clear_pending_verifications()
def _seed_hashed_content(session: Session, path: Path, data: bytes) -> AssetContent:
path.write_bytes(data)
snapshot = snapshot_hash(str(path))
assert snapshot is not None
digest, verified_stat = snapshot
return create_content(
session,
str(path),
to_stored_hash(digest),
verified_stat.st_size,
verified_stat.st_mtime_ns,
)
def _bump_mtime(path: Path) -> os.stat_result:
before = path.stat()
bumped = before.st_mtime_ns + 5_000_000_000
os.utime(path, ns=(bumped, bumped))
after = path.stat()
assert after.st_mtime_ns != before.st_mtime_ns
assert after.st_size == before.st_size
return after
def _rewrite_same_size(path: Path, data: bytes) -> os.stat_result:
before = path.stat()
assert len(data) == before.st_size
assert data != path.read_bytes()
path.write_bytes(data)
bumped = before.st_mtime_ns + 5_000_000_000
os.utime(path, ns=(bumped, bumped))
after = path.stat()
assert after.st_size == before.st_size
assert after.st_mtime_ns != before.st_mtime_ns
return after
def test_split_record_is_enrich_candidate_in_hash_mode(
session: Session, temp_dir: Path
) -> None:
path = temp_dir / "split.safetensors"
content = create_content(session, str(path), hash="blake3:deadbeef")
record = create_record(session, content.id, path.name)
session.commit()
record_id = record.id
candidates = _candidates_under(session, temp_dir, compute_hashes=True)
assert record_id in candidates
def test_enriched_record_is_not_enrich_candidate_in_hash_mode(
session: Session, temp_dir: Path
) -> None:
path = temp_dir / "enriched.safetensors"
content = create_content(session, str(path), hash="blake3:deadbeef")
record = create_record(
session, content.id, path.name, system_metadata={"architecture": "flux"}
)
session.commit()
record_id = record.id
candidates = _candidates_under(session, temp_dir, compute_hashes=True)
assert record_id not in candidates
def test_same_size_mtime_bump_does_not_split(session: Session, temp_dir: Path) -> None:
path = temp_dir / "touched.safetensors"
content = create_content(session, str(path), hash=None, size_bytes=100, mtime_ns=1000)
record = create_record(
session, content.id, path.name, tags=["keepme"], system_metadata={"k": "v"}
)
session.commit()
content_id, record_id = content.id, record.id
detect_content_change(
session, content, _FakeStat(st_size=100, st_mtime_ns=2000), hashing_is_enabled=False
)
session.commit()
session.expire_all()
live = session.get(AssetContent, content_id)
assert live is not None and live.is_missing is False
rows_at_path = list(
session.scalars(select(AssetContent).where(AssetContent.path == str(path)))
)
assert len(rows_at_path) == 1
surviving = session.get(Asset, record_id)
assert surviving.system_metadata == {"k": "v"}
tags = fetch_record_tags(session, record_id)
assert "keepme" in tags
assert "missing" not in tags
assert live.mtime_ns == 2000
assert live.size_bytes == 100
def test_accepted_mtime_bump_drops_the_unverifiable_hash(
session: Session, temp_dir: Path
) -> None:
path = temp_dir / "synced.safetensors"
content = _seed_hashed_content(session, path, b"rsynced bytes")
record = create_record(
session, content.id, path.name, tags=["keepme"], system_metadata={"k": "v"}
)
session.commit()
content_id, record_id, stored_hash = content.id, record.id, content.hash
assert lookup_for_view(session, stored_hash) is not None
observed = _bump_mtime(path)
detect_content_change(session, content, observed, hashing_is_enabled=False)
session.commit()
session.expire_all()
live = session.get(AssetContent, content_id)
assert live.hash is None
assert lookup_for_view(session, stored_hash) is None
assert live.is_missing is False
assert live.mtime_ns == observed.st_mtime_ns
assert live.size_bytes == observed.st_size
surviving = session.get(Asset, record_id)
assert surviving is not None and surviving.content_id == content_id
assert surviving.system_metadata == {"k": "v"}
assert "keepme" in fetch_record_tags(session, record_id)
listed, _, _ = list_records_page(session, RecordPageSpec(limit=100))
assert record_id in {row.id for row in listed}
def test_same_size_content_change_is_never_served_under_the_old_hash(
session: Session, temp_dir: Path
) -> None:
path = temp_dir / "overwritten.safetensors"
content = _seed_hashed_content(session, path, b"AAAA")
create_record(session, content.id, path.name, tags=["keepme"])
session.commit()
old_hash = content.hash
assert lookup_for_view(session, old_hash) is not None
observed = _rewrite_same_size(path, b"BBBB")
detect_content_change(session, content, observed, hashing_is_enabled=False)
session.commit()
session.expire_all()
assert path.read_bytes() == b"BBBB"
assert lookup_for_view(session, old_hash) is None
def test_accepted_mtime_bump_is_not_re_detected_by_the_next_scan(
session: Session, temp_dir: Path
) -> None:
path = temp_dir / "resynced.safetensors"
content = _seed_hashed_content(session, path, b"cloud synced bytes")
create_record(session, content.id, path.name, tags=["keepme"])
session.commit()
content_id = content.id
detect_content_change(session, content, _bump_mtime(path), hashing_is_enabled=False)
session.commit()
detect_content_change(session, content, path.stat(), hashing_is_enabled=True)
assert drain_pending_verifications(session) == 0
detect_content_change(session, content, path.stat(), hashing_is_enabled=False)
session.commit()
session.expire_all()
rows_at_path = list(
session.scalars(select(AssetContent).where(AssetContent.path == str(path)))
)
assert len(rows_at_path) == 1
assert rows_at_path[0].id == content_id
assert rows_at_path[0].is_missing is False
def test_dropped_hash_is_refilled_in_place_by_a_later_hash_mode_pass(
session: Session, temp_dir: Path
) -> None:
path = temp_dir / "refilled.safetensors"
content = _seed_hashed_content(session, path, b"cloud synced bytes")
record = create_record(
session, content.id, path.name, tags=["keepme"], system_metadata={"k": "v"}
)
session.commit()
content_id, record_id = content.id, record.id
assert record_id not in _candidates_under(session, temp_dir, compute_hashes=True)
detect_content_change(session, content, _bump_mtime(path), hashing_is_enabled=False)
session.commit()
session.expire_all()
assert session.get(AssetContent, content_id).hash is None
assert record_id in _candidates_under(session, temp_dir, compute_hashes=True)
enqueue_transition_work(session, "off_to_on")
drain_transition_queue(session)
session.commit()
session.expire_all()
snapshot = snapshot_hash(str(path))
assert snapshot is not None
live = session.get(AssetContent, content_id)
assert live.is_missing is False
assert live.hash == to_stored_hash(snapshot[0])
assert lookup_for_view(session, live.hash).id == content_id
rows_at_path = list(
session.scalars(select(AssetContent).where(AssetContent.path == str(path)))
)
assert len(rows_at_path) == 1
assert "keepme" in fetch_record_tags(session, record_id)
assert session.get(Asset, record_id).system_metadata == {"k": "v"}
assert record_id not in _candidates_under(session, temp_dir, compute_hashes=True)
def test_mtime_and_size_change_splits_with_null_metadata(
session: Session, temp_dir: Path
) -> None:
path = temp_dir / "grown.safetensors"
content = create_content(session, str(path), hash=None, size_bytes=100, mtime_ns=1000)
create_record(
session, content.id, path.name, tags=["oldtag"], system_metadata={"k": "v"}
)
session.commit()
old_content_id = content.id
detect_content_change(
session, content, _FakeStat(st_size=200, st_mtime_ns=2000), hashing_is_enabled=False
)
session.commit()
session.expire_all()
assert session.get(AssetContent, old_content_id).is_missing is True
live = session.scalar(
select(AssetContent).where(
AssetContent.path == str(path), AssetContent.is_missing.is_(False)
)
)
assert live is not None and live.id != old_content_id
assert live.size_bytes == 200 and live.mtime_ns == 2000
new_record = session.scalar(select(Asset).where(Asset.content_id == live.id))
assert new_record is not None
assert new_record.system_metadata is None
assert _raw_system_metadata(session, new_record.id) is None
assert "oldtag" not in fetch_record_tags(session, new_record.id)
assert new_record.id in _candidates_under(session, temp_dir, compute_hashes=True)
def test_mtime_unchanged_size_changed_does_not_split(
session: Session, temp_dir: Path
) -> None:
path = temp_dir / "weird.safetensors"
content = create_content(session, str(path), hash=None, size_bytes=100, mtime_ns=1000)
create_record(session, content.id, path.name, tags=["keepme"])
session.commit()
content_id = content.id
detect_content_change(
session, content, _FakeStat(st_size=999, st_mtime_ns=1000), hashing_is_enabled=False
)
session.commit()
session.expire_all()
assert session.get(AssetContent, content_id).is_missing is False
rows_at_path = list(
session.scalars(select(AssetContent).where(AssetContent.path == str(path)))
)
assert len(rows_at_path) == 1
def test_transition_drain_split_replacement_has_null_metadata(
session: Session, temp_dir: Path
) -> None:
path = temp_dir / "changed.bin"
path.write_bytes(b"old bytes")
old_snapshot = snapshot_hash(str(path))
assert old_snapshot is not None
old_digest, _ = old_snapshot
stat = path.stat()
old_content = create_content(
session, str(path), to_stored_hash(old_digest), stat.st_size, stat.st_mtime_ns
)
old_content_id = old_content.id
create_record(
session, old_content_id, "changed.bin", tags=["oldtag"], system_metadata={"k": "v"}
)
path.write_bytes(b"different new bytes")
enqueue_transition_work(session, "off_to_on")
drain_transition_queue(session)
session.commit()
session.expire_all()
assert session.get(AssetContent, old_content_id).is_missing is True
live = session.scalar(
select(AssetContent).where(
AssetContent.path == str(path), AssetContent.is_missing.is_(False)
)
)
assert live is not None and live.id != old_content_id
new_record = session.scalar(select(Asset).where(Asset.content_id == live.id))
assert new_record is not None
assert new_record.system_metadata is None
assert _raw_system_metadata(session, new_record.id) is None
assert "oldtag" not in fetch_record_tags(session, new_record.id)