mirror of
https://github.com/Comfy-Org/ComfyUI.git
synced 2026-09-28 00:47:44 -05:00
* review-stack 1/4: code (37 files, +3217/-3958) Review-and-land stack for synap5e/feat/asset-record-content-split, generated by review-stack.py. Once approved, merges DOWN into the layer below (a fast-forward); only the bottom layer squash-merges into the real base. See ~/adocs/review-stack.md. Rule: path not under tests-unit/ or tests/ Question: Is the logic change right? Source tip:7007d18582Merge-base:783545f689* review-stack 2/4: tests-removed (24 files, +274/-8220) Review-and-land stack for synap5e/feat/asset-record-content-split, generated by review-stack.py. Once approved, merges DOWN into the layer below (a fast-forward); only the bottom layer squash-merges into the real base. See ~/adocs/review-stack.md. Rule: test file deleted, or modified with deleted/(added+deleted) >= 0.9 Question: For each dropped assertion: obsolete by a ruling, or covered by a tests-new test? Source tip:7007d18582Merge-base:783545f689* review-stack 3/4: tests-changed (13 files, +1043/-1218) Review-and-land stack for synap5e/feat/asset-record-content-split, generated by review-stack.py. Once approved, merges DOWN into the layer below (a fast-forward); only the bottom layer squash-merges into the real base. See ~/adocs/review-stack.md. Rule: remaining modified test files (incl. conftest.py / helpers) Question: Did the edits weaken an existing check? Source tip:7007d18582Merge-base:783545f689* review-stack 4/4: tests-new (46 files, +8601/-0) Review-and-land stack for synap5e/feat/asset-record-content-split, generated by review-stack.py. Once approved, merges DOWN into the layer below (a fast-forward); only the bottom layer squash-merges into the real base. See ~/adocs/review-stack.md. Rule: test file added Question: Is the code layer well covered? Source tip:7007d18582Merge-base:783545f689* review-stack 5/6: code (13 files, +351/-104) Review-and-land stack for synap5e/feat/assets-di, generated by review-stack.py. Once approved, merges DOWN into the layer below (a fast-forward); only the bottom layer squash-merges into the real base. See ~/adocs/review-stack.md. Rule: path not under tests-unit/ or tests/ Question: Is the logic change right? Source tip:eca2c74bffMerge-base:20d59d2a5f* review-stack 6/6: tests (8 files, +753/-238) Review-and-land stack for synap5e/feat/assets-di, generated by review-stack.py. Once approved, merges DOWN into the layer below (a fast-forward); only the bottom layer squash-merges into the real base. See ~/adocs/review-stack.md. Rule: every changed file under tests-unit/ or tests/ (added, modified, or deleted) Question: Is the code layer well covered, and did any edit weaken an existing check? Source tip:eca2c74bffMerge-base:20d59d2a5f* review-stack 7/8: ported-fixes (42 files, +1361/-180) Review-and-land stack for synap5e/feat/assets-di-v2, generated by review-stack.py conventions (hand-built continuation layer; see the PR body). Once approved, merges DOWN into the layer below (a fast-forward); only the bottom layer squash-merges into the real base. See ~/adocs/review-stack.md. Rule: the 11 base-branch fix/docs commits 595cd6e4..94d7185b cherry-picked across the DI refactor (7efdd1d7excluded, superseded by layer 8) Question: was each base fix ported faithfully across the DI refactor? Source tip: 6841881069284803b902b4a9e33bdcda13126771 Merge-base:7fdfb40f4b* review-stack 8/8: defensive-parity (4 files, +36/-3) Review-and-land stack for synap5e/feat/assets-di-v2, generated by review-stack.py conventions (hand-built continuation layer; see the PR body). Once approved, merges DOWN into the layer below (a fast-forward); only the bottom layer squash-merges into the real base. See ~/adocs/review-stack.md. Rule: match-or-improve master's dependency defenses — NoAssets selection when DB deps unavailable (7efdd1d7's outcome via the DI seam), requirements warning before assets imports, blake3 in the guarded dependency set Question: does each degradation path now match or improve master's behavior? Source tip:ebc2cfeebcMerge-base:7fdfb40f4b* fix(assets): only discard content rows this operation actually inserted CR-9: Enumerated all six create_content call sites. Only scanner seeding and the three ingest registration paths track IDs for failure cleanup. * fix(assets): reject hash-only uploads with FEATURE_DISABLED when hashing is off CodeRabbit finding CR-2: reject hash-only multipart uploads before create_from_hash when hashing is disabled. * fix(assets): seed persists the stat it verified CR-7: persist the fresh seed-time restat instead of walk-time spec values. * fix(assets): route database lock failures to the lock guidance CR-16: route file-lock startup failures through the existing lock guidance and exit path. * fix(assets): drop the inaccurate temp-cleanup claim from the shutdown warning References CR-10. * fix(assets): walk the output root after execution so undeclared outputs register promptly Custom nodes that write files into the output directory without declaring them in output_ui only became assets when the next full walk happened - a frontend GET /object_info or a restart. Headless and API-only sessions never trigger either, so those files never converged into the asset database. The post-execution hook now requests a FULL scan of the output root instead of an enrich-only pass. The seeder's pending-request queue was generalised from enrich-specific to carrying a scan phase, so the request starts immediately when the seeder is idle and coalesces (escalating to FULL on a phase mismatch) when a scan is already running. queue_output_enrichment is renamed to queue_output_scan across the protocol, the NoAssets no-op and the call site. References FIX-6. * chore(assets): remove seeder paths orphaned by the output-scan change 45c2f96e rerouted both former enrich call sites to start()/enqueue_scan(), leaving two seeder methods that look live but are not. Review round F2 raised this along with four smaller items; the user's disposition was to fix all six here. - Delete start_enrich: zero callers repo-wide after 45c2f96e. - Delete enqueue_enrich: no production callers; its ~18 call sites in tests/test_asset_seeder.py move to enqueue_scan(phase=ScanPhase.ENRICH) with their semantics unchanged. The deletion forces the half-done class renames (TestEnqueueEnrich* -> TestEnqueueScan*, consistent with the already-renamed TestPendingScanDrain) and restores the module docstring that was dropped rather than reworded. - Document at manager.queue_output_scan that ScanPhase.FULL per debounce window is the deliberate, user-ratified trade, so it is not optimised back to ENRICH without revisiting the decision. - Document that SeedAssetSpec.size_bytes/mtime_ns are walk-time diagnostics only - production persists the seed-time restat since CR-7. - Export create_content_reporting_insert from the queries facade and fold scanner.py's direct-module import into the existing facade block. - Harden test_queue_output_scan_does_not_duplicate_declared_output against a vacuous pass: it now asserts the seeder finished without errors and that an undeclared sibling written into the same directory WAS registered by the same scan, proving the walk actually ran. No production behaviour changes beyond the two deletions. References F2-cleanup. * chore: comment cleanup Comment-Gate: 18 quarantined * fix(assets): preserve pause across the seeder's pending-scan drain pause() runs before every prompt, while pending-scan enqueue and resume only run inside the debounced gc-interval gate. If the active scan finishes just after the next prompt's pause, its finally block resets the seeder to idle and the pending drain starts a replacement with the run gate open, so resume becomes a no-op. Capture pausedness under the lock before resetting to idle, then start the drained scan already paused. Setting the state and gate before launching the thread avoids the start-then-reclear window and lets resume release the existing scan checkpoints. Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent) Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai> * test(assets): pin job_id absence for scan-discovered assets Owner ruling, recorded 2026-09-03 in the stack-9-hardening planning notepad: scan-discovered assets — including undeclared outputs found by the post-execution walk — carry job_id = None, always; only emission-time registration (output_ui declaration) attributes a job; attributing walk finds to the most recent prompt would be a temporal-correlation guess that is wrong exactly when prompts interleave; None is honest provenance. Do NOT add proximity-based attribution heuristics to the scanner. Ratified against Jacob Segal's cross-job-attribution concern (2026-09-08 review meeting) — a wrongly-attributed asset could mean one user's cloud job sees another user's asset. Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent) Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai> * [review-stack 10/10] assets-tests (#16218) * test(execution): run the battery with assets enabled and assert asset-system health at teardown * test(execution): cover list-shaped outputs registering assets Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent) Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai> --------- Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai> * [review-stack 11/11] review-fixes (#16261) * fix(assets): only exit on database file-lock timeout when assets are enabled * test(assets): pin live_contents_under_prefixes path-filtering semantics * perf(assets): push live-content prefix filtering into SQL * test(assets): declare per-entry intent in the path-prefix corpus * test(assets): normalize POSIX-literal path expectations for Windows * test(assets): force observable stat changes and close-before-mutate on Windows-sensitive rewrites * test(assets): force an observable mtime change in the hash-mode split test --------- Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai> Co-authored-by: guill <jacob.e.segal@gmail.com>
206 lines
6.9 KiB
Python
206 lines
6.9 KiB
Python
"""Reconciles catalogued content against what is actually on disk: retiring rows
|
|
whose file is gone, splitting a row whose bytes changed, and recovering one
|
|
whose file came back. Recovery fires only when the returning file's hash
|
|
identifies exactly one missing row and no live row already occupies that path,
|
|
so a restored file can never leave two live rows describing one location.
|
|
"""
|
|
|
|
from __future__ import annotations
|
|
|
|
import os
|
|
from pathlib import Path
|
|
from typing import Literal
|
|
|
|
import sqlalchemy as sa
|
|
from sqlalchemy.orm import Session
|
|
|
|
from app.assets.database.models import AssetContent
|
|
from app.assets.database.queries.records import (
|
|
create_content,
|
|
create_record,
|
|
mark_content_missing,
|
|
unset_content_missing,
|
|
)
|
|
from app.assets.helpers import sql_path_under_prefix, to_stored_hash
|
|
from app.assets.services.path_utils import compute_loader_path, get_name_and_tags_from_asset_path
|
|
from app.assets.services.snapshot_hash import snapshot_hash
|
|
|
|
_pending_verification_ids: list[str] = []
|
|
_pending_recovery_paths: list[str] = []
|
|
|
|
|
|
def clear_pending_verifications() -> None:
|
|
_pending_verification_ids.clear()
|
|
_pending_recovery_paths.clear()
|
|
|
|
|
|
def queue_pending_verification(content_id: str) -> None:
|
|
if content_id not in _pending_verification_ids:
|
|
_pending_verification_ids.append(content_id)
|
|
|
|
|
|
def pending_recovery_count() -> int:
|
|
return len(_pending_recovery_paths)
|
|
|
|
|
|
def recover_missing_content(
|
|
session: Session, path: str, stat_result: os.stat_result, hashing_is_enabled: bool
|
|
) -> Literal["recovered", "no_match", "unstable"]:
|
|
if not hashing_is_enabled:
|
|
return "no_match"
|
|
occupied = session.scalar(
|
|
sa.select(AssetContent.id)
|
|
.where(AssetContent.path == path, AssetContent.is_missing.is_(False))
|
|
.limit(1)
|
|
)
|
|
if occupied is not None:
|
|
return "no_match"
|
|
snapshot = snapshot_hash(path)
|
|
if snapshot is None:
|
|
if path not in _pending_recovery_paths:
|
|
_pending_recovery_paths.append(path)
|
|
return "unstable"
|
|
digest, verified_stat = snapshot
|
|
stored_hash = to_stored_hash(digest)
|
|
matches = list(
|
|
session.scalars(
|
|
sa.select(AssetContent).where(
|
|
AssetContent.path == path,
|
|
AssetContent.is_missing.is_(True),
|
|
AssetContent.hash == stored_hash,
|
|
)
|
|
)
|
|
)
|
|
if len(matches) == 1:
|
|
recovered = matches[0]
|
|
unset_content_missing(session, recovered.id)
|
|
recovered.size_bytes = verified_stat.st_size
|
|
recovered.mtime_ns = verified_stat.st_mtime_ns
|
|
return "recovered"
|
|
if len(matches) > 1:
|
|
return "no_match"
|
|
null_hash_matches = list(
|
|
session.scalars(
|
|
sa.select(AssetContent).where(
|
|
AssetContent.path == path,
|
|
AssetContent.is_missing.is_(True),
|
|
AssetContent.hash.is_(None),
|
|
)
|
|
)
|
|
)
|
|
if len(null_hash_matches) != 1:
|
|
return "no_match"
|
|
candidate = null_hash_matches[0]
|
|
if (candidate.size_bytes, candidate.mtime_ns) != (
|
|
verified_stat.st_size,
|
|
verified_stat.st_mtime_ns,
|
|
):
|
|
return "no_match"
|
|
unset_content_missing(session, candidate.id)
|
|
candidate.hash = stored_hash
|
|
candidate.size_bytes = verified_stat.st_size
|
|
candidate.mtime_ns = verified_stat.st_mtime_ns
|
|
return "recovered"
|
|
|
|
|
|
def is_path_under_prefixes(path: str, prefixes: list[str]) -> bool:
|
|
candidate = Path(os.path.abspath(path))
|
|
return any(candidate.is_relative_to(os.path.abspath(prefix)) for prefix in prefixes)
|
|
|
|
|
|
def split_content(session: Session, content: AssetContent, stat_result: os.stat_result, hash_value: str | None) -> AssetContent:
|
|
mark_content_missing(session, content.id)
|
|
name, tags = get_name_and_tags_from_asset_path(content.path)
|
|
replacement = create_content(
|
|
session,
|
|
path=content.path,
|
|
hash=hash_value,
|
|
size_bytes=stat_result.st_size,
|
|
mtime_ns=stat_result.st_mtime_ns,
|
|
)
|
|
create_record(
|
|
session,
|
|
content_id=replacement.id,
|
|
name=name,
|
|
loader_path=compute_loader_path(content.path),
|
|
tags=tags,
|
|
)
|
|
return replacement
|
|
|
|
|
|
def detect_content_change(
|
|
session: Session,
|
|
content: AssetContent,
|
|
stat_result: os.stat_result,
|
|
hashing_is_enabled: bool,
|
|
) -> None:
|
|
if content.mtime_ns == stat_result.st_mtime_ns:
|
|
# Ruling #10: size drift with unchanged mtime is undefined behavior.
|
|
return
|
|
if hashing_is_enabled:
|
|
queue_pending_verification(content.id)
|
|
return
|
|
if content.size_bytes == stat_result.st_size:
|
|
# User identity rule: a same-size mtime bump (rsync, cloud sync, backup restore) is the
|
|
# same file — never split, or the record's tags and metadata are destroyed.
|
|
# The stored hash goes with the refreshed stat: OFF mode cannot prove the bytes, and a
|
|
# refreshed stat alone would re-qualify the row to be served under a digest it may no
|
|
# longer match.
|
|
content.size_bytes = stat_result.st_size
|
|
content.mtime_ns = stat_result.st_mtime_ns
|
|
content.hash = None
|
|
return
|
|
split_content(session, content, stat_result, hash_value=None)
|
|
|
|
|
|
def drain_pending_verifications(session: Session, limit: int | None = None) -> int:
|
|
queued_count = min(len(_pending_verification_ids), limit or len(_pending_verification_ids))
|
|
processed = 0
|
|
for _ in range(queued_count):
|
|
content_id = _pending_verification_ids.pop(0)
|
|
content = session.get(AssetContent, content_id)
|
|
if content is None or content.is_missing:
|
|
continue
|
|
try:
|
|
os.stat(content.path, follow_symlinks=True)
|
|
except FileNotFoundError:
|
|
mark_content_missing(session, content.id)
|
|
processed += 1
|
|
continue
|
|
except OSError:
|
|
queue_pending_verification(content_id)
|
|
continue
|
|
|
|
try:
|
|
snapshot = snapshot_hash(content.path)
|
|
except OSError:
|
|
queue_pending_verification(content_id)
|
|
continue
|
|
if snapshot is None:
|
|
queue_pending_verification(content_id)
|
|
continue
|
|
digest, verified_stat = snapshot
|
|
stored_hash = to_stored_hash(digest)
|
|
|
|
if content.hash == stored_hash or content.hash is None:
|
|
content.hash = stored_hash
|
|
content.size_bytes = verified_stat.st_size
|
|
content.mtime_ns = verified_stat.st_mtime_ns
|
|
else:
|
|
split_content(session, content, verified_stat, hash_value=stored_hash)
|
|
processed += 1
|
|
return processed
|
|
|
|
|
|
def live_contents_under_prefixes(session: Session, prefixes: list[str]) -> list[AssetContent]:
|
|
if not prefixes:
|
|
return []
|
|
return list(
|
|
session.scalars(
|
|
sa.select(AssetContent).where(
|
|
AssetContent.is_missing.is_(False),
|
|
sa.or_(*(sql_path_under_prefix(AssetContent.path, prefix) for prefix in prefixes)),
|
|
)
|
|
)
|
|
)
|