* fix(assets): don't take the database lock when assets are off
An assets-off ComfyUI no longer initialises the database or takes its lock, so it
can't block a later assets-on start on the same install. If another process holds
the lock, it prints a startup warning saying a future version will refuse to
start the second instance, and keeps running.
* fix(assets): never let the lock check stop startup; tighten the lock tests
* fix(assets): keep the lock check's release inside its guard
* fix(assets): reword the shared-database startup warning
seeder.scan_completed reported only elapsed_ms, which is wall-clock time
and includes every pause taken while prompts ran. It now also carries:
- cpu_ms: the scan thread's CPU time (time.thread_time()).
- paused_ms: time blocked at the pause gate, summed across pauses.
- dirs_listed_count: directories the input/output walk or the output
rescan's folder listing listed.
- files_statted_count: os.stat calls on files in the scan's per-file
loops (reference sync, the listing check, discovery, admission, seed,
watch list and enrich).
Scan failure events carried only the exception class, so every SQLite
failure read as OperationalError. seeder.scan_failed, batch_insert_failed
and scanner.{fast_scan,temp_sync,mark_missing,stat,watch_stat,watch_seed}
_failed now also carry error_kind, one of a closed set:
expression_tree_too_large, too_many_variables, database_locked,
disk_full, disk_io, unable_to_open, database_corrupt, permission_denied,
file_locked, read_only, other.
It is a closed enum rather than a scrubbed message because str() of a
SQLAlchemy error includes the statement and its bound parameters, which
are file paths. Classification reads SQLite's result code where the
driver exposes it (Python 3.11+), then SQLite's fixed message on the
driver exception (exc.orig), then errno and, on Windows, winerror.
Co-authored-by: guill <jacob.e.segal@gmail.com>
* Disable pinned memory automatically on AMD APUs
APU VRAM is carved out of system RAM, so pinning host memory only takes
RAM away from the GPU. Detect AMD integrated GPUs and treat them as if
--disable-pinned-memory was passed. dGPUs and the flag are unchanged.
* mm: generalize is_integrated pin disable check to all cuda
---------
Co-authored-by: tvukovic-amd <tvukovic@amd.com>
* fix(assets): batch prefix filters so scans work with many model folders
Every scan prefix (one per model folder, including each extra_model_paths
base) adds terms to a single OR, and SQLite rejects an expression tree
deeper than 1000. From about 500 prefixes every models scan failed with
"Expression tree is too large" and the catalog stayed empty.
Run the prefix filter in batches of at most 200 prefixes and merge the
results, deduping rows that nested or overlapping prefixes put in more than
one batch. With 200 prefixes or fewer the statement is unchanged. The
enrichment candidate query merges each batch's keyset page, which yields
the same page a single statement would.
* fix(assets): filter enrich candidates in one pass above one prefix batch
Paging each prefix batch separately made every page scan to the end of
the table for any batch with few matches, which is quadratic in catalog
size. Above one batch, read the candidates in id order once and apply the
same prefix test in Python, stopping at the page limit. At 200 prefixes
or fewer the statement is unchanged.
Run the many-prefix tests under a 999 bound-variable cap too, the limit
of SQLite before 3.32, which each 200-prefix batch stays under.
* fix(assets): yield the GIL while filtering enrich candidates in Python
Above one prefix batch the enrich candidates are filtered in a Python loop,
which could hold the GIL across many rows outside the prefixes. Yield it per
row as the other scan loops do, so the UI stays responsive.
* Revert "fix(assets): yield the GIL while filtering enrich candidates in Python"
Measured, it bought nothing: with a 60k-asset catalogue at 450 and 2000
prefixes, p99 lateness of a 1 ms sleeper thread is the same within noise
either way, because the Python work between 500-row fetches is short and
sqlite3 releases the GIL during each fetch. The yield made the loop 30-90%
slower.
* test(assets): import db at module level in the many-prefix scan test
* [Partner Nodes] feat(Anthropic): add Claude Sonnet 5.5 to the Claude node
Signed-off-by: bigcat88 <bigcat88@icloud.com>
* [Partner Nodes] feat(Anthropic): allow max_tokens down to 1024 for Opus 5.5 and Sonnet 5.5
Signed-off-by: bigcat88 <bigcat88@icloud.com>
---------
Signed-off-by: bigcat88 <bigcat88@icloud.com>
* fix(assets): recover records when a drive comes back with hashing off
A scan that runs while a drive is offline marks every row on it missing. With
hashing off (the default), nothing could recover those rows when the drive
returned, so the scan created new records and the user's names, tags, metadata
and job links stayed on the orphaned originals.
- With hashing off, a file that reappears at a path no live row occupies
recovers the missing row there whose size and mtime match exactly. If several
match, the newest recovers. Hashing on is unchanged.
- A reference stat that fails with an I/O error other than not-found leaves the
row live, as the output listing rescan already does.
- seeder.marked_missing is emitted per root when a fast scan marks rows missing,
and scan_completed carries missing_marked_count and recovered_count.
* fix(assets): read recovery candidates before the write transaction; count recoveries only once committed
* test(assets): a pruned model folder recovers its records when it is registered again
* fix(assets): never recover a missing row whose records were all deleted
* docs(assets): scope the multiple-match rule to hashing on
* docs(assets): only new content splits a same-path edit
* test(assets): scan_completed carries nonzero missing and recovered counts
* perf(assets): let the partial live-path index serve the hashing-off recovery check
Written as IS 0, the check that no live row occupies the path scanned
asset_contents once per recovered file, inside the batch's write transaction.
At 50k rows that stretched each batch's lock window to about 2.5s and made
concurrent output registrations time out.
* ci: run the checks against master
* test(assets): import folder_paths at module level
folder_paths.get_filename_list raises when a folder it cached earlier has gone
away (or on an OSError other than not-found), and collect_models_files let that
abort the whole asset scan, so nothing in models, input or output was
catalogued. Fall back to a fresh listing of the category, which skips a folder
that's gone; if that raises too, skip the category with a warning and scan the
rest.
With --enable-assets, the scans that keep the asset database in sync could starve the web server's event loop and burn CPU on large libraries. Three changes:
- First scan yields the GIL. The loop that derives names and tags for new files now sleeps briefly every 2 ms of work, so the event loop keeps running. At 51k files, time to a usable UI on the first run drops from 4.7 s to 2.0 s (1.5 s with assets off); the first scan takes about 3% longer. (was #16546)
- Output rescans list folders instead of stat'ing every file. The post-prompt output rescan compares directory listings against the database instead of stat'ing each catalogued file. That per-file stat was 70–95% of the rescan's cost: ~5.8 s per rescan at 200k outputs on local SSD, ~57 s at 10k on NFS. Any row the listing can't vouch for gets an individual stat before it is retired, and only "file not found" retires it: case-insensitive or normalizing filesystems, hidden folders, folders that fail to list, symlink aliases, and permission or I/O errors keep their rows. The one behaviour left to the next full scan is a file overwritten in place under the same name. (was #16599)
- Rescan loops pause. The listing and comparison loops pause every 10 ms of work, so prompts start and the UI stays responsive during a long rescan. At 200k outputs, without this a prompt could wait ~1 s to start. This costs 7–15% more rescan time. (was #16600)
Assets-only: none of this runs with assets off.
ComfyUI failed to start when the assets database packages (sqlalchemy, alembic, blake3) were missing, e.g. in a venv not re-synced after requirements.txt changed. That happened with assets off too, although that mode never needs them.
- The asset imports in lifecycle.py, manager.py and server.py are guarded by db.py's existing dependencies_available() check. Without the packages, NoAssets registers no asset routes and temp cleanup still runs.
- --enable-assets without the packages logs which packages are missing, prints the install command, and continues with assets disabled.
- /view resolves blake3: filenames only when assets are enabled. Trade-off: if assets are turned off on an install whose saved workflows hold blake3: widget values, those previews return 404. Running such workflows already fails, since execution never resolves hashes.
* fix(db): wait briefly for the database lock at startup
A relaunch can start while the previous process is still exiting and
holding the lock. Wait up to 5 seconds for it to be released before
treating the database as in use, and log how long the wait took. The
lock is never taken from a holder.
* fix(db): log the lock wait at info without guessing its cause
* fix(db): say when startup waits for the database lock
Log once before the wait instead of after it, so a lock that stays held
explains the pause before the error. Document the wait in the docstring,
and make the wait test release the lock only after the wait has started.