Eleven vendor-shaped Protocols are replaced by one IntegrationsDomain with
`call(integration, operation, **params)` and `describe()`. The published API
was growing by one class per third-party service while the enforcement point
— `integrations.<name>`, checked at the broker — never varied. The hand-kept
public stub had already fallen behind that growth: Luma, ImgBB and SenseNova
existed in the implementation but had never been published.
Three defects surfaced and are fixed here:
- LlamaCppModelRef.generate reached the vendor by attribute, so every pack
using the returned handle broke at the wire, not at the call site.
- The in-process path exposed vendors as public attributes and had no
`call`, so a node could work unsandboxed and fail once sandboxed. The
vendors are now private and reached only by name, giving both paths the
same surface.
- `describe` was declared sync in the Protocol but implemented async in the
guest, which must round-trip to the host.
Also removes a test for load_onnx_image_classifier, deleted from core in
2076700f without its tests; the surviving _validate_onnx_weight_file test
is kept.
The twenty handlers for IP-Adapter, SAM, CLIPSeg, image classification,
advanced ControlNet, transparent VAE, segmentation, inpainting, image
preprocessing and interpolation state used no instance state: they were
module functions parked on a class. They now live in _vendor_ops and
register through the same table, which leaves InProcessOps holding the
engine primitives and collects every pack-specific operation in one file
that can move out to the packs that own them.
They reach the SDK through the module rather than by importing names, so
substitution still resolves at call time.
A stock install — no overlay, no SDK markers on the node — now takes the
original invocation path: no ExecutionPlan, no ref table, no runtime
binding, and the pre-existing async task semantics. The seam engages per
node when it declares SDK_REFS/SDK_PERMISSIONS/SDK_REQUIRED_WEIGHTS, or
globally once any default provider is replaced, so a registered backend
still sees every node and provenance-based sandboxing cannot be bypassed
by a node simply declaring nothing.
The pipeline patches a live model, so its operations stay in core. What
leaves is the typed facade over them: IpAdapterRef and IpAdapterEmbedsRef
were published contract, so every one of their methods was a permanent
commitment. The same ergonomics now ship with the packs that want them,
built on the generic dispatch, and core keeps the operations without
keeping the surface.
ipadapter.apply, apply_tiled, encode and ipadapter_embeds.combine all
remain available by name. Handles cross as generic Refs carrying the
IPADAPTER_PIPE and IPADAPTER_EMBEDS kinds, which is what the validation and
the marshaller now check.
Ref gains _wrap so a handle with no dedicated class can still be re-typed
by the marshaller; the base keeps the resolver's kind while a subclass that
declares KIND asserts its own.
op() was defined only on ImageRef, so a pack holding any other handle had
no way to call a named operation on it and had to wait for core to grow a
typed method. The dispatch itself was always generic: the broker's wire
parameter is named image but passes straight through to
ops.apply(op, subject, params), and core's own vendor wrappers already send
non-image handles along it.
Moving op() to the base Ref is what lets a pack build its own typed
accessor over a capability core knows nothing about, so the operation
vocabulary can keep growing while the API does not. ImageRef keeps its
override to narrow the return type.
Removes OnnxDetectorRef, ctx.models.load_onnx_detector, the
onnx_detector.detect operation, its entry type, loader and cache, and the
secure_kind inference branch that mapped to the removed type.
The detector is one model family, and running an untrusted ONNX graph
belongs in the sandbox rather than in the trusted host process, so the packs
now load and execute it themselves through the generic asset broker.
Completes the removal: the hand-maintained public surface still declared
MattingModelRef and load_vitmatte after the implementation moved to the
LayerStyle pack.
ViTMatte is one pack's model family, so the architecture, the loader and the
refinement move to the pack that uses it. Removes MattingModelRef,
ctx.models.load_vitmatte, the matting.refine operation, the ViTMatte entry
type, its loader and its cache from core.
The pack asks only for generic primitives: resolve a declared weight and read
its state dict. It selects and offloads its own device, because the model is
its own and runs in its own process.
Ports 31d76988 from api-v2-runtime so both cores carry the same cache
behaviour. Seven of the thirteen copy-pasted caches evicted with
dict.popitem(), dropping the most recently inserted entry rather than the
least recently used, which pinned whichever model loaded first.
Ports the core half of the LayerStyle BLIP move to this branch, which had
only received the pack half. Removes VqaModelRef, ctx.models.load_vqa, the
VQA cache and the 30,522-token BERT vocabulary from core; LayerStyle now
carries that implementation and its own vocabulary and asks for weights
through the generic asset broker.
Brings this core back into step with api-v2-runtime, where this landed as
f1510546.
Preserves in-progress work so it is not lost: a new _cloud_media module
plus the SDK, public-surface, and model-transform changes that reference
it. All four files compile; committed as a recovery point rather than a
validated release.
Pack schemas use it for input bounds, and the module that has always
defined it (nodes.py) is a host module a sandboxed pack cannot import.
Same value by definition: 16384 is frozen into every workflow that
ever serialized a bound.
The V2 additions the KJNodes completion needed on the core side:
- _model_transforms: the closed, core-owned transform vocabulary
behind ModelRef.patch — 29 named transforms, declaratively
parameterized, validated host-side, immutable and stacking. No
function ever crosses the boundary; a pack cannot register one.
- structured-vs-live split: value()/from_value() only on structured
data refs (LATENT, AUDIO, TRACKS...). MODEL/CLIP/VAE/asset refs are
handles in every execution mode — in-process identity resolution no
longer hands a live model to node code.
- preview overrides (tiny-VAE, LTX factors), triton VAE seam, memory
attention, and profiling surfaces backing the corresponding closed
brokers in the overlay.
- torch_compile/model_patcher/model_management: compiled-view
aliasing recognized by the model manager (no double-counted
weights); shared state-dict loading path so the native loader and
the V2 broker cannot drift.
VaeRef.decode/encode, ClipRef.tokenize/encode_from_tokens_scheduled/encode,
CondRef.combine/concat, with in-process implementations behind the existing
named-op registry so an overlay extends the vocabulary without touching the
contract.
These keep the old API's shape on purpose — you still write vae.decode(latent)
— and change only what the call means: the node holds a handle, awaits, and the
decode runs on the trusted plane against weights it never sees. That is what
lets a node DECLARE a VAE input and still be sandboxable. Compatible in shape so
conversion is mechanical across a corpus nobody here maintains; different in
substance so a converted node is sandboxable by construction.
Two mirroring defects fixed before shipping: encode() sliced its input to three
channels (core's VAEEncode does not, so it silently dropped alpha), and ClipRef
offered only a combined encode(text) (CLIPTextEncodeSDXL builds one token dict
from two prompts, and the ACE nodes pass a dozen tokenizer kwargs, so the
collapse made both inexpressible).
InProcessRefResolver now records a ref's kind at creation and checks it at
resolve. Ref tokens cross as {kind, id, cls} and the host rebuilt from what
arrived, so the holder chose its own ref's type: an ImageRef id could be
presented as a VaeRef. Unguessable ids already stopped a guest reaching a handle
it was never given; possessing a handle is not the same as labelling it.
Release is explicit, not collected. InProcessRefResolver.clear() drops the
table's strong references at a known point, called from execution.py in a
finally so it also covers the path where the node raised or its guest died —
which is exactly when nothing else will run. A ref table can hold multi-gigabyte
tensors, and refcount timing neither crosses a process boundary nor is bounded
under reference cycles.
OpsProvider.apply is annotated (op, subject: Ref, params) -> Any; it claimed
ImageRef in and out while handle ops take a VaeRef and may return a LatentRef,
a CondRef, or a plain token dict.
Two generic gaps that blocked an out-of-process node from being a sampler.
1. wrap_inputs only wrapped tensors and latents. A MODEL or CONDITIONING is a
live engine object, so it passed straight through — and an out-of-process
backend cannot serialize a ModelPatcher, so any node taking one simply could
not run in a guest. The rule is now by capability rather than an enumerated
type list: a value that can cross as data does, and anything else becomes a
handle. That is the correct boundary rule anyway — objects do not cross,
handles do — and it is what lets a node take a MODEL and still run isolated.
2. ExecutionPlan.permissions was never populated, so the seam could not tell a
backend what a node needs. A node now declares SDK_PERMISSIONS and the seam
copies it onto the plan. Declaring is not granting: the backend decides, and
an out-of-process one still gates every call at the wire. Nodes that declare
nothing — the overwhelming majority — get nothing.
Behaviour-preserving in-process: 54 core seam tests and 26 overlay tests pass.
`unwrap_outputs` rebuilds a node's NodeOutput in order to swap output refs back
for real objects. It rebuilt it from the results alone — `NodeOutput(*resolved)`
— silently discarding `ui`, `expand` and `block_execution`.
The practical effect: no SDK_REFS node could be an output node. ComfyUI only
emits the `executed` websocket event, the one that delivers a node's results to
the frontend, for nodes that return ui data (`if len(output_ui) > 0`). So a
converted PreviewImage-style node executed perfectly and then displayed
nothing, with no error anywhere to explain it. `expand` (subgraph expansion)
and `block_execution` were lost the same way.
Resolving refs is a transport concern and has no business changing what the
node said. Generic fix, not specific to any backend: it applies equally to the
in-process path.
Replaces enumerated invert/scale methods on OpsProvider with generic dispatch:
apply(op, image, params) + supports(op) + a built-in registry {invert, scale}
+ register_op. ImageRef.op(name, **params) is the untyped transport seam;
invert()/scale() remain as built-in convenience. Adds OpNotSupported (carries
the capability name) so a node can fall back to the raw tier. An overlay now
extends the op vocabulary without touching OSS core.
Statically-typed op methods live on the secure-lib side (per guidance), not
here — core stays a generic seam.
Generic out-of-process enablement: the ExecutionPlan now ships the node's
module spec and ref-wrapped inputs, and dispatch receives the per-node host
runtime (refs/ctx/ops) so an external backend can execute the node elsewhere
and broker guest calls against the same ref table. In-process default ignores
all of it; legacy nodes unaffected (seam tests green).
Nodes operate on assets (image.invert()) and never receive buffers; compute
runs on the trusted plane via the OpsProvider seam. Raw buffer access becomes
a permissioned, discouraged escape hatch (raw(); forces dedicated tier under
the overlay). The execution seam wraps heavy inputs as refs for SDK_REFS
nodes and resolves output refs for downstream legacy nodes. The .pyi contract
no longer imports torch.
POC stand-ins (interface debt, ledgered in the overlay repo DEBT.md):
invert/scale enumerated on OpsProvider; duck-typed wrap_inputs; SDK_REFS
class-attr opt-in.
- execution.py: route V3 node dispatch through providers.execution_backend
with a behavior-preserving local_call closure (exact sync/async-task
semantics retained); bind per-node ctx+refs inside the invocation scope so
the concurrent-async path is correct. V1 nodes untouched. Default backend =
in-process => byte-identical to today.
- comfy_api/latest/_sdk_public.pyi: authoritative type contract for the secure
SDK (backend analog of the frontend comfy-api.d.ts): refs, ctx + domains,
ctx() accessor; host/overlay seam separated.
- custom_nodes/comfy_sdk_poc: SandboxInvert POC node authored to the SDK.
- tests-unit: seam regression (sync+async SDK nodes through the real engine;
provider-swap intercept). Verified PASS.
Establishes the open-source 'key Python API' for secure custom nodes:
- comfy_api/latest/_sdk.py: opaque typed refs (ImageRef/ModelRef/AssetRef...),
a brokered ctx surface (assets/progress/scratch/events/storage + stubs),
and a Providers registry (ExecutionBackend/CtxProvider/RefResolver) with
in-process DEFAULTS so OSS behaves exactly as today (a ref wraps the real
object; ctx is a passthrough; zero-copy, zero-overhead).
- load_overlay(): env-var (COMFY_OVERLAY_MODULE) path loader that lets a
separable, proprietary cloud overlay register isolated implementations at
the seam. Unset => pure OSS. Wired guarded into main.startup.
- comfy_api/v0_0_3: new API version exposing sdk; registered in version_list.
Nothing isolation-specific lives in OSS: this is the seam, not the engine.
* chore(assets): drop the unused asset_meta table from migration 0007
* fix(assets): drop asset_meta when downgrading a database that already created it
* docs(assets): clarify asset schema docstring
* test(assets): assert alembic and ORM index parity for the surviving asset tables
* test(assets): cover asset system state index parity
* refactor(tests): hoist migration-0007 test imports to module scope
* fix(assets): guard the hashing dependency and chain the real import error
* fix(assets): always resume background scanning when prompt handling fails
* fix(api): derive the assets feature flag from the selected manager
* fix(assets): paginate enrichment by id cursor so failures cannot starve or overflow the query
* refactor(assets): extract prompt_worker so its resume contract is testable in-process
* refactor(assets): test the blake3 import guard in-process instead of via subprocess
* fix(assets): advance the enrichment cursor only past rows the batch attempted
* fix(assets): track the scan pause across prompt worker iterations
* chore(assets): address review follow-ups in the hashing guard, feature flags, and pagination pin
* docs(assets): describe enrichment rows as attempted rather than selected
The cursor holds at the last row a batch actually attempted, so a pause ends a batch early and the rows behind it are selected again when the scan resumes. Only the attempt is capped at once per pass.
* test(api): derive the expected assets flag from the manager under test
The assertion asked for the flag with no argument, so it read the parameter default rather than anything the manager reported - in a test whose subject is the two agreeing. Passing the manager's own state keeps it honest if the setup ever yields an enabled manager.
* test(assets): let a broken prompt worker import fail instead of skipping
The fixture wrapped importlib.import_module in a bare except that called
pytest.skip, so a circular import, a missing dependency or a syntax error in
app/prompt_worker.py would retire all four resume-contract tests while CI
stayed green. The whole premise of extracting the module is that main.py can
import it, so an import failure has to be a collection error.
The CPU guard is genuinely load-bearing — comfy.model_management selects its
device at import time and a CUDA build with no driver raises there — so it is
kept, but as a precondition rather than an exception handler, matching the
args.cpu-before-import convention already used by the comfy_test and
comfy_api_nodes_test modules. Nothing is caught now.
* docs(assets): name the unattempted rows instead of the ones behind the cursor
The cursor moves forward through ascending ids, so 'the rows behind it' reads as the rows already passed - the opposite of what is selected again.
* fix(assets): only absorb duplicate-path races when seeding scanned assets
* fix(db): copy the legacy database inside the process lock
* fix(assets): drop watch-list entries on stat errors instead of aborting the scan
* fix(assets): clean up temp uploads on validation failures
Multipart parsing writes the uploaded bytes to a temporary file before it
validates the remaining form fields, so a request rejected after its file
part had already been read left the temp file and its uuid directory on
disk.
Routing those removals through delete_temp_file_if_exists also changes the
success path. The previous helper returned early when the temp file was
already gone, so it never reached the parent rmdir; the shared helper
attempts the rmdir unconditionally. Moving the upload to its destination
leaves the temp path absent, so a successful upload now also discards its
empty uuid directory, closing a pre-existing leak.
* fix(assets): keep updated_at stable on no-op renames
* docs(assets): make module docstrings and the rebuild warning truthful
* fix(assets): correct event-log status snapshots and failure telemetry
* test(assets): make the keyset tie-breaker and temp-exclusion tests falsifiable
* chore: comment cleanup
Comment-Gate: 3 quarantined
* fix(assets): keep unreadable filesystem metadata from failing a whole scan batch
* fix(assets): remove temporary uploads on non-UploadError failures
* fix(assets): preserve successful specs when a scan batch fault propagates
* test(db): drop the inert legacy-copy patch from the path preparation tests
prepare_file_db_path no longer copies the legacy database - that moved inside the process lock in _init_file_db - so patching copy_legacy_default_db here did nothing. Leaving it implied a side effect the function does not have, and would have masked one if it were reintroduced.
* fix(assets): report specs committed before a batch fault and preserve the fault itself
* fix(assets): distinguish partial batch insert failures
* refactor(assets): collapse the duplicated batch fault deferral into seed_asset_specs
* fix(assets): reject duplicate file parts instead of stranding the first upload
* refactor(assets): reap empty upload directories without importing the API layer
* fix(assets): emit the invalid-mtime event once per scan
Every other per-file emit on this scan path is gated -- mark_emitted(
"stat_failed:enrich"), "hash_discarded_modified", "hash_failed",
"enrich_failed" -- but scanner.invalid_mtime fired per file, so a restored
archive or a FAT volume of pre-epoch mtimes put one structured event per file
into the stream the closed vocabulary exists to keep parseable.
Counted and emitted once, carrying the count. seed_asset_specs receives no
_ScanProgress object and neither does insert_asset_specs above it, so routing
this through mark_emitted would mean changing both signatures plus the seeder
call site; the count form needs neither and the event is now strictly more
informative than N identical fieldless lines. The per-file logging.warning is
unchanged, and the emit stays inside seed_asset_specs so the static call-site
manifest still matches.
test_seed_skips_negative_fresh_mtime_with_warning_and_telemetry now pins the
full list of invalid_mtime lines to exactly ["... count=1"] instead of
asserting one such line exists -- a strictly stronger assertion, and the only
change the new field required.
* fix(assets): keep spec construction failures from wedging the watch list
get_name_and_tags_from_asset_path raises ValueError by contract when a path
stops resolving to a configured root, and it sat outside the guard, as did
compute_loader_path and mimetypes.guess_type. An escape skipped the
_WATCH_LIST[:] = remaining write at the end, so drained entries stayed on the
list and were re-attempted every tick while entries past the fault never
reached the increment _WATCH_SCAN_RETRIES needs to retire them. The list
wedged permanently.
Spec construction is now inside a guard that drops just the offending entry,
and the list write moved into a finally so no future escape can skip it. The
loop walks an iterator rather than the list, so the finally can put back the
entries it never reached instead of discarding them.
New event name rather than reusing one: scanner.watch_seed_failed is emitted
only when seed_asset_specs returns an error, and widening it to also mean
"never got as far as seeding" would make it lie -- a consumer treating it as a
database-health signal would get false positives from what is really a path
layout problem. scanner.watch_spec_failed is registered in ALLOWED_EVENTS and
in the static call-site manifest.
* refactor(assets): export the live-path conflict check as public API
scanner.py reached past the package's own re-export surface to import
_is_live_path_conflict directly out of records.py. The underscore said
module-private while the import said otherwise, and records.py deliberately
publishes its public names through app.assets.database.queries -- which the
same import block three lines above was already using.
The use is correct and unchanged; only the name and the route change. Renamed
to is_live_path_conflict, listed in the package __init__ import and __all__
alongside its siblings, and scanner.py now takes it from the package like
everything else it imports from there.
* docs(db): restore the rationale for locking before migration
Commit 1dbcdcd7 and the comment-cleanup pass 8205022f reduced this to "All
database reads and writes, including the legacy import, run under the lock",
dropping the part that did the work: upstream master locks after migrating and
justifies it with "Alembic uses its own connection, so we must wait until it's
done before locking -- otherwise our own lock blocks the migration". That is
false, the lock is on a separate <db>.lock file, and the surviving sentence
said nothing to stop a contributor "fixing" the ordering back.
Restored and adapted rather than pasted: the legacy copy and the db_exists
probe now happen inside the lock, which the original text predates, so both
are named in the list of things the ordering makes mutually exclusive.
* fix(assets): stop a scan on memory exhaustion instead of deferring it
MemoryError is an Exception, so the per-spec and per-batch handlers stored it alongside ordinary faults and carried on - allocating for every remaining spec and then every remaining batch while the process was already out of memory. Both handlers now let it through, and the scan records a failure and stops.
* docs(assets): document the prune failure response and its None result
The route gained a 500 PRUNE_FAILED branch and the seeder method gained a None return, both so a prune that did not run cannot be reported as a clean one. Neither contract was written down.
* test(assets): assert the surviving spec count after a propagated fault
* docs(db): shorten the lock-ordering comment while keeping its rationale
* docs(tests): drop the cross-module justification from the import-order comment
* test(assets): restore the cpu flag after the guarded prompt worker import
* test(assets): restore the cpu flag even when the prompt worker import fails
* fix(assets): drop asset_meta in a new migration instead of editing 0007
0007 shipped in v0.36.0, v0.37.0 and v0.37.1, so editing it would leave two
installs at that revision with different schemas depending on when they
upgraded. Restore 0007 to its released form and drop the unused asset_meta
table in 0008 instead.
Nothing reads or writes asset_meta; asset metadata lives in the JSON columns
on assets. 0008 downgrades by recreating the table and its four indexes
exactly as 0007 creates them.
* refactor(assets): keep prompt_worker in main.py
Custom nodes may reference main.prompt_worker, and tests can already
import main (test_db_init_locking does), so the function stays where it
was. Its body keeps the resume-on-failure handling and the pause flag
that persists across loop iterations; the tests now call
main.prompt_worker.
* fix(assets): report the selected manager through the existing feature flags
* fix(assets): record a failed prune in the scan status
A failed prune left the scan's error list empty, so the run looked clean.
Record it instead and let discovery continue as before.
* test(assets): test the feature flag API directly instead of copying the startup write
Both parity tests wrote SERVER_FEATURE_FLAGS["assets"] themselves, so they
passed whatever startup did. Pin the one real claim - a no-argument
get_server_features() reports the flag - in the feature flags tests, and drop
the disabled case, which default_asset_manager already covers.
* Bring #16486's watch-list batching and seed logging into this branch
The merge before this commit is `git merge -X ours origin/master`: in
conflicting hunks it keeps this branch's side. This commit ports what
#16486 changed in those hunks.
- tick_watch_list takes no session. It collects settled entries and seeds
them in one batch through insert_asset_specs, keeping this branch's
per-entry stat and spec handling and the finally that always rewrites
the watch list. A failed seed is reported once per batch, since the
batch only returns its first error.
- seed_asset_specs no longer warns again for a skipped spec;
observe_asset_specs already logged why.
- Tests call tick_watch_list() without a session, bind the write session
where they seed for real, and fake insert_asset_specs with its
(created, error) return. The mid-drain fault now comes from stat,
because seeding runs after the loop.
* Note where the assets core flag is set and why prompt_worker catches BaseException
---------
Co-authored-by: guill <jacob.e.segal@gmail.com>
* Take the SQLite write lock up front for scan and output-registration writes
The scanner's seeding and reference sync, and executed-output registration,
write through a separate engine whose transactions open with BEGIN
IMMEDIATE, and the database runs in WAL mode. On those paths stat, hashing
and metadata extraction now happen before the write transaction opens, and
reference-sync results are applied only to rows unchanged since they were
observed. busy_timeout stays at pysqlite's 5s default.
A fast-scan batch now commits as one transaction, so an unexpected error
partway through discards the whole batch; the next scan recreates it.
Enrichment, verification, uploads and tagging still write through the
existing sessions.
Migration backups use SQLite's backup API, and a legacy database is
checkpointed before it is relocated, since in WAL mode committed pages can
live in the -wal file that a plain file copy misses.
* Skip relocating a legacy database whose WAL cannot be checkpointed
The checkpoint can report busy without raising; moving the file then would
leave committed pages behind in the -wal. Also keep the source's file mode
on SQLite backups, as the plain copy did.
* Replace run_write_txn with a create_write_session factory
Write paths open the write engine's session the same way the rest of the
code opens create_session(), and commit explicitly. Executed-output
registration reads the new record's fields before committing, so expiry
does not reload them in a second write transaction.
* Document the scanner's pre-transaction observation types
* Document that write sessions must not nest
* Warn that a nested write session looks like lock contention
* Seed a hashed spec with the stat its hash was verified against
* Bound the SQLite backup so a locked destination cannot hang startup
* Time out a backup only while it is blocked
* Commit each drained entry before hashing the next
drain_pending_verifications and drain_transition_queue ran every entry in
one session, so an entry that wrote (marking a vanished file missing, say)
left a transaction open while the next entry's file was stat'ed and
hashed. Commit at the top of each entry instead, so the hash runs with no
transaction open.
* Seed settled watch-list entries through insert_asset_specs
tick_watch_list seeded each settled file through seed_asset_specs in the
caller's session, where the enrich phase still held drain_pending's
writes open, and each seed ran in a deferred savepoint that reads before
it writes. Collect the settled specs and hand them to insert_asset_specs,
which stats and hashes before opening one write session. The seeder
commits the pending verifications before ticking, so no transaction is
open while the watch list stats or waits for the write lock.
* Read upload metadata before the claim transaction
_create_upload_record read the file for system metadata after the
content claim had opened a write transaction. Callers now extract the
metadata before opening their session (or, when reusing content, before
claiming it) and pass it in. The claim's own stat re-check stays inside
the transaction: that is what makes the claim sound.
* Keep the upgrade error when restoring the backup also fails
If restoring the pre-upgrade backup raised, that exception replaced the
upgrade's, and the backup's location was never logged. Log the upgrade
error first, log where the pre-upgrade copy is kept if the restore or its
cleanup fails, and re-raise the upgrade error either way.
* Note the drains' session requirement and word the restore log for either failure
The drains' per-entry commit only leaves no transaction open on a
create_session() session. The restore log now covers a failed backup
removal as well as a failed restore. The watch-list admission test
patches insert_asset_specs, the seam tick_watch_list now calls.
* Log the real error when a seed spec cannot be read
observe_asset_specs treated every OSError as a vanished file, so a
permission error or an I/O error on a file that still exists was
reported only as "Skipping vanished asset during scan". A missing file
is still handled as before; any other OSError is now also logged through
_log_scan_error before the spec is skipped. The vanished-path test's
fake now raises FileNotFoundError, the error a vanished file produces.
* Report an unreadable seed spec once, without also calling it vanished
* Skip a pending verification whose row changed while it was hashed
Committing per entry means no lock is held while the file is hashed, so
another writer can retire or replace the row in that time. The drain
would then store a hash on a missing row, or split it and attach a
record to the other writer's row. Re-read the row after hashing and skip
it unless it is still live with the hash, size and mtime it was loaded
with, as apply_reference_observations does.
* Remove the stored upload file in the reupload claim test
The test cleaned up its temp files but left the first upload's stored
file in the output directory.
* Take the SQLite write lock up front for scan and output-registration writes
The scanner's seeding and reference sync, and executed-output registration,
write through a separate engine whose transactions open with BEGIN
IMMEDIATE, and the database runs in WAL mode. On those paths stat, hashing
and metadata extraction now happen before the write transaction opens, and
reference-sync results are applied only to rows unchanged since they were
observed. busy_timeout stays at pysqlite's 5s default.
A fast-scan batch now commits as one transaction, so an unexpected error
partway through discards the whole batch; the next scan recreates it.
Enrichment, verification, uploads and tagging still write through the
existing sessions.
Migration backups use SQLite's backup API, and a legacy database is
checkpointed before it is relocated, since in WAL mode committed pages can
live in the -wal file that a plain file copy misses.
* Skip relocating a legacy database whose WAL cannot be checkpointed
The checkpoint can report busy without raising; moving the file then would
leave committed pages behind in the -wal. Also keep the source's file mode
on SQLite backups, as the plain copy did.
* Replace run_write_txn with a create_write_session factory
Write paths open the write engine's session the same way the rest of the
code opens create_session(), and commit explicitly. Executed-output
registration reads the new record's fields before committing, so expiry
does not reload them in a second write transaction.
* Document the scanner's pre-transaction observation types
* Document that write sessions must not nest
* Warn that a nested write session looks like lock contention
* Seed a hashed spec with the stat its hash was verified against
* Bound the SQLite backup so a locked destination cannot hang startup
* Time out a backup only while it is blocked
Upscale models loaded via Spandrel expect exactly 3 input channels, so a 4-channel IMAGE crashed in the model's first conv. Split off the alpha channel before upscaling, then resize it to the output resolution and concatenate it back so transparency survives the upscale.
Fixes#16499