mirror of
https://github.com/qdrant/qdrant.git
synced 2026-10-02 19:07:48 -05:00
* feat: global quota API Memory and disk are node-wide resources, so configuring their thresholds per collection through strict mode makes little sense. Move them behind a single cluster-wide `QuotaManager`. The quota config is seeded from `storage.quotas` in the settings (and so from env vars), overridden by `quota.json` in the storage directory, and updated cluster-wide through a new `SetQuotaConfig` consensus operation which rewrites that file on every peer. Raft snapshots carry it too, so a peer that joins by snapshot picks it up. Quotas are enforced wherever the strict mode memory and disk checks used to run, but no longer gated behind `strict_mode.enabled`: a value set in an enabled strict mode config still wins per resource, the quota is the default. Rejections name both the condition that tripped and the config that governs it. `GET /quotas` reports the config plus current utilization to global read users; `PUT /quotas` replaces it for global manage users. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: cover the quota endpoints in the API consistency checks `test_all_rest_endpoints_are_covered` and the OpenAPI endpoint count both break on any new REST endpoint. Add `GET`/`PUT /quotas` to `ACTION_ACCESS` with their JWT access tests, and bump the expected API count. The quota endpoints stay out of `REST_ENDPOINT_WHITELIST`: that list is for data-plane endpoints reported per-endpoint in metrics. Also add a Raft snapshot CBOR compatibility test — snapshots are exchanged between peers of different versions during a rolling upgrade, so `quota_config` must be absent-tolerant in both directions. Review feedback: persist through `SaveOnDisk`, which already implements the write-before-swap protocol this was doing by hand; validate the config at both persistence boundaries, since a hand-edited quota file or a config arriving through consensus does not pass the REST handler's validation, and a `0%` limit would reject every update forever. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test: assert seeding a quota from invalid settings persists nothing Follow-up to review feedback claiming `SaveOnDisk::load_or_init` writes the init value before it is validated. It does not — only `SaveOnDisk::new` persists — but the property matters: were seeding to persist first, invalid settings would leave a `quota.json` that fails validation on every subsequent start, and the node could only be recovered by deleting it by hand. Pin it down with a test. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * refactor: make QuotaManager the single reader of memory and disk The quota checks measured memory and disk themselves, while the optimizer and the WAL disk watcher each called `fs4::available_space` behind their own ad-hoc caches. Fold all of it into QuotaManager: it owns the readings, the freshness policy, and the limits they are compared against. Moves the module to `lib/shard`, since the optimizer sits below `storage` and has to reach it; `storage::quota` re-exports it, so consensus, the `/quotas` API and StorageConfig are unchanged. The manager is installed as a process singleton by TableOfContent, ahead of loading any collection. - Callers hand in QuotaLimits overrides instead of a StrictModeConfig, and an override can now only tighten. A collection-level admin could raise `max_disk_usage_percent` past a cluster-wide limit that needed global manage rights to set; ties resolve to the quota so the rejection names the knob that actually has to change. - Measurements are cached for 5s, but a reading at or above its limit is never reused: a rejected client retries, and freeing the resource has to take effect on the next request rather than a TTL later. - `fits_on_disk` sizes an optimization against physical free space only, never the configured limits. Optimizations are what free a full disk, so the quota must not be what stops one. - `percent_of` widens to u128 instead of saturating the multiply, which under-reported utilization (failing open) above ~184 PB. - StorageConfig::quotas is optional; absent means no quota is enforced. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat: don't recover dead replicas onto a node at a resource limit Recovering a dead replica pulls a whole copy of its shard onto this node. If it is already at its memory or disk quota that transfer cannot finish, and starting it only pushes the node further past the limit. Skip it and reconsider on a later sync, once the resource frees up. Adds QuotaManager::check_capacity for work that lands bytes here without being an update. Unlike fits_on_disk the configured limits do apply: taking on a replica is not what frees a full node, so there is no deadlock to avoid by letting it through. The check is hoisted out of the per-shard loop because a node over its limit re-measures on every call, so checking per dead shard would cost a statvfs each. It is free when no quota is configured. Also trims the comments across the quota module, which had grown well past what the code needs. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test: drop trivial and duplicated quota tests Six tests removed, ~140 lines, with no loss of coverage: - a_rejection_names_the_knob_that_has_to_change asserted that a format! contains its own literals; the message is covered end-to-end by the override test and by test_global_quota.py. - a_node_over_its_quota_has_no_capacity_to_take_on_a_replica was 30 lines for check_capacity, a one-line delegation to check_update the test above it already calls. - a_rejecting_measurement_is_never_served_from_the_cache duplicated the meter test, which proves the same rule with an injected reader instead of inferring it from the real filesystem. - free_space_is_reported_without_enforcing_anything covered a one-line accessor, and its point is what the fits_on_disk test is for. - The two resolve tests and the three meter tests each collapse into one. DiskFit::Unknown keeps its coverage as two lines inside the fits_on_disk test rather than its own fixture. The snapshot compat pair becomes one test: the second only asserted cluster_metadata.is_empty(), which says nothing about quotas — the real check was the deserialize, now an expect that states it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: re-measure free space as the disk fills, and drop a Windows-only assert Two CI failures, both from this branch. e2e test_low_disk: the DiskUsageWatcher I replaced escalated to checking on every call once free space fell below 512 MB. Folding it into the quota manager lost that — available_bytes passed no limit, so a reading was reused for the full 5s however little space was left. On a disk filling as fast as that test fills it, 5s blind is enough to actually run out and the WAL write dies instead of returning "No space left on device". available_bytes now takes a watch_below level and never reuses a reading under it, which is what the old ladder was expressing. The watcher passes max(min_free, 512 MB), so the escalation point is back; above it the 5s cache still costs fewer syscalls than the old 128-call ladder. fits_on_disk gets the same rule by passing required_bytes, so a merge that does not fit re-checks rather than sitting on a stale sample. Windows: fits_on_disk on a missing path was asserted to be Unknown, but GetDiskFreeSpaceEx resolves up to the containing drive and succeeds — as common::disk_usage's own test documents. Dropped; the branch is a two-line else and is not portably reachable. Also renames an_optimization_is_sized_against_the_disk_not_the_quota, which needed explaining to be understood. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat: report the quota config in telemetry Reads it from the quota manager rather than the settings, so it is the config the node is actually enforcing: a peer that missed a consensus update reports what it is applying, not what the cluster agreed on. Gated on global access, the same access `GET /quotas` requires, and left out of `PeerTelemetry` — a quota is per-node state, so each peer reports its own. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: regenerate OpenAPI, and cover the quota in the telemetry key sets Two CI failures from the previous commit. Referencing QuotaConfig from TelemetryData moves its definition earlier in `components/schemas`, because TelemetryData is generated ahead of QuotaStatus. Regenerated rather than hand-patched, so the schema is a pure move. test_telemetry_detail asserts the exact set of top-level telemetry keys. The quota is reported at every level, including 0 — it is three scalars, it is the default the endpoint serves, and it is what explains an update being rejected — so both key sets gain it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat: remove `max_disk_usage_percent` from strict mode Disk is a node-wide resource, so a per-collection percentage of it never meant anything a caller could act on: the limit describes how full the *node* is, and which collection the write happens to target has nothing to do with it. The global quota is where it belongs. It shipped in 1.18.2 without documentation, so this drops it outright rather than deprecating. Removal is soft in every direction: StrictModeConfig has no `deny_unknown_fields`, so a client still sending it gets it ignored rather than a 400, and the same struct deserializes the persisted collection config, so collections created on 1.18.2+ keep loading. Proto field 22 is reserved so the number is never reused. The e2e test becomes a quota test — the fixture and the timing are the interesting parts and they carry over unchanged; only how the threshold is configured differs. `max_resident_memory_percent` was documented and stays for now. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * refactor: enforce the strict mode memory limit outside the quota `max_resident_memory_percent` was folded into the quota as an override, which meant the quota check had to know about strict mode, and retiring the setting would mean unpicking `EffectiveLimit` and `LimitSource` from the resolution logic. It is now a check of its own in `verification/mod.rs`, next to the strict mode checks it belongs with, borrowing only the measurement from the quota manager — which stays the node's single reader of process memory, so both checks still share one reading. Deleting the setting later is deleting one function and its one caller. `QuotaManager::check_update` takes no arguments and consults the quota alone. A collection can still only tighten the limit for itself, because its own check runs in addition rather than in place of the quota's, and each rejection now names the config that has to change without having to carry a `LimitSource` to say so. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * refactor: enforce the quota on the update path, not in strict mode The quota check sat inside `check_strict_mode_toc_batch` only because that was the one place holding the collection's strict mode config. It doesn't need one any more, and the placement had a real cost: coverage depended on each handler remembering to ask for a strict mode check, and four of the internal update RPCs do — `sync_internal`, which moves the most bytes onto a node, does not. It now runs in `Collection::update_from_client` and `update_from_peer`, which every update passes through. `update_from_client` checks ahead of the shard split, so an operation is accepted or refused whole rather than landing on some shards and being refused by others. Classification moves with it, from ~10 `consumes_memory` impls on request DTOs to one exhaustive `CollectionUpdateOperations::consumes_quota`. The internal enum has variants — raw upserts, conditional upserts, the syncs — that have no client-facing request type, so per-DTO impls structurally could not classify them. Shard-transfer syncs stay excluded, as they are today: a transfer is sized up once before it starts, and refusing its batches partway abandons work that is nearly done only for it to restart from the beginning. Index and named-vector creation reach shards through consensus, past this check — a peer must not refuse what the cluster agreed to — so they keep their pre-consensus check, now against the quota manager directly. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat: deprecate `max_resident_memory_percent` in strict mode Same reason the disk threshold went: memory is node-wide, so a per-collection percentage of it caps how full the *node* is, which has nothing to do with which collection is being written to. The node-wide quota caps it once for everything. Unlike the disk threshold this one shipped documented, in 1.18.0, so it keeps working — as a limit a collection can tighten for itself, never lift — and gets the usual markers: `#[deprecated]` on both Rust structs, `[deprecated = true]` on proto field 21, and `deprecated: true` in the OpenAPI schema, which schemars derives from the attribute. The note names 1.21 as the removal. Recording a version matters here: the audit in docs/plans/overdue-deprecations.md found that this repo has never written a removal deadline down, and members of the 1.15.0 deprecation batch are still in tree. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: reconcile the quota readers with #9891 #9891 landed effective (cgroup) figures in telemetry while this branch was making QuotaManager the single reader of memory and disk. Two collisions, neither of which git sees. `segment::utils::mem::total_memory_bytes` is now a shared accessor with a 5s TTL, so a cgroup resize is picked up. The quota module had its own `OnceLock` copy that froze the value at startup — exactly what #9891 set out to fix — so it delegates to the shared one instead. Telemetry's new `disk_size` called `common::disk_usage::disk_usage` directly. That reader lost its TTL cache on this branch when the caching moved into the quota manager's meter, so it would have taken an uncached `statvfs` on every telemetry request, and it put a second disk reader back in the tree. It goes through `QuotaManager::disk_capacity_bytes` now, sharing the reading the quota check already takes. Verified it still reports the storage filesystem, matching `df`. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * refactor: split the quota manager by what each half does `manager.rs` had grown to 450 lines holding three separate jobs: owning the config and its file, taking the readings, and comparing one against the other. - `manager/store.rs` — the `Store` enum, `QUOTA_CONFIG_FILE`, and config validation, which is now the store's own business rather than something every caller has to remember to do first. - `manager/measure.rs` — every reading, and `DiskFit`. The "nothing else calls `statvfs` or reads process RSS" claim is now checkable by looking at one file. - `manager/enforce.rs` — `check_update` / `check_capacity` and the threshold comparison. - `manager/mod.rs` — the struct, its construction, and the config accessors: what a reader needs to see first. Tests move with their subject. No behaviour change: `set_config` used to validate before delegating to the store, and now the store validates on write, which is the same order of operations from the outside. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: count copy-on-write deletes toward the quota Dropping a vector or a payload key does not free anything on its own: copy-on-write rewrites the point to produce the version without that field, so storage grows first and is only reclaimed once the optimizer gets to it. Gating those as if they were reclaiming space let a full node keep taking writes that make it fuller. Deleting whole points stays exempt. That is the one operation that has to work on a node at its limit, or there is no way back under it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * docs: say that the quota's reported usage is per node `GET /quotas` returns one cluster-wide config and one set of utilization figures, which reads as though both describe the cluster. They do not: memory and disk are node-local, so `usage` is whatever the peer that served the request is seeing, and a peer under its limit says nothing about the others. Also corrects `resident_memory_percent`, which claimed to be a share of total system memory. It is a share of the memory available to the process, which under a cgroup is the limit rather than the host's RAM. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat: treat a node over its quota as a failed replica, not a bad request A quota rejection described the request as invalid (400) and was classified non-transient, which is how the replica set recognises errors that every replica would produce alike. A quota is the opposite: the input is fine and the answer depends on which machine you ask. On the default `wait=false` path that combination silently dropped the write — `update.rs` only deactivates transient failures when nothing completed — leaving the replica Active and permanently missing data its co-replicas had. It is now `InsufficientStorage`, transient, HTTP 507 / gRPC `ResourceExhausted`. So a node that is out of room is handled like one that is offline: - last active replica, or every replica over quota: nothing could take the write, and the client is told the cluster is out of room. - more than one replica: the full node is deactivated through the same path a dead peer takes, and the update stands if enough replicas accepted it. `check_capacity` already keeps recovery off that node until it has room. The check also moves off `update_from_client`, which applied the coordinator's own limit to the whole operation even when it held no replica of the shards being written. Each replica set now gates its own local write and records the refusal as a failure of this peer, so a node only ever answers for itself. `ResourceExhausted` is shared with rate limiting, and the reverse conversion mapped it straight to `RateLimitExceeded` — a forwarded rejection came back as 429. Statuses now carry a marker so the two stay distinguishable across the wire. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat: report quota pressure per node, and across the cluster A quota is node-local, so finding out which node has hit one meant asking each of them in turn — and nothing at all showed up in monitoring. `/metrics` gains a `quota_exceeded` gauge for the local node. It is emitted only while the quota is enabled: with it off the value would be a constant 0 that says nothing about the node, and an alert built on it would go quiet rather than fire if someone disabled the quota. Telemetry's `quota` field carries the same verdict alongside the config, since that is where the metric is derived from. `GET /quotas` now answers for the whole cluster. A new `GetQuotaUsage` RPC on the internal `QdrantInternal` service returns what one peer is using, and the handler fans it out to every known peer in parallel. Peers that do not answer are left out rather than failing the request — the nodes that are out of room are exactly the ones most likely to time out, and a partial answer still names them. Outside distributed mode the field is absent rather than a map of one. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat: report the quota metric per resource `quota_exceeded` was one flag for the whole node, which does not say what to go and fix — disk is freed by deleting or optimizing, memory by unloading. It now carries a `resource` label: quota_exceeded{resource="memory"} 0 quota_exceeded{resource="disk"} 1 A resource with no limit gets no series at all, for the same reason the metric is absent while the quota is disabled: a series that can never reach 1 reads as healthy and would quietly carry an alert that cannot fire. `QuotaManager::exceeded` returns the per-resource verdict, with `None` for a resource this node does not cap. Telemetry reports the same breakdown, since the metric is derived from it. The peer usage RPC keeps a single flag — it sits next to both percentages, so it only has to answer "is this peer refusing writes". Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: drop a no-op error conversion the linter caught `check_global_access` already returns a `StorageError`, so mapping it through `StorageError::from` converted the type to itself and tripped `clippy::useless_conversion`. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat: hold a tripped quota until usage clears a release margin A resource resting on its limit crosses it in both directions on the noise between two readings, and each crossing is expensive: the node refuses a write, its replica is deactivated, usage dips, recovery starts sending a whole shard copy back, and the arriving data pushes it over again. The loop sustains itself, and every lap costs a shard transfer. A limit now trips at its configured value but only clears once usage has fallen 5 percentage points below it, so the crossing has to be real. The margin is floored at 1%, since a limit smaller than the margin would otherwise be impossible to fall back under and would strand the node. The verdict is carried on the manager rather than recomputed, which makes it the thing reporting shows: expect `exceeded` to be set while the utilization next to it is already back under the limit. Rejections say so too, rather than claiming a limit that is no longer exceeded: Disk usage is at 87% of total capacity. It reached the configured limit of 90% and has to fall below 85% before this node takes writes again. Changing the config clears the verdicts. New limits are a deliberate act, and should not be held back by the margin of a limit that no longer exists. Both resources are now evaluated on every check instead of stopping at the first failure, so a verdict is never left behind reporting a reading that has since been superseded. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat: make the quota release margin configurable 5 points is a guess about how noisy a deployment's usage is, which is not something one number can be right about: a node whose disk moves in gigabyte steps needs a wider margin than one that creeps, and an operator who wants the old flip-on-every-reading behaviour should be able to ask for it. `release_margin_percent` joins the rest of the quota config, so it seeds from `QDRANT__STORAGE__QUOTAS__RELEASE_MARGIN_PERCENT`, replicates through consensus, and changes with `PUT /quotas`. Defaults to 5 and is filled in when a request omits it, so it always answers with the margin actually in force rather than leaving the caller to assume one. `0` releases as soon as usage is back under the limit. `QuotaConfig` grows a hand-written `Default` for it, since deriving one would have quietly defaulted the margin to 0 and disabled the hysteresis for anyone constructing a config in code. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * refactor: leave the release margin unset by default, and hold verdicts in atomics `release_margin_percent` is `null` unless someone sets it, rather than materialising 5 into every config. A quota written today then does not pin a number a later release may want to revise, and `{"enabled": false}` still round-trips as itself. `QuotaConfig::limits` resolves it, next to `enabled`, so enforcement never sees the unset case. The verdicts move from a `Mutex<QuotaExceeded>` to one `AtomicBool` per resource. They are judged independently and nothing reads them as a pair, so the lock only added contention to the path every update takes; a verdict that races a concurrent check is re-decided by the next one from a fresh reading. That also drops the tri-state. Only "was this over its limit" has to survive between checks — whether a resource is enforced at all follows from the config and the reading, so it is derived when reporting rather than stored. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * docs: drop a comment arguing with a design that was never here Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
473 lines
18 KiB
YAML
473 lines
18 KiB
YAML
log_level: INFO
|
|
|
|
# Logging configuration
|
|
# Qdrant logs to stdout. You may configure to also write logs to a file on disk.
|
|
# Be aware that this file may grow indefinitely.
|
|
# logger:
|
|
# # Logging format, supports `text` and `json`
|
|
# format: text
|
|
# on_disk:
|
|
# enabled: true
|
|
# log_file: path/to/log/file.log
|
|
# log_level: INFO
|
|
# # Logging format, supports `text` and `json`
|
|
# format: text
|
|
# buffer_size_bytes: 1024
|
|
|
|
storage:
|
|
# Where to store all the data
|
|
storage_path: ./storage
|
|
|
|
# Where to store snapshots
|
|
snapshots_path: ./snapshots
|
|
|
|
snapshots_config:
|
|
# "local" or "s3" - where to store snapshots
|
|
snapshots_storage: local
|
|
# s3_config:
|
|
# bucket: ""
|
|
# region: ""
|
|
# access_key: ""
|
|
# secret_key: ""
|
|
|
|
# Where to store temporary files
|
|
# If null, temporary snapshots are stored in: storage/snapshots_temp/
|
|
temp_path: null
|
|
|
|
# Deprecated: use `payload.memory` instead.
|
|
# If true - point payloads will not be stored in memory.
|
|
# It will be read from the disk every time it is requested.
|
|
# This setting saves RAM by (slightly) increasing the response time.
|
|
# Note: those payload values that are involved in filtering and are indexed - remain in RAM.
|
|
#
|
|
# Default: true
|
|
on_disk_payload: true
|
|
|
|
# Default payload storage configuration for newly created collections.
|
|
# Overrides the deprecated `on_disk_payload` flag if both are set.
|
|
# payload:
|
|
# # Memory placement of the payload storage: cold or cached.
|
|
# memory: cold
|
|
|
|
# Load-time memory mode. Only affects how segments are loaded on startup;
|
|
# does not modify any persisted configuration. Intended as a recovery knob
|
|
# when a node crash-loops on out-of-memory.
|
|
#
|
|
# Options:
|
|
# - disabled (default): load segments as persisted.
|
|
# - no_resident: downgrade components to their on-disk variants where
|
|
# possible — quantization loads as if always_ram=false, payload field
|
|
# indexes as if on_disk=true, payload storage as mmap (not populated).
|
|
# - no_populate: same as no_resident, plus skip mmap prefault on load for
|
|
# vectors, HNSW graph and payload storage.
|
|
#low_memory_mode: disabled
|
|
|
|
# Maximum number of concurrent updates to shard replicas
|
|
# If `null` - maximum concurrency is used.
|
|
update_concurrency: null
|
|
|
|
# Write-ahead-log related configuration
|
|
wal:
|
|
# Size of a single WAL segment
|
|
wal_capacity_mb: 32
|
|
|
|
# Number of WAL segments to create ahead of actual data requirement
|
|
wal_segments_ahead: 0
|
|
|
|
# Normal node - receives all updates and answers all queries
|
|
node_type: "Normal"
|
|
|
|
# Listener node - receives all updates, but does not answer search/read queries
|
|
# Useful for setting up a dedicated backup node
|
|
# node_type: "Listener"
|
|
|
|
performance:
|
|
# Number of parallel threads used for search operations. If 0 - auto selection.
|
|
max_search_threads: 0
|
|
|
|
# CPU budget, how many CPUs (threads) to allocate for an optimization job.
|
|
# If 0 - auto selection, keep 1 or more CPUs unallocated depending on CPU size
|
|
# If negative - subtract this number of CPUs from the available CPUs.
|
|
# If positive - use this exact number of CPUs.
|
|
optimizer_cpu_budget: 0
|
|
|
|
# Prevent DDoS of too many concurrent updates in distributed mode.
|
|
# One external update usually triggers multiple internal updates, which breaks internal
|
|
# timings. For example, the health check timing and consensus timing.
|
|
# If null - auto selection.
|
|
update_rate_limit: null
|
|
|
|
# Limit for number of incoming automatic shard transfers per collection on this node, does not affect user-requested transfers.
|
|
# The same value should be used on all nodes in a cluster.
|
|
# Default is to allow 1 transfer.
|
|
# If null - allow unlimited transfers.
|
|
#incoming_shard_transfers_limit: 1
|
|
|
|
# Limit for number of outgoing automatic shard transfers per collection on this node, does not affect user-requested transfers.
|
|
# The same value should be used on all nodes in a cluster.
|
|
# Default is to allow 1 transfer.
|
|
# If null - allow unlimited transfers.
|
|
#outgoing_shard_transfers_limit: 1
|
|
|
|
# Enable async scorer which uses io_uring when rescoring.
|
|
# Only supported on Linux, must be enabled in your kernel.
|
|
# See: <https://qdrant.tech/articles/io_uring/#and-what-about-qdrant>
|
|
#async_scorer: false
|
|
|
|
# Whether components readable through either a memory mapping or io_uring should use
|
|
# io_uring. Only has an effect on Linux.
|
|
#
|
|
# - unset (default): the immutable vector storages follow `async_scorer`, nothing else
|
|
# uses io_uring.
|
|
# - "disabled": no component uses io_uring.
|
|
# - "auto": use io_uring for components with a `cold` memory placement, where reads hit
|
|
# the disk. Components meant to sit in RAM keep using mmap, which is faster there.
|
|
#io_uring: disabled
|
|
|
|
# Maximum number of collections to load concurrently.
|
|
#max_concurrent_collection_loads: 1
|
|
# Maximum number of local shards to load concurrently when loading a collection.
|
|
#max_concurrent_shard_loads: 1
|
|
# Maximum number of segments to load concurrently when loading a local shard.
|
|
#max_concurrent_segment_loads: 8
|
|
|
|
optimizers:
|
|
# The minimal fraction of deleted vectors in a segment, required to perform segment optimization
|
|
deleted_threshold: 0.2
|
|
|
|
# The minimal number of vectors in a segment, required to perform segment optimization
|
|
vacuum_min_vector_number: 1000
|
|
|
|
# Target amount of segments optimizer will try to keep.
|
|
# Real amount of segments may vary depending on multiple parameters:
|
|
# - Amount of stored points
|
|
# - Current write RPS
|
|
#
|
|
# It is recommended to select default number of segments as a factor of the number of search threads,
|
|
# so that each segment would be handled evenly by one of the threads.
|
|
# If `default_segment_number = 0`, will be automatically selected by the number of available CPUs
|
|
default_segment_number: 0
|
|
|
|
# Do not create segments larger this size (in KiloBytes).
|
|
# Large segments might require disproportionately long indexation times,
|
|
# therefore it makes sense to limit the size of segments.
|
|
#
|
|
# If indexation speed have more priority for your - make this parameter lower.
|
|
# If search speed is more important - make this parameter higher.
|
|
# Note: 1Kb = 1 vector of size 256
|
|
# If not set, will be automatically selected considering the number of available CPUs.
|
|
max_segment_size_kb: null
|
|
|
|
# Maximum size (in KiloBytes) of vectors allowed for plain index.
|
|
# Default value based on experiments and observations.
|
|
# Note: 1Kb = 1 vector of size 256
|
|
# To explicitly disable vector indexing, set to `0`.
|
|
# If not set, the default value will be used.
|
|
indexing_threshold_kb: 10000
|
|
|
|
# Interval between forced flushes.
|
|
flush_interval_sec: 5
|
|
|
|
# Max number of threads (jobs) for running optimizations per shard.
|
|
# Note: each optimization job will also use `max_indexing_threads` threads by itself for index building.
|
|
# If null - have no limit and choose dynamically to saturate CPU.
|
|
# If 0 - no optimization threads, optimizations will be disabled.
|
|
max_optimization_threads: null
|
|
|
|
# This section has the same options as 'optimizers' above. All values specified here will overwrite the collections
|
|
# optimizers configs regardless of the config above and the options specified at collection creation.
|
|
#optimizers_overwrite:
|
|
# deleted_threshold: 0.2
|
|
# vacuum_min_vector_number: 1000
|
|
# default_segment_number: 0
|
|
# max_segment_size_kb: null
|
|
# indexing_threshold_kb: 10000
|
|
# flush_interval_sec: 5
|
|
# max_optimization_threads: null
|
|
|
|
# Default parameters of HNSW Index. Could be overridden for each collection or named vector individually
|
|
hnsw_index:
|
|
# Number of edges per node in the index graph. Larger the value - more accurate the search, more space required.
|
|
m: 16
|
|
|
|
# Number of neighbours to consider during the index building. Larger the value - more accurate the search, more time required to build index.
|
|
ef_construct: 100
|
|
|
|
# Minimal size threshold (in KiloBytes) below which full-scan is preferred over HNSW search.
|
|
# This measures the total size of vectors being queried against.
|
|
# When the maximum estimated amount of points that a condition satisfies is smaller than
|
|
# `full_scan_threshold_kb`, the query planner will use full-scan search instead of HNSW index
|
|
# traversal for better performance.
|
|
# Note: 1Kb = 1 vector of size 256
|
|
full_scan_threshold_kb: 10000
|
|
|
|
# Number of parallel threads used for background index building.
|
|
# If 0 - automatically select.
|
|
# Best to keep between 8 and 16 to prevent likelihood of building broken/inefficient HNSW graphs.
|
|
# On small CPUs, less threads are used.
|
|
max_indexing_threads: 0
|
|
|
|
# Deprecated: use `memory` instead.
|
|
# Store HNSW index on disk. If set to false, index will be stored in RAM. Default: false
|
|
on_disk: false
|
|
|
|
# Memory placement of the HNSW index: cold, cached or pinned.
|
|
# Overrides the deprecated `on_disk` flag if both are set.
|
|
# memory: cached
|
|
|
|
# Custom M param for hnsw graph built for payload index. If not set, default M will be used.
|
|
payload_m: null
|
|
|
|
# Default shard transfer method to use if none is defined.
|
|
# If null - don't have a shard transfer preference, choose automatically.
|
|
# If stream_records, snapshot or wal_delta - prefer this specific method.
|
|
# More info: https://qdrant.tech/documentation/guides/distributed_deployment/#shard-transfer-method
|
|
shard_transfer_method: null
|
|
|
|
# Default parameters for collections
|
|
collection:
|
|
# Number of replicas of each shard that network tries to maintain
|
|
replication_factor: 1
|
|
|
|
# How many replicas should apply the operation for us to consider it successful
|
|
write_consistency_factor: 1
|
|
|
|
# Default parameters for vectors.
|
|
vectors:
|
|
# Deprecated: use `memory` instead.
|
|
# Whether vectors should be stored in memory or on disk.
|
|
on_disk: null
|
|
|
|
# Memory placement of the vector storage: cold or cached.
|
|
# Overrides the deprecated `on_disk` flag if both are set.
|
|
# memory: null
|
|
|
|
# shard_number_per_node: 1
|
|
|
|
# Default quantization configuration.
|
|
# More info: https://qdrant.tech/documentation/guides/quantization
|
|
quantization: null
|
|
|
|
# Default strict mode parameters for newly created collections.
|
|
#strict_mode:
|
|
# Whether strict mode is enabled for a collection or not.
|
|
#enabled: false
|
|
|
|
# Max allowed `limit` parameter for all APIs that don't have their own max limit.
|
|
#max_query_limit: null
|
|
|
|
# Max allowed `timeout` parameter.
|
|
#max_timeout: null
|
|
|
|
# Allow usage of unindexed fields in retrieval based (eg. search) filters.
|
|
#unindexed_filtering_retrieve: null
|
|
|
|
# Allow usage of unindexed fields in filtered updates (eg. delete by payload).
|
|
#unindexed_filtering_update: null
|
|
|
|
# Max HNSW value allowed in search parameters.
|
|
#search_max_hnsw_ef: null
|
|
|
|
# Whether exact search is allowed or not.
|
|
#search_allow_exact: null
|
|
|
|
# Max oversampling value allowed in search.
|
|
#search_max_oversampling: null
|
|
|
|
# Maximum number of collections allowed to be created
|
|
# If null - no limit.
|
|
max_collections: null
|
|
|
|
# Cluster-wide resource quotas.
|
|
#
|
|
# Memory and disk are node-wide resources, so their limits are configured once for the whole
|
|
# cluster instead of per collection. Quotas reject updates that would consume more of a resource,
|
|
# regardless of whether strict mode is enabled for the collection being written to. The deprecated
|
|
# `max_resident_memory_percent` in a collection's strict mode config can tighten the memory limit
|
|
# for that collection, but never lift it.
|
|
#
|
|
# A limit that is reached has to be cleared by `release_margin_percent` before the node accepts
|
|
# writes again. Without that margin a resource resting on its limit would put the node in and out
|
|
# of service on the noise between two readings, restarting a shard recovery every time.
|
|
#
|
|
# These values only seed the quota manager on the first start. Afterwards the quota config is read
|
|
# from (and written to) `quota.json` in the storage directory, which is kept in sync across the
|
|
# cluster through consensus and can be changed with the `PUT /quotas` API.
|
|
#
|
|
# If this section is absent - no quota is enforced.
|
|
#quotas:
|
|
# Whether the limits below are enforced.
|
|
#enabled: false
|
|
|
|
# Reject memory-consuming updates once process resident memory reaches this percentage of total
|
|
# system memory (or of the cgroup limit, if one applies).
|
|
# If null - resident memory is not capped.
|
|
#max_resident_memory_percent: null
|
|
|
|
# Reject disk-consuming updates once the filesystem hosting the storage directory is filled to
|
|
# this percentage of its capacity.
|
|
# If null - disk usage is not capped.
|
|
#max_disk_usage_percent: null
|
|
|
|
# How many percentage points below its limit a resource has to fall before this node starts
|
|
# accepting work again. Raise it where usage is volatile; 0 releases as soon as usage is back
|
|
# under the limit.
|
|
# If null - the built-in default of 5 applies.
|
|
#release_margin_percent: null
|
|
|
|
service:
|
|
# Maximum size of POST data in a single request in megabytes
|
|
max_request_size_mb: 32
|
|
|
|
# Number of parallel workers used for serving the api. If 0 - equal to the number of available cores.
|
|
# If missing - Same as storage.max_search_threads
|
|
max_workers: 0
|
|
|
|
# Host to bind the service on
|
|
host: 0.0.0.0
|
|
|
|
# HTTP(S) port to bind the service on
|
|
http_port: 6333
|
|
|
|
# gRPC port to bind the service on.
|
|
# If `null` - gRPC is disabled. Default: null
|
|
# Comment to disable gRPC:
|
|
grpc_port: 6334
|
|
|
|
# Enable CORS headers in REST API.
|
|
# If enabled, browsers would be allowed to query REST endpoints regardless of query origin.
|
|
# More info: https://developer.mozilla.org/en-US/docs/Web/HTTP/CORS
|
|
# Default: true
|
|
enable_cors: true
|
|
|
|
# Enable HTTPS for the REST and gRPC API
|
|
enable_tls: false
|
|
|
|
# Check user HTTPS client certificate against CA file specified in tls config
|
|
verify_https_client_certificate: false
|
|
|
|
# Set an api-key.
|
|
# If set, all requests must include a header with the api-key.
|
|
# example header: `api-key: <API-KEY>`
|
|
#
|
|
# If you enable this you should also enable TLS.
|
|
# (Either above or via an external service like nginx.)
|
|
# Sending an api-key over an unencrypted channel is insecure.
|
|
#
|
|
# Uncomment to enable.
|
|
# api_key: your_secret_api_key_here
|
|
|
|
# Set an api-key for read-only operations.
|
|
# If set, all requests must include a header with the api-key.
|
|
# example header: `api-key: <API-KEY>`
|
|
#
|
|
# If you enable this you should also enable TLS.
|
|
# (Either above or via an external service like nginx.)
|
|
# Sending an api-key over an unencrypted channel is insecure.
|
|
#
|
|
# Uncomment to enable.
|
|
# read_only_api_key: your_secret_read_only_api_key_here
|
|
|
|
# Uncomment to enable JWT Role Based Access Control (RBAC).
|
|
# If enabled, you can generate JWT tokens with fine-grained rules for access control.
|
|
# Use generated token instead of API key.
|
|
#
|
|
# jwt_rbac: true
|
|
|
|
# Enforce API key / JWT authentication on the internal (p2p) gRPC API.
|
|
# The regular api_key is always forwarded on internal requests, but the
|
|
# receiving side only verifies it when this flag is enabled. Keep disabled
|
|
# during a rolling upgrade so older peers (that do not attach a key yet)
|
|
# can still reach newer peers, then turn it on once every node is upgraded.
|
|
#
|
|
# enforce_internal_auth: true
|
|
|
|
# Hardware reporting adds information to the API responses with a
|
|
# hint on how many resources were used to execute the request.
|
|
#
|
|
# Warning: experimental, this feature is still under development and is not supported yet.
|
|
#
|
|
# Uncomment to enable.
|
|
# hardware_reporting: true
|
|
#
|
|
# Uncomment to enable.
|
|
# Prefix for the names of metrics in the /metrics API.
|
|
# metrics_prefix: qdrant_
|
|
|
|
# Allow snapshot recovery from remote HTTP/HTTPS URLs.
|
|
# If disabled, snapshot recovery will only work with local files and uploads.
|
|
# Disabling this can mitigate SSRF risks in environments where the Qdrant node
|
|
# has access to internal resources that should not be reachable by users.
|
|
# Default: true
|
|
# enable_snapshot_url_recovery: true
|
|
|
|
cluster:
|
|
# Use `enabled: true` to run Qdrant in distributed deployment mode
|
|
enabled: false
|
|
|
|
# Configuration of the inter-cluster communication
|
|
p2p:
|
|
# Port for internal communication between peers
|
|
port: 6335
|
|
|
|
# Use TLS for communication between peers
|
|
enable_tls: false
|
|
|
|
# Configuration related to distributed consensus algorithm
|
|
consensus:
|
|
# How frequently peers should ping each other.
|
|
# Setting this parameter to lower value will allow consensus
|
|
# to detect disconnected nodes earlier, but too frequent
|
|
# tick period may create significant network and CPU overhead.
|
|
# We encourage you NOT to change this parameter unless you know what you are doing.
|
|
tick_period_ms: 100
|
|
|
|
# Compact consensus operations once we have this amount of applied
|
|
# operations. Allows peers to join quickly with a consensus snapshot without
|
|
# replaying a huge amount of operations.
|
|
# If 0 - disable compaction
|
|
compact_wal_entries: 128
|
|
|
|
# Set to true to prevent service from sending usage statistics to the developers.
|
|
# Read more: https://qdrant.tech/documentation/guides/telemetry
|
|
telemetry_disabled: false
|
|
|
|
# TLS configuration.
|
|
# Required if either service.enable_tls or cluster.p2p.enable_tls is true.
|
|
tls:
|
|
# Server certificate chain file
|
|
cert: ./tls/cert.pem
|
|
|
|
# Server private key file
|
|
key: ./tls/key.pem
|
|
|
|
# Certificate authority certificate file.
|
|
# This certificate will be used to validate the certificates
|
|
# presented by other nodes during inter-cluster communication.
|
|
#
|
|
# If verify_https_client_certificate is true, it will verify
|
|
# HTTPS client certificate
|
|
#
|
|
# Required if cluster.p2p.enable_tls is true.
|
|
ca_cert: ./tls/cacert.pem
|
|
|
|
# TTL in seconds to reload certificate from disk, useful for certificate rotations.
|
|
# Only works for HTTPS endpoints. Does not support gRPC (and intra-cluster communication).
|
|
# If `null` - TTL is disabled.
|
|
cert_ttl: 3600
|
|
|
|
# Audit logging configuration.
|
|
# When enabled, Qdrant writes structured JSON audit log entries for every
|
|
# access-checked API request.
|
|
#
|
|
# audit:
|
|
# enabled: false
|
|
# dir: ./storage/audit
|
|
# rotation: daily
|
|
# max_log_files: 7
|
|
# # If true, use X-Forwarded-For header to determine client IP in audit logs.
|
|
# # Only enable this when running behind a trusted reverse proxy or load balancer.
|
|
# # WARNING: Enabling this without a trusted proxy allows clients to spoof their IP.
|
|
# # Default: false
|
|
# trust_forwarded_headers: false
|