mirror of
https://github.com/qdrant/qdrant.git
synced 2026-09-29 09:27:53 -05:00
Every vector candidate now carries a distance metric instead of the hardcoded Dot, threaded into the fixture schema, the CreateVectorName generator and the read-back prediction. The model predicts Cosine read-backs exactly by mirroring the engine's ingestion preprocessing: metric_preprocess follows NamedVectors::preprocess_dense_vector's per-datatype dispatch and calls Distance::preprocess_vector itself, so predictions track the engine by construction (including the identity preprocess of the byte metric, which stores Uint8 vectors un-normalized). Stored vectors are preprocessed exactly once (optimizer and CoW moves transfer raw bytes), so predictions stay exact across moves, including Cosine + Float16. New candidates: "e" (dense Cosine), "n" (multi-dense Cosine, per-row normalization), "x" (dense Cosine + Float16), "o" (dense Cosine + Turbo4, padding-free dim), "j" (dense Euclid), "k" (dense Manhattan). Euclid/Manhattan preprocess is an identity, so their value is engine side: Order::SmallBetter comparator coverage. The startup predictability check now also rejects sparse + non-Dot (sparse schemas carry no distance) and Turbo4 + Euclid/Manhattan (TQ's L1/L2 modes store lengths differently from Dot/Cosine and their copy-on-write re-quantization fixed point is not soak-validated yet). Soak-validated on seeds 1/2/4/5/6/7/8 (30k ops), including two restart runs (restart probability 0.002) with the optimizer enabled. Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Collection
Crate, which implements all functions required for operations with a single collection of points. Points within a collection should share the same payload schema and have same vector size. So that search requests could be performed over all points of a single collection.

