* Raise ValueError for non-positive batch_size in iter_batch
iter_batch passed batch_size straight into islice with no lower-bound
check. A batch_size of 0 made every islice call return an empty list
immediately, so embed and rerank returned an empty result with no error
across every modality that shares this helper. A negative batch_size hit
islice's own argument validation and raised a cryptic ValueError instead
of a clear one.
Validate batch_size >= 1 once in iter_batch itself, since every embed
and rerank entrypoint across dense text, sparse text, image,
late-interaction and cross-encoder rerank funnels through it.
Added test_iter_batch_rejects_non_positive_size and
test_iter_batch_accepts_positive_size to tests/test_common.py. The new
rejection test fails on unmodified code (batch_size=0 returns silently
instead of raising) and passes with the fix.
Fixes#719
* fix: fix batch size check for parallel execution
---------
Co-authored-by: George Panchuk <george.panchuk@qdrant.tech>
* feat: support pluggable stemmer and inline stopwords for BM25
Allow passing a custom stemmer (any object with a stem_word(word) -> str
method, per the new Stemmer protocol) to Bm25, overriding the default
SnowballStemmer. When a custom stemmer is provided, the supported-languages
check is skipped, enabling languages without a Snowball algorithm such as
Polish, Czech, Ukrainian, Slovak, Bulgarian, or Vietnamese.
Also allow passing stopwords inline, overriding the per-language stopwords
file shipped with the model. Both parameters are forwarded to parallel
workers. Default behavior is unchanged when neither parameter is given.
Fixes#654
* test: cover pluggable stemmer and inline stopwords for BM25
- custom stemmer callable is applied to tokens
- default Snowball path is unchanged
- unsupported language (Polish) works with a custom stemmer and still
raises without one
- inline stopwords override the file-based stopwords
* fix: refine BM25 custom stemmer and stopword handling
---------
Co-authored-by: George Panchuk <george.panchuk@qdrant.tech>
* fix: re-download model when cached files fail verification
An interrupted or externally truncated download can leave a corrupt
file in the Hugging Face cache. On the next load the offline-first
probe sees the file on disk and returns it, so the model fails with
INVALID_PROTOBUF and the only fix is to delete the cache by hand.
This makes the cache self-heal. The probe now raises when a model
file's size does not match the recorded metadata, so download_model
falls through to an online retry with force_download=True
(huggingface_hub's cache check is existence-only and will not
re-fetch a present-but-truncated blob).
Verification also walks the repo tree recursively and matches files
by their repo-relative path. Without that, weights kept in a
subdirectory (onnx/model.onnx, more than half the models) were never
recorded in the metadata, so neither the new check nor the existing
one ever looked at them. A mismatch on an auxiliary config file is
still left to best-effort loading, as before.
* fix: merge force_download into kwargs on the recovery retry
If a caller passes force_download through kwargs, the explicit
force_download=True on the recovery retry would raise a duplicate
keyword error. Merge it into kwargs so the recovery value wins
without the collision.
* fix: only force a re-download when a cached model file is corrupt
---------
Co-authored-by: George Panchuk <george.panchuk@qdrant.tech>
* fix(image): normalize batched input along the channel axis
normalize() advertises 4D (N, C, H, W) support via its num_channels
branch and the channel-count validation, but the actual math used
((image.T - mean) / std).T. Transpose reverses every axis, so on 4D
input the channels no longer line up with mean/std: it raises when
N != C and silently normalizes along the batch axis when N == C.
Reshape mean/std to broadcast on the real channel axis instead; the
(C, H, W) path is unchanged.
* test(image): cover channel-wise normalize for 3D and batched input
* refactor
---------
Co-authored-by: George Panchuk <george.panchuk@qdrant.tech>
* fix: support optional tokenizer metadata files
* fix: pad id fallback chain and additional_special_tokens lists
* refactor: remove redundant tests and comments
---------
Co-authored-by: George Panchuk <george.panchuk@qdrant.tech>
* fix: pass (width, height) to Pillow in the Resize transform
`resize()` handed a tuple size straight to `PIL.Image.resize()`. fastembed
keeps sizes as (height, width) — `Transform.from_config` builds the tuple
as `(size["height"], size["width"])` — while Pillow takes (width, height),
so a non-square image processor configuration produced a transposed image:
Resize(size=(100, 200))(Image.new("RGB", (300, 300)))[0].size
# (100, 200), expected (200, 100)
Square sizes are unaffected, which is why this went unnoticed. The int
branch of `resize()` already emits Pillow order and is untouched, as are
`resize_ndarray()`'s callers, which pass (width, height) explicitly.
`Resize.__call__` is the only caller of this function and always supplies
fastembed's height-first order, so converting here is safe.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* tests: simplify tests
---------
Co-authored-by: George Panchuk <george.panchuk@qdrant.tech>
* new: drop python3.9, replace optional and union with |
* new: remove python 3.9 from pyproject
* refactor: replace remaining union and optional with |
* new: remove optional and union in dataclasses
* fix: add typealias to numpy type
* new: replace union with | in token count
* tests: introduce model cache to tests
* fix: fix not cached model deletion
* new: do not run CI tests on mac os and windows on python 3.10-3.12
* fix: lowercase cache keys, bm25 caching
* tests: do not run parallel processing on all cpus in sparse text embed
* fix: fix models to cache names, do not run parallel=0
* fix: fix sparse embedding tests
* fix: bm42 language by lower case model name
* Custom rerankers support
* Test for reranker_custom_model
* test fix
* Model description type fix
* Test fix
* fix: fix naming
* fix: remove redundant arg from tests
* new: update readme
---------
Co-authored-by: George Panchuk <george.panchuk@qdrant.tech>
* chore: Trigger CI test
* chore: Trigger CI test
* chore: Trigger CI test
* chore: Trigger CI test
* chore: Trigger CI test
* chore: Trigger CI test
* chore: Trigger CI test
* chore: Trigger CI test
* chore: Trigger CI test
* Trigger CI
* Trigger CI
* Trigger CI
* Trigger CI test
* Trigger CI test
* Trigger CI test
* Trigger CI test
* Trigger CI test
* Trigger CI test
* Trigger CI test
* Trigger CI test
* Trigger CI test
* Trigger CI test
* Trigger CI test
* Trigger CI test
* Trigger CI test
* Trigger CI test
* Trigger CI test
* Trigger CI test
* Trigger CI test
* Trigger CI test
* Trigger CI test
* Trigger CI test
* Trigger CI test
* Trigger CI test
* Trigger CI test
* Trigger CI test
* Trigger CI test
* Trigger CI test
* new: Added on workflow dispatch
* tests: Updated tests
* fix: Fix CI
* fix: Fix CI
* fix: Fix CI
* improve: Prevent stop iteration error caused by next
* fix: Fix variable might be referenced before assignment
* refactor: Revised the way of getting models to test
* fix: Fix test in image model
* refactor: Call one model
* fix: Fix ci
* fix: Fix splade model name
* tests: Updated tests
* chore: Remove cache
* tests: Update multi task tests
* tests: Update multi task tests
* tests: Updated tests
* refactor: refactor utils func, add comments, conditions refactor
---------
Co-authored-by: George Panchuk <george.panchuk@qdrant.tech>
* Migration of models to dataclasses
* Model description file
* Test fix
* kw_only support
* Multitask embeddings test fix
* list_supported_models type fix
* Dim fix for sparsemodels
* Dim fix for sparsemodels
* Dim fix for sparsemodels (x2)
* Model management type fix
* Interface docstring fixes
* Mypy fixes
* Typing fix again
* Typing fix again
* Special cast to SparseModelDescription
* Special cast to SparseModelDescription
* Special cast to SparseModelDescription
* typing fix for colpali
* typing fix for colpali
* typing fix for colpali
* typing fix for colpali
* Let's try generic typing for ModelManagment
* wip: dataclass idea, small fixes (#475)
* wip: dataclass idea, small fixes
* fix: fix exception message in base model description
* remove custom model descriptions
* make license, description and size in gb mandatory in model description
* fix: introduce _list_supported_models which returns model description objects
* test: add test for list supported models
* fix: fix list supported models usage in tests
---------
Co-authored-by: George <george.panchuk@qdrant.tech>
* wip: design draft
* Operators fix
* Fix model inputs
* Import from fastembed.late_interaction_multimodal
* Fixed method misspelling
* Tests, which do not run in CI
Docstring improvements
* Fix tests
* Bump colpali to version v1.3
* Remove colpali v1.2
* Remove colpali v1.2 from tests
* partial fix of change requests:
descriptions
docs
black
* query_max_length
* black colpali
* Added comment for EMPTY_TEXT_PLACEHOLDER
* Review fixes
* Removed redundant VISUAL_PROMPT_PREFIX
* type fix + model info
* new: add specific model path to colpali
* fix: revert accidental renaming
* fix: remove max_length from encode_batch
* refactoring: remove redundant QUERY_MAX_LENGTH variable
* refactoring: remove redundant document marker token id
* fix: fix type hints, fix tests, handle single image path embed, rename model, update description
* license: add gemma to NOTICE
* fix: do not run colpali test in ci
* fix: fix colpali test
---------
Co-authored-by: d.rudenko <dmitrii.rudenko@qdrant.com>
* new: Added type stub
* chore: Updated stubs
* chore: device_id type hint
* chore: add -> none to init without args
* new: Added workflow type check
* chore: Revert added type checkers
* chore: Added type hints
* fix: Fix generic type
* fix: Fix generic type
* new: Add type hints for parallel processor
* fix: Revert queue sub type as its not supported
* fix: Revert queue sub type as its not supported
* fix: Fixed type hints
* chore: Updated type hints
* fix: Update task id to be public
* chore: Updated type hints
* chore: Updated type hints
* fix: minor reverts in parallel processor and onnx text model
* chore: Add missing type hints in functions
* add missing import, small type refactor
---------
Co-authored-by: George Panchuk <george.panchuk@qdrant.tech>
* fix: Fix minilm paraphrase by adding it to pool models
* tests: Updated minilm paraphrase canonical vector
* chore: Added a warning message for updating the model
* chore: Added version where model will be removed
* new: Added jina embedding v3
* refactor: Changed dim to int value
* new: Updated notice
* new: Extended text embedding with query embed and passage embed
* fix: Fix lazy load in query and passage embed
* tests: Added test for multitask embeddings
* nit: Remove cache dir from tests
* tests: Updated tests
* improve: Improve task selection
* fix: Fix ci
* fix: Update fastembed/text/multitask_embedding.py
Co-authored-by: George <george.panchuk@qdrant.tech>
* Update fastembed/text/multitask_embedding.py
Co-authored-by: George <george.panchuk@qdrant.tech>
* fix: Pass task id using kwargs to parallel processor
* tests: Added test for task assignment
* prefer enums over ints
* tests: Added test for parallel
* improve: Updated model description
* fix: Fix ci
* fix: Fix ci
* refactor: Refactor query_embed and passage_embed
* tests: Added task propagation to parallel
* refactor: Set default task as retrieval passage
* chore: Update default task in tests
---------
Co-authored-by: George <george.panchuk@qdrant.tech>
* Merge master
* rerank_pairs interface + parallelism support
* remove test notebook
* Removed unused code
* New tests for cross encoders and new interface
* Importing Self fix. We will need it for mypy support in newer versions
* Removed Self typing
* Removed non-needed changes from text
* Isort + black
* wip: start reviewing (#420)
Co-authored-by: Dmitrii Ogn <dimitriy_rudenko@mail.ru>
* Test fix
* Update fastembed/rerank/cross_encoder/text_cross_encoder.py
Co-authored-by: George <george.panchuk@qdrant.tech>
* Update fastembed/rerank/cross_encoder/text_cross_encoder.py
Co-authored-by: George <george.panchuk@qdrant.tech>
* Update fastembed/rerank/cross_encoder/text_cross_encoder.py
Co-authored-by: George <george.panchuk@qdrant.tech>
* Update fastembed/rerank/cross_encoder/text_cross_encoder_base.py
Co-authored-by: George <george.panchuk@qdrant.tech>
* Test for parallel processing + bugfix of PosixPath passing
* Removed non-needed import and added docstring
* Typing fix + argument passing
* Test parametrization
Moved to selected models set to test
* Run base test on all models
* Typing fix + improvement of input_names check
* nit: fix post process, update docstring, update tokenize, remove redundant imports
---------
Co-authored-by: George <george.panchuk@qdrant.tech>