* new: Add missing type hints
* refactor: Removed type ignore
* fix: fix mypy complaints
* fix: remove redundant type coercion, fix skip list type
* new: more precise type for sparse embedding inference, a small revert for parallel processor
---------
Co-authored-by: George Panchuk <george.panchuk@qdrant.tech>
* chore: Added type hints
* fix: Fix generic type
* fix: Fix generic type
* new: Add type hints for parallel processor
* fix: Revert queue sub type as its not supported
* fix: Revert queue sub type as its not supported
* fix: Fixed type hints
* chore: Updated type hints
* fix: Update task id to be public
* chore: Updated type hints
* chore: Updated type hints
* fix: minor reverts in parallel processor and onnx text model
* fix: Fix return of onnx_embed
* fix: Fix type hint of start method in worker class
* fix: Fix not passing kwargs in _preprocess_onnx_input and tokenize as base class
* fix: Fix not passing kwargs in _preprocess_onnx_input as base class
* fix: change tokenize in simpleTokenizer to classmethod
* chore: Changed query argument to Iterable to match base class
* chore: changed mask token id and pad token id to be int
* review suggestions (#398)
---------
Co-authored-by: George <george.panchuk@qdrant.tech>
* feat: Added multi gpu support for text embedding
* feat: Add support for multi-gpu for special text models
* fix: Fix lazy_load to load the model to child processes when parallel is not none
* feat: Added lazy_load and multi-gpu to colbert
* feat: Add lazy_load and multi gpu to image models
* feat: Support lazy_load and multi-gpu to sparse models (except BM25)
* fix: Fixed BM25 not working
* refactor: Remove redundant GPUParallelProcessor
* refactor: Refactor _embed_*_parallel
* feat: Add cuda argument
refactor: Refactor how worker assign device
* fix: Fix if providers and cuda are None
* fix: Fix providers and cuda are none
* WIP: Multi gpu support review (#361)
* WIP: review
* wip: review
* refactor: refactor images
* refactor: refactor sparse
* refactor: refactor late interaction
* add model loading
* add tests
* fix: uncomment models in tests
* fix: fix variable declaration order
* fix: fix device id assignment
* tests: add multi gpu tests
* fix: fix device id assignment for sparse embeddings
* tests: update multi gpu tests
---------
Co-authored-by: George Panchuk <george.panchuk@qdrant.tech>
* refactor: remove redundant declarations
* fix: rollback redundant changes
* fix: remove num workers device ids dep, fix type hint
* fix: fix post process for sparse models
* fix: remove redundant model loading
* new: add lazy load and new gpu support to cross encoders
* fix: add rerankers to multi gpu tests
* fix: unlock multilingual test
* fix: fix gpu test with cross encoder
---------
Co-authored-by: Andrey Vasnetsov <andrey@vasnetsov.com>
Co-authored-by: George Panchuk <george.panchuk@qdrant.tech>
* fix: Fix deadlock when child gets kill -9 sig
* chore: Better cleanup for resources
* chore: changed place of processes.clear
* fix: Added cancle_join_thread for emergency shutdown
* MiniLM fix
* Added MiniLM to text embedding
Fixed MiniLM source destination
Black + isort for repo
* Fixed model all-MiniLM-L6-v2 description
Recomputed canonical vector for all-MiniLM-L6-v2 in test
---------
Co-authored-by: d.rudenko <dimitriyrudenk@gmail.com>