Dmitrii Rudenko
4e699cc9fa
Isort for image tests
2024-07-30 19:37:44 +03:00
Dmitrii Rudenko
90a430cc40
Added selfish logo test as requests input
2024-07-30 19:23:01 +03:00
Dmitrii Rudenko
f59ba2e89d
Additional tests for image embeddings
2024-07-30 18:44:37 +03:00
Anush
0e258ab875
feat: Added jina-embeddings-v2-base-code ( #301 )
...
* feat: Added jina-embeddings-v2-base-code
* fix: test embeddings for "hello world" not "Hello"
* docs: Updated supported models
2024-07-18 18:17:46 +05:30
Dmitrii Ogn
f0ff09c546
Oml zoo ( #291 )
...
* Support of Qdrant/Unicom-ViT-B-16 and Qdrant/Unicom-ViT-B-32
2024-07-10 16:43:54 +03:00
Dmitrii Ogn
d09af55edd
Nomic-embeddings-support ( #280 )
...
* Nomic-embeddings-support
* Jina models moved to pooled-normalized embeddings
* Canonical vector for nomic-ai/nomic-embed-text-v1.5-Q
* Moved all nomics to pooled_embeddings
---------
Co-authored-by: d.rudenko <dimitriyrudenk@gmail.com >
2024-07-10 11:45:30 +03:00
Andrey Vasnetsov
f820c36656
fix + test for empty from_dict ( #285 )
2024-07-05 13:30:14 +02:00
George
e1ecfe9c2f
fix: fix None cache dir in parallel mode ( #277 )
2024-06-14 17:19:53 +02:00
Dmitrii Ogn
331207976e
MiniLM fix ( #275 )
...
* MiniLM fix
* Added MiniLM to text embedding
Fixed MiniLM source destination
Black + isort for repo
* Fixed model all-MiniLM-L6-v2 description
Recomputed canonical vector for all-MiniLM-L6-v2 in test
---------
Co-authored-by: d.rudenko <dimitriyrudenk@gmail.com >
2024-06-14 16:42:31 +03:00
Klaus Hueck
fd0b26f009
Add support for jinaai/jina-embeddings-v2-base-de ( #270 )
...
* feat: add support for SOTA german embedding model with long context length jinaai/jina-embeddings-v2-base-de
* Fix jina de model weight
---------
Co-authored-by: George <panchuk.george@outlook.com >
2024-06-14 13:39:13 +02:00
George
5461012ab1
new: add bm25, fix param propagation in parallel mode, fix bm42 parallel ( #274 )
...
* new: add bm25, fix param propagation in parallel mode, fix bm42 parallel
* refactoring: remove redundant example
* fix: fix mp start method in bm25
* refactoring: refactor token id generation
* new: replace model repository
2024-06-13 19:52:48 +02:00
George
c8fff66b18
Colbert ( #248 )
...
* new: add late interaction embedding, colbert
* new: update imports
* new: add comments
* fix: rollback mp methods
* fix: restore existing padding after embed query
* fix: fix OnnxOutputContext in onnx embed, fix preprocessing for colbert
2024-05-31 17:06:43 +02:00
Dmitrii Ogn
85aaae4c08
Add resnet ( #246 )
...
* Resnet support added
* Tests fixed
Shapes matching for Resnet50-onnx
Example of Resnet50 to onnx conversion (basic)
* Removed optional conversion from PIL to np.ndarray and now it it's made default
Fixed test accordingly
* Refactoring of pil2ndarray
* Partial support of convnext preprocessing
Resize logic
* normalize canonical value
* Style changes for review
* new: update resnet repo
---------
Co-authored-by: d.rudenko <dimitriyrudenk@gmail.com >
Co-authored-by: George Panchuk <george.panchuk@qdrant.tech >
2024-05-31 16:56:13 +02:00
Andrey Vasnetsov
dfd25d41c9
Attention sparse embeddings ( #235 )
...
* WIP: sparse embeddings using attention
* support for stopwords
* apply stopwords
* proceed implementation of sparse attention embeddings (#234 )
* complete inference
* query embed + comment
* use simpler weights formula instead of sorting of words
* update tests
* fix: fix bm42 usage, add query_embed to SparseTextEmbedding, update tests
---------
Co-authored-by: George <george.panchuk@qdrant.tech >
2024-05-24 15:25:34 +02:00
George
99164c9050
Clip ( #219 )
...
* wip: init image embeddings
* new: add clip
* new: fix clip text embedding
* fix: fix image parallel
* fix: add test images
* fix: add PIL
* fix: fix generics
* fix: fix sparse worker
* fix: fix image test path
* fix: replace models repo
* new: follow-up for onnx providers and local_files_only option
* fix: add types, refactor a bit
* refactoring: move onnxprovider type alias to types
* fix: fix type alias import
2024-05-09 11:12:25 +02:00
Anush
ab7a99a748
feat: Quantized models ( #201 )
...
* feat: Quantized models
* refactor: use model_file for GCS
* refactoring: refactor model downloading (#209 )
* refactoring: refactor model downloading
* refactor: update docstring
Co-authored-by: Anush <anushshetty90@gmail.com >
* Update fastembed/common/model_management.py
Co-authored-by: George <george.panchuk@qdrant.tech >
* fix: model_file for Snowflake models
---------
Co-authored-by: George <george.panchuk@qdrant.tech >
2024-04-26 17:27:16 +02:00
Anush
466886a317
ci: Schedule python-tests.yml ( #211 )
...
* ci: Schedule python-tests.yml
* ci: use emojis
* ci: Bump action versions python-tests.yml
* ci: python-tests.yml
2024-04-26 10:28:37 +05:30
Anush
cc4112d859
feat: Snowflake models ( #207 )
...
* feat: Snowflake models
* Added snowflake/snowflake-arctic-embed-m
* docs: snowflake/snowflake-arctic-embed-m
2024-04-19 19:00:10 +05:30
George
8b7b8476a5
fix: fix model sizes in supported models lists ( #167 )
...
* fix: fix model sizes in supported models lists
* fix: remove redundant comment
* fix: fix test
* fix: update supported models notebook
* Consistentcy around quantization in supported_onnx_models
---------
Co-authored-by: Nirant <NirantK@users.noreply.github.com >
Co-authored-by: Nirant Kasliwal <nirant.bits@gmail.com >
2024-04-01 16:18:49 +05:30
George
ce98631b9a
Fix spladepp parallelism ( #169 )
...
* fix: add get_worker_class implementation to spladepp
* fix: add tests for parallel embed for spladepp
2024-04-01 10:22:24 +05:30
Yuvraj Wale
bc1e23849b
feat: support mixedbread-ai/mxbai-embed-large-v1 ( #158 )
...
* add: support mixedbread-ai/mxbai-embed-large-v1
* refactor: canonical vector
* Update fastembed/text/onnx_embedding.py
---------
Co-authored-by: Nirant <NirantK@users.noreply.github.com >
2024-03-22 06:23:54 +05:30
Nirant
96a2a9097a
Fix model name typo + Add SPLADE notebook ( #155 )
...
* Rename model + Add SPLADE notebook
* Update docs/examples/SPLADE_with_FastEmbed.ipynb
Co-authored-by: Anush <anushshetty90@gmail.com >
* Update docs/examples/SPLADE_with_FastEmbed.ipynb
Co-authored-by: Anush <anushshetty90@gmail.com >
* Update docs/examples/SPLADE_with_FastEmbed.ipynb
Co-authored-by: Anush <anushshetty90@gmail.com >
* Update CANONICAL_COLUMN_VALUES in test_sparse_embeddings.py
---------
Co-authored-by: Anush <anushshetty90@gmail.com >
2024-03-20 18:15:53 +05:30
Andrey Vasnetsov
361f674e47
avoid changing output dimentionality for a single input ( #148 )
2024-03-13 23:38:51 +05:30
Nirant
d817da2e01
Add Splade v1 ( #144 )
...
* Add SPLADE v1
* WIP SPLADE Export errors
* add ONNX model to HF hub and use that
* Update sentences in Converting_SPLADE_to_ONNX.ipynb
* Remove unnecessary files and directories
* Rename var in TextEmbedding class to use EMBEDDING_MODEL_TYPE
* Add SPLADE to list of text embeddings
* Add SPLADE model support for text embedding
* Fix deprecation warning in embedding.py
* Add test for batch embedding with sparse embeddings
* Refactor import statement in test_sparse_embeddings.py
* Rename nbs
* Update vocab size in SPLADE model
* Fix canonical vector lookup in test_text_onnx_embeddings.py
* review refactoring
* restore list_supported_models in OnnxTextEmbedding
* Remove unused method _preprocess_onnx_input() in SpladePP class
* Update SPLADE_PP_en_v1 source in splade_pp.py
* Refactor onnx_model.py to change base model behavior
* extend tests to sparse values as well as indicies
* chore: pre-commit hooks
---------
Co-authored-by: generall <andrey@vasnetsov.com >
Co-authored-by: Anush008 <anushshetty90@gmail.com >
2024-03-13 18:04:44 +05:30
Anush
1e298a00b3
feat: Added gte-large, nomic-text 1.5, cleanup ( #130 )
2024-02-21 17:54:32 +05:30
Armaghan
406f432edc
feat: Support sentence-transformers/paraphrase-multilingual-MiniLM-L12-v2 ( #129 )
...
* feat: Support sentence-transformers/paraphrase-multilingual-MiniLM-L12-v2
* test: Include sentence-transformers/paraphrase-multilingual-MiniLM-L12-v2
* docs: supported models update
2024-02-21 13:35:49 +05:30
Anush
98141cc8d3
feat: Added nomic-embed-text-v1 support + formatting changes + import fixes ( #118 )
...
* feat: Added nomic-embed-text-v1 support
* chore: xenova/nomic-embed-text-v1 -> nomic-ai/nomic-embed-text-v1
---------
Co-authored-by: Nirant <NirantK@users.noreply.github.com >
2024-02-19 14:17:33 +05:30
generall
fcdc5690b9
rename models
2024-02-02 15:51:19 +01:00
generall
7883fa3c41
new multilingual models
2024-02-02 15:32:42 +01:00
generall
8b800da7bc
ruff
2024-02-02 15:12:57 +01:00
generall
4813b18854
refactoring
...
Co-authored-by: George Panchuk <george.panchuk@qdrant.tech >
2024-02-02 15:12:57 +01:00
Anush
96f7d83d33
feat: Support xenova/multilingual-e5-large, xenova/paraphrase-multili… ( #103 )
...
* feat: Support xenova/multilingual-e5-large, xenova/paraphrase-multilingual-mpnet-base-v2
* chore: updated exclude_token_type_ids check
* docs: supported models update
2024-02-02 15:16:35 +05:30
Anush
ede507e2cf
feat: HuggingFace download support for FlagEmbedding ( #94 )
...
* feat: HF support for FlagEmbedding
* chore: update docstring embedding.py
* refactor: GCS URLs models.json
* chore: toLower() models.json
* chore: update tqdm declarative
* chore: exclude keys list_supported_models
* chore: review changes
2024-01-23 12:47:55 +05:30
Anush
9b63427118
chore: pre-commit formatting ( #91 )
...
* chore: formatting
* chore: formatting
* chore: remove other hooks
* Update poetry lock
---------
Co-authored-by: Nirant Kasliwal <nirant.bits@gmail.com >
2024-01-16 15:06:54 +05:30
Joan Fontanals
f222d7cd87
add JinaEmbeddings class ( #67 )
...
* add JinaEmbeddings class
* fix tests dimensions
---------
Co-authored-by: Joan Fontanals Martinez <joan.fontanals.martinez@jina.ai >
2023-11-20 16:37:23 +05:30
NirantK
f28087c71a
* fix(embedding.py): change dim value from 384 to 512 for the "BAAI/bge-small-zh-v1.5" model
...
* fix(test_onnx_embeddings.py): add canonical vector values for the "BAAI/bge-small-zh-v1.5"
2023-10-16 18:23:02 +05:30
Nirant
f79de09ff7
Merge branch 'streaming-inference' into add_model_size
2023-10-16 17:54:33 +05:30
NirantK
e24ea64e21
* test(test_onnx_embeddings.py): skip specific model if size_in_GB is greater than 1
2023-10-16 17:52:35 +05:30
generall
35ec40b3a3
review fixes
2023-10-16 14:19:58 +02:00
Nirant
c953a083cc
Merge branch 'main' into streaming-inference
2023-10-16 17:33:36 +05:30
NirantK
9e5d37846c
* feat(embedding.py): add support for BAAI/bge-small-en-v1.5 and BAAI/bge-base-en-v1.5 models
...
* feat(embedding.py): change default model to v1.5
2023-10-16 17:08:58 +05:30
generall
c719fc696d
disable large models on non-ubuntu CI
2023-10-16 13:27:28 +02:00
generall
24dc24b02d
implement data-parallel inference and up version
2023-10-16 13:07:07 +02:00
generall
8c6d4d2b52
remove optimum + more test + fix batching embed + ci on other machines
2023-10-03 21:19:58 +02:00
NirantK
4d6d27cffb
* test(test_onnx_embeddings.py): remove unnecessary list conversion in embeddings assignment
2023-09-25 16:50:52 +05:30
NirantK
3149694c04
Utility script version of the notebook
2023-09-25 15:33:44 +05:30
NirantK
a335c8898f
Remove attention pooling
2023-09-18 07:45:12 +05:30
NirantK
03b23af0a0
* refactor(test_onnx_embeddings.py): rename test_onnx_embeddigns.py to test_onnx_embeddings.py
...
* chore(test_onnx_embeddings.py): remove unnecessary line breaks and whitespace
* test(test_onnx_inference): add test case for onnx inference
2023-08-25 00:37:31 +05:30
NirantK
0c7574e95a
* fix(test_onnx_embeddigns.py): change method name from encode to embed
2023-08-23 09:51:44 +05:30
generall
770beb1023
refactoring
2023-08-18 01:36:36 +02:00