Nirant
c09909773e
Add Execution Counts ( #177 )
...
* Update notebooks
* Update FastEmbed usage across docs
* Refactor code for better readability and maintainability
* Clean outputs
* Change dataset
* Add numbers inline in output
* Remove inline outputs since I used :memory:
* Fix syntax error in Hindi_Tamil_RAG_with_Navarasa7B.ipynb
2024-04-01 20:48:26 +05:30
Nirant
3d2254d215
Bump version to 0.2.6 in pyproject.toml ( #175 )
2024-04-01 17:16:07 +05:30
George
e11f66bb55
refactoring: update imports in notebooks ( #173 )
...
* new: simplify imports
* refactoring: update import
* refactoring: update imports in notebooks
* fix: fix notebook output
* Re-run notebook with revised imports
---------
Co-authored-by: Nirant Kasliwal <nirant.bits@gmail.com >
2024-04-01 17:02:13 +05:30
George
8b7b8476a5
fix: fix model sizes in supported models lists ( #167 )
...
* fix: fix model sizes in supported models lists
* fix: remove redundant comment
* fix: fix test
* fix: update supported models notebook
* Consistentcy around quantization in supported_onnx_models
---------
Co-authored-by: Nirant <NirantK@users.noreply.github.com >
Co-authored-by: Nirant Kasliwal <nirant.bits@gmail.com >
2024-04-01 16:18:49 +05:30
Nirant
e3d2e1dc44
Hybrid Search Tutorial ( #165 )
...
* Re-organize docs
* Rename notebooks
* Move nbs
* Working Sparse and Dense Search
* Add RRF
* Refactor code to improve performance and readability
* Add ESCI label for the RRF results
* Update docs/examples/Hybrid_Search.ipynb
Co-authored-by: Anush <anushshetty90@gmail.com >
* Update docs/examples/Hybrid_Search.ipynb
Co-authored-by: Anush <anushshetty90@gmail.com >
* Remove unnecessary code and update vector format
---------
Co-authored-by: Anush <anushshetty90@gmail.com >
2024-03-29 21:10:14 +05:30
Nirant Kasliwal
c651b2b539
Update SPLADE notebook with new sections
2024-03-22 15:18:34 +05:30
Yuvraj Wale
bc1e23849b
feat: support mixedbread-ai/mxbai-embed-large-v1 ( #158 )
...
* add: support mixedbread-ai/mxbai-embed-large-v1
* refactor: canonical vector
* Update fastembed/text/onnx_embedding.py
---------
Co-authored-by: Nirant <NirantK@users.noreply.github.com >
2024-03-22 06:23:54 +05:30
Nirant
96a2a9097a
Fix model name typo + Add SPLADE notebook ( #155 )
...
* Rename model + Add SPLADE notebook
* Update docs/examples/SPLADE_with_FastEmbed.ipynb
Co-authored-by: Anush <anushshetty90@gmail.com >
* Update docs/examples/SPLADE_with_FastEmbed.ipynb
Co-authored-by: Anush <anushshetty90@gmail.com >
* Update docs/examples/SPLADE_with_FastEmbed.ipynb
Co-authored-by: Anush <anushshetty90@gmail.com >
* Update CANONICAL_COLUMN_VALUES in test_sparse_embeddings.py
---------
Co-authored-by: Anush <anushshetty90@gmail.com >
2024-03-20 18:15:53 +05:30
Nirant
d817da2e01
Add Splade v1 ( #144 )
...
* Add SPLADE v1
* WIP SPLADE Export errors
* add ONNX model to HF hub and use that
* Update sentences in Converting_SPLADE_to_ONNX.ipynb
* Remove unnecessary files and directories
* Rename var in TextEmbedding class to use EMBEDDING_MODEL_TYPE
* Add SPLADE to list of text embeddings
* Add SPLADE model support for text embedding
* Fix deprecation warning in embedding.py
* Add test for batch embedding with sparse embeddings
* Refactor import statement in test_sparse_embeddings.py
* Rename nbs
* Update vocab size in SPLADE model
* Fix canonical vector lookup in test_text_onnx_embeddings.py
* review refactoring
* restore list_supported_models in OnnxTextEmbedding
* Remove unused method _preprocess_onnx_input() in SpladePP class
* Update SPLADE_PP_en_v1 source in splade_pp.py
* Refactor onnx_model.py to change base model behavior
* extend tests to sparse values as well as indicies
* chore: pre-commit hooks
---------
Co-authored-by: generall <andrey@vasnetsov.com >
Co-authored-by: Anush008 <anushshetty90@gmail.com >
2024-03-13 18:04:44 +05:30
Nirant Kasliwal
5603fbe1fb
Fix naming typo
2024-03-06 14:30:54 +05:30
Nirant Kasliwal
9ed4486d9c
Rename typo in notebook
2024-03-06 14:24:48 +05:30
Nirant Kasliwal
e36a39c388
Update author information in notebook
2024-03-06 14:18:48 +05:30
Nirant Kasliwal
97c359c1b9
Add Navrasa LLM download and explanation sections
2024-03-06 14:17:41 +05:30
Nirant Kasliwal
9fd51425fe
Add author information and Colab link to notebook
2024-03-06 14:12:04 +05:30
Nirant Kasliwal
0dec22d02d
Remove inline outputs
2024-03-06 14:06:46 +05:30
Nirant Kasliwal
8a3b746b71
Rename notebook
2024-03-06 14:05:13 +05:30
Nirant
337ad9c93f
Hindi RAG with Qdrant and FastEmbed ( #135 )
...
* Add workingnb
* Add A100 Colab
* Remove old checkpoint
* Refactor code to separate HF Token
2024-03-06 14:02:32 +05:30
Anush
1e298a00b3
feat: Added gte-large, nomic-text 1.5, cleanup ( #130 )
2024-02-21 17:54:32 +05:30
Armaghan
406f432edc
feat: Support sentence-transformers/paraphrase-multilingual-MiniLM-L12-v2 ( #129 )
...
* feat: Support sentence-transformers/paraphrase-multilingual-MiniLM-L12-v2
* test: Include sentence-transformers/paraphrase-multilingual-MiniLM-L12-v2
* docs: supported models update
2024-02-21 13:35:49 +05:30
Anush
98141cc8d3
feat: Added nomic-embed-text-v1 support + formatting changes + import fixes ( #118 )
...
* feat: Added nomic-embed-text-v1 support
* chore: xenova/nomic-embed-text-v1 -> nomic-ai/nomic-embed-text-v1
---------
Co-authored-by: Nirant <NirantK@users.noreply.github.com >
2024-02-19 14:17:33 +05:30
Nirant
81bab0cd1d
Make 0.2.1 Release + Update docs ( #116 )
...
* Update version from 0.2.0 (yanked) to 0.2.1
* Update text embedding to include prefix for passages and queries
* Update supported models to use the latest API
* * fix(text_embedding_base.py): remove unnecessary prefix from texts in embed method
* feat(text_embedding_base.py): update query_embed method to updated instruction for the v1.5 model
* Remove comparison, since the ranking is identical even with varying embedding
* Refactor text embedding query handling
2024-02-08 09:06:00 +05:30
Nirant
973da354ae
Fix query to align with Qdrant mixin usage ( #115 )
...
* fix: query in text_embedding_base to work with both Iterable and str as users might supply both
* Fix Qdrant query to align with future usage
* * refactor(text_embedding_base.py): change query parameter type from str to Union[str, Iterable[str]] in query_embed method
* Update return type of query_embed method
* Update return type in TextEmbeddingBase
2024-02-07 22:01:31 +05:30
Anush
96f7d83d33
feat: Support xenova/multilingual-e5-large, xenova/paraphrase-multili… ( #103 )
...
* feat: Support xenova/multilingual-e5-large, xenova/paraphrase-multilingual-mpnet-base-v2
* chore: updated exclude_token_type_ids check
* docs: supported models update
2024-02-02 15:16:35 +05:30
Anush
9b63427118
chore: pre-commit formatting ( #91 )
...
* chore: formatting
* chore: formatting
* chore: remove other hooks
* Update poetry lock
---------
Co-authored-by: Nirant Kasliwal <nirant.bits@gmail.com >
2024-01-16 15:06:54 +05:30
Nirant
55379539ef
* chore(docs): update Getting Started.ipynb with progressbar + New Models
...
* * feat(Supported_Models.ipynb): add support for BAAI/bge-small-zh-v1.5 model
* feat(Supported_Models.ipynb): add support for jinaai/jina-embeddings-v2-base-en model
* feat(Supported_Models.ip
* * chore(docs): update Getting Started.ipynb with progressbar
2023-12-13 14:17:25 +05:30
NirantK
299042d592
* docs(examples): add explanation of Qdrant Client usage with FastEmbed library and Qdrant API
2023-10-30 23:01:26 +05:30
Nirant
f8f8316fea
Merge pull request #38 from qdrant/explain_cossim
...
* docs(examples): update FastEmbed_vs_HF_Comparison.ipynb
2023-10-19 22:56:04 +05:30
NirantK
4999fa17b5
* docs(examples): update FastEmbed_vs_HF_Comparison.ipynb
...
with cosine similarity values for BAAI/bge-small-en and BAAI/bge-small-en-v1.5 embeddings
2023-10-19 22:49:25 +05:30
NirantK
ad297c4f13
* chore(Usage_With_Qdrant.ipynb): remove unnecessary outputs in code cells
2023-10-18 23:03:52 +05:30
NirantK
c1fdaf3303
* chore(Supported_Models.ipynb): update supported models table
...
* feat(Supported_Models.ipynb): add size_in_GB column to supported models table
2023-10-18 23:02:50 +05:30
NirantK
b61f8a48cc
Update to v1.5 model
2023-10-18 20:49:31 +05:30
NirantK
d2bdfee4e0
* docs(examples): update comparison notebook with more accurate description of embeddings similarity
2023-10-10 17:56:25 +05:30
NirantK
ac4375516f
Rename nbs; add Cosine similarity check
2023-10-10 17:55:37 +05:30
NirantK
bc402694bd
Update code to handle Generator
2023-10-05 19:17:18 +05:30
NirantK
4949158eff
* chore(Usage_With_Qdrant.ipynb): update notebook title and remove experimental note
...
* docs(Usage_With_Qdrant.ipynb): add support for qdrant-client[fastembed] installation
2023-09-27 17:53:05 +05:30
NirantK
dd2a9e9bcc
* docs(Supported_Models.ipynb): update model descriptions
2023-09-27 17:41:39 +05:30
NirantK
71289926c0
* feat(Supported_Models.ipynb): add example notebook for supported models
2023-09-27 17:40:40 +05:30
NirantK
547130a4c3
Move throughput comparison to docs
2023-09-25 17:40:38 +05:30
NirantK
b72cae1e9b
* chore(docs): rename HF_vs_FastEmbed.ipynb to 02_HF_vs_FastEmbed.ipynb
2023-09-18 08:46:13 +05:30
NirantK
ecfaba9e60
Fix yield/iterator issue
2023-09-18 08:42:34 +05:30
NirantK
5afa103d45
Remove attention pooling
2023-09-18 07:35:52 +05:30
NirantK
9c5d32f271
move nbs
2023-09-07 15:25:32 +05:30
NirantK
1292ded017
Add Qdrant Usage back to the example
2023-09-07 15:19:46 +05:30
NirantK
5a16ce8c63
* chore(docs): rename notebook files
...
* - Rename '03_HF_vs_FastEmbed.ipynb' to 'HF_vs_FastEmbed.ipynb'
* - Rename '01_Retrieval_with_FastEmbed.ipynb' to 'Retrieval_with_FastEmbed
2023-08-25 13:41:32 +05:30
NirantK
743a095a93
* chore(docs): update HF_vs_FastEmbed.ipynb and 04_1M_Embedding_Creation.ipynb
...
*
* HF_vs_FastEmbed.ipynb:
* - Add code to calculate embeddings using mean pooling with attention weighting
* - Fix variable name error in line 4
*
*
2023-08-25 12:59:22 +05:30
NirantK
1bd639b817
* chore(docs): rename 02_Usage_with_Qdrant.ipynb to FastEmbed_Usage_with_Qdrant.ipynb in the experimental folder
2023-08-25 12:56:38 +05:30
NirantK
ac0f3315ed
Cleaner notebook
2023-08-25 01:06:24 +05:30
NirantK
135a4524ac
* refactor(docs/examples): rename files for better organization and numbering
...
* rename(docs/examples): Rename 'Retrieval with FastEmbed.ipynb' to '01_Retrieval_with_FastEmbed.ipynb'
* rename(docs/examples): Rename 'Usage with Qdrant.ipynb
2023-08-25 00:57:39 +05:30
NirantK
2f0e9a8507
* feat(examples): add 1M Embedding Creation notebook
2023-08-25 00:54:23 +05:30
NirantK
7479640a7b
* (feat) HFvsFastEmbed.ipynb: Add new nb example
2023-08-25 00:54:18 +05:30