Commit Graph

3 Commits

Author SHA1 Message Date
Thomas E Lackey
eaf7a36c3d Add a configurable LangExtract recognizer for use with any provider. (#1815)
* Add a basic, configurable LangExtract-based recognizer class for use with any provider.

* Add a basic, configurable LangExtract-based recognizer class for use with any provider.

* Address comments (#4)

* Address comments

* Replace ollama_langextract_recognizer with basic_langextract_recognizer.

* Replace ollama_langextract_recognizer with basic_langextract_recognizer.

* Replace ollama_langextract_recognizer with basic_langextract_recognizer.

* Working so far

* Working so far

* Working so far

* remove dead code

* bad comment

---------

Co-authored-by: Kassymkhan Bekbolatov <kbekbolatov@solidcore.ai>

* Address comment in telackey/lellm (#5)

* Replace ollama_langextract_recognizer with basic_langextract_recognizer.

* Fix LangExtract error

---------

Co-authored-by: Kassymkhan Bekbolatov <kbekbolatov@solidcore.ai>

* Remove changes not required

* docstring

* update docs

* ruff

---------

Co-authored-by: Kassymkhan Bekbolatov <kasymhan007@gmail.com>
Co-authored-by: Kassymkhan Bekbolatov <kbekbolatov@solidcore.ai>
2026-02-03 10:38:06 +02:00
Dor Lugasi-Gal
dd1623fadc Add Azure OpenAI support for LangExtract recognizer (#1801)
* Add Azure OpenAI support for LangExtract recognizer

* ruff

* add redundnant tests to achieve 100% coverage

* Add error handling and tests for Azure OpenAI provider initialization

* Fix Microsoft Defender secret scanning false positives

Replace test API keys with obviously fake placeholders:
- test-api-key → PLACEHOLDER_NOT_A_REAL_KEY
- test-key-123 → PLACEHOLDER_NOT_A_REAL_KEY
- env-key → PLACEHOLDER_FROM_ENV
- key → PLACEHOLDER_KEY

These are unit test mock values, not real secrets. Using
placeholder patterns that won't trigger security scanners
while maintaining test validity. All 29 tests passing.

* remove bandit

* Add bandit tool to Microsoft Security DevOps workflow

* Remove bandit from Microsoft Security DevOps workflow tools

* Update AzureOpenAILangExtractRecognizer to use deployment name from environment variable

* Refactor Azure OpenAI integration: remove legacy provider, update recognizer, and adjust tests

* Improve error handling during Azure OpenAI provider registration by logging as error and raising exception

* Refactor Azure OpenAI provider initialization logging for improved readability

* Refactor test imports in Azure OpenAI recognizer tests for improved clarity and organization

* Refactor Azure OpenAI provider tests to remove unnecessary variable assignments for improved clarity

* Refactor langextract configuration: reorder supported entities and update entity mappings for consistency

* Refactor Azure OpenAI integration: enhance documentation, improve endpoint validation, and streamline provider registration

* Refactor Azure OpenAI LangExtract Recognizer: remove unused import and clean up code formatting

* Refactor Azure OpenAI provider tests: update imports to use the correct module and remove obsolete test for langextract availability

* Refactor Azure authentication handling: consolidate credential management into azure_auth_helper and update related recognizers and tests

* Refactor Azure OpenAI recognizers: enhance module imports for registration, streamline model ID handling, and improve test coverage for credential selection

* Refactor AHDS Surrogate operator: streamline error handling by mocking Azure credentials and client, and improve code readability

* Refactor AHDS Recognizer tests: replace multiple credential mocks with a single get_azure_credential mock for improved clarity and maintainability
2025-12-04 12:11:19 +02:00
Ron Shakutai
08219e7663 Language models integration (LangExtract) (#1775)
* Add LangExtract recognizer for PII extraction

- Introduced LangExtract recognizer to enhance PII detection capabilities.
- Added configuration files for LangExtract prompts and examples.
- Implemented LangExtractRecognizer class to handle PII extraction using LangExtract.
- Created tests for LangExtract recognizer to ensure functionality and reliability.
- Added a simple standalone test script for quick validation of LangExtract setup.
- Updated pyproject.toml to include langextract as a dependency.

* refine the docs

* narrow support for oollama only

* Refactor LangExtract tests to use Ollama; remove API key dependency

* adding first draft of docker compose

* Update model_id in tests to use 'gemma2:2b' instead of 'gemini-2.5-flash'

* Refactor LangExtract documentation to focus on Ollama support; remove references to other LLM providers

* Update README to remove Ollama setup instructions and clarify integration guide reference

* Enhance Ollama installation script with progress messages and error handling; update model download method for better user feedback

* auto ruff fixes

* Enhance LangExtractRecognizer tests with real Ollama integration

- Updated `langextract_recognizer_class` fixture to create a test-specific configuration for LangExtractRecognizer, enabling it for testing.
- Refactored tests in `test_langextract_recognizer.py` to utilize the new configuration and validate the recognizer's behavior with real Ollama.
- Removed mock-based tests for LangExtract and replaced them with integration tests that check the recognizer's functionality against a running Ollama instance.
- Added tests to verify the recognizer's initialization, entity detection, and error handling when the Ollama server is unreachable.
- Ensured that only requested entities are returned and that results include analysis explanations.

* Remove unnecessary line breaks in LLM-based PII detection section of README

* Add LangExtract LLM-based PII detection test and configuration

* Improve Ollama availability check with setup attempt message

* Increase wait time for services and update healthcheck parameters for Ollama service

* Add Ollama setup for Analyzer tests and improve availability check

* Set timeout for Ollama setup in Analyzer tests to 8 minutes

* Enhance Ollama setup for Analyzer tests with improved installation and readiness checks

* Update Ollama model references from gemma2:2b to llama3.2:1b across configuration and test files

* Update model references from llama3.2:1b to gemma2:2b across configuration, scripts, and tests

* Remove 'enabled' configuration from LangExtract settings in YAML files and update tests accordingly

* Update Ollama service configuration: change port mapping and modify healthcheck command

* Refactor logging in LangExtractRecognizer: reduce verbosity and improve clarity of extraction results

* Update Ollama service configuration: modify port mapping and healthcheck command

* Update CI workflow and tests: reduce sleep duration and add environment variable for LangExtract recognizer

* Reduce sleep duration in CI workflow from 150 to 60 seconds

* Update LangExtract model references from gemma2:2b to gemma3:1b and remove obsolete installation script

* docs and prompt fixes

* finalizing the pr

* Update Ollama image to latest version and add LangExtract PII/PHI extraction examples and prompts

* fix bad example

* fix unit-tests

* refactor: clean up .env file and simplify skip_engine logic in tests

* chore: add a new line to .env file for better readability

* chore: remove unnecessary blank line from .env file

* revert .env

* intial commit

* Remove skip marker for spacy_nlp_engine fixture

* Remove skip markers for stanza and transformers NLP engine fixtures

* Remove Ollama recognizer test and update default recognizers configuration

* Remove unused Ollama recognizer configuration and update prompt file references

* Add end-to-end tests for API anonymization and redaction features

- Implemented tests for the anonymization API in `test_api_anonymizer.py`, covering various scenarios including valid requests, empty inputs, malformed requests, and custom anonymizers.
- Created integration tests in `test_api_e2e_integration_flows.py` to validate the analyze and anonymize workflow with PII detection.
- Added tests for image redaction functionality in `test_api_image_redactor.py`, ensuring proper handling of image data and error responses.
- Developed package-level tests in `test_package_e2e_integration_flows.py` to verify the functionality of the analyzer and anonymizer engines, including support for third-party recognizers.

* Remove unused Ollama recognizer imports and related tests

* Update requirements and improve Ollama recognizer availability checks in e2e tests

* Fix formatting in requirements.txt for analyzer and anonymizer dependencies

* Update Ollama model ID from gemma3:1b to gemma2:2b in configuration and tests

* gemma2:2b

* finalizing the pr

* Remove unused ABC import from lm_recognizer.py

* Fix indentation in docker-compose.yml for volumes section

* Fix line break for clarity in adding_recognizers.md

* Add timeout settings for Ollama recognizer and test cases

* Refactor timeout comment for clarity in OllamaLangExtractRecognizer

* Update Ollama model version and add configuration for LangExtract recognizer tests

* Update examples_file path in configuration for Ollama recognizer

* Remove timeout decorator from Ollama recognizer

* Add rerun settings to unit and E2E tests for improved stability

* Remove rerun settings from unit and E2E test commands for simplification

* Set max-parallel to 2 for local build and E2E tests

* Remove max-parallel setting from local build and E2E tests

* move poetry cache dir

* pr changes

* code review changes

* ruff check

* remove unused json import in test_ollama_recognizer.py

* refactor test names for clarity and consistency in test_ollama_recognizer.py

* finalizing the PR

* self code review fixes

* Refactor LangExtractRecognizer to use yaml for configuration loading

* ruff fixes

* Update error messages in OllamaLangExtractRecognizer tests for clarity

* CR comment addressed

* Remove unused variables from Jinja2 prompt rendering in LangExtractRecognizer

* exporting functionality to helpers enlarging composition

* composition

* Refactor entity mapper and langextract utilities for improved clarity and consistency

* Refactor tests to use get_langextract_module for mocking LangExtract availability

* Refactor langextract utilities to improve clarity and error handling; remove deprecated functions and update tests accordingly

* Refactor LLM utilities by simplifying docstrings and consolidating imports for improved readability and maintainability

* Update error message for missing Jinja2 installation to include poetry installation instructions

* Add Ollama recognizer configuration and tests for YAML integration

- Introduced `test_ollama_enabled_recognizers.yaml` to define recognizers including OllamaLangExtractRecognizer.
- Enhanced `test_package_e2e_integration_flows.py` with a test to validate loading of Ollama recognizer from YAML configuration.
- Updated `OllamaLangExtractRecognizer` to support configuration path and language parameters.
- Improved handling of relative paths for configuration files.

* Refactor docstrings in OllamaLangExtractRecognizer for improved clarity and formatting

* Enhance OllamaLangExtractRecognizer initialization docstring to clarify kwargs usage

* pr comments

* Refactor OllamaLangExtractRecognizer to streamline config path handling and remove redundant comments

* Refactor Ollama recognizer test to improve clarity and enhance entity detection validation

* Refactor tests for Ollama recognizer and LMRecognizer to improve exception handling and configuration validation

* Update config path for Ollama recognizer in test configuration

* Update config paths for Ollama recognizer and add test configuration for LangExtract

* Remove test configuration for Ollama LangExtract

* Remove test configuration for Ollama LangExtract recognizer

* Fix formatting in resolve_config_path function for improved readability

* Enable UsLangExtractRecognizer and update its config path

* change all configs to use gemma3:1b

* Disable Ollama LangExtract recognizer and update its configuration path

* Update langextract configuration paths to use absolute paths for prompt and examples files

* Remove test script for Ollama recognizer configuration loading

* Refactor config loading in examples and prompt loaders to use resolve_config_path; update logging level in LMRecognizer; add langextract availability check in OllamaLangExtractRecognizer.

* Refactor parameter description in load_yaml_examples and clean up imports in prompt_loader

* Update langextract paths to use repo-root-relative paths in tests and prompt loader

* Enhance documentation for Ollama setup and improve __init__.py imports for clarity and maintainability

* code review changes

* pr comments & align to main

---------

Co-authored-by: Tamir Kamara <26870601+tamirkamara@users.noreply.github.com>
Co-authored-by: Sharon Hart <sharonh.dev@gmail.com>
2025-11-30 15:22:10 +02:00