mirror of
https://github.com/data-privacy-stack/presidio.git
synced 2026-07-31 23:30:46 -05:00
copilot/fix-attribute-error-json-analysis
3 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
eaf7a36c3d |
Add a configurable LangExtract recognizer for use with any provider. (#1815)
* Add a basic, configurable LangExtract-based recognizer class for use with any provider. * Add a basic, configurable LangExtract-based recognizer class for use with any provider. * Address comments (#4) * Address comments * Replace ollama_langextract_recognizer with basic_langextract_recognizer. * Replace ollama_langextract_recognizer with basic_langextract_recognizer. * Replace ollama_langextract_recognizer with basic_langextract_recognizer. * Working so far * Working so far * Working so far * remove dead code * bad comment --------- Co-authored-by: Kassymkhan Bekbolatov <kbekbolatov@solidcore.ai> * Address comment in telackey/lellm (#5) * Replace ollama_langextract_recognizer with basic_langextract_recognizer. * Fix LangExtract error --------- Co-authored-by: Kassymkhan Bekbolatov <kbekbolatov@solidcore.ai> * Remove changes not required * docstring * update docs * ruff --------- Co-authored-by: Kassymkhan Bekbolatov <kasymhan007@gmail.com> Co-authored-by: Kassymkhan Bekbolatov <kbekbolatov@solidcore.ai> |
||
|
|
dd1623fadc |
Add Azure OpenAI support for LangExtract recognizer (#1801)
* Add Azure OpenAI support for LangExtract recognizer * ruff * add redundnant tests to achieve 100% coverage * Add error handling and tests for Azure OpenAI provider initialization * Fix Microsoft Defender secret scanning false positives Replace test API keys with obviously fake placeholders: - test-api-key → PLACEHOLDER_NOT_A_REAL_KEY - test-key-123 → PLACEHOLDER_NOT_A_REAL_KEY - env-key → PLACEHOLDER_FROM_ENV - key → PLACEHOLDER_KEY These are unit test mock values, not real secrets. Using placeholder patterns that won't trigger security scanners while maintaining test validity. All 29 tests passing. * remove bandit * Add bandit tool to Microsoft Security DevOps workflow * Remove bandit from Microsoft Security DevOps workflow tools * Update AzureOpenAILangExtractRecognizer to use deployment name from environment variable * Refactor Azure OpenAI integration: remove legacy provider, update recognizer, and adjust tests * Improve error handling during Azure OpenAI provider registration by logging as error and raising exception * Refactor Azure OpenAI provider initialization logging for improved readability * Refactor test imports in Azure OpenAI recognizer tests for improved clarity and organization * Refactor Azure OpenAI provider tests to remove unnecessary variable assignments for improved clarity * Refactor langextract configuration: reorder supported entities and update entity mappings for consistency * Refactor Azure OpenAI integration: enhance documentation, improve endpoint validation, and streamline provider registration * Refactor Azure OpenAI LangExtract Recognizer: remove unused import and clean up code formatting * Refactor Azure OpenAI provider tests: update imports to use the correct module and remove obsolete test for langextract availability * Refactor Azure authentication handling: consolidate credential management into azure_auth_helper and update related recognizers and tests * Refactor Azure OpenAI recognizers: enhance module imports for registration, streamline model ID handling, and improve test coverage for credential selection * Refactor AHDS Surrogate operator: streamline error handling by mocking Azure credentials and client, and improve code readability * Refactor AHDS Recognizer tests: replace multiple credential mocks with a single get_azure_credential mock for improved clarity and maintainability |
||
|
|
08219e7663 |
Language models integration (LangExtract) (#1775)
* Add LangExtract recognizer for PII extraction - Introduced LangExtract recognizer to enhance PII detection capabilities. - Added configuration files for LangExtract prompts and examples. - Implemented LangExtractRecognizer class to handle PII extraction using LangExtract. - Created tests for LangExtract recognizer to ensure functionality and reliability. - Added a simple standalone test script for quick validation of LangExtract setup. - Updated pyproject.toml to include langextract as a dependency. * refine the docs * narrow support for oollama only * Refactor LangExtract tests to use Ollama; remove API key dependency * adding first draft of docker compose * Update model_id in tests to use 'gemma2:2b' instead of 'gemini-2.5-flash' * Refactor LangExtract documentation to focus on Ollama support; remove references to other LLM providers * Update README to remove Ollama setup instructions and clarify integration guide reference * Enhance Ollama installation script with progress messages and error handling; update model download method for better user feedback * auto ruff fixes * Enhance LangExtractRecognizer tests with real Ollama integration - Updated `langextract_recognizer_class` fixture to create a test-specific configuration for LangExtractRecognizer, enabling it for testing. - Refactored tests in `test_langextract_recognizer.py` to utilize the new configuration and validate the recognizer's behavior with real Ollama. - Removed mock-based tests for LangExtract and replaced them with integration tests that check the recognizer's functionality against a running Ollama instance. - Added tests to verify the recognizer's initialization, entity detection, and error handling when the Ollama server is unreachable. - Ensured that only requested entities are returned and that results include analysis explanations. * Remove unnecessary line breaks in LLM-based PII detection section of README * Add LangExtract LLM-based PII detection test and configuration * Improve Ollama availability check with setup attempt message * Increase wait time for services and update healthcheck parameters for Ollama service * Add Ollama setup for Analyzer tests and improve availability check * Set timeout for Ollama setup in Analyzer tests to 8 minutes * Enhance Ollama setup for Analyzer tests with improved installation and readiness checks * Update Ollama model references from gemma2:2b to llama3.2:1b across configuration and test files * Update model references from llama3.2:1b to gemma2:2b across configuration, scripts, and tests * Remove 'enabled' configuration from LangExtract settings in YAML files and update tests accordingly * Update Ollama service configuration: change port mapping and modify healthcheck command * Refactor logging in LangExtractRecognizer: reduce verbosity and improve clarity of extraction results * Update Ollama service configuration: modify port mapping and healthcheck command * Update CI workflow and tests: reduce sleep duration and add environment variable for LangExtract recognizer * Reduce sleep duration in CI workflow from 150 to 60 seconds * Update LangExtract model references from gemma2:2b to gemma3:1b and remove obsolete installation script * docs and prompt fixes * finalizing the pr * Update Ollama image to latest version and add LangExtract PII/PHI extraction examples and prompts * fix bad example * fix unit-tests * refactor: clean up .env file and simplify skip_engine logic in tests * chore: add a new line to .env file for better readability * chore: remove unnecessary blank line from .env file * revert .env * intial commit * Remove skip marker for spacy_nlp_engine fixture * Remove skip markers for stanza and transformers NLP engine fixtures * Remove Ollama recognizer test and update default recognizers configuration * Remove unused Ollama recognizer configuration and update prompt file references * Add end-to-end tests for API anonymization and redaction features - Implemented tests for the anonymization API in `test_api_anonymizer.py`, covering various scenarios including valid requests, empty inputs, malformed requests, and custom anonymizers. - Created integration tests in `test_api_e2e_integration_flows.py` to validate the analyze and anonymize workflow with PII detection. - Added tests for image redaction functionality in `test_api_image_redactor.py`, ensuring proper handling of image data and error responses. - Developed package-level tests in `test_package_e2e_integration_flows.py` to verify the functionality of the analyzer and anonymizer engines, including support for third-party recognizers. * Remove unused Ollama recognizer imports and related tests * Update requirements and improve Ollama recognizer availability checks in e2e tests * Fix formatting in requirements.txt for analyzer and anonymizer dependencies * Update Ollama model ID from gemma3:1b to gemma2:2b in configuration and tests * gemma2:2b * finalizing the pr * Remove unused ABC import from lm_recognizer.py * Fix indentation in docker-compose.yml for volumes section * Fix line break for clarity in adding_recognizers.md * Add timeout settings for Ollama recognizer and test cases * Refactor timeout comment for clarity in OllamaLangExtractRecognizer * Update Ollama model version and add configuration for LangExtract recognizer tests * Update examples_file path in configuration for Ollama recognizer * Remove timeout decorator from Ollama recognizer * Add rerun settings to unit and E2E tests for improved stability * Remove rerun settings from unit and E2E test commands for simplification * Set max-parallel to 2 for local build and E2E tests * Remove max-parallel setting from local build and E2E tests * move poetry cache dir * pr changes * code review changes * ruff check * remove unused json import in test_ollama_recognizer.py * refactor test names for clarity and consistency in test_ollama_recognizer.py * finalizing the PR * self code review fixes * Refactor LangExtractRecognizer to use yaml for configuration loading * ruff fixes * Update error messages in OllamaLangExtractRecognizer tests for clarity * CR comment addressed * Remove unused variables from Jinja2 prompt rendering in LangExtractRecognizer * exporting functionality to helpers enlarging composition * composition * Refactor entity mapper and langextract utilities for improved clarity and consistency * Refactor tests to use get_langextract_module for mocking LangExtract availability * Refactor langextract utilities to improve clarity and error handling; remove deprecated functions and update tests accordingly * Refactor LLM utilities by simplifying docstrings and consolidating imports for improved readability and maintainability * Update error message for missing Jinja2 installation to include poetry installation instructions * Add Ollama recognizer configuration and tests for YAML integration - Introduced `test_ollama_enabled_recognizers.yaml` to define recognizers including OllamaLangExtractRecognizer. - Enhanced `test_package_e2e_integration_flows.py` with a test to validate loading of Ollama recognizer from YAML configuration. - Updated `OllamaLangExtractRecognizer` to support configuration path and language parameters. - Improved handling of relative paths for configuration files. * Refactor docstrings in OllamaLangExtractRecognizer for improved clarity and formatting * Enhance OllamaLangExtractRecognizer initialization docstring to clarify kwargs usage * pr comments * Refactor OllamaLangExtractRecognizer to streamline config path handling and remove redundant comments * Refactor Ollama recognizer test to improve clarity and enhance entity detection validation * Refactor tests for Ollama recognizer and LMRecognizer to improve exception handling and configuration validation * Update config path for Ollama recognizer in test configuration * Update config paths for Ollama recognizer and add test configuration for LangExtract * Remove test configuration for Ollama LangExtract * Remove test configuration for Ollama LangExtract recognizer * Fix formatting in resolve_config_path function for improved readability * Enable UsLangExtractRecognizer and update its config path * change all configs to use gemma3:1b * Disable Ollama LangExtract recognizer and update its configuration path * Update langextract configuration paths to use absolute paths for prompt and examples files * Remove test script for Ollama recognizer configuration loading * Refactor config loading in examples and prompt loaders to use resolve_config_path; update logging level in LMRecognizer; add langextract availability check in OllamaLangExtractRecognizer. * Refactor parameter description in load_yaml_examples and clean up imports in prompt_loader * Update langextract paths to use repo-root-relative paths in tests and prompt loader * Enhance documentation for Ollama setup and improve __init__.py imports for clarity and maintainability * code review changes * pr comments & align to main --------- Co-authored-by: Tamir Kamara <26870601+tamirkamara@users.noreply.github.com> Co-authored-by: Sharon Hart <sharonh.dev@gmail.com> |