Add Ruff linter + Apply Ruff fix (#1379)

* Add Ruff linter + Apply Ruff fix

* Move up linting

* Move up linting

* Move up linting

* docs

* docs
This commit is contained in:
Sharon Hart
2024-05-12 10:16:47 +03:00
committed by GitHub
parent 2348fff508
commit cb0184afa2
28 changed files with 163 additions and 134 deletions

View File

@@ -3,14 +3,6 @@ parameters:
- name: WORKING_FOLDER - name: WORKING_FOLDER
steps: steps:
- task: Bash@3
displayName: 'Linting: ${{ parameters.SERVICE }}'
inputs:
targetType: 'inline'
workingDirectory: ${{ parameters.WORKING_FOLDER }}
script: |
set -eux # fail on error
pipenv run flake8
- task: Bash@3 - task: Bash@3
displayName: 'Unit tests: ${{ parameters.SERVICE }}' displayName: 'Unit tests: ${{ parameters.SERVICE }}'

View File

@@ -11,6 +11,21 @@ stages:
steps: steps:
- template: ./security-analysis.yml - template: ./security-analysis.yml
- job: Linting
displayName: Linting
pool:
vmImage: 'ubuntu-latest'
steps:
- task: Bash@3
displayName: 'Linting: Presidio for $(python.version)'
inputs:
targetType: 'inline'
script: |
set -eux # fail on error
pip install ruff
ruff check
- job: TestAnalyzer - job: TestAnalyzer
displayName: Test Analyzer displayName: Test Analyzer
pool: pool:

View File

@@ -1,29 +1,7 @@
repos: repos:
- repo: https://github.com/ambv/black - repo: https://github.com/astral-sh/ruff-pre-commit
rev: 22.3.0 rev: v0.4.3
hooks: hooks:
- id: black - id: ruff
language_version: python3 args: [ --fix ]
exclude: ^tests/ - id: ruff-format
- repo: https://github.com/pycqa/flake8
rev: 3.9.0
hooks:
- id: flake8
additional_dependencies: [
'pep8-naming',
'flake8-docstrings',
]
args: ['--max-line-length=88',
'--docstring-convention=numpy',
# 'PEP8 Rules' to ignore in tests. Ignore documentation rules for all tests
# and ignore long lines / whitespaces for e2e-tests where we define jsons in-code.
'--per-file-ignores=**/tests/**.py:D docs/**.py:D e2e-tests/**.py:D,E501,W291,W293 docs/samples/deployments/spark/notebooks/*.py:E501,F821,D103',
'--extend-ignore=
E203,
D100,
D202,
ANN101,
ANN102,
ANN204,
ANN203'
]

View File

@@ -38,7 +38,7 @@ To get started, refer to the documentation for [setting up a development environ
### How to test? ### How to test?
For Python, Presidio leverages `pytest` and `flake8`. See [this tutorial](docs/development.md#testing) on more information on testing presidio modules. For Python, Presidio leverages `pytest` and `ruff`. See [this tutorial](docs/development.md#testing) on more information on testing presidio modules.
### Adding new recognizers for new PII types ### Adding new recognizers for new PII types

View File

@@ -56,7 +56,7 @@ Follow these steps when starting to work on a Presidio service with Pipenv:
4. To run arbitrary scripts within the virtual env, start the command with 4. To run arbitrary scripts within the virtual env, start the command with
`pipenv run`. For example: `pipenv run`. For example:
1. `pipenv run flake8` 1. `pipenv run ruff check`
2. `pipenv run pip freeze` 2. `pipenv run pip freeze`
3. `pipenv run python -m spacy download en_core_web_lg` 3. `pipenv run python -m spacy download en_core_web_lg`
@@ -233,39 +233,36 @@ run.bat
### Linting ### Linting
Presidio services are PEP8 compliant and continuously enforced on style guide issues during the build process using `flake8`. Presidio services are PEP8 compliant and continuously enforced on style guide issues during the build process using `ruff`, in turn running `flake8` and other linters.
Running flake8 locally, using `pipenv run flake8`, you can check for those issues prior to committing a change. Running ruff locally, using `pipenv run ruff check`, you can check for those issues prior to committing a change.
In addition to the basic `flake8` functionality, Presidio uses the following extensions: Ruff runs linters in addition to the basic `flake8` functionality, Presidio uses linters as part as ruff such as:
- _pep8-naming_: To check that variable names are PEP8 compliant. - _pep8-naming_: To check that variable names are PEP8 compliant.
- _flake8-docstrings_: To check that docstrings are compliant. - _flake8-docstrings_: To check that docstrings are compliant.
### Automatically format code and check for code styling ### Automatically format code and check for code styling
To make the linting process easier, you can use pre-commit hooks to verify and automatically format code upon a git commit, using `black`: To make the linting process easier, you can use pre-commit hooks to verify and automatically format code upon a git commit, using `ruff-format`:
1. [Install pre-commit package manager locally.](https://pre-commit.com/#install) 1. [Install pre-commit package manager locally.](https://pre-commit.com/#install)
2. From the project's root, enable pre-commit, installing git hooks in the `.git/` directory by running: `pre-commit install`. 2. From the project's root, enable pre-commit, installing git hooks in the `.git/` directory by running: `pre-commit install`.
3. Commit non PEP8 compliant code will cause commit failure and automatically 3. Commit non PEP8 compliant code will cause commit failure and automatically
format your code using `black`, as well as checking code formatting using `flake8` format your code using, as well as checking code formatting using `ruff`
```sh ```sh
>git commit -m 'autoformat' presidio-analyzer/presidio_analyzer/predefined_recognizers/us_ssn_recognizer.py [INFO] Initializing environment for https://github.com/astral-sh/ruff-pre-commit.
[INFO] Installing environment for https://github.com/astral-sh/ruff-pre-commit.
black....................................................................Failed [INFO] Once installed this environment will be reused.
- hook id: black [INFO] This may take a few minutes...
- files were modified by this hook ruff.....................................................................Passed
ruff-format..............................................................Failed
reformatted presidio-analyzer/presidio_analyzer/predefined_recognizers/us_ssn_recognizer.py - hook id: ruff-format
All done! - files were modified by this hook
1 file reformatted. 5 files reformatted, 4 files left unchanged
```
flake8...................................................................Passed
```
4. Committing again will finish successfully, with a well-formatted code. 4. Committing again will finish successfully, with a well-formatted code.

View File

@@ -20,8 +20,6 @@ azure-core = "*"
[dev-packages] [dev-packages]
pytest = "*" pytest = "*"
pytest-mock = "*" pytest-mock = "*"
flake8= {version = ">=3.7.9"} ruff = "*"
pep8-naming = "*"
flake8-docstrings = "*"
pre_commit = "*" pre_commit = "*"
python-dotenv = "*" python-dotenv = "*"

View File

@@ -187,7 +187,7 @@ class AnalyzerEngine:
>>> results = analyzer.analyze(text='My phone number is 212-555-5555', entities=['PHONE_NUMBER'], language='en') # noqa D501 >>> results = analyzer.analyze(text='My phone number is 212-555-5555', entities=['PHONE_NUMBER'], language='en') # noqa D501
>>> print(results) >>> print(results)
[type: PHONE_NUMBER, start: 19, end: 31, score: 0.85] [type: PHONE_NUMBER, start: 19, end: 31, score: 0.85]
""" """ # noqa: E501
all_fields = not entities all_fields = not entities

View File

@@ -19,7 +19,6 @@ class BatchAnalyzerEngine:
""" """
def __init__(self, analyzer_engine: Optional[AnalyzerEngine] = None): def __init__(self, analyzer_engine: Optional[AnalyzerEngine] = None):
self.analyzer_engine = analyzer_engine self.analyzer_engine = analyzer_engine
if not analyzer_engine: if not analyzer_engine:
self.analyzer_engine = AnalyzerEngine() self.analyzer_engine = AnalyzerEngine()
@@ -42,10 +41,10 @@ class BatchAnalyzerEngine:
texts = self._validate_types(texts) texts = self._validate_types(texts)
# Process the texts as batch for improved performance # Process the texts as batch for improved performance
nlp_artifacts_batch: Iterator[ nlp_artifacts_batch: Iterator[Tuple[str, NlpArtifacts]] = (
Tuple[str, NlpArtifacts] self.analyzer_engine.nlp_engine.process_batch(
] = self.analyzer_engine.nlp_engine.process_batch( texts=texts, language=language
texts=texts, language=language )
) )
list_results = [] list_results = []
@@ -127,7 +126,7 @@ class BatchAnalyzerEngine:
@staticmethod @staticmethod
def _validate_types(value_iterator: Iterable[Any]) -> Iterator[Any]: def _validate_types(value_iterator: Iterable[Any]) -> Iterator[Any]:
for val in value_iterator: for val in value_iterator:
if val and not type(val) in (int, float, bool, str): if val and type(val) not in (int, float, bool, str):
err_msg = ( err_msg = (
"Analyzer.analyze_iterator only works " "Analyzer.analyze_iterator only works "
"on primitive types (int, float, bool, str). " "on primitive types (int, float, bool, str). "

View File

@@ -27,7 +27,8 @@ class AzureAILanguageRecognizer(RemoteRecognizer):
azure_ai_endpoint: Optional[str] = None, azure_ai_endpoint: Optional[str] = None,
): ):
""" """
Wrapper for the PII detection in Azure AI Language Wrap the PII detection in Azure AI Language.
:param supported_entities: List of supported entities for this recognizer. :param supported_entities: List of supported entities for this recognizer.
If None, all supported entities will be used. If None, all supported entities will be used.
:param supported_language: Language code to use for the recognizer. :param supported_language: Language code to use for the recognizer.
@@ -36,8 +37,8 @@ class AzureAILanguageRecognizer(RemoteRecognizer):
:param azure_ai_key: Azure AI for language key :param azure_ai_key: Azure AI for language key
:param azure_ai_endpoint: Azure AI for language endpoint :param azure_ai_endpoint: Azure AI for language endpoint
For more info, see https://learn.microsoft.com/en-us/azure/ai-services/language-service/personally-identifiable-information/overview # noqa For more info, see https://learn.microsoft.com/en-us/azure/ai-services/language-service/personally-identifiable-information/overview
""" """ # noqa E501
super().__init__( super().__init__(
supported_entities=supported_entities, supported_entities=supported_entities,
@@ -73,7 +74,7 @@ class AzureAILanguageRecognizer(RemoteRecognizer):
@staticmethod @staticmethod
def __get_azure_ai_supported_entities() -> List[str]: def __get_azure_ai_supported_entities() -> List[str]:
"""Return the list of all supported entities for Azure AI Language.""" """Return the list of all supported entities for Azure AI Language."""
from azure.ai.textanalytics._models import PiiEntityCategory # noqa from azure.ai.textanalytics._models import PiiEntityCategory # noqa
return [r.value.upper() for r in PiiEntityCategory] return [r.value.upper() for r in PiiEntityCategory]

View File

@@ -334,9 +334,9 @@ class RecognizerRegistry:
:example: :example:
>>> registry = RecognizerRegistry() >>> registry = RecognizerRegistry()
>>> recognizer = { "name": "Titles Recognizer", "supported_language": "de","supported_entity": "TITLE", "deny_list": ["Mr.","Mrs."]} # noqa: E501 >>> recognizer = { "name": "Titles Recognizer", "supported_language": "de","supported_entity": "TITLE", "deny_list": ["Mr.","Mrs."]}
>>> registry.add_pattern_recognizer_from_dict(recognizer) >>> registry.add_pattern_recognizer_from_dict(recognizer)
""" """ # noqa: E501
recognizer = PatternRecognizer.from_dict(recognizer_dict) recognizer = PatternRecognizer.from_dict(recognizer_dict)
self.add_recognizer(recognizer) self.add_recognizer(recognizer)

View File

@@ -1,10 +0,0 @@
[flake8]
max-line-length = 88
exclude =
.git,
__pycache__,
build,
dist,
tests
docstring-convention = numpy
extend-ignore = E203 D100 D202 ANN101 ANN102 ANN204 ANN203 TC

View File

@@ -1,4 +1,5 @@
"""Setup.py for Presidio Analyzer.""" """Setup.py for Presidio Analyzer."""
import os.path import os.path
from os import path from os import path
@@ -27,7 +28,7 @@ setuptools.setup(
"presidio_analyzer": ["py.typed", "conf/*"], "presidio_analyzer": ["py.typed", "conf/*"],
}, },
trusted_host=["pypi.org"], trusted_host=["pypi.org"],
tests_require=["pytest", "flake8>=3.7.9"], tests_require=["pytest", "ruff"],
install_requires=[ install_requires=[
"spacy>=3.4.4, <4.0.0", "spacy>=3.4.4, <4.0.0",
"regex", "regex",

View File

@@ -9,7 +9,5 @@ pycryptodome = ">=3.10,<4.0.0"
[dev-packages] [dev-packages]
pytest = "*" pytest = "*"
flake8 = { version = ">=3.7.9" } ruff = "*"
pep8-naming = "*"
flake8-docstrings = "*"
pre_commit = "*" pre_commit = "*"

View File

@@ -1,6 +0,0 @@
[flake8]
max-line-length = 88
docstring-convention = numpy
per-file-ignores =
tests/*: D
extend-ignore = E203 D100 D202 ANN101 ANN102 ANN204 ANN203

View File

@@ -5,7 +5,7 @@ from os import path
from setuptools import setup, find_packages from setuptools import setup, find_packages
test_requirements = ["pytest>=3", "flake8==3.7.9"] test_requirements = ["pytest>=3", "ruff"]
__version__ = "" __version__ = ""
this_directory = path.abspath(path.dirname(__file__)) this_directory = path.abspath(path.dirname(__file__))

View File

@@ -5,7 +5,7 @@ pathspec = "*"
[dev-packages] [dev-packages]
pytest = ">=6" pytest = ">=6"
flake8= {version = ">=3.7"} ruff = "*"
pre_commit = ">=2" pre_commit = ">=2"
pytest-cov = "*" pytest-cov = "*"
pytest-mock = "*" pytest-mock = "*"

View File

@@ -0,0 +1,6 @@
[tool.ruff]
line-length = 120
# To be fixed:
[tool.ruff.lint]
ignore = ["D205", "D400", "E721"]

View File

@@ -1,12 +1,6 @@
[bdist_wheel] [bdist_wheel]
universal = 1 universal = 1
[flake8]
import-order-style = pep8
application-import-names = presidio_cli
ignore = E203,W503
max-line-length = 120
[build_sphinx] [build_sphinx]
all-files = 1 all-files = 1
source-dir = docs source-dir = docs

View File

@@ -31,5 +31,5 @@ setup(
], ],
install_requires=["presidio-analyzer>=2.2", "pyyaml", "pathspec"], install_requires=["presidio-analyzer>=2.2", "pyyaml", "pathspec"],
trusted_host=["pypi.org"], trusted_host=["pypi.org"],
tests_require=["pytest", "flake8>=3.7.9"], tests_require=["pytest", "ruff"],
) )

View File

@@ -19,6 +19,4 @@ azure-ai-formrecognizer = ">=3.3.0,<4.0.0"
[dev-packages] [dev-packages]
pytest = "*" pytest = "*"
pytest-mock = "*" pytest-mock = "*"
flake8 = { version = ">=3.9.2" } ruff = "*"
pep8-naming = "*"
flake8-docstrings = "*"

View File

@@ -0,0 +1,6 @@
[tool.ruff]
line-length = 120
# To be fixed:
[tool.ruff.lint]
ignore = ["F401", "F841", "E402", "E721", "F541", "E712", "F811", "E722"]

View File

@@ -1,10 +0,0 @@
[flake8]
max-line-length = 88
exclude =
.git,
__pycache__,
build,
dist,
tests
docstring-convention = numpy
extend-ignore = E203 D100 D202 D407 ANN101 ANN102 ANN204 ANN203

View File

@@ -1,4 +1,5 @@
"""Setup.py for Presidio Image Redactor.""" """Setup.py for Presidio Image Redactor."""
import os.path import os.path
from os import path from os import path
@@ -12,10 +13,10 @@ requirements = [
"pydicom>=2.3.0", "pydicom>=2.3.0",
"pypng>=0.20220715.0", "pypng>=0.20220715.0",
"azure-ai-formrecognizer>=3.3.0,<4.0.0", "azure-ai-formrecognizer>=3.3.0,<4.0.0",
"opencv-python>=4.0.0,<5.0.0" "opencv-python>=4.0.0,<5.0.0",
] ]
test_requirements = ["pytest>=3", "pytest-mock>=3.10.0", "flake8>=3.7.9"] test_requirements = ["pytest>=3", "pytest-mock>=3.10.0", "ruff"]
__version__ = "" __version__ = ""
this_directory = path.abspath(path.dirname(__file__)) this_directory = path.abspath(path.dirname(__file__))

View File

@@ -11,7 +11,5 @@ pandas = ">=1.5.2"
[dev-packages] [dev-packages]
pytest = "*" pytest = "*"
flake8 = { version = ">=3.7.9" } ruff = "*"
pep8-naming = "*"
flake8-docstrings = "*"
pre_commit = "*" pre_commit = "*"

View File

@@ -0,0 +1,6 @@
[tool.ruff]
line-length = 120
# To be fixed:
[tool.ruff.lint]
ignore = ["F841"]

View File

@@ -1,10 +0,0 @@
[flake8]
max-line-length = 88
exclude =
.git,
__pycache__,
build,
dist,
tests
docstring-convention = numpy
extend-ignore = E203 D100 D202 ANN101 ANN102 ANN204 ANN203 TC

View File

@@ -5,7 +5,7 @@ from os import path
from setuptools import setup, find_packages from setuptools import setup, find_packages
test_requirements = ["pytest>=3", "flake8==3.7.9"] test_requirements = ["pytest>=3", "ruff"]
__version__ = "" __version__ = ""
this_directory = path.abspath(path.dirname(__file__)) this_directory = path.abspath(path.dirname(__file__))
@@ -15,9 +15,7 @@ with open(path.join(this_directory, "README.md"), encoding="utf-8") as f:
long_description = f.read() long_description = f.read()
try: try:
with open( with open(os.path.join(parent_directory, "PRESIDIO-STRUCTURED-VERSION")) as version_file:
os.path.join(parent_directory, "PRESIDIO-STRUCTURED-VERSION")
) as version_file:
__version__ = version_file.read().strip() __version__ = version_file.read().strip()
except Exception: except Exception:
__version__ = os.environ.get("PRESIDIO_STRUCTURED_VERSION", "0.0.1-alpha") __version__ = os.environ.get("PRESIDIO_STRUCTURED_VERSION", "0.0.1-alpha")

79
pyproject.toml Normal file
View File

@@ -0,0 +1,79 @@
[tool.ruff]
exclude = [
# Ruff recommended:
".bzr",
".direnv",
".eggs",
".git",
".git-rewrite",
".hg",
".ipynb_checkpoints",
".mypy_cache",
".nox",
".pants.d",
".pyenv",
".pytest_cache",
".pytype",
".ruff_cache",
".svn",
".tox",
".venv",
".vscode",
"__pypackages__",
"_build",
"buck-out",
"build",
"dist",
"node_modules",
"site-packages",
"venv",
# Project specific:
"docs/samples",
"e2e-tests/",
"*/tests/*"
]
# Same as Black.
line-length = 88
indent-width = 4
[tool.ruff.lint]
select = ["E", "F", "I", "D", "N", "W",
# To be added:
# "SIM", "UP", "ANN", "B"
]
ignore = ["E203", "D100", "D202", "D407", "ANN101", "ANN102", "ANN204",
# To be fixed:
"I001", "E721", "N818"]
fixable = ["ALL"]
## Allow unused variables when underscore-prefixed.
#dummy-variable-rgx = "^(_+|(_+[a-zA-Z0-9_]*[a-zA-Z0-9]+?))$"
[tool.ruff.format]
# Like Black, use double quotes for strings.
quote-style = "double"
# Like Black, indent with spaces, rather than tabs.
indent-style = "space"
# Like Black, respect magic trailing commas.
skip-magic-trailing-comma = false
# Like Black, automatically detect the appropriate line ending.
line-ending = "auto"
## 3. Avoid trying to fix flake8-bugbear (`B`) violations.
#unfixable = ["B"]
## 4. Ignore `E402` (import violations) in all `__init__.py` files, and in select subdirectories.
#[tool.ruff.lint.per-file-ignores]
#"__init__.py" = ["E402"]
#"**/{tests,docs,tools}/*" = ["E402"]
[tool.ruff.lint.pydocstyle]
convention = "numpy"