csv_metadata_quality/app.py: Update help text

Use DCTERMS fields where possible.
CHANGELOG.md: Add note about requests cache
2025-12-16 10:24:11 +01:00 · 2021-03-14 10:52:58 +02:00 · 2021-03-14 09:13:51 +02:00 · 2021-03-14 09:07:35 +02:00 · 2021-03-13 12:59:45 +02:00 · 2021-03-13 11:56:52 +02:00
9 changed files with 81 additions and 26 deletions
--- a/.github/workflows/python-app.yml
+++ b/.github/workflows/python-app.yml
@@ -16,10 +16,10 @@ jobs:

    steps:
    - uses: actions/checkout@v2
-    - name: Set up Python 3.8
+    - name: Set up Python 3.9
      uses: actions/setup-python@v2
      with:
-        python-version: 3.8
+        python-version: 3.9
    - name: Install dependencies
      run: |
        python -m pip install --upgrade pip
--- a/CHANGELOG.md
+++ b/CHANGELOG.md
@@ -4,6 +4,15 @@ All notable changes to this project will be documented in this file.
 The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.0.0/),
 and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).

+## Unreleased
+### Changed
+- Fixing invalid multi-value separators like `|` and `|||` is no longer class-
+ified as "unsafe" as I have yet to see a case where this was intentional
+
+### Added
+- Configurable directory for AGROVOC requests cache (to allow running the web
+version from Google App Engine where we can only write to /tmp)
+
 ## [0.4.6] - 2021-03-11
 ### Added
 - Validation of dcterms.license field against SPDX license identifiers 
--- a/README.md
+++ b/README.md
@@ -1,7 +1,7 @@
 # DSpace CSV Metadata Quality Checker ![GitHub Actions](https://github.com/ilri/csv-metadata-quality/workflows/Build%20and%20Test/badge.svg) [![Build Status](https://ci.mjanja.ch/api/badges/alanorth/csv-metadata-quality/status.svg)](https://ci.mjanja.ch/alanorth/csv-metadata-quality)
 A simple, but opinionated metadata quality checker and fixer designed to work with CSVs in the DSpace ecosystem (though it could theoretically work on any CSV that uses Dublin Core fields as columns). The implementation is essentially a pipeline of checks and fixes that begins with splitting multi-value fields on the standard DSpace "||" separator, trimming leading/trailing whitespace, and then proceeding to more specialized cases like ISSNs, ISBNs, languages, unnecessary Unicode, AGROVOC terms, etc.

-Requires Python 3.7 or greater (3.8 recommended). CSV and Excel support comes from the [Pandas](https://pandas.pydata.org/) library, though your mileage may vary with Excel because this is much less tested.
+Requires Python 3.7.1 or greater (3.8+ recommended). CSV and Excel support comes from the [Pandas](https://pandas.pydata.org/) library, though your mileage may vary with Excel because this is much less tested.

 If you use the DSpace CSV metadata quality checker please cite:

@@ -13,13 +13,14 @@ If you use the DSpace CSV metadata quality checker please cite:
 - Validate languages against ISO 639-1 (alpha2) and ISO 639-3 (alpha3)
 - Experimental validation of titles and abstracts against item's Dublin Core language field
 - Validate subjects against the AGROVOC REST API (see the `--agrovoc-fields` option)
+- Validation of licenses against the list of [SPDX license identifiers](https://spdx.org/licenses)
 - Fix leading, trailing, and excessive (ie, more than one) whitespace
- Fix invalid and unnecessary multi-value separators (`|`) using `--unsafe-fixes`
+- Fix invalid and unnecessary multi-value separators (`|`)
 - Fix problematic newlines (line feeds) using `--unsafe-fixes`
+- Perform [Unicode normalization](https://withblue.ink/2019/03/11/why-you-need-to-normalize-unicode-strings.html) on strings using `--unsafe-fixes`
 - Remove unnecessary Unicode like [non-breaking spaces](https://en.wikipedia.org/wiki/Non-breaking_space), [replacement characters](https://en.wikipedia.org/wiki/Specials_(Unicode_block)#Replacement_character), etc
 - Check for "suspicious" characters that indicate encoding or copy/paste issues, for example "foreˆt" should be "forêt"
 - Remove duplicate metadata values
- Perform [Unicode normalization](https://withblue.ink/2019/03/11/why-you-need-to-normalize-unicode-strings.html) on strings using `--unsafe-fixes`

 ## Installation
 The easiest way to install CSV Metadata Quality is with [poetry](https://python-poetry.org):
@@ -54,14 +55,14 @@ To validate and clean a CSV file you must specify input and output files using t
 $ csv-metadata-quality -i data/test.csv -o /tmp/test.csv
 ```

-## Unsafe Fixes
-You can enable several "unsafe" fixes with the `--unsafe-fixes` option. Currently this will attempt to fix invalid multi-value separators and remove newlines.
-
-### Invalid Multi-Value Separators
-This is considered "unsafe" because it is *theoretically* possible for a single `|` character to be used legitimately in a metadata value, though in my experience it is always a typo. For example, if a user mistakenly writes `Kenya|Tanzania` when attempting to indicate two countries, the result will be one metadata value with the literal text `Kenya|Tanzania`. The `--unsafe-fixes` option will correct the invalid multi-value separator so that there are two metadata values, ie `Kenya||Tanzania`.
+## Invalid Multi-Value Separators
+While it is *theoretically* possible for a single `|` character to be used legitimately in a metadata value, in my experience it is always a typo. For example, if a user mistakenly writes `Kenya|Tanzania` when attempting to indicate two countries, the result will be one metadata value with the literal text `Kenya|Tanzania`. This utility will correct the invalid multi-value separator so that there are two metadata values, ie `Kenya||Tanzania`.

 This will also remove unnecessary trailing multi-value separators, for example `Kenya||Tanzania||`.

+## Unsafe Fixes
+You can enable several "unsafe" fixes with the `--unsafe-fixes` option. Currently this will remove newlines and perform Unicode normalization.
+
 ### Newlines
 This is considered "unsafe" because some systems give special importance to vertical space and render it properly. DSpace does not support rendering newlines in its XMLUI and has, at times, suffered from parsing errors that cause the import process to fail if an input file had newlines. The `--unsafe-fixes` option strips Unix line feeds (U+000A).

--- a/csv_metadata_quality/app.py
+++ b/csv_metadata_quality/app.py
@@ -17,7 +17,7 @@ def parse_args(argv):
    parser.add_argument(
        "--agrovoc-fields",
        "-a",
-        help="Comma-separated list of fields to validate against AGROVOC, for example: dc.subject,cg.coverage.country",
+        help="Comma-separated list of fields to validate against AGROVOC, for example: dcterms.subject,cg.coverage.country",
    )
    parser.add_argument(
        "--experimental-checks",
@@ -48,7 +48,7 @@ def parse_args(argv):
    parser.add_argument(
        "--exclude-fields",
        "-x",
-        help="Comma-separated list of fields to skip, for example: dc.contributor.author,dc.identifier.citation",
+        help="Comma-separated list of fields to skip, for example: dc.contributor.author,dcterms.bibliographicCitation",
    )
    args = parser.parse_args()

@@ -111,10 +111,9 @@ def run(argv):
        df[column] = df[column].apply(check.suspicious_characters, field_name=column)

        # Fix: invalid and unnecessary multi-value separators
-        if args.unsafe_fixes:
-            df[column] = df[column].apply(fix.separators, field_name=column)
-            # Run whitespace fix again after fixing invalid separators
-            df[column] = df[column].apply(fix.whitespace, field_name=column)
+        df[column] = df[column].apply(fix.separators, field_name=column)
+        # Run whitespace fix again after fixing invalid separators
+        df[column] = df[column].apply(fix.whitespace, field_name=column)

        # Fix: duplicate metadata values
        df[column] = df[column].apply(fix.duplicates, field_name=column)
--- a/csv_metadata_quality/check.py
+++ b/csv_metadata_quality/check.py
@@ -1,3 +1,4 @@
+import os
 import re
 from datetime import datetime, timedelta

@@ -242,7 +243,11 @@ def agrovoc(field, field_name):

    # enable transparent request cache with thirty days expiry
    expire_after = timedelta(days=30)
-    requests_cache.install_cache("agrovoc-response-cache", expire_after=expire_after)
+    # Allow overriding the location of the requests cache, just in case we are
+    # running in an environment where we can't write to the current working di-
+    # rectory (for example from csv-metadata-quality-web).
+    REQUESTS_CACHE_DIR = os.environ.get("REQUESTS_CACHE_DIR", ".")
+    requests_cache.install_cache(f"{REQUESTS_CACHE_DIR}/agrovoc-response-cache", expire_after=expire_after)

    # prune old cache entries
    requests_cache.core.remove_expired_responses()
--- a/csv_metadata_quality/version.py
+++ b/csv_metadata_quality/version.py
@@ -1 +1 @@
-VERSION = "0.4.6"
+VERSION = "0.4.6-dev"
--- a/poetry.lock
+++ b/poetry.lock
@@ -212,6 +212,7 @@ optional = false
 python-versions = "!=3.0.*,!=3.1.*,!=3.2.*,!=3.3.*,>=2.7"

 [package.dependencies]
+importlib-metadata = {version = "*", markers = "python_version < \"3.8\""}
 mccabe = ">=0.6.0,<0.7.0"
 pycodestyle = ">=2.6.0a1,<2.7.0"
 pyflakes = ">=2.2.0,<2.3.0"
@@ -224,6 +225,22 @@ category = "main"
 optional = false
 python-versions = ">=2.7, !=3.0.*, !=3.1.*, !=3.2.*, !=3.3.*"

+[[package]]
+name = "importlib-metadata"
+version = "3.7.2"
+description = "Read metadata from Python packages"
+category = "dev"
+optional = false
+python-versions = ">=3.6"
+
+[package.dependencies]
+typing-extensions = {version = ">=3.6.4", markers = "python_version < \"3.8\""}
+zipp = ">=0.5"
+
+[package.extras]
+docs = ["sphinx", "jaraco.packaging (>=8.2)", "rst.linker (>=1.9)"]
+testing = ["pytest (>=3.5,!=3.7.3)", "pytest-checkdocs (>=1.2.3)", "pytest-flake8", "pytest-cov", "pytest-enabler", "packaging", "pep517", "pyfakefs", "flufl.flake8", "pytest-black (>=0.3.7)", "pytest-mypy", "importlib-resources (>=1.3)"]
+
 [[package]]
 name = "iniconfig"
 version = "1.1.1"
@@ -449,12 +466,15 @@ category = "dev"
 optional = false
 python-versions = ">=2.7, !=3.0.*, !=3.1.*, !=3.2.*, !=3.3.*"

+[package.dependencies]
+importlib-metadata = {version = ">=0.12", markers = "python_version < \"3.8\""}
+
 [package.extras]
 dev = ["pre-commit", "tox"]

 [[package]]
 name = "prompt-toolkit"
-version = "3.0.16"
+version = "3.0.17"
 description = "Library for building powerful interactive command lines in Python"
 category = "dev"
 optional = false
@@ -539,6 +559,7 @@ python-versions = ">=3.6"
 atomicwrites = {version = ">=1.0", markers = "sys_platform == \"win32\""}
 attrs = ">=19.2.0"
 colorama = {version = "*", markers = "sys_platform == \"win32\""}
+importlib-metadata = {version = ">=0.12", markers = "python_version < \"3.8\""}
 iniconfig = "*"
 packaging = "*"
 pluggy = ">=0.12,<1.0.0a1"
@@ -770,10 +791,22 @@ category = "main"
 optional = false
 python-versions = ">=2.7, !=3.0.*, !=3.1.*, !=3.2.*, !=3.3.*"

+[[package]]
+name = "zipp"
+version = "3.4.1"
+description = "Backport of pathlib-compatible object wrapper for zip files"
+category = "dev"
+optional = false
+python-versions = ">=3.6"
+
+[package.extras]
+docs = ["sphinx", "jaraco.packaging (>=8.2)", "rst.linker (>=1.9)"]
+testing = ["pytest (>=4.6)", "pytest-checkdocs (>=1.2.3)", "pytest-flake8", "pytest-cov", "pytest-enabler", "jaraco.itertools", "func-timeout", "pytest-black (>=0.3.7)", "pytest-mypy"]
+
 [metadata]
 lock-version = "1.1"
-python-versions = "^3.8"
-content-hash = "6a9ee0f26b50f361d7e0e6a2275f0e3174dee1c89fbd460583c4ea3d873857b8"
+python-versions = "^3.7.1"
+content-hash = "e60e882e091af667b968c00521fd378e1220c1836d394d90bbc783920e38bb62"

 [metadata.files]
 agate = [
@@ -853,6 +886,10 @@ idna = [
    {file = "idna-2.10-py2.py3-none-any.whl", hash = "sha256:b97d804b1e9b523befed77c48dacec60e6dcb0b5391d57af6a65a312a90648c0"},
    {file = "idna-2.10.tar.gz", hash = "sha256:b307872f855b18632ce0c21c5e45be78c0ea7ae4c15c828c20788b26921eb3f6"},
 ]
+importlib-metadata = [
+    {file = "importlib_metadata-3.7.2-py3-none-any.whl", hash = "sha256:407d13f55dc6f2a844e62325d18ad7019a436c4bfcaee34cda35f2be6e7c3e34"},
+    {file = "importlib_metadata-3.7.2.tar.gz", hash = "sha256:18d5ff601069f98d5d605b6a4b50c18a34811d655c55548adc833e687289acde"},
+]
 iniconfig = [
    {file = "iniconfig-1.1.1-py2.py3-none-any.whl", hash = "sha256:011e24c64b7f47f6ebd835bb12a743f2fbe9a26d4cecaa7f53bc4f35ee9da8b3"},
    {file = "iniconfig-1.1.1.tar.gz", hash = "sha256:bc3af051d7d14b2ee5ef9969666def0cd1a000e121eaea580d4a313df4b37f32"},
@@ -969,8 +1006,8 @@ pluggy = [
    {file = "pluggy-0.13.1.tar.gz", hash = "sha256:15b2acde666561e1298d71b523007ed7364de07029219b604cf808bfa1c765b0"},
 ]
 prompt-toolkit = [
-    {file = "prompt_toolkit-3.0.16-py3-none-any.whl", hash = "sha256:62c811e46bd09130fb11ab759012a4ae385ce4fb2073442d1898867a824183bd"},
-    {file = "prompt_toolkit-3.0.16.tar.gz", hash = "sha256:0fa02fa80363844a4ab4b8d6891f62dd0645ba672723130423ca4037b80c1974"},
+    {file = "prompt_toolkit-3.0.17-py3-none-any.whl", hash = "sha256:4cea7d09e46723885cb8bc54678175453e5071e9449821dce6f017b1d1fbfc1a"},
+    {file = "prompt_toolkit-3.0.17.tar.gz", hash = "sha256:9397a7162cf45449147ad6042fa37983a081b8a73363a5253dd4072666333137"},
 ]
 ptyprocess = [
    {file = "ptyprocess-0.7.0-py2.py3-none-any.whl", hash = "sha256:4b41f3967fce3af57cc7e94b888626c18bf37a083e3651ca8feeb66d492fef35"},
@@ -1191,3 +1228,7 @@ xlrd = [
    {file = "xlrd-1.2.0-py2.py3-none-any.whl", hash = "sha256:e551fb498759fa3a5384a94ccd4c3c02eb7c00ea424426e212ac0c57be9dfbde"},
    {file = "xlrd-1.2.0.tar.gz", hash = "sha256:546eb36cee8db40c3eaa46c351e67ffee6eeb5fa2650b71bc4c758a29a1b29b2"},
 ]
+zipp = [
+    {file = "zipp-3.4.1-py3-none-any.whl", hash = "sha256:51cb66cc54621609dd593d1787f286ee42a5c0adbb4b29abea5a63edc3e03098"},
+    {file = "zipp-3.4.1.tar.gz", hash = "sha256:3607921face881ba3e026887d8150cca609d517579abe052ac81fc5aeffdbd76"},
+]
--- a/pyproject.toml
+++ b/pyproject.toml
@@ -1,6 +1,6 @@
 [tool.poetry]
 name = "csv-metadata-quality"
-version = "0.4.6"
+version = "0.4.6-dev"
 description="A simple, but opinionated CSV quality checking and fixing pipeline for CSVs in the DSpace ecosystem."
 authors = ["Alan Orth <alan.orth@gmail.com>"]
 license="GPL-3.0-only"
@@ -11,7 +11,7 @@ homepage = "https://github.com/ilri/csv-metadata-quality"
 csv-metadata-quality = 'csv_metadata_quality.__main__:main'

 [tool.poetry.dependencies]
-python = "^3.8"
+python = "^3.7.1"
 pandas = "^1.0.4"
 python-stdnum = "^1.13"
 xlrd = "^1.2.0"
--- a/setup.py
+++ b/setup.py
@@ -14,7 +14,7 @@ install_requires = [

 setuptools.setup(
    name="csv-metadata-quality",
-    version="0.4.6",
+    version="0.4.6-dev",
    author="Alan Orth",
    author_email="aorth@mjanja.ch",
    description="A simple, but opinionated CSV quality checking and fixing pipeline for CSVs in the DSpace ecosystem.",
Author	SHA1	Message	Date
Alan Orth	c9c277f8df	csv_metadata_quality/app.py: Update help text All checks were successful continuous-integration/drone/push Build is passing Details Use DCTERMS fields where possible.	2021-03-14 10:52:58 +02:00
Alan Orth	fb35afd937	CHANGELOG.md: Add note about requests cache	2021-03-14 09:13:51 +02:00
Alan Orth	0e9176f0a6	csv_metadata_quality/check.py: requests cache Allow overriding the directory for the requests cache. In the case of csv-metadata-quality-web, which currently runs on Google's App Engine, we can only write to /tmp.	2021-03-14 09:07:35 +02:00
Alan Orth	1008acf35e	Always fix invalid multi-value separators All checks were successful continuous-integration/drone/push Build is passing Details This is no longer class-ified as "unsafe" as I have yet to see a case where this was intentional, and it always causes issues when you import the data in a DSpace repository.	2021-03-13 12:59:45 +02:00
Alan Orth	f00a07e2cd	README.md: Reorganize unsafe functionality All checks were successful continuous-integration/drone/push Build is passing Details	2021-03-13 11:56:52 +02:00
Alan Orth	46098861ed	poetry.lock: Run poetry update All checks were successful continuous-integration/drone/push Build is passing Details	2021-03-11 22:45:32 +02:00
Alan Orth	fa84cfa440	Bump version to 0.4.6-dev	2021-03-11 22:44:36 +02:00
Alan Orth	6cc1401f88	pyproject.toml: Minimum Python is technically 3.7.1 All checks were successful continuous-integration/drone/push Build is passing Details See: https://pandas.pydata.org/pandas-docs/stable/whatsnew/v1.2.0.html	2021-03-11 13:41:58 +02:00
Alan Orth	ad2cda8a41	README.md: Add note about SPDX license identifiers All checks were successful continuous-integration/drone/push Build is passing Details	2021-03-11 12:21:34 +02:00
Alan Orth	dc6920802e	.github/workflows/python-app.yml: Use Python 3.9 I now use this version in my development environment. Eventually I should add a matrix of versions to use, but I don't know the GitHub Actions syntax well enough yet.	2021-03-11 12:17:57 +02:00
Alan Orth	6ca449d8ed	README.md: Update note about Python 3.8 to 3.8+ Currently the lower bound on Python version support is 3.7 because of Pandas 1.2.0 requiring it, but I use 3.9 on my development box.	2021-03-11 12:16:07 +02:00