cgspace-notes/2024-10.md at 512848fc738a6ae6cab6901cebd2454b402ef6b4

alanorth/cgspace-notes

Fork 0

mirror of https://github.com/alanorth/cgspace-notes.git synced 2024-10-23 01:13:01 +02:00

Alan Orth 512848fc73

Add notes for 2024-10-03

2024-10-03 11:51:44 +03:00

1.3 KiB

Raw Blame History

title

date

author

2024-10-03

I had an idea to get abstracts from OpenAlex
- For copyright reasons they don't include plain abstracts, but the pyalex library can convert them on the fly

I filtered for journal articles that were Creative Commons and missing abstracts:

$ csvcut -c 'id,dc.title[en_US],dcterms.abstract[en_US],cg.identifier.doi[en_US],dcterms.type[en_US],dcterms.language[en_US],dcterms.license[en_US]' ~/Downloads/2024-09-30-cgspace.csv | csvgrep -c 'dcterms.type[en_US]' -r '^Journal Article$' | csvgrep -c 'cg.identifier.doi[en_US]' -r '^.+$' | csvgrep -c 'dcterms.license[en_US]' -r '^CC-' | csvgrep -c 'dcterms.abstract[en_US]' -r '^$' | csvgrep -c 'dcterms.language[en_US]' -r '^en$' | grep -v "||" | grep -v -- '-ND' | grep -v -E 'https://doi.org/10.(2499|4160|17528)/' > /tmp/missing-abstracts.csv

Then wrote a script to get them from OpenAlex
- After inspecting and cleaning a few dozen up in OpenRefine (removing "Keywords:" and copyright, and HTML entities, etc) I managed to get about 440

1.3 KiB Raw Blame History

2024-10-03

1.3 KiB

Raw Blame History