mirror of
https://github.com/alanorth/cgspace-notes.git
synced 2025-01-27 05:49:12 +01:00
Update notes for 2022-04-04
This commit is contained in:
@ -21,5 +21,11 @@ sys 3m43.037s
|
||||
- Start a full harvest on AReS
|
||||
- Help Marianne with submit/approve access on a new collection on CGSpace
|
||||
- Go back in Gaia's batch reports to find records that she indicated for replacing on CGSpace (ie, those with better new copies, new versions, etc)
|
||||
- Looking at the Solr statistics for 2022-03 on CGSpace
|
||||
- I see 54.229.218.204 on Amazon AWS made 49,000 requests, some of which with this user agent: `Apache-HttpClient/4.5.9 (Java/1.8.0_322)`, and many others with a normal browser agent, so that's fishy!
|
||||
- The DSpace agent pattern `http.?agent` seems to have caught the first ones, but I'll purge the IP ones
|
||||
- I see 40.77.167.80 is Bing or MSN Bot, but using a normal browser user agent, and if I search Solr for `dns:*msnbot* AND dns:*.msn.com.` I see over 100,000, which is a problem I noticed a few months ago too...
|
||||
- I extracted the MSN Bot IPs from Solr using an IP facet, then used the `check-spider-ip-hits.sh` script to purge them
|
||||
-
|
||||
|
||||
<!-- vim: set sw=2 ts=2: -->
|
||||
|
Reference in New Issue
Block a user