PaperSuitorBot
The crawler of PaperSuitor
What it does
PaperSuitorBot keeps PaperSuitor up to date. PaperSuitor is a free guide to journals in the humanities and social sciences, starting with law, that helps authors decide where to submit. For each journal it shows the length limits, the review process and other conditions for authors, and links to the page where the journal states them.
What it reads
- robots.txt, before anything else on a site.
- Journals' home pages and their pages for authors, such as author guidelines and submission instructions, as web pages or PDFs. It finds them from the home page, the site's sitemap or a DOI.
- Up to 12 recent article pages of a journal, for the submission and publication dates they show. It reads these only when the publisher does not deposit those dates with Crossref.
- The OAI-PMH feed of journals on Open Journal Systems or Digital Commons, for the titles and abstracts of recent articles. It reads a feed only when other sources hold too few of the journal's articles.
It does not log in, fill in forms, or follow links to accounts, shopping carts or social media.
How often, and how fast
It runs once a week. A journal's pages for authors are read again after 60 days, and its article pages and feed after 28 days. A journal whose pages could not be read or found is tried again the next week.
It sends one request at a time to each organisation's servers, at least one second apart. If your robots.txt sets a Crawl-delay, it waits the full delay between requests. After a 429 (Too Many Requests) answer, it waits at least 10 seconds between requests to that organisation for the rest of the run, or the Retry-After time if that is longer. If Retry-After asks for more than a minute, it sends nothing more to that organisation until the next week's run.
When a site says no
It follows robots.txt. Rules written for PaperSuitorBot apply to it; if there are none, it follows the rules for all robots (*). It reads robots.txt afresh at every run. Before following a redirect to another site, it checks that site's robots.txt too. To keep it off your site:
User-agent: PaperSuitorBot Disallow: /
It never tries to get around bot protection and never solves CAPTCHAs. Some pages refuse it with 401, 403 or 503, or show a bot check. It leaves such a page alone until its next scheduled visit, 60 days later.
Checking that a request is ours
Its user-agent names PaperSuitorBot and gives the address of this page:
Mozilla/5.0 (compatible; PaperSuitorBot/1.0; +https://(this website)/crawler.html)
Anyone can copy a user-agent, and the crawler has no fixed IP addresses. It does have a signing key, and the public half is published at https://(this website)/.well-known/http-message-signatures-directory. Once its registration with Cloudflare's verified bots programme is complete, it signs every request with that key, using HTTP Message Signatures as the Web Bot Auth protocol describes. A signed request also carries the header Signature-Agent: "https://(this website)". A request whose signature verifies against that key comes from this crawler.
What it keeps
It keeps facts, not copies of pages:
- a journal's stated limits, with the sentence that states them, the page's address and the date it was read;
- how long the journal takes from submission to publication, as a median over recent articles;
- the titles and abstracts of recent articles, used to match manuscripts to journals.
Contact
Questions, problems, or a request that it stop visiting your site: crawler@papersuitor.com