PaperSuitorBot

The crawler of PaperSuitor

What it does

PaperSuitorBot keeps PaperSuitor up to date. PaperSuitor is a free guide to journals in the humanities and social sciences, starting with law, that helps authors decide where to submit. For each journal it shows the length limits, the review process and other conditions for authors, and links to the page where the journal states them.

What it reads

It does not log in, fill in forms, or follow links to accounts, shopping carts or social media.

How often, and how fast

It runs once a week. A journal's pages for authors are read again after 60 days, and its article pages and feed after 28 days. A journal whose pages could not be read or found is tried again the next week.

It sends one request at a time to each organisation's servers, at least one second apart. If your robots.txt sets a Crawl-delay, it waits the full delay between requests. After a 429 (Too Many Requests) answer, it waits at least 10 seconds between requests to that organisation for the rest of the run, or the Retry-After time if that is longer. If Retry-After asks for more than a minute, it sends nothing more to that organisation until the next week's run.

When a site says no

It follows robots.txt. Rules written for PaperSuitorBot apply to it; if there are none, it follows the rules for all robots (*). It reads robots.txt afresh at every run. Before following a redirect to another site, it checks that site's robots.txt too. To keep it off your site:

User-agent: PaperSuitorBot
Disallow: /

It never tries to get around bot protection and never solves CAPTCHAs. Some pages refuse it with 401, 403 or 503, or show a bot check. It leaves such a page alone until its next scheduled visit, 60 days later.

Checking that a request is ours

Its user-agent names PaperSuitorBot and gives the address of this page:

Mozilla/5.0 (compatible; PaperSuitorBot/1.0; +https://(this website)/crawler.html)

Anyone can copy a user-agent, and the crawler has no fixed IP addresses. It does have a signing key, and the public half is published at https://(this website)/.well-known/http-message-signatures-directory. Once its registration with Cloudflare's verified bots programme is complete, it signs every request with that key, using HTTP Message Signatures as the Web Bot Auth protocol describes. A signed request also carries the header Signature-Agent: "https://(this website)". A request whose signature verifies against that key comes from this crawler.

What it keeps

It keeps facts, not copies of pages:

Contact

Questions, problems, or a request that it stop visiting your site: crawler@papersuitor.com

Back to PaperSuitor