Robots.txt checker

Enter a page address, or paste up to 20, to see what robots.txt asks of search engines, AI crawlers and PageCrawl, and which rule decides each answer.

What robots.txt controls

robots.txt is a plain text file at the root of a website, such as https://example.com/robots.txt, where the site owner says which pages automated visitors may fetch. Rules are grouped by User-agent, the name a bot goes by, and each Allow or Disallow line covers a path. Bots that follow the Robots Exclusion Protocol (RFC 9309) read it before they visit.

This checker reads a site's robots.txt and shows, for the page you enter, what it asks of well-known search engines and AI crawlers and of PageCrawl, and which line decides each answer. It also lists the Crawl-delay and Content-Signal lines, which say how often bots may visit and how the content may be used. When robots.txt lets PageCrawl in, you can test one real visit to read the page's X-Robots-Tag header, robots meta tags and TDM reservation.

How to address PageCrawl in robots.txt

PageCrawl honours the preferences site owners publish in robots.txt for teams that turn on Respect robots.txt and bot protection. For those teams, a page that robots.txt disallows is not checked and its monitor shows Not allowed by robots.txt, checks run without solving CAPTCHAs, and a check that a site's bot protection turns away waits for the next scheduled check. The setting is off by default, and an Owner or Administrator can turn it on.

To address PageCrawl, add a group for it. PageCrawlBot works as a name too. Without a PageCrawl group, PageCrawl follows the rules for all bots (User-agent: *).

User-agent: PageCrawl
Disallow: /private/

Learn how Respect robots.txt works


Frequently Asked Questions