# Respect robots.txt and Bot Protection

Source: PageCrawl.io Help Center
URL: https://pagecrawl.io/help/features/article/respect-robots-txt
Published: 7 October, 2026
Last updated: 8 October, 2026

---

**Respect robots.txt and bot protection** is a team setting for organisations whose policy asks them to honour the preferences site owners publish. When it is on, PageCrawl skips pages a site's robots.txt does not allow, checks pages only over its own connections, and stops when a site's bot protection turns a check away instead of trying another way.

The setting is off by default. PageCrawl checks the pages you choose, on the schedule you set, the way a person asked to see them, while robots.txt was written for automated crawlers. The team owner can turn it on for the whole team at any time.

Run a website and found PageCrawlBot in your logs? See [how a site owner can allow or block PageCrawl](#how-can-a-site-owner-allow-or-block-pagecrawl).

### What does the setting do?

With the setting on, every monitor in the team follows the rules each site publishes. PageCrawl skips pages the site's robots.txt does not allow, uses only its own connections and the standard browser, makes no second attempt after a bot-protection refusal, waits out Crawl-delay, and does not check pages the site keeps out of AI use.

- **robots.txt.** A page that robots.txt disallows, or whose robots.txt cannot be read, is not checked. The monitor shows **Not allowed by robots.txt**.
- **Connections.** Checks go out over PageCrawl's own connections only, and pages are checked without solving CAPTCHAs.
- **No signing in.** Saved website logins and HTTP Basic credentials are not used, so only pages open to any visitor are checked.
- **Bot protection.** When a site's bot protection turns a check away, there is no second attempt. The monitor shows **Stopped by bot protection**, and the next scheduled check tries once more.
- **Crawl-delay.** Your team's checks to one site are spaced by its Crawl-delay, up to 300 seconds.
- **AI use.** Where robots.txt says `Content-Signal: ai-input=no`, the page is not checked and shows "Not allowed by robots.txt", because every check's changes are read by AI.
- **Page Discovery.** [Page Discovery](/help/features/article/page-discovery.md) follows robots.txt and Crawl-delay too, and finds pages from sitemaps and ordinary links rather than opening them in a browser.
- **Live preview.** The preview you use to set up a monitor follows the same rules, so a page that robots.txt disallows does not open there either.

Note: robots.txt states a site owner's preferences. The setting makes PageCrawl honour them; whether your team needs it is a decision for your own policy.

### How do I turn it on?

The team owner turns it on once, and it applies to every workspace and monitor in the team. In a team with one workspace, open Settings > Workspace > General. In a team with several workspaces, open Settings > Team > More > Team Settings. Then switch it on and choose Update.

1. Open **[Settings > Workspace > General](/app/settings/workspace#respect-robots)** if your team has one workspace, or **[Settings > Team > More > Team Settings](/app/settings/team#respect-robots)** if it has several.
2. Find the **Robots.txt and bot protection** section.
3. Switch on **Respect robots.txt and bot protection**.
4. Choose **Update**.

Each monitor follows the new rules from its next check.

### What do "Not allowed by robots.txt" and "Stopped by bot protection" mean?

Both statuses mean your team's setting stopped a check, not that the monitor is broken. **Not allowed by robots.txt** means the site's robots.txt does not allow PageCrawl to check the page, or could not be read. **Stopped by bot protection** means the site's bot protection turned the check away and PageCrawl did not try another way.

| Status | What happened | What happens next |
|---|---|---|
| **Not allowed by robots.txt** | The site's robots.txt disallows the page for PageCrawl, or the file could not be read. PageCrawl did not visit the page. | Each scheduled check reads the rules again, and checks resume once the site allows the page. |
| **Stopped by bot protection** | The site's bot protection turned the check away. PageCrawl made no second attempt through another connection or browser. | The next scheduled check makes one more ordinary visit. |

The usual ways to unblock a page, such as Stealth Mode, residential connections or PageCrawl Relay, are not used while the setting is on. If you have permission to monitor the page, ask the site owner to allow PageCrawl in robots.txt, or ask the team owner to turn the setting off.

### Which robots.txt rules does PageCrawl follow?

PageCrawl follows the robots.txt standard, [RFC 9309](https://www.rfc-editor.org/rfc/rfc9309.html). It obeys the group for `User-agent: PageCrawl` or `User-agent: PageCrawlBot` when the file has one, and otherwise the group for `User-agent: *`. A site without a robots.txt allows everything, and a robots.txt that cannot be read counts as a refusal.

- **Which group applies.** Groups naming PageCrawl or PageCrawlBot apply, in any capitalisation. When there are none, the `*` group applies. A PageCrawl group replaces the `*` rules rather than adding to them.
- **Which rule wins.** The most specific matching rule decides, and `Allow` wins a tie, as with search engines.
- **No robots.txt.** When the site answers "not found" or with another client error, there is no file, so every page is allowed.
- **Unreadable robots.txt.** When the site answers with a server error or does not answer at all, the file counts as unreadable, and the standard treats that as a refusal of every page. Like Google, PageCrawl counts a 429 Too Many Requests answer as unreadable too.
- **How it is read.** PageCrawl reads robots.txt as `PageCrawlBot/2.1` and reads it again about every hour.

### How can a site owner allow or block PageCrawl?

Add a group for PageCrawl to the site's robots.txt. `Disallow: /` asks every team that respects robots.txt to skip the whole site, and an empty `Disallow:` allows everything. `PageCrawlBot` works as the name too. Without a PageCrawl group, the rules for `User-agent: *` apply.

To ask PageCrawl to skip the whole site:

```
User-agent: PageCrawl
Disallow: /
```

To allow PageCrawl while asking other bots to stay away:

```
User-agent: *
Disallow: /

User-agent: PageCrawl
Allow: /
```

A PageCrawl group replaces your `*` rules for PageCrawl, so repeat any `*` rules you still want it to follow. These rules apply to teams that turned on Respect robots.txt and bot protection; other teams check the pages they chose the way a person asked to see them. PageCrawl reads robots.txt as PageCrawlBot, and the check itself opens the page in a standard browser. If you have a question about visits from PageCrawl, [contact us](/contact-us).

### What changes about connections and the browser?

Checks use only PageCrawl's own connections. Proxy pools, custom proxies, PageCrawl Relay and residential proxies, including Web Unblocker, are not used. Stealth Mode is paused, so pages open in the standard browser. Pages are also checked without solving CAPTCHAs, so the CAPTCHA integration is not used. Website logins are not used either: checks and previews do not sign in.

Your monitors keep their saved **Location** and **Engine**. A monitor set to one of PageCrawl's own locations keeps using it, and the other choices apply again when the setting is turned off. Some pages only load through the options the setting pauses, so expect those monitors to show **Stopped by bot protection** while it is on.

### How does PageCrawl handle Crawl-delay?

Crawl-delay is not part of the robots.txt standard, but many sites use it to ask for a pause between requests. With the setting on, PageCrawl leaves at least that many seconds between your team's checks to the same site, up to 300 seconds. A larger value counts as 300 seconds.

The delay is counted per team. Your checks to one site wait their turn, so on a site with many monitors and a long Crawl-delay, some checks run later than their schedule. Page Discovery makes one request at a time on that site, spaced by the same delay.

### What happens when a site opts out of AI use?

Some sites add a Content-Signal line to robots.txt to say how their content may be used. Where it says `ai-input=no` for a page, PageCrawl does not check that page and the monitor shows "Not allowed by robots.txt". PageCrawl reads every change with AI to summarise and score it, so honouring the signal means not monitoring the page. Checks start again once the site allows AI use.

```
User-agent: *
Content-Signal: search=yes, ai-input=no
Allow: /
```

PageCrawl never uses monitored content to train AI models, so `ai-train` needs no action. The newer `use=` value is not acted on yet.

### What does the setting not cover?

The robots.txt decision is made for the address you monitor. A few things fall outside it:

- **Pages reached through the monitor's own actions,** such as a link clicked in [Perform Actions](/help/features/article/perform-actions.md), are not checked against robots.txt separately.
- **Redirects** to another page are followed without a second robots.txt check.
- **Images, scripts, styles and other files** the page loads come with the page, as in any browser.
- **A site's terms of use.** robots.txt is a request from the site owner, not permission. The setting does not read a site's terms, so review them yourself for the pages you monitor.
- **Rule changes on the site** take up to 24 hours to apply, because robots.txt is read again once a day, as the robots.txt standard allows.

### How can I check a URL before adding it?

Use the free [robots.txt checker](/robots-txt-checker). Enter a page address to see whether the site's robots.txt allows PageCrawl to check it, which line decides, and whether the site sets a Crawl-delay or opts out of AI use. It also shows what the file says to well-known search and AI crawlers. To check several pages, paste up to 20 addresses, one per line: each gets a robots.txt verdict and a quick visit that shows whether the page opens or turns automated visits away, and you can open any of them for the full report.

### What happens when I turn it off?

Switch off **Respect robots.txt and bot protection** in the same place and choose **Update**. Monitors stopped by the setting return to pending and are checked again, and their saved Location and Engine choices apply from the next check.

---

Need more? The complete PageCrawl.io help center, with every article, is available as a single document at https://pagecrawl.io/llms-full.txt. Read it for context on anything this page does not cover.
