AI Disclosure Statement
Effective date: September 2026
What PageCrawl.io does
PageCrawl.io is a website change monitoring service. Customers choose the pages and the parts of a page they want watched. PageCrawl captures each page on a schedule, compares the new capture with the previous one, records what changed and sends an alert.
Whether a page has changed is decided without AI: the same captures always give the same result. AI comes in afterwards, for a small number of tasks that describe, score or extract what the comparison already found. Each of those runs on the requesting customer's own data, and none of them makes a decision about a person. The one processing that looks across customers is the model comparison described under Data and training, and nothing it produces is shown to a customer.
Defaults. AI features are on by default for content monitors. They can be switched off for the whole team or for a workspace, which stops every AI call for it, and the written summary can be switched off for a single monitor. Change detection itself is deterministic in every tracking mode except AI Extract, where the value the AI reads from the page is what gets compared.
The providers and the models. We may use Anthropic, Google Gemini, OpenAI, or another model available through OpenRouter. Whichever is used, it must not train on your content, and the request is encrypted in transit. Which model answers a request changes over time as more capable ones become available. These are general-purpose models operated by third parties; PageCrawl trains, fine-tunes and hosts none of them. The providers we contract with are named in the sub-processor annex to our Data Processing Agreement, and a change to that list carries the notice and objection rights the agreement sets out.
What is sent to the AI providers
What is sent. What goes to a provider depends on the feature in use. It can include: the address and title of the monitored page, its previous and current content, cropped screenshots when a visual change is analysed, and the customer's own instructions and label names. PageCrawl never adds the customer's account, billing or sign-in details to a request, though a monitored page can itself display them.
The other features. A written overview in a report or digest is produced from the monitor names and the change summaries of the period it covers. Setup suggestions send the content of the page being set up and the description the customer typed in their own words. Product matching sends the names, brands and addresses of the pages being compared. Learning from a customer's feedback sends the summaries of the changes that customer marked, or the changed text where there is no summary.
Personal data on monitored pages. A monitored page can contain anything its publisher put there, including personal data, and a customer can monitor a page behind a login they have saved with us. PageCrawl does not add the customer's account details, billing data or sign-in credentials to an AI request. It does not remove personal data from page content either: a page captured while signed in, a screenshot of it, or a forwarded verification email can carry the customer's own name or email address, or the personal data of other people shown on the page. It is the customer, as controller, who decides which pages to monitor and on what lawful basis.
Data and training
No training on customer content. Content sent for analysis is not used to train or fine-tune any model. Every request carries an instruction that the content must not be collected, and we use only providers that accept it. A provider may hold a request for a limited period for abuse monitoring under its own terms, typically up to 30 days.
No cross-customer use. An AI request contains only the requesting customer's data, and one customer's content is never used to produce output for another customer.
Choosing the models. Separately from those features, PageCrawl evaluates candidate models on a small sample of real checks. Those requests go to the same providers, the sample is used for that evaluation and nothing else, and nothing it produces is shown to a customer.
What PageCrawl keeps. A summary, an importance score and an extracted value are stored with the check they belong to and are deleted with it, under the retention period that applies to the customer's plan. Guidance kept from a customer's feedback is removed when that feedback is cleared or the workspace is deleted.
Using your own API key
Customers can connect their own API key for OpenAI, Anthropic or Google Gemini, in which case the request goes to that provider under the customer's own agreement. If the customer's key fails twice in a row, PageCrawl falls back to its own AI providers and the included credits in six hour windows until the key works again, and emails the account owner. Where the customer's provider refuses the key for quota, use of that key is suspended for a day, then a week, then a month if it keeps happening, and requests fall back the same way for as long as the suspension lasts. Keys are stored encrypted and are never returned by the API.
While a fallback is running. The content of those requests goes to PageCrawl's AI sub-processors under our agreements with them, not to the customer's own provider, and the customer's included AI credits are used. The account owner is told by email and in the application. Where the key is failing, PageCrawl tries it again after each six hour window and returns to it as soon as it works. Where the provider is refusing the key for quota, the cover lasts longer: one day the first time, one week the second and one month after that, and it ends at once if the customer updates their AI settings.
Your provider's terms are yours. Whether a provider retains or uses what is submitted depends on the customer's own agreement with it. OpenAI and Anthropic exclude API traffic from training by default under their standard commercial API terms. Google applies the same default to paid, billing-enabled use of the Gemini API; its terms for unpaid use may allow submitted content to be used to improve Google products, so a billing-enabled key is the safer choice for confidential content.
Importance scores, alerts and human oversight
What a score does. Every change is recorded. An importance score decides whether that change also raises an alert: a change scored below the notification threshold is saved without one. A default threshold applies until the customer changes it, and scoring is on unless it is switched off. Where a customer has written alert rules that decide a check, those rules decide it and the score is not applied. Where no score is produced, the alert is sent.
Nothing is hidden. A change that scored below the threshold stays in the change history with its status, and can be listed by filtering on that status. The underlying record is always there: the text difference, the visual difference and the stored captures. Labels are added and removed only from the label rules the customer set up.
Human oversight (GDPR Article 22). AI output in PageCrawl is an aid to reading and sorting changes, never a decision about a person. Importance scores rate a change to a web page, not an individual. The service carries out no automated decision-making producing legal effects concerning a person or similarly significantly affecting them within the meaning of Article 22 GDPR. Where a customer's own process depends on a value the AI read from a page, we recommend confirming it against the stored capture before relying on it.
Controls available to customers
- Switch AI off for the whole team, or for a single workspace. That stops every AI call for it, including importance scoring, AI Extract and the description of visual changes.
- Switch the written summary off for a single monitor. Importance scoring still runs unless it is switched off as well.
- Switch importance scoring off, or change the notification threshold, for a workspace, a template or a single monitor.
- Stop the surrounding page content being sent with a change, so that only the changed text goes to the provider.
- Choose whether AI runs on the first check of a new monitor.
- Write your own instructions for what matters on a page, and your own rules for automatic labels.
- Switch change feedback off, so nothing is learned from your team's marks.
- Set daily and monthly limits on how much AI a workspace may use.
- Connect your own AI provider key instead of using the included AI credits.
Data ownership and deletion
All data collected through PageCrawl.io belongs to the customer. We do not sell or license it, and the only third parties that receive it are the sub-processors listed in our Data Processing Agreement, which process it on our documented instructions.
A summary, an importance score and an extracted value are stored with the check they belong to and are deleted with it. Guidance kept from a customer's feedback is removed when that feedback is cleared or the workspace is deleted. What an AI provider holds in its own abuse-monitoring window is governed by that provider's terms, described above. How long monitoring data is kept in general, and what happens when an account is closed, is set out in our privacy policy and in our Data Processing Agreement.
The full AI Disclosure Statement and our Data Processing Agreement are available on request from hey@pagecrawl.io.
