# PageCrawl.io > AI-powered website change monitoring and change detection platform PageCrawl.io monitors web pages for changes and alerts you when a scheduled check detects an update, with checks as frequent as every 2 minutes depending on plan. Track prices, availability, text, documents, and more using real browser rendering with AI-powered change summaries. ## Company Facts - **Product**: PageCrawl.io, a website change monitoring and change detection SaaS - **Website**: https://pagecrawl.io - **Launched**: 2018 - **Founder**: Laurynas Sakalauskas - **Category**: Website change monitoring / change detection. Not an uptime monitor: uptime tools answer "is the site up", PageCrawl answers "what changed on the page" - **Contact**: hey@pagecrawl.io (sales), help_me@pagecrawl.io (support) ## What PageCrawl Does PageCrawl.io is a cloud-based website monitoring service that detects changes on any web page and notifies you through your preferred channels. It uses real browsers (with full JavaScript support) to render pages exactly as a human would see them, then compares snapshots to detect what changed. AI analyzes every change to provide plain-language summaries and importance scoring, so you only get notified when it matters. ## Core Capabilities - **Website Change Monitoring**: Detect text, visual, and structural changes on any webpage - **Real Browser Rendering**: JavaScript-enabled monitoring for dynamic sites and SPAs - **AI-Powered Change Summaries**: Get plain-language summaries of what changed, included on all plans - **AI Importance Scoring**: Every change scored 0-100 by importance, set thresholds to filter noise - **AI Pattern Learning**: Dismiss changes and AI learns, similar low-value changes are auto-filtered - **Multi-Channel Alerts**: Email, Slack, Discord, Microsoft Teams, Telegram, webhooks, web push notifications - **Element-Level Tracking**: Track specific elements with CSS/XPath selectors or monitor full pages - **File Monitoring**: Track changes in PDFs, Word (.docx), Excel (.xlsx), PowerPoint (.pptx), CSV files - **Cloud Document Monitoring**: Google Sheets, Google Drive, Microsoft SharePoint documents - **Visual Comparison**: Side-by-side screenshot diffs and historical screenshot archives - **Automatic Page Discovery**: Find new pages on any website via sitemap scanning, deep crawl, or URL scanning - **Check Frequencies**: From every 2 minutes (Ultimate) to monthly schedules - **Smart False-Positive Reduction**: Filter dates, ads, popups, dynamic content, combined with AI noise filtering - **Protected-Site Monitoring**: Reliably monitors sites that use bot protection - **Password-Protected Pages**: Monitor pages behind login forms - **Browser Actions**: Click buttons, type text, remove overlays, scroll, wait for elements before capture - **Team Collaboration**: Workspaces, user roles (owner, admin, standard, viewer), shared monitors - **Review Boards**: Kanban-like board for compliance teams to review and flag changes - **Bulk Management**: Import multiple URLs, bulk edit configurations, reusable page templates - **Web Archiving**: Save complete web archives in WACZ format (Ultimate plan) - **Monitoring Reports**: Daily, weekly, or monthly digest reports with AI summaries - **Data Exports**: Export change history to Excel spreadsheets ## Tracking Modes - **Full Page**: Monitor the entire visible text content of a page - **Content Only**: Ignores navigation, headers, footers, focuses on main content - **Reader Mode**: Uses Readability algorithm for clean article extraction - **Price**: Auto-detects prices and tracks availability for e-commerce pages - **Specific Text**: Track a specific element using CSS or XPath selector - **Specific Number**: Track a numeric value on a page using CSS or XPath selector - **Visual/Screenshot**: Detect visual layout changes via screenshot comparison - **Links**: Track all links on a page for changes - **HTML**: Monitor raw HTML source changes - **File Hash**: Track binary file changes via SHA-256 hash ## Integrations - Slack, Microsoft Teams, Discord, Telegram - Email with Cc support and digest reports - Web push notifications (Chrome, Firefox, Edge, Safari 16+) - Webhooks for custom integrations - Zapier (connect to thousands of apps) - n8n (community node with full API support) - Google Sheets sync (auto-sync monitored data) - Dropbox sync (auto-save screenshots) - RSS feeds - Home Assistant (via webhooks and REST API) - RESTful API for programmatic access - Browser extension for Chrome and Firefox - iOS Safari bookmarklet - MCP Server for AI assistants ## MCP Server (AI Assistant Integration) PageCrawl.io provides an MCP (Model Context Protocol) server that lets AI assistants manage website monitors through natural conversation. It works on every plan, including Free: list and read monitors, create monitors, organize them into folders, manage labels (tags), mark changes reviewed, and enable or disable monitors. Plan monitor-count and monthly check limits still apply. - **MCP Server URL**: https://pagecrawl.io/mcp - **Authentication**: OAuth 2.0 - **Compatible with**: Claude.ai, ChatGPT, Claude Code, Claude Desktop, Cursor, Windsurf, and any MCP-compatible client ### MCP Setup **Claude (Web/Desktop):** Settings > Connectors > Add custom connector > Name: "PageCrawl", URL: `https://pagecrawl.io/mcp` > Authorize **Claude Code:** Add to `.mcp.json`: ```json { "mcpServers": { "pagecrawl": { "url": "https://pagecrawl.io/mcp" } } } ``` **ChatGPT:** Settings > Connected apps > Add > URL: `https://pagecrawl.io/mcp` ### MCP Tools | Tool | Description | |------|-------------| | list-workspaces | List all accessible workspaces | | add-page-monitor | Create a new monitor with URL, tracking mode, frequency, notifications | | list-monitors | Search and list monitors across all workspaces by URL, domain, or name | | get-monitor-details | Get full configuration for a monitor (supports batch via monitor_ids) | | get-monitor-history | Retrieve historical checks and detected changes (supports batch, since filter) | | get-latest-values | Get current tracked values for monitors (supports batch, lightest-weight) | | get-check-diff | View formatted text diffs showing additions and removals for a check | | get-changes-since | Get all changes across all monitors since a given date | | trigger-check | Force an immediate check on a monitor (paid plans) | | manage-tags | List workspace tags, add or remove tags from monitors | | manage-folders | Organize monitors into hierarchical folders: list, create, rename, nest, and move monitors between folders | | set-monitor-status | Enable or disable a monitor; disabling stops checks but keeps the monitor and its history | | mark-changes-seen | Mark detected changes as reviewed on one or all monitors | | list-templates | List templates available in a workspace | | update-monitor-defaults | View or update default settings for new monitors per workspace | | create-data-source | Create a push data source: a monitor that receives values you send via API, MCP, or email instead of being crawled | | record-data | Record a value on a push data source; it is stored in history and triggers notifications and AI summaries like any check | ### MCP Usage Guide - "What is the current price?" -> use get-latest-values - "What changed in the last check?" -> use get-check-diff - "Show me the change history" -> use get-monitor-history - "What changed today/this week?" -> use get-changes-since - "Find monitors for example.com" -> use list-monitors - "Monitor this URL" -> use add-page-monitor - frequency is in minutes: 60 = hourly, 1440 = daily, 10080 = weekly - get-monitor-details, get-monitor-history, and get-latest-values support batch requests via monitor_ids array (max 50) ## Common Use Cases - **E-commerce & Price Monitoring**: Track competitor prices, product availability, inventory restocks - **Compliance & Legal**: Monitor terms of service, privacy policies, regulatory filings, legal documents - **Competitive Intelligence**: Track competitor websites for pricing, product, hiring, and content changes - **Stock and Restock Alerts**: GPU stock (NVIDIA RTX), gaming consoles (PS5), networking equipment (UniFi), Tesla inventory, Raspberry Pi - **Job Hunting**: Get alerts for new job postings on career pages and job boards - **SEC Filing Monitoring**: Track EDGAR filings (10-K, 10-Q, 8-K, S-1, Form 4, 13-F) - **Software Release Tracking**: Monitor release notes, changelogs, and version updates - **Academic Research**: Track research papers, news, and competitor publications - **Real Estate**: Monitor new property listings and price changes - **Content Monitoring**: Track news sites, blogs, government pages for updates - **Smart Home Automation**: Integrate with Home Assistant for automated responses to web changes ## Key Differentiators - **Real browser rendering**: Unlike simple HTTP scrapers, PageCrawl uses real browsers that execute JavaScript, handle SPAs, and render pages exactly as humans see them - **AI included on all plans**: Every plan includes AI credits for change summaries and smart filtering, no separate AI add-on needed - **BYOK (Bring Your Own Key)**: Connect your own API key from OpenAI, Gemini, Anthropic, or OpenRouter for unlimited AI usage. OpenRouter gives access to any AI model available on the market, so you can choose the model that best fits your needs and budget - **No data selling**: PageCrawl does not sell user data, offer data packs, or use customer data for AI training. Your data stays private. - **File monitoring built-in**: Track PDFs, Excel, Word, PowerPoint, CSV, Google Sheets, and SharePoint documents natively - **Automatic page discovery**: Find and start monitoring new pages automatically as they appear on any website - **MCP Server for AI assistants**: Manage monitors through AI tools like Claude and ChatGPT - **Browser extension**: One-click monitoring from Chrome or Firefox - **Cloud-based, no installation**: Entirely browser-based, nothing to install - **Since 2018**: Serving thousands of organizations including Microsoft, New York Times, Deloitte, Goldman Sachs, HubSpot, Yahoo, EY, and Stanford ## Pricing PageCrawl.io plans as of July 2026 (current details: https://pagecrawl.io/pricing): - **Free Forever**: $0/month, 6 pages, 220 checks/mo, 60-min frequency, 90-day history, 15 AI credits, all integrations, MCP server - **Standard**: $8/month (or $80/year), 100 pages, 15,000 checks/mo, 15-min frequency, 12-month history, 500 AI credits, MCP server - **Enterprise** (Most Popular): $30/month (or $300/year), 500 pages, 100,000 checks/mo, 5-min frequency, unlimited history, 2,500 AI credits, SSO, premium proxy pool, MCP server - **Ultimate**: $99/month (or $999/year), 1,000 pages, 100,000 checks/mo, 2-min frequency, unlimited history, 10,000 AI credits with Pro-tier models, web archiving (WACZ), dedicated account manager, MCP server - **Custom plans**: Available for higher volumes, 1-minute checks, and enterprise needs. Contact hey@pagecrawl.io - All paid plans scale with quantity multipliers for more pages and checks ## Enterprise Features - SAML 2.0 Single Sign-On (Azure AD, Google Workspace, Okta, OneLogin) - User access roles and permissions (owner, admin, standard, viewer) - Premium proxy pool for difficult-to-access sites - CAPTCHA solving service (optional, additional cost) - Higher quality screenshots - MCP Server for AI assistant integration - Invoice and PO billing (Ultimate, annual) - Dedicated account manager (Ultimate) ## API PageCrawl.io offers a RESTful API for programmatic access to: - Create and manage monitored pages - Retrieve change history and diffs (PNG, HTML, Markdown, Patch formats) - Get screenshots and visual comparisons - Configure webhooks that fire when a check detects a change - Trigger manual checks - Manage folders, tags, and workspaces - **Interactive API Reference**: https://pagecrawl.io/developers - **OpenAPI 3.0 Spec**: https://pagecrawl.io/api/openapi.yaml - **Authentication**: Bearer token via `Authorization: Bearer ` - **Quick Track Endpoint**: `POST /api/track-simple` - create a monitor with just a URL - **Webhook Payload Fields**: 25+ customizable fields including AI summary, priority score, markdown diff, screenshot URL - **Rate Limits**: 60 requests/minute on Free, 300 requests/minute on paid plans - **API access**: available on every plan, including Free ## Support - Free plan: Community support only - Standard: Email support within 72 hours - Enterprise: Premium email support within 48 hours - Ultimate: Priority support within 24 hours with dedicated account manager - Contact: hey@pagecrawl.io (sales), help_me@pagecrawl.io (support) ## Frequently Asked Questions ### What is PageCrawl.io? PageCrawl.io is a website change monitoring service, launched in 2018, that checks web pages on a schedule, detects text, visual, number, and file changes, summarizes each change with AI, and delivers alerts to email, Slack, Discord, Microsoft Teams, Telegram, webhooks, and browser push notifications. ### How does PageCrawl detect changes on a website? PageCrawl loads each monitored page in a real browser with JavaScript enabled, captures the tracked content, and compares it with the previous check. Differences are shown as highlighted text diffs and side-by-side screenshots, and AI scores each change from 0 to 100 by importance so noise can be filtered out. ### How quickly will I know when a page changes? An alert is sent when the next scheduled check detects the change. Check frequency depends on plan: hourly on Free, every 15 minutes on Standard, every 5 minutes on Enterprise, every 2 minutes on Ultimate, and down to every minute on custom plans. ### Does PageCrawl work on JavaScript-heavy or protected sites? Yes. Pages are rendered in real browsers with full JavaScript execution, so single-page applications and dynamic sites are captured as a visitor sees them. Sites that use bot protection are monitored reliably, with optional premium residential proxies for the hardest cases. ### Can PageCrawl monitor PDFs and other documents? Yes. PageCrawl tracks text changes in PDF, Word, Excel, PowerPoint, and CSV files, as well as Google Sheets, Google Docs, and Microsoft SharePoint documents. Any binary file can also be tracked by its SHA-256 checksum. ### Is there a free plan? Yes. The Free Forever plan includes real browser monitoring, AI change summaries, and every notification channel and integration, including the REST API and MCP server. Paid plans start at $8 per month and add more pages, more checks, and faster frequencies. ### Can AI assistants like Claude or ChatGPT manage PageCrawl? Yes. PageCrawl's MCP server at https://pagecrawl.io/mcp (OAuth 2.0) works with Claude, ChatGPT, Cursor, and any MCP-compatible client on every plan. Assistants can create monitors, read diffs and history, organize folders and tags, and record data in natural language. ## Resources - [Website](https://pagecrawl.io): Main site - [Help Center](https://pagecrawl.io/help): All help articles - [Blog](https://pagecrawl.io/blog): Monitoring use cases, tutorials, and product news - [Pricing](https://pagecrawl.io/pricing): Plans and pricing - [API Documentation](https://pagecrawl.io/help/features/article/api-webhooks-for-custom-integrations.md): REST API, webhooks, and RSS feeds - [Contact](https://pagecrawl.io/contact-us): Sales and support - [About](https://pagecrawl.io/about): Company information - [Privacy Policy](https://pagecrawl.io/privacy-policy): Privacy practices - [Security](https://pagecrawl.io/security-statement): Security practices - [GDPR](https://pagecrawl.io/gdpr): GDPR compliance - [Full LLM reference](https://pagecrawl.io/llms-full.txt): Complete content of every help article ## Optional Lower-priority resources that can be skipped under context pressure. - [Blog](https://pagecrawl.io/blog): Monitoring use cases, tutorials, and product news. Full URL index of all published posts is included in llms-full.txt. --- ## Knowledge Base / Help Center The following sections contain the full content of all PageCrawl.io help center articles. ### Account & Settings ### Cancel or Upgrade Account URL: https://pagecrawl.io/help/account-settings/article/cancel-or-upgrade-account ### Changing plan or billing interval If you would like to change or upgrade your plan, just go to your [Subscription settings](/app/settings/subscription) and choose a plan you want to switch to. [Image: Subscription settings with the Choose Your Plan section and a monthly or yearly billing toggle] **Upgrades take effect immediately.** When you upgrade, a new billing cycle starts on the day of the upgrade. You are charged the full price of the new plan for the new cycle, and the unused time on your old plan is applied as a credit on that invoice. Your renewal date moves to one billing period from the day you upgrade, and your usage limits reset right away. For example, you are halfway through a month on the $8/mo plan and upgrade to the $30/mo plan. Roughly $4 of unused time is credited, so the invoice today comes to about $26, and your next renewal is one month from the upgrade date. The exact amount, including any tax or discounts, is shown before you confirm. **Downgrades take effect at the end of your billing period.** When you switch to a cheaper plan, you keep your current plan and all its limits until your renewal date, so no credit or partial refund is needed. At renewal, the new plan starts and is billed at its price. Until then the downgrade is only scheduled, and you can cancel it from your [Subscription settings](/app/settings/subscription) to stay on your current plan. Note: switching your billing interval works the same way. Moving from monthly to yearly is treated as an upgrade and applies immediately, while moving from yearly to monthly is treated as a downgrade and applies at renewal. ### Canceling or Suspending your account You can cancel your subscription by going to your [Subscription settings](/app/settings/subscription) and clicking on the red **"Downgrade to Free"** button. This will open a multi-step confirmation modal where you can optionally provide feedback about why you are canceling. To complete the cancellation, you will need to type "CANCEL" and confirm. Once confirmed, your subscription will not end immediately. You will retain full access to your paid features until the end of your current billing period (grace period). After that date, your account will automatically downgrade to the Free plan. --- ### How to Change Email Address URL: https://pagecrawl.io/help/account-settings/article/how-to-change-email-address Unfortunately, for security and to prevent service abuse, email addresses cannot be changed directly by users. [Image: Your Account section in Account Settings showing the read-only email field with a note to contact support to change it] To change your email address please contact support at [help_me@pagecrawl.io](mailto:help_me@pagecrawl.io) from your originally registered email address. We will verify the information and get back to you as soon as possible. _Email address for 'Free Forever' plan users cannot be changed to prevent service abuse._ --- ### How to Delete My Account URL: https://pagecrawl.io/help/account-settings/article/how-to-delete-my-account **Deletion of your account will result in loss of ALL data associated with it.** To delete your account go to the **Account Settings**, scroll to the bottom of the page, press **Permanently delete my account**, and proceed with the instructions. [Image: Account Deletion section at the bottom of Account Settings with the Permanently delete my account button] --- ### SAML SSO Configuration in PageCrawl [Audience: Developer / technical] URL: https://pagecrawl.io/help/account-settings/article/saml-sso-configuration This guide covers the PageCrawl side of SSO setup: importing your identity provider's metadata, enabling SSO, configuring enforcement and user provisioning. For step-by-step instructions on configuring your identity provider (Azure AD, Google Workspace, Okta, etc.), see the [Identity Provider Setup Guide](/help/account-settings/article/set-up-identity-provider-for-saml-sso.md). Single Sign-On (SSO) allows your team members to securely access PageCrawl using your organization's identity provider, such as Azure AD, Google Workspace, Okta, or OneLogin. ## Requirements To use SAML SSO, your team must meet the following requirements: - **Enterprise or Ultimate Plan** subscription - **Corporate email domain** - The team owner must use a verified corporate email address (free email providers like Gmail, Yahoo, Outlook, and iCloud are not supported) - **Identity Provider** that supports SAML 2.0 standard ## How to Configure SAML SSO ### 1. Access SSO Settings Navigate to **Settings → Team → Auth & SSO** in your PageCrawl account. You must be a team Owner or Administrator to access these settings. [Image: Single Sign-On via SAML 2.0 section in Auth and SSO settings with the Enable SSO toggle] When you first access the SSO settings page, PageCrawl automatically generates a unique identifier (UUID) and creates an initial SSO configuration for your team. This UUID is immediately available and used to create your Entity ID and Metadata URL. ### 2. Get Service Provider Information Before configuring your Identity Provider, copy the **Metadata URL** displayed in the blue information box at the top of the SSO settings page. [Image: Service Provider Information box showing the Metadata URL, Entity ID, Reply URL (ACS), Sign on URL, and Logout URL] The URL will look like: `https://pagecrawl.io/sso/saml/abc-123-def-456/metadata` **Important:** Copy the actual URL shown in PageCrawl, not this example. Most Identity Providers can automatically import all necessary configuration (Entity ID, ACS URL, Logout URL, etc.) from this metadata URL. **Note:** If your IdP requires manual entry, the individual URLs are also displayed in the same box: - Reply URL (Assertion Consumer Service URL) - Sign on URL - Logout URL ### 3. Configure Your Identity Provider Follow the instructions in our [Identity Provider Setup Guide](./set-up-identity-provider-for-saml-sso) for your specific IdP (Azure AD, Google Workspace, Okta, etc.). You'll need to create a SAML application in your IdP and provide the ACS URL and Entity ID from step 2. ### 4. Import Identity Provider Metadata into PageCrawl You have three options to configure your IdP: [Image: Metadata import tabs in PageCrawl SSO settings: Metadata URL, Metadata XML, and Manual Entry] #### Option A: Metadata URL (Recommended) - Enter your IdP's metadata URL - Click "Parse Metadata from URL" - PageCrawl will automatically extract all required settings #### Option B: Metadata XML - Copy your IdP's metadata XML - Paste it into the metadata XML field - Click "Parse Metadata XML" #### Option C: Manual Entry - Manually enter Entity ID, SSO URL, SLO URL, and X.509 Certificate - This option is useful for custom configurations ### 5. Enable SSO Features Configure the following settings based on your needs: #### Enable SSO Turn on SAML authentication for your domain. #### Enforce SSO When enabled, password login will be disabled for users with your email domain. Users must authenticate via your identity provider. #### Exempt Users (Breakglass) When "Enforce SSO" is enabled, an **Exempt Users (Breakglass)** field appears. Add one or more team members here to keep an emergency way in if your identity provider ever becomes unavailable. Exempt users are a safety net for outages: if your IdP goes down, a certificate expires, or a misconfiguration locks everyone out, an exempt admin can still get in and fix or disable SSO. Exempt team members: - Can still log in with their PageCrawl email and password while SSO is enforced for everyone else. - Can still use the password reset flow (which is otherwise disabled for enforced domains). **Notes:** - Only existing team members can be added as exempt users. The email you enter must belong to a current member of the team, or it will be rejected. - Keep the exempt list short (typically one or two trusted administrators) and make sure those accounts have strong passwords and their own two-factor authentication enabled, since they bypass the SSO requirement. - Exempt users are not exempt from SSO, they can still choose to log in via SSO. Enforcement is simply not applied to them. #### Just-in-Time (JIT) Provisioning [Image: Just-in-Time provisioning settings: Enable Automatic Account Creation, Default Role, Create Personal Workspace, and Default Workspaces] **Enable Automatic Account Creation** - **Enabled**: New users logging in via SSO will automatically get accounts created - **Disabled**: Only existing users can log in via SSO. New users must be manually added first. When JIT provisioning is enabled, you can configure: **Default Role for New SSO Users** - Administrator - Standard User - Viewer **Default Workspaces** - Leave empty to assign all workspaces - Select specific workspaces to limit access **Auto-Create Personal Workspace** - When enabled, each new SSO user gets a personal workspace - Note: Your account has a workspace limit based on your subscription - If the limit is reached, no personal workspaces will be created ## Workspace Limits Personal workspace creation depends on your [subscription plan](/pricing): If you enable "Auto-Create Personal Workspace" and have reached your limit, new SSO users will be assigned to default workspaces instead of creating personal workspaces. ## SSO Login Flow Once configured, users with your email domain will: 1. Go to PageCrawl login page 2. Enter their email address 3. Be redirected to your identity provider 4. Authenticate with their corporate credentials 5. Be redirected back to PageCrawl and logged in automatically If JIT provisioning is enabled and they're a new user, an account will be created automatically with the configured role and workspace assignments. ## Troubleshooting Common Issues ### "Team has reached member limit" **Error:** "Unable to provision SSO user: Team has reached its member limit." **Solution:** - Check your subscription plan in **Settings → Team → Subscription** - Either upgrade to a plan with more seats or remove inactive members - Once you have available seats, the user can try logging in again ### "Automatic account creation is disabled" **Error:** "Automatic account creation is disabled. Please ask your team administrator to enable JIT provisioning." **Solution:** - Enable **"Enable Automatic Account Creation"** in **Settings → Team → Auth & SSO** - Or manually add the user in **Settings → Team → Users** before they log in ### User Not Assigned in Identity Provider **Symptoms:** User gets error after authenticating at IdP. **Solution:** - **Azure AD:** Go to Enterprise Applications → PageCrawl → Users and groups → Add user/group - **Google Workspace:** Admin Console → PageCrawl app → User access → Enable for user's org unit - **Okta:** Applications → PageCrawl → Assignments → Assign to People ### Certificate Expired or Invalid **Symptoms:** "Invalid signature" or authentication fails at final step. **Solution:** 1. In PageCrawl SSO settings, update the metadata: - Click **Parse Metadata from URL** to refresh, or - Download fresh XML from IdP and paste it, then click **Parse Metadata XML** 2. Most IdPs rotate certificates every 1-3 years ### Metadata Import Errors **Common Issues:** - **EntitiesDescriptor Format:** PageCrawl requires `EntityDescriptor` format, not `EntitiesDescriptor` - **Invalid XML:** Ensure you copied the entire XML including ` Primary email - **Signed response**: Leave unchecked (PageCrawl requires signed assertions, which is the industry standard default) 2. Click **Continue** 3. Click **Finish** (skip attribute mapping) ### Step 4: Import Metadata to PageCrawl 1. Open the downloaded metadata XML file 2. In PageCrawl SSO settings, paste the content into **Metadata XML** field 3. Click **Parse Metadata XML** ### Step 5: Turn On the App 1. In Google Admin, click on your PageCrawl app 2. Click **User access** 3. Select **ON for everyone** or specific organizational units 4. Click **Save** --- [Image: Okta logo] ### Step 1: Add Application 1. Sign in to your [Okta Admin Console](https://admin.okta.com) 2. Go to **Applications → Applications** 3. Click **Create App Integration** 4. Select **SAML 2.0** and click **Next** ### Step 2: General Settings 1. Enter "PageCrawl" as the **App name** 2. (Optional) Upload a logo. You can [download the PageCrawl logo](https://pagecrawl.io/images/logo-dark.png) to use here. 3. Click **Next** ### Step 3: Configure SAML 1. In the **SAML Settings** section, enter: - **Single sign-on URL**: Paste your Reply URL from PageCrawl (e.g., `https://pagecrawl.io/sso/saml/abc-123.../acs`) - **Audience URI (SP Entity ID)**: Paste your Entity ID from PageCrawl (e.g., `https://pagecrawl.io/sso/saml/abc-123.../metadata`) - **Name ID format**: EmailAddress - **Application username**: Email 2. Leave other settings as default 3. Click **Next** ### Step 4: Feedback 1. Select **I'm an Okta customer adding an internal app** 2. Click **Finish** ### Step 5: Get Metadata URL 1. On the **Sign On** tab, scroll to **SAML Signing Certificates** 2. Click **Actions** next to the active certificate 3. Click **View IdP metadata** 4. Copy the URL from your browser's address bar 5. In PageCrawl SSO settings, paste this URL in the **Metadata URL** field 6. Click **Parse Metadata from URL** ### Step 6: Assign Users 1. Go to the **Assignments** tab 2. Click **Assign** and select **Assign to People** or **Assign to Groups** 3. Assign users who should have access to PageCrawl 4. Click **Done** --- [Image: OneLogin logo] ### Step 1: Add Application 1. Sign in to your [OneLogin Admin Console](https://app.onelogin.com/admin) 2. Go to **Applications → Applications** 3. Click **Add App** 4. Search for "SAML Test Connector (Advanced)" and select it ### Step 2: Configure Application 1. Enter "PageCrawl" as the **Display Name** 2. Click **Save** ### Step 3: Configure SAML Settings 1. Go to the **Configuration** tab 2. Enter the following: - **Audience (Entity ID)**: Paste your Entity ID from PageCrawl (e.g., `https://pagecrawl.io/sso/saml/abc-123.../metadata`) - **Recipient**: Paste your Reply URL from PageCrawl (e.g., `https://pagecrawl.io/sso/saml/abc-123.../acs`) - **ACS (Consumer) URL Validator**: Use regex pattern `https://pagecrawl\.io/sso/saml/[^/]+/acs` - **ACS (Consumer) URL**: Paste your Reply URL from PageCrawl (e.g., `https://pagecrawl.io/sso/saml/abc-123.../acs`) 3. Click **Save** ### Step 4: Get Metadata URL 1. Go to the **More Actions** menu 2. Select **SAML Metadata** 3. Copy the metadata URL 4. In PageCrawl SSO settings, paste this URL in the **Metadata URL** field 5. Click **Parse Metadata from URL** ### Step 5: Assign Users 1. Go to the **Users** tab 2. Select users who should have access 3. Click **Save** --- ## Custom SAML 2.0 Provider If your identity provider isn't listed above but supports SAML 2.0, you can configure it manually: ### Step 1: Configure Your Identity Provider In your IdP, create a new SAML application with these settings: - **Entity ID**: Paste your Entity ID from PageCrawl (you copied this in the first section above, e.g., `https://pagecrawl.io/sso/saml/abc-123.../metadata`) - **ACS URL**: Paste your Reply URL from PageCrawl (e.g., `https://pagecrawl.io/sso/saml/abc-123.../acs`) - **NameID Format**: Email Address - **Binding**: HTTP-POST for ACS, HTTP-Redirect for SSO ### Step 2: Get IdP Information From your identity provider, collect: - **Entity ID** (IdP Issuer) - **SSO URL** (Sign-on URL) - **SLO URL** (Sign-out URL) - Optional - **X.509 Certificate** ### Step 3: Manual Configuration in PageCrawl 1. In PageCrawl SSO settings, select the **Manual Entry** tab 2. Enter the collected information: - Entity ID - SSO URL - SLO URL (optional) - X.509 Certificate (paste the full certificate including BEGIN/END markers) 3. Enable SSO and configure JIT provisioning settings 4. Click **Save Changes** --- ## Validation After configuration, test your SSO: 1. Open an incognito/private browser window 2. Go to PageCrawl login page 3. Enter a test user's email address with your domain 4. Verify you're redirected to your IdP 5. Complete authentication 6. Verify you're logged into PageCrawl successfully If you encounter issues, check: - User is assigned to the PageCrawl application in your IdP - Email domain matches your configured domain - Metadata was imported correctly - X.509 certificate is valid and not expired --- ## Notes - **Metadata XML Format**: PageCrawl does not support the `EntitiesDescriptor` element. Use `EntityDescriptor` format. - **Multiple IdPs**: Each team supports one identity provider. Organizations that need more than one IdP can set up multiple teams (each with its own identity provider) on a custom plan. Contact [support@pagecrawl.io](mailto:support@pagecrawl.io) to set this up. - **Certificate Rotation**: When your IdP certificate expires, update the metadata in PageCrawl SSO settings. ## Support For assistance with your specific identity provider, contact [support@pagecrawl.io](mailto:support@pagecrawl.io). --- ### User Access Roles and Permissions URL: https://pagecrawl.io/help/account-settings/article/user-access-roles PageCrawl uses role-based access control to manage what each team member can do. There are four roles, each with different permission levels. Manage members and their roles under **Settings > Team > Users**. [Image: Users settings page listing team members with their email, assigned workspaces, and role, plus the Invite User button] ### Available Roles | Role | Manage Team | Manage Workspaces | Edit Pages | View Pages | |------|:-----------:|:-----------------:|:----------:|:----------:| | **Owner** | Yes | Yes | Yes | Yes | | **Administrator** | Yes | Yes | Yes | Yes | | **Standard User** | No | No | Yes | Yes | | **Viewer** | No | No | No | Yes | ### Owner Each team has exactly one Owner (the account creator). The Owner has full control over all team settings, billing, and member management. Ownership cannot be transferred or removed. ### Administrator Administrators can manage the team on behalf of the Owner: - Invite and remove team members - Change member roles - Assign workspace access to members - Create and delete workspaces - Edit all team and workspace settings (notifications, integrations, AI, etc.) - Full access to all workspaces ### Standard User Standard Users can work within their assigned workspaces: - View and edit monitored pages in assigned workspaces - Create new pages and tracked elements - Review changes and leave feedback - Access all monitoring features within their workspaces Standard Users cannot invite members, change roles, or access workspaces they haven't been assigned to. ### Viewer Viewers have read-only access to their assigned workspaces: - View monitored pages and detected changes - Browse change history and reports - Cannot create, edit, or delete pages - Cannot modify any settings ### Managing Team Members To manage roles and access: 1. Go to **Settings** > **Team** > **Users** 2. View the member list showing name, email, workspaces, and role 3. Click a member's role to change it (Owner and Administrator only) 4. Click **Update** in the Workspaces column to assign or revoke workspace access ### Inviting New Members 1. Go to **Settings** > **Team** > **Users** 2. Click **Invite User** 3. Enter their email address and select a role 4. The invite expires after 2 weeks. You can resend it if needed. ### Workspace Access Members only see workspaces they've been assigned to. Administrators can assign workspace access per user. If all workspace access is removed from a user, they are removed from the team entirely. This means you can have team members who only see specific projects, clients, or departments without exposure to other workspaces. --- ### Why Was My Password Rejected? URL: https://pagecrawl.io/help/account-settings/article/why-was-my-password-rejected When you create a PageCrawl.io account or change your password, we sometimes ask you to pick a different one. This page explains why that happens and how to choose a password that will be accepted. ## What we check Our password rules follow NIST SP 800-63B, the digital identity guideline published by the US National Institute of Standards and Technology. In line with that guideline, a password only needs to meet two conditions: - It must be at least 8 characters long. Spaces are fine, and you can use a whole sentence if you like. - It must not appear on public lists of commonly used passwords. That is all. The same guideline advises against requiring uppercase letters, numbers, or special characters, because length protects an account far better than forced symbols. We would rather you pick a longer password you can actually remember. ## Why some passwords are declined Over the years, many websites around the world have had security incidents, and large lists of the passwords used there now circulate publicly. Attackers try those lists first whenever they attempt to break into accounts. If a password appears on those lists, it can be guessed in seconds, no matter how personal it feels. When we decline a password for this reason, it simply means many other people have used the same password before and it is now publicly known. ## Does this mean my account was hacked? No. Nothing has happened to your PageCrawl.io account or your data. The check looks at the password itself, not at you. It only tells us that the same combination of characters appears on public lists. PageCrawl.io has not been breached, and no one is targeting your account. Declining the password is purely a precaution, so that your account never depends on a password attackers already know. ## How the check works Your password is never shared with anyone. The check uses only the first few characters of a scrambled fingerprint of your password (called a hash) and compares that tiny fragment against Have I Been Pwned, a well known public safety database used by many major online services. It is impossible to reconstruct your password from the fragment used in the lookup. ## How to pick a strong password - Longer is stronger. A short sentence such as "my dog naps on the blue sofa" is easy to remember and very hard to guess. - Avoid single common words, names, and simple patterns like "12345678". - Use a different password for every website. A password manager makes this effortless and can generate and remember passwords for you. If your password was declined, try a longer and more personal phrase. And if you have used the declined password on other websites, changing it there as well is a good idea. --- ### Features ### Advanced Configuration Options for Power Users [Audience: Developer / technical] URL: https://pagecrawl.io/help/features/article/advanced-configuration PageCrawl offers advanced configuration options for users who need fine-grained control over their monitoring setup. This guide covers the key power-user features. ### Power User Mode When editing a monitored page, you can enable **Power User** mode using the toggle in the page settings. This reveals additional settings that are hidden by default to keep the interface clean for everyday use. With Power User mode enabled, you get access to: - **Engine selection** - Choose between the default browser engine, Stealth Mode (for sites that block bots), or Fast mode (optimized for static pages) - **Intelligent Reconnect** - Automatically retry failed checks with a different approach - **Custom User Agent** - Set a specific browser user agent string - **Custom Headers** - Add custom HTTP headers to requests - **Custom JavaScript** - Run JavaScript code before or after page load - **Device emulation** - Emulate specific device viewports Power User settings are marked with a special icon throughout the edit form so you can easily identify them. [Image: Power User Settings enabled in the page editor, revealing Engine, Intelligent Reconnect, Device Simulation, User-Agent, Request Headers, and Custom Proxies options] ### Advanced Mode vs Simple Mode PageCrawl offers two ways to add and edit monitored pages: **Simple Mode** (default) guides you through setup step by step. It auto-detects the best settings, shows a live preview, and covers the most common use cases. Best for getting started quickly. **Advanced Mode** gives you full control over every setting in a single form. Use it when you need to: - Track multiple elements on the same page simultaneously - Configure complex action sequences - Set up templates or apply existing ones - Fine-tune notification conditions per element - Work with custom selectors, thresholds, and comparison methods You can switch to Advanced Mode from the Simple Mode page by clicking the "Advanced setup" link at the bottom. If you prefer to always use Advanced Mode, check the "Always show Advanced Setup" option. ### Multiple Tracked Elements Each monitored page can track multiple elements simultaneously, each with its own comparison method: | Type | What It Tracks | |------|---------------| | **Full Page** | Entire page text content | | **Text** | Text content of a specific element (by CSS/XPath selector) | | **Number** | Numeric values with configurable change thresholds | | **Price** | Price values with currency detection | | **Availability** | In-stock/out-of-stock status | | **Links** | All outgoing links on the page | | **Visual** | Visual screenshot comparison with diff percentage | | **HTML** | Raw HTML structure of an element | | **Boolean** | Presence or absence of an element | | **Feed/List** | RSS, Atom, or other feed content | | **Rating** | Star ratings or review scores | | **Reviews** | Customer review text and metadata | | **JavaScript** | Values extracted by running custom JavaScript | | **SEO Tags** | Meta tags, Open Graph data, and structured data | | **PDF** | Text content extracted from PDF files | | **Word** | Text content extracted from Word documents | | **Excel** | Data extracted from Excel spreadsheets | | **CSV** | Data extracted from CSV files | | **PowerPoint** | Text content extracted from PowerPoint presentations | Each tracked element can have its own set of [actions](/help/features/article/perform-actions.md) and comparison settings. [Image: Multiple tracked elements configured on a single page, each with its own type and comparison method] ### Templates Templates let you save a monitoring configuration and apply it to multiple pages automatically. This is especially useful when combined with [Page Discovery](/help/features/article/page-discovery.md) for auto-monitoring newly discovered pages. To create a template: 1. Go to **Settings** > **Workspace** > **Templates** 2. Enter a sample URL to auto-fill settings 3. Configure tracked elements, actions, check frequency, and notifications 4. Save the template [Image: Template Details form with a template label, sample URL, check frequency, and proxy location] Templates can also define URL filters for page discovery, so new pages matching your criteria are automatically monitored with the template's settings. ### Bulk Editing Edit settings across multiple pages at once: 1. Select pages from your page list using the checkboxes 2. Click **Bulk Edit** in the toolbar 3. Choose what to change: check frequency, engine, proxy, actions, notifications, tags, or folder 4. Apply changes to all selected pages [Image: Tracked pages with rows selected and the Bulk actions menu open] Available on paid plans. ### AI Configuration Configure AI-powered change analysis per workspace: 1. Go to **Settings** > **Workspace** > **Integrations** > **AI** 2. Choose your AI provider (OpenAI, Gemini, or Anthropic) 3. Select a model 4. Optionally set focus areas to guide the AI on what changes matter most [Image: AI Features Configuration in workspace integrations with provider, model, and credit usage] Each plan includes monthly AI credits. You can also bring your own API key (BYOK) for unlimited usage. See [AI BYOK Setup](/help/integrations/article/ai-byok-setup-guide.md) for details. ### Custom Check Scheduling Control exactly when PageCrawl checks your pages: 1. Go to **Settings** > **Workspace** > **Schedule** 2. Set active monitoring hours (e.g., business hours only) 3. Choose which days of the week to run checks 4. Set the workspace timezone [Image: Workspace schedule settings with quick presets, day selection, an active-hours slider, and a schedule summary] This helps reduce unnecessary checks during off-hours and keeps your check quota focused on the times that matter. ### Global Filters Apply text filters across all pages in a workspace: 1. Go to **Settings** > **Workspace** > **General** 2. Add global ignored text patterns 3. These patterns are excluded from change detection on every page in the workspace Useful for filtering out dynamic content like timestamps, ad copy, or session IDs that appear across many pages. ### Proxy Configuration Choose where PageCrawl checks your pages from: - **Default** - Automatic server selection - **Proxy Pool** - Use your own proxies (managed under Settings → Proxy Pools) for pages behind firewalls or geo-restrictions - **Location-specific** - Select from available proxy locations (London, New York, San Francisco, Toronto, Frankfurt, Tel Aviv) - **Residential** - Use residential IP addresses for pages that block datacenter IPs Configure per page or apply via bulk edit. --- ### AI-Powered Change Detection and Smart Filtering URL: https://pagecrawl.io/help/features/article/ai-powered-change-detection PageCrawl.io includes AI-powered analysis for all users. Every plan comes with monthly AI credits that work automatically with zero setup. When a page changes, AI summarizes what happened and scores how important the change is, so you only get notified about what matters. For users who need more, you can also bring your own API key (BYOK) for unlimited AI usage and full model control. ## AI Credits Every plan includes monthly AI credits, visible in the workspace AI settings: [Image: AI Features Configuration with the AI Credits usage meter and a summary of what AI does] Every plan includes monthly AI credits: | Plan | Monthly Credits | |------|----------------| | **Free** | 15 | | **Standard** | 200 (scales with quantity) | | **Enterprise** | 1,000 (scales with quantity) | | **Ultimate** | 10,000 (scales with quantity, includes Pro tier) | Credits are based on page size. Each 4,000-token block costs 1 credit on Basic tier or 10 credits on Pro tier (Ultimate plan only). A typical blog post uses 1-2 credits. Credits reset monthly. When credits run out, page monitoring continues normally, but AI summaries and importance filtering pause until the next billing cycle. You can also switch to BYOK at any time for unlimited usage. ## Getting Started No setup is required. AI features are enabled by default for all workspaces: 1. Add pages to monitor as usual 2. When changes are detected, AI automatically summarizes them and assigns importance scores 3. View your credit usage in **Settings > Workspace > Integrations > AI** **Workspace-specific**: AI features are configured per workspace. You can have some workspaces with AI enabled and others without. [Image: What matters AI panel in the page editor with the AI summarization toggle and focus instructions] ## How AI Features Work | Feature | Process | |---------|---------| | **Summarization** | Change detected > Content sent to AI > Human-readable summary generated > Included in notification | | **Importance Scoring** | Change detected > AI analyzes content > Priority score assigned (0-100) > Low-priority changes filtered | | **[Label Automation](#ai-label-automation)** | Change detected > AI evaluates your label rules > Labels automatically added or removed | ## Configuration ### Available for All Users | Setting | Description | |---------|-------------| | **Custom Instructions** | Teach AI what matters for your monitoring (max 2,000 chars) | | **Summary Language** | Generate summaries in 19 languages: English, Spanish, French, German, Italian, Portuguese, Dutch, Polish, Ukrainian, Russian, Japanese, Korean, Chinese, Arabic, Hindi, Turkish, Lithuanian, Latvian, and Estonian | | **Minimum importance** | The lowest importance score (0-100) that sends a notification. Changes scoring below it are still tracked, they just stay quiet, and you can reveal them in the change history with the status filter. Set it for the whole workspace, or override it on a single page or template under Power User settings. | ### Additional BYOK Settings These settings are available when using your own API key: | Setting | Description | |---------|-------------| | **Deep Analysis** | Send full page content to AI for better context. Uses more tokens but provides more accurate analysis. When disabled, only the changed text (diff) is sent. | | **Run on First Check** | Get AI analysis on the initial page check, before any changes are detected | | **AI Requests Per Month** | Set a monthly cap to control costs. When the limit is reached, AI features pause until the next month. Leave empty for unlimited. | | **Per Page Per Day** | Limit how many AI analyses a single page can trigger in 24 hours. Prevents noisy pages from consuming your entire budget. Default: 10. | | **Max Tokens** | Limit content size per request. If content exceeds this limit, AI analysis is skipped for that change. | ### Understanding Tokens A **token** is roughly 4 characters or about 3/4 of a word. With included credits, each 4,000-token block counts as 1 credit. | Page Type | Typical Tokens | |-----------|---------------| | Simple (blog, article) | ~1,000-2,000 | | Medium (product, news) | ~2,000-5,000 | | Large (documentation) | ~5,000-10,000 | ## Using Your Own API Key (BYOK) If your included credits are not enough, or you want full control over model selection, you can connect your own API key from OpenAI, Google Gemini, Anthropic, or OpenRouter. 1. Go to **Settings > Workspace > Integrations > AI** 2. Select your AI provider and enter your API key 3. Click **Test Connection** to verify 4. Choose your preferred model and save [Image: BYOK API Configuration with AI provider, model selection, Quick Select model tiers, and a confirmed API key] When using BYOK, AI credits are not consumed and you pay your AI provider directly. See the [BYOK Setup Guide](/help/integrations/article/ai-byok-setup-guide.md) for detailed instructions. ## Best Practices ### Start Small - AI is enabled by default, so monitor your credit usage for the first few weeks - Check usage statistics in **Settings > Workspace > Integrations > AI** - If you need more credits, upgrade your plan or connect your own API key ### Optimize Credit Usage - Use Custom Instructions to help AI focus on what matters - A daily cap of 10 analyses per page prevents noisy pages from consuming your budget - For high-volume monitoring, consider BYOK with a budget model like Gemini Flash-Lite ### Choose the Right Mode | Scenario | Recommendation | |----------|---------------| | Getting started | Use included credits (no setup needed) | | High-volume pages | Enable Importance Scoring to filter noise | | Technical pages | Enable Summarization for readable changes | | Need unlimited AI | Connect your own API key (BYOK) | | Critical pages | Use BYOK with premium models (GPT-5.5, Claude Sonnet) | ## AI Label Automation AI can automatically apply or remove labels on detected changes based on rules you define. Instead of manually categorizing changes, the AI reads each change and decides which labels to add or remove according to your instructions. ### How to Set It Up 1. Go to **Settings > Workspace > Labels** 2. Scroll to the **AI Label Automation** section 3. Click **Add Rule** to create a label/instruction pair 4. For each rule, choose a label name and write a plain-language instruction explaining when the AI should apply it 5. Click **Save Changes** [Image: AI Label Automation section with a label/instruction rule pairing the Price Drop label to an instruction, plus the Add Rule button] You can configure up to 10 label rules per workspace. ### How It Works Each time a change is detected and AI analysis runs, the AI evaluates the change against your label rules and decides which labels to add or remove. The AI receives the current labels on the page, so it can remove labels that no longer apply (e.g., removing "Out of Stock" when a product is back in stock). Labels are applied to the change record, making them available for filtering on the [Review Board](/help/features/article/review-board.md) and in your page list. ### Example Rules | Label | Instruction | |-------|-------------| | Breaking News | Apply when urgent or breaking news appears | | Policy Update | Apply when terms, policies, or legal text changes | | New Event | Apply when a new conference or event is announced | | Job Posted | Apply when new job listings are added | | Content Removed | Apply when significant content is deleted from the page | ### Important Notes - AI can only manage labels defined in your automation rules. Manually applied labels are never touched. - Label names have a maximum of 50 characters; instructions have a maximum of 500 characters. - Labels are created automatically if they do not already exist in your workspace. - AI Label Automation requires AI to be configured for the workspace (either included credits or BYOK). - Label decisions run as part of the standard AI analysis, so no additional credits are used beyond the normal change analysis. ## Security and Privacy | Aspect | Details | |--------|---------| | **Included credits** | Content is processed through PageCrawl's managed AI infrastructure | | **BYOK mode** | Content is sent directly to your chosen AI provider | | **Storage** | AI summaries stored in PageCrawl.io for your reference | | **Security** | All transmission via HTTPS, API keys encrypted at rest | | **Provider policies** | Review your AI provider's data usage and retention policies when using BYOK | ## Related Articles - [How PageCrawl Uses AI](/help/features/article/how-pagecrawl-uses-ai.md) - An overview of where AI summaries, scoring, and labels appear, and how to control them - [AI Integration Setup Guide (BYOK)](/help/integrations/article/ai-byok-setup-guide.md) - Step-by-step guide to configure your own API keys for unlimited AI usage - [Choosing the Right AI Model for Website Monitoring](/help/tutorials/article/choosing-best-ai-model-website-monitoring.md) - Compare models and pricing for BYOK users --- ### API and Webhooks for Custom Integrations [Audience: Developer / technical] URL: https://pagecrawl.io/help/features/article/api-webhooks-for-custom-integrations PageCrawl provides three ways to integrate page monitoring into your own applications and workflows: a **REST API** to manage monitors programmatically, **webhooks** for real-time change notifications, and **RSS feeds** for lightweight consumption. [Image: Interactive API reference at pagecrawl.io/developers with endpoints, authentication, and copy-paste code examples] Note: The REST API and webhooks are available on every plan, including Free. ### Authentication All API requests require a Bearer token. Go to **Settings > API > API Tokens** and click **Create Token**, then copy it immediately (it is not shown again). [Image: API Tokens card with the tokens table, Create Token button, and the Bearer authorization note] Include the token in the `Authorization` header on every request: ``` Authorization: Bearer YOUR_API_TOKEN ``` For the full reference with every endpoint, parameter, and response schema, see [pagecrawl.io/developers](/developers). ### Quick Start #### Step 1: Create a monitor The simplest way to start monitoring is the `/api/track-simple` endpoint. It only requires a URL. ```bash curl -X POST "https://pagecrawl.io/api/track-simple" \ -H "Authorization: Bearer YOUR_API_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "url": "https://example.com/pricing", "tracking_mode": "fullpage", "ai_page_focus": "Alert me if the price drops or the item goes out of stock; skip header, footer, and cookie-banner changes" }' ``` ```python import requests response = requests.post( "https://pagecrawl.io/api/track-simple", headers={"Authorization": "Bearer YOUR_API_TOKEN"}, json={"url": "https://example.com/pricing", "tracking_mode": "fullpage", "ai_page_focus": "Alert me if the price drops or the item goes out of stock; skip header, footer, and cookie-banner changes"}, ) page = response.json() print(f"Monitoring: {page['name']} (ID: {page['id']})") ``` ```javascript const response = await fetch("https://pagecrawl.io/api/track-simple", { method: "POST", headers: { "Authorization": "Bearer YOUR_API_TOKEN", "Content-Type": "application/json", }, body: JSON.stringify({ url: "https://example.com/pricing", tracking_mode: "fullpage", ai_page_focus: "Alert me if the price drops or the item goes out of stock; skip header, footer, and cookie-banner changes" }), }); const page = await response.json(); console.log(`Monitoring: ${page.name} (ID: ${page.id})`); ``` **Tracking modes:** `fullpage` (all visible text, default), `content_only` (text without navigation/headers/footers), `reader` (reader-mode content), `price` (auto-detect prices), `specific_text` (requires `selector`), `specific_number` (requires `selector`). **`ai_page_focus`** (optional) is free text telling the AI what matters most on this page, for example `"Alert me if the price drops or the item goes out of stock; skip header, footer, and cookie-banner changes"`. It sharpens change summaries and priority scoring. **Frequency:** an optional `frequency` field (in minutes) controls how often the page is checked: `1440` for daily, `60` for hourly, `15` for every 15 minutes (depends on your plan). When omitted, it defaults to daily. #### Step 2: Set up a webhook ```bash curl -X POST "https://pagecrawl.io/api/hooks" \ -H "Authorization: Bearer YOUR_API_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "target_url": "https://your-server.com/webhook", "match_type": "all", "events": ["change_detected"] }' ``` **Match types:** `all` (every page), `monitors`, `tags`, `folders`, `domains`. **Events:** `change_detected`, `error`, `price_change_detected`. #### Step 3: Handle webhook payloads When a change is detected, PageCrawl POSTs a JSON payload to your endpoint. ```python from flask import Flask, request app = Flask(__name__) @app.route("/webhook", methods=["POST"]) def handle_change(): data = request.json print(f"Change detected: {data['title']}") print(f"Difference: {data['human_difference']}") if data.get("ai_summary"): print(f"AI Summary: {data['ai_summary']}") return "", 200 ``` **Key payload fields:** `title`, `contents` (current value), `difference` (0-100), `human_difference`, `ai_summary`, `ai_priority_score`, `markdown_difference`, `page_screenshot_image`. ### API Endpoints | Method | Endpoint | Description | |--------|----------|-------------| | `GET` | `/api/pages` | List all monitored pages | | `POST` | `/api/pages` | Create a new monitored page | | `GET` | `/api/pages/{slug}` | Get page details and latest values | | `PUT` | `/api/pages/{id}` | Update page settings | | `DELETE` | `/api/pages/{id}` | Delete a monitored page | | `PUT` | `/api/pages/{id}/check` | Trigger an immediate check | | `PUT` | `/api/pages/{id}/status` | Enable or disable a page | | `GET` | `/api/pages/{id}/history` | Get check history for a page | | `GET` | `/api/pages/{id}/checks/{checkId}/diff.markdown` | Get a text diff as markdown | For the complete endpoint list with parameters and schemas, see [pagecrawl.io/developers](/developers). ### Webhooks Webhooks send HTTP POST requests with a JSON body to your endpoint whenever a page change is detected or an error occurs. Configure them in **Settings** > **Webhooks**. | Setting | Description | |---------|-------------| | **Target URL** | The HTTP endpoint that receives the POST request | | **Event triggers** | Change detected, error, or both | | **Page filter** | Limit to specific pages, tags, folders, or a domain, or fire for all pages | | **Payload fields** | Select which fields to include (all by default) | Available payload fields include page ID, title, change summary, diff data (markdown and HTML), screenshots, AI summary, AI priority score, and per-element values. See the [Webhook Integration guide](/help/integrations/article/webhook-integration.md) for the full field reference and example payloads. **Reliable delivery (automatic retries):** If your endpoint is briefly unreachable, for example your server was offline for a few minutes or returned a temporary error, PageCrawl automatically retries the webhook with a backoff delay rather than dropping the event. This means a short outage on your side does not cost you the change data: once your server is back online, the queued events are delivered. Return a `2xx` status code to acknowledge receipt; any other response (or a timeout) is treated as a failure and scheduled for retry. ### RSS Feeds Prefer to consume changes in an RSS reader or automation tool? PageCrawl can generate an Atom feed of detected changes scoped to all pages, tags, folders, a domain, or specific monitors. See the [RSS Feeds guide](/help/features/article/page-monitoring-rss-feeds.md) for setup. ### Download the OpenAPI Spec The full specification is available as an OpenAPI 3.0 file you can import into Postman, Insomnia, or any API client: ``` https://pagecrawl.io/api/openapi.yaml ``` ### Common Use Cases - **Custom dashboards** - Pull change data into your own monitoring dashboard via API - **Automation workflows** - Trigger actions in n8n, Make, Zapier, or custom scripts via webhooks - **Database logging** - Store all detected changes in your own database - **Alerting systems** - Forward high-priority changes to PagerDuty, Opsgenie, or similar ### Related Articles - [Full API Reference](/developers) - Interactive OpenAPI reference with every endpoint and schema - [Webhook Integration](/help/integrations/article/webhook-integration.md) - Detailed webhook setup, payload reference, and testing - [Advanced Integrations](/help/tutorials/article/reference-implementations.md) - Copy-paste polling and webhook code in Python, Node.js, and PHP - [Zapier Integration](/help/integrations/article/pagecrawl-zapier-integration.md) - Connect PageCrawl to 5,000+ apps - [n8n Integration](/help/integrations/article/pagecrawl-n8n-integration.md) - Open-source workflow automation - [RSS Feeds](/help/features/article/page-monitoring-rss-feeds.md) - Subscribe to changes via RSS --- ### Automatically Discover New Pages To Track URL: https://pagecrawl.io/help/features/article/page-discovery PageCrawl is designed to make website change monitoring and management seamless. Page Discovery automatically finds new pages on a website and starts monitoring them for you, so your coverage stays up to date as the site grows. The fastest way to get started is to add a whole website from the **Discovered Pages** page, then fine-tune how it works from the template that gets created for you. ### Add a Website to Discover Start from the [Discovered Pages](/app/discovered-pages) page and click **Add Website**. This is the quickest way to begin: enter a website and PageCrawl handles the rest. [Image: Add Website for Discovery dialog with the website URL, what to track options, check frequency, and notification settings] In the dialog, set: 1. **Website URL** - the site you want to monitor (for example, a competitor's store or your own marketing site). 2. **What to track** - choose to track every discovered page, only top-level pages (like `/pricing` and `/about`), or review pages yourself before anything is monitored. 3. **Check Frequency** - how often discovered pages are checked for changes. 4. **Notify me via** - the channels that should receive change alerts. Click **Add Website**. PageCrawl scans the site, lists everything it finds on the Discovered Pages page, and creates a **template** with sensible defaults that controls how this website is monitored. ### Adjust Advanced Configuration via the Template Adding a website automatically creates a configured **template** for it. The template is where you fine-tune scanning, filters, and what gets tracked. Open it from [Templates settings](/app/settings/workspace/templates) (or the "Templates settings" link in the Add Website dialog) and edit the template for your website. #### Choose a Scanning Method The template's **Discover New Pages** setting controls how PageCrawl looks for new pages. The default is **Automatic (recommended)**, which combines methods to find pages using the best approach for each website: * **Automatic (recommended)**: Combines sitemap and link discovery to find pages using the best method for the website. This is the default and recommended setting. * **Homepage Links Only**: Discover new links by following links on the homepage. Available as a daily or weekly check. Useful if you want to focus on pages directly linked from the main page. * **Sitemap Only**: Discover pages listed in the website's sitemap. Most websites have a sitemap to help search engines find their pages, making this an efficient method for large sites. * **Follow Links 2 Levels Deep**: Follows links on the homepage, then follows links on those pages too. Available as a weekly check. Note: Only available on Enterprise and Ultimate plans. * **Follow Links 3 Levels Deep**: Follows links on the homepage, then follows links two more levels deep. Available as a weekly check. Note: Only available on Enterprise and Ultimate plans. * **Deep Scan**: Conduct a comprehensive analysis by visiting every accessible page on your website. This ensures that no new links go unnoticed, even on deeply nested pages. Note: Only available on Enterprise and Ultimate plans. #### Apply Include and Exclude Filters Use the template's filters to control exactly which discovered pages are monitored, so you avoid tracking irrelevant pages. [Image: Discovery filter controls in the template: include and exclude rules, Track All Pages, and a tracked page limit] * **Include rules**: Specify keywords or patterns that a page must match to be tracked. Useful for tracking specific types of content, such as `/product/` pages only. * **Exclude rules**: Define keywords or patterns that should be skipped. Ideal for ignoring pages you do not care about, such as `/tag/` or `/archive/` URLs. * **Track All Pages**: Track every discovered page without filtering. * **Tracked Page Limit**: Cap how many pages are auto-tracked to keep usage under control. #### Configure Tracked Elements The template also defines what is tracked on each discovered page. You can monitor all pages, or only those with a specific structure (for example, only product pages). 1. To monitor whole pages, set the tracked element to **Full-page Text**. 2. To monitor pages with a specific layout, configure multiple tracked elements, such as product title, price, and description. If these elements do not exist on a discovered page, that page is simply skipped. Save the template, and PageCrawl will keep discovering and monitoring matching pages automatically. If too many irrelevant pages are discovered, tighten the include/exclude filters and remove the pages you do not want. --- ### Available Tracked Element Types URL: https://pagecrawl.io/help/features/article/available-tracked-monitoring-types [Image: What to Track panel with the element type options (Full Page Text, Specific Area, Visual, Price, Feed, Parse) highlighted] When monitoring changes on a webpage, the type of tracked element selected defines what kind of content will be tracked and how updates are detected. The six standard types below appear in the standard editor under **What to Track**. Everything else is available in **Advanced mode**, where you can also track several elements on the same page. Note: You can configure multiple tracked elements on a single page and mix types freely. For example, track a product's Price, Availability, and Rating all at once, each with its own threshold, conditions, and notifications. Add as many as you need in Advanced mode. ### Standard Element Types These appear in the standard editor under **What to Track** as soon as you add a page. #### Full Page Text - **Description:** Tracks all visible text on the entire webpage. - **Use Case:** Useful for capturing comprehensive textual content. [Image: Full Page Text example: a web page on the left, and PageCrawl's text diff showing the changed line on the right] #### Text (Specific Area) - **Description:** Monitors text changes in a specified area of a webpage. - **Important Note:** Only the first element matching the selector is tracked. - **Use Case:** Ideal for tracking text in specific areas, like headlines or descriptions. [Image: Text example: a status page on the left, and PageCrawl's red/green text diff of the tracked text on the right] #### Visual - **Description:** Monitors and alerts on visual changes in a specified area. - **Note:** This is a beta feature; report any issues encountered. - **Use Case:** Ideal for tracking visual changes like layout updates or style changes. [Image: Visual example: a web page on the left, and PageCrawl's Visually Compare control with a difference-percentage progress bar on the right] #### Price - **Description:** Detects and extracts the first price found on the page. - **Limitation:** May not work well on pages with multiple prices. - **Use Case:** Monitoring product prices on e-commerce websites. [Image: Price example: a product page on the left, and PageCrawl showing the old price struck through, a down arrow, the new price, the percent change, and a sparkline] #### Feed / List - **Description:** Tracks entries from RSS or Atom feeds, detecting new, removed, or changed items. - **Use Case:** Monitoring blog feeds, news feeds, or any structured list for new entries and updates. [Image: Feed example: a blog feed on the left, and PageCrawl listing the newly added feed items on the right] #### Parse (AI Extract) - **Description:** Uses AI to extract the specific information you describe in plain language (for example "the event date" or "the lowest price"), even when there is no clean selector to target. - **Use Case:** Pulling a single fact or field out of an unstructured page; the extracted value is then tracked for changes. [Image: Parse example: an event page on the left, and PageCrawl's AI-extracted event date shown as a diff on the right] ### Advanced Mode Element Types Switch to **Advanced mode** (the link at the top of the page editor) to open the full **TYPE** dropdown, track several elements on one page, and use the additional types below. [Image: Tracked Elements section in Advanced mode with the TYPE, ELEMENT, and THRESHOLD selectors] #### Number - **Description:** Extracts and monitors numeric values in a specific webpage area. - **Features:** Provides basic statistical analysis and visual graphs. - **Use Case:** Useful for tracking numbers, such as stock levels or scores. - **Also in the standard editor:** You can enable Number tracking from **Text** mode in the standard create flow, without switching to Advanced mode. [Image: Number example: a star count on the left, and PageCrawl showing the old value struck through, an up arrow, the new value, the percent change, and a sparkline] #### Availability - **Description:** Tracks the availability status of a product on the page. - **Use Case:** Monitoring whether a product is in stock, out of stock, or on pre-order. - **Also in the standard editor:** You can enable Availability from **Price** mode in the standard create flow, without switching to Advanced mode. [Image: Availability example: a product page on the left, and PageCrawl showing 'Out of Stock (was: In Stock)' on the right] #### Rating - **Description:** Tracks the product rating displayed on the page. - **Use Case:** Monitoring changes to product ratings on review sites or e-commerce platforms. [Image: Rating example: a product's star rating on the left, and PageCrawl showing the old rating, a down arrow, the new rating, and the percent change on the right] #### Reviews - **Description:** Tracks the review count displayed on the page. - **Use Case:** Monitoring how many reviews a product has received over time. [Image: Reviews example: a review count on the left, and PageCrawl showing the old count, an up arrow, the new count, the percent change, and a sparkline on the right] #### Links - **Description:** Tracks internal and external links originating from a webpage. - **Use Case:** Ideal for monitoring link changes on resource-heavy websites. [Image: Links example: a navigation menu on the left, and PageCrawl listing added and removed links on the right] #### Iframes - **Description:** Monitors embedded content within `