openapi: 3.0.3
info:
  title: PageCrawl API
  description: |
    The PageCrawl API lets you interact with the PageCrawl.io web monitoring service programmatically.
    Create, update, and delete monitored pages, view change history, and retrieve screenshots.

    ## Quickstart

    Adding a page to monitor can be a one-liner. The shortest possible call needs only a URL:

    ```bash
    curl -X POST "https://pagecrawl.io/api/track-simple?url=https://example.com/" \
      -H "Authorization: Bearer YOUR_API_TOKEN"
    ```

    That's it — PageCrawl will start checking the page on the default schedule and notify you when it changes.

    **Add an AI focus to get smarter alerts.** `ai_page_focus` is a free-text hint that tells PageCrawl what actually matters on this page, so noisy edits (footer rotations, A/B copy tweaks) are deprioritised and the changes you care about float to the top:

    ```bash
    curl -X POST "https://pagecrawl.io/api/track-simple" \
      -H "Authorization: Bearer YOUR_API_TOKEN" \
      -H "Content-Type: application/json" \
      -d '{
        "url": "https://example.com/pricing",
        "ai_page_focus": "Focus on pricing tiers and plan limits; ignore testimonials and footer changes."
      }'
    ```

    Need more control? The same [Quick track a page](#operation/trackSimple) endpoint accepts every common option as you grow into them — check frequency, what to track (full page, a specific element, a price, an AI-extracted value), folders, tags, notification channels, custom rules, headers, proxies, and more. Start simple and add fields when you need them.

    For full configuration (multiple tracked elements, advanced auth, templates, etc.) use [Create a new page](#operation/createPage) instead.

    ## Authentication

    All API requests require authentication via a Bearer token or query parameter.

    **Bearer token (recommended):**
    ```
    Authorization: Bearer your_api_token
    ```

    **Query parameter:**
    ```
    GET /api/pages?api_token=your_api_token
    ```

    API tokens can be created in Settings > API. Do not expose tokens publicly.

    ## Workspace Selection

    If your account has multiple workspaces, use the `workspace_id` query parameter:
    ```
    GET /api/pages?workspace_id=123
    ```

    ## Rate Limiting

    Limits are applied per endpoint and per account, not as one global quota. Listing pages, reading history and listing recordings (`/api/pages`, `/api/pages/:id/history`, `/api/pages/:id/recordings`) allow 300 requests per minute on the Free plan and 1,200 per minute on paid plans. Diff images allow 60 per minute (Free) and 600 per minute (paid), exports 10 per hour (Free) and 60 per hour (paid), and pushing values into a data source 60 per hour (Free) and 1,800 per hour (paid). Screenshots, the timeline, websites, the summary and AI summaries allow 600 requests per minute. A few actions have their own smaller limits (webhook tests, on-demand checks per page). Exceeding a limit returns HTTP 429 with a `Retry-After` header.

    ## Validation Errors

    Failed validation returns HTTP 422 with error details. If you exceed your page limit, new pages are saved as disabled.

    With new page track creation, the Page object will be returned with a 201 HTTP status code.
    When a page is updated, the Page object will be returned with a 200 HTTP status code.

    ## Common Workflows

    **Webhooks (most effective):** If your objective is to store data related to detected changes, use the webhook functionality. When a change is detected, a webhook will be immediately sent to your server. Webhooks can be configured in Settings > API > Webhooks.

    **API Polling (simple):** Configure pages in the UI and poll `/api/pages?simple=1` to retrieve all tracked pages with their checks.

    **API Polling (advanced):** Retrieve the list of pages via `/api/pages`, then for each page call `/api/pages/:id/history?simple=1` to get check data.

    **Push data (data sources):** Not every value you want to track lives on a web page. Create a data source with [Create a data source](#operation/createDataSource), then push values into it with a one-liner against its secret `ingest_url` (`curl -d value=42 INGEST_URL`, no Authorization header, see [Push a value](#operation/ingestValueByUrl)), with the [authenticated ingest endpoint](#operation/ingestValue), or by emailing the value to the data source's private `ingest_email` address. Pushed values get the same change comparison, history, charts, notifications, and AI summaries as monitored pages.

    **Reference implementations:** Copy-paste polling, webhook verification, and hybrid push-plus-poll examples in Python, Node.js, and PHP are in the [Reference Implementations](#section/Reference-Implementations) section below.

    ## Reference Implementations

    This guide shows three ways to connect your own application to PageCrawl and provides working code for each in Python, Node.js, and PHP. These are the same patterns the official Home Assistant integration uses, distilled into minimal examples you can adapt.

    Pick the pattern that fits your needs:

    - **Polling** is the simplest. You read the API on a timer. Best for dashboards and reports that do not need instant updates.
    - **Webhooks (push)** deliver changes to your server the moment they happen. Best for real-time automation and alerting.
    - **Hybrid** combines a webhook for instant updates with a slow reconcile poll that catches anything missed. This is the most robust option and what the Home Assistant integration runs.

    ### Authentication

    All API requests use a bearer token in the `Authorization` header:

    ```
    Authorization: Bearer YOUR_TOKEN
    ```

    You can use an API token (Settings > API) or an OAuth access token. Free accounts can use the API. Treat the token like a password and keep it server-side.

    ### Rate Limits

    - Free accounts: 60 requests per minute.
    - Paid accounts: 300 requests per minute.

    When you exceed the limit the API responds with HTTP `429`. Honor the `Retry-After` response header (seconds to wait) before retrying. Choose a poll interval that stays well under your limit, especially if you paginate across many monitors.

    ### Polling

    Poll `GET /api/pages?simple=1` on an interval. Each page object includes a `latest` snapshot and a `checks` array. Read `latest.contents` for the primary tracked element, and read per-element values from `checks[0].elements`, keyed by `element_id` so each value maps to a stable tracked element in your own system. Use pagination if your workspace returns multiple pages of results.

    **Python**

    ```python
    import time
    import requests

    BASE = "https://pagecrawl.io"
    TOKEN = "YOUR_TOKEN"
    SESSION = requests.Session()
    SESSION.headers["Authorization"] = f"Bearer {TOKEN}"


    def fetch_pages():
        """Fetch all monitors, following pagination and honoring 429."""
        pages, url = [], f"{BASE}/api/pages?simple=1"
        while url:
            resp = SESSION.get(url, timeout=30)
            if resp.status_code == 429:
                wait = int(resp.headers.get("Retry-After", "5"))
                time.sleep(wait)
                continue
            resp.raise_for_status()
            body = resp.json()
            pages.extend(body.get("data", []))
            url = body.get("links", {}).get("next")
        return pages


    def poll_once():
        for page in fetch_pages():
            latest = page.get("latest") or {}
            print(page["id"], page.get("title"), "->", latest.get("contents"))

            checks = page.get("checks") or []
            elements = checks[0].get("elements", []) if checks else []
            for el in elements:
                # element_id is stable across every check; use it as your key.
                print("  ", el.get("element_id"), el.get("label"), el.get("contents"))


    if __name__ == "__main__":
        while True:
            poll_once()
            time.sleep(300)  # stay well under the rate limit
    ```

    **Node.js**

    ```js
    const BASE = "https://pagecrawl.io";
    const TOKEN = "YOUR_TOKEN";
    const HEADERS = { Authorization: `Bearer ${TOKEN}` };

    const sleep = (ms) => new Promise((r) => setTimeout(r, ms));

    async function fetchPages() {
      const pages = [];
      let url = `${BASE}/api/pages?simple=1`;
      while (url) {
        const resp = await fetch(url, { headers: HEADERS });
        if (resp.status === 429) {
          const wait = parseInt(resp.headers.get("Retry-After") || "5", 10);
          await sleep(wait * 1000);
          continue;
        }
        if (!resp.ok) throw new Error(`HTTP ${resp.status}`);
        const body = await resp.json();
        pages.push(...(body.data || []));
        url = body.links?.next || null;
      }
      return pages;
    }

    async function pollOnce() {
      for (const page of await fetchPages()) {
        const latest = page.latest || {};
        console.log(page.id, page.title, "->", latest.contents);

        const elements = page.checks?.[0]?.elements || [];
        for (const el of elements) {
          // element_id is stable across every check; use it as your key.
          console.log("  ", el.element_id, el.label, el.contents);
        }
      }
    }

    async function main() {
      while (true) {
        await pollOnce();
        await sleep(300_000); // stay well under the rate limit
      }
    }

    main();
    ```

    ### Webhooks (Push)

    Create a hook so PageCrawl POSTs to your server the instant a change is detected, then verify every delivery.

    **1. Create the hook**

    ```
    POST /api/hooks
    Authorization: Bearer YOUR_TOKEN
    Content-Type: application/json

    {
      "target_url": "https://your-server.example.com/pagecrawl",
      "match_type": "all",
      "event_type": "change_detected"
    }
    ```

    The response includes a `signing_secret`. Store it securely. You will use it to verify deliveries. (You can also create hooks in the UI under Settings > API > Webhooks.)

    **2. Verify each delivery**

    Every webhook includes two headers:

    - `X-PageCrawl-Signature: sha256=<hmac>`
    - `X-PageCrawl-Timestamp: <unix>`

    The HMAC is `HMAC_SHA256(signing_secret, "{timestamp}.{body}")` where `{body}` is the exact raw request body. Compute the same value, compare it in constant time, and reject deliveries whose timestamp is too old (to prevent replay). Always verify against the raw bytes, not a re-serialized object.

    **Python**

    ```python
    import hashlib
    import hmac
    import time

    MAX_AGE = 300  # seconds


    def verify_signature(secret: str, timestamp: str, raw_body: bytes, header: str) -> bool:
        if not secret or not timestamp or not header:
            return False
        try:
            ts = int(timestamp)
        except (TypeError, ValueError):
            return False
        if abs(time.time() - ts) > MAX_AGE:
            return False  # stale, possible replay

        expected = hmac.new(
            secret.encode("utf-8"),
            f"{timestamp}.".encode("utf-8") + raw_body,
            hashlib.sha256,
        ).hexdigest()

        provided = header[len("sha256="):] if header.startswith("sha256=") else header
        return hmac.compare_digest(expected, provided)
    ```

    A minimal Flask receiver:

    ```python
    from flask import Flask, request, abort

    app = Flask(__name__)
    SIGNING_SECRET = "YOUR_SIGNING_SECRET"


    @app.post("/pagecrawl")
    def receive():
        sig = request.headers.get("X-PageCrawl-Signature")
        ts = request.headers.get("X-PageCrawl-Timestamp")
        if not verify_signature(SIGNING_SECRET, ts, request.get_data(), sig):
            abort(401)
        payload = request.get_json()
        print("change on", payload.get("id"), payload.get("short_summary"))
        return "", 204
    ```

    **Node.js**

    ```js
    const crypto = require("crypto");
    const express = require("express");

    const SIGNING_SECRET = "YOUR_SIGNING_SECRET";
    const MAX_AGE = 300; // seconds

    function verifySignature(secret, timestamp, rawBody, header) {
      if (!secret || !timestamp || !header) return false;
      const ts = parseInt(timestamp, 10);
      if (Number.isNaN(ts)) return false;
      if (Math.abs(Date.now() / 1000 - ts) > MAX_AGE) return false; // stale

      const expected = crypto
        .createHmac("sha256", secret)
        .update(`${timestamp}.${rawBody}`)
        .digest("hex");

      const provided = header.startsWith("sha256=") ? header.slice(7) : header;
      const a = Buffer.from(expected);
      const b = Buffer.from(provided);
      return a.length === b.length && crypto.timingSafeEqual(a, b);
    }

    const app = express();
    // Capture the raw body exactly as received so the HMAC matches.
    app.use(express.raw({ type: "*/*" }));

    app.post("/pagecrawl", (req, res) => {
      const sig = req.get("X-PageCrawl-Signature");
      const ts = req.get("X-PageCrawl-Timestamp");
      const raw = req.body.toString("utf8");
      if (!verifySignature(SIGNING_SECRET, ts, raw, sig)) {
        return res.sendStatus(401);
      }
      const payload = JSON.parse(raw);
      console.log("change on", payload.id, payload.short_summary);
      res.sendStatus(204);
    });

    app.listen(8080);
    ```

    **PHP**

    ```php
    <?php

    function verify_signature(string $secret, ?string $timestamp, string $rawBody, ?string $header): bool
    {
        $maxAge = 300; // seconds
        if ($secret === '' || $timestamp === null || $header === null) {
            return false;
        }
        if (! ctype_digit($timestamp)) {
            return false;
        }
        if (abs(time() - (int) $timestamp) > $maxAge) {
            return false; // stale, possible replay
        }

        $expected = hash_hmac('sha256', "{$timestamp}.{$rawBody}", $secret);
        $provided = str_starts_with($header, 'sha256=') ? substr($header, 7) : $header;

        return hash_equals($expected, $provided);
    }

    $signingSecret = 'YOUR_SIGNING_SECRET';
    $rawBody = file_get_contents('php://input');
    $sig = $_SERVER['HTTP_X_PAGECRAWL_SIGNATURE'] ?? null;
    $ts = $_SERVER['HTTP_X_PAGECRAWL_TIMESTAMP'] ?? null;

    if (! verify_signature($signingSecret, $ts, $rawBody, $sig)) {
        http_response_code(401);
        exit;
    }

    $payload = json_decode($rawBody, true);
    error_log('change on '.$payload['id'].' '.($payload['short_summary'] ?? ''));
    http_response_code(204);
    ```

    ### Hybrid (Push Plus Reconcile)

    The most robust integration uses a webhook for instant updates and a slow background poll that reconciles state. The webhook keeps you current in real time. The reconcile poll catches anything a webhook might miss (for example if your server was briefly offline) and refreshes monitors that did not change. This is the model the Home Assistant integration runs: push updates the in-memory snapshot, and a slow loop re-fetches the full list on a long interval.

    **Python (sketch)**

    ```python
    import threading
    import time

    state = {}  # element_id -> latest value, shared between push and poll
    lock = threading.Lock()


    def on_webhook(payload):
        """Called from your verified webhook receiver. Instant update."""
        with lock:
            for el in payload.get("page_elements", []):
                state[el["element_id"]] = el.get("contents")


    def reconcile_loop():
        """Slow safety net. Re-reads everything on a long interval."""
        while True:
            for page in fetch_pages():  # from the polling example above
                checks = page.get("checks") or []
                for el in (checks[0].get("elements", []) if checks else []):
                    with lock:
                        state[el["element_id"]] = el.get("contents")
            time.sleep(3600)  # reconcile hourly; the webhook handles real time


    threading.Thread(target=reconcile_loop, daemon=True).start()
    ```

    **Node.js (sketch)**

    ```js
    const state = new Map(); // element_id -> latest value

    function onWebhook(payload) {
      // Called from your verified webhook receiver. Instant update.
      for (const el of payload.page_elements || []) {
        state.set(el.element_id, el.contents);
      }
    }

    async function reconcileLoop() {
      // Slow safety net. Re-reads everything on a long interval.
      while (true) {
        for (const page of await fetchPages()) {
          // fetchPages from the polling example
          for (const el of page.checks?.[0]?.elements || []) {
            state.set(el.element_id, el.contents);
          }
        }
        await new Promise((r) => setTimeout(r, 3_600_000)); // reconcile hourly
      }
    }

    reconcileLoop();
    ```

    Keep the reconcile interval long (hourly or slower) so the webhook does the real-time work and the poll stays comfortably within your rate limit.

    ## Primary Tracked Element

    The "Primary Tracked Element" is the first element you track on a page (the first item in the `elements` array). Several fields in the API refer to it:

    - `latest.contents`, `latest.difference`, `latest.human_difference` (on Page)
    - `history` (on Check)

    If you only track one element per page, you can simplify your code by reading these primary fields directly.
    If you track multiple elements per page, you can ignore mentions of "primary" and use the per-element data in `elements` instead.

    ## Need additional API endpoints?

    If you need API endpoints that are not documented here, please contact support. While undocumented internal API endpoints may be available, they may break or change in future updates without notice.

    ## MCP Server (AI Integrations)

    PageCrawl provides a Model Context Protocol (MCP) server for integration with AI assistants like Claude, Cursor, and Windsurf. The MCP server allows AI tools to interact with your monitors programmatically.

    **Available MCP tools:**

    | Tool | Description |
    |------|-------------|
    | `list-workspaces` | List all accessible workspaces |
    | `add-page-monitor` | Add a new page monitor with simplified tracking modes |
    | `list-monitors` | List and search monitors across all workspaces |
    | `get-monitor-details` | Get detailed configuration for a specific monitor |
    | `get-latest-values` | Get latest tracked values for a monitor |
    | `get-monitor-history` | Retrieve historical monitoring data and detected changes |
    | `get-check-diff` | View formatted text diffs showing what changed in a check |
    | `trigger-check` | Trigger an immediate check on a monitor |
    | `manage-tags` | List workspace tags, add or remove tags from monitors |
    | `mark-changes-seen` | Mark detected changes as seen/reviewed |
    | `list-templates` | List templates available in a workspace |
    | `update-monitor-defaults` | View and update default monitor settings per workspace |

    **Setup with Claude (Web/Desktop):** Go to Settings > Connectors > Add custom connector. Set the URL to `https://mcp.pagecrawl.io/mcp` and authorize access.

    **Setup with Claude Code:** Add to your `.mcp.json` file:
    ```json
    { "mcpServers": { "pagecrawl": { "url": "https://mcp.pagecrawl.io/mcp" } } }
    ```

    **Setup with ChatGPT:** Go to Settings > Apps & Connectors > Create. Enter `https://mcp.pagecrawl.io/mcp` as the MCP server URL, set Authentication to OAuth, and set the server transport to Streamable HTTP (not SSE). Requires a Plus, Pro, Team, Enterprise, or Edu plan.

    For full setup instructions, see the [MCP Server guide](https://pagecrawl.io/help/integrations/article/mcp-server-ai-tools).
  version: '1.0'
  contact:
    name: PageCrawl Support
    url: https://pagecrawl.io
    email: support@pagecrawl.io

servers:
  - url: https://pagecrawl.io/api
    description: Production

security:
  - bearerAuth: []
  - apiTokenQuery: []

tags:
  - name: Pages
    description: Create, read, update, and delete monitored pages
  - name: History
    description: Retrieve check history and text diffs for monitored pages
  - name: Screenshots
    description: Retrieve full-page screenshots and visual diffs
  - name: Recordings
    description: |
      Screenshots and web archives of checks that found no change. Available on
      the **Ultimate plan**.

      PageCrawl normally keeps a check's screenshot and web archive only when it
      detects a change. With **Record every check** it also keeps them for the
      checks in between, so you have a complete record of the page over time.
      Checks that detected a change are in the [check history](#operation/getPageHistory);
      recordings are the checks in between.

      Turn it on for every page of a workspace in Settings > Workspace, or per
      page with `every_check_screenshots` and `every_check_archives` when you
      [create](#operation/createPage) or [update](#operation/updatePage) a page.
      Per-page values apply only while the workspace setting for that medium is
      "Selected monitors and templates" or "All monitors"; while it is Off,
      nothing is recorded, whatever the page says.
      Web archives of every check need a check frequency of 6 hours or slower.

      Recordings use your archive storage (500 GB per Ultimate unit). When it is
      full, recording pauses and nothing is deleted.
  - name: Review Boards
    description: |
      Move pages between Review Boards (To Review / Reviewed / custom).

      A **review card** is created on the "To Review" board the first time a
      change is detected on a page. Any further changes detected on that page
      before it is marked reviewed are bundled into the same card (the card's
      check range expands to cover them). Once you move the card to "Reviewed",
      the next detected change starts a fresh card.

      > **Limited availability.** The full Reviews API (listing boards, fetching
      > review cards, reading comments, retrieving review history) is not yet
      > publicly exposed. Only the move endpoints below are currently documented.
      > If you need to read reviews programmatically, contact support@pagecrawl.io
      > and we will scope it with you.
  - name: Data Sources
    description: |
      Push values into PageCrawl instead of having PageCrawl fetch a page.
      Every data source gets a secret ingest URL, so pushing a value is a
      one-liner with no Authorization header:

      ```bash
      curl -d value=42 https://pagecrawl.io/api/ingest/data-a1b2c3d4e5f6g7h8i9j0k1l2
      ```

      The long random token in the URL is the credential, so treat the URL
      like a password. Pushed values get the same change comparison, history,
      charts, multi-channel notifications, and AI summaries as monitored
      pages. Available on every plan, including Free.

      Create a data source, then send values via the secret ingest URL, the
      authenticated ingest endpoint, or by email. The secret URL is returned
      as `ingest_url` when the data source is created and shown in the data
      source's "Send data" panel in the app.

      **Email ingestion.** Every data source has a private inbound email
      address, returned as `ingest_email` on the created object and shown in
      the app. The address local part is the same secret as the ingest URL,
      e.g. `data-a1b2c3d4e5f6g7h8i9j0k1l2@ingest.pagecrawl.io`. Email a value
      and the message body becomes the stored value
      (if the body is empty, the subject is used instead). Lines like
      `price: 9.99` that match existing field names map to those fields;
      otherwise the whole body is stored as the single value. One-liner:

      ```bash
      echo "42" | mail data-a1b2c3d4e5f6g7h8i9j0k1l2@ingest.pagecrawl.io
      ```

      The address
      itself is the credential, so keep it secret. An optional sender filter
      (an exact address or an exact domain) can restrict who may email values
      in. Useful for cron jobs, no-code tools, and systems that can only
      send email.

paths:
  /track-simple:
    post:
      tags: [Pages]
      summary: Quick track a page
      description: |
        Simplified page creation for most common use cases. You only need to provide a URL and optionally a CSS/XPath selector.

        **Example:**
        ```bash
        curl -X POST "https://pagecrawl.io/api/track-simple?url=https://example.com/&frequency=60" \
          -H "Authorization: Bearer YOUR_API_KEY"
        ```
      operationId: trackSimple
      requestBody:
        required: true
        content:
          application/json:
            schema:
              $ref: '#/components/schemas/TrackSimpleRequest'
      responses:
        '201':
          description: Page created
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/Page'
        '200':
          description: Duplicate page found (when ignore_duplicates is true)
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/Page'
        '422':
          $ref: '#/components/responses/ValidationError'

  /pages:
    get:
      tags: [Pages]
      summary: List all pages
      description: Retrieves all monitored pages in the current workspace with their latest check data.
      operationId: listPages
      parameters:
        - name: simple
          in: query
          description: Return simplified response without configuration options.
          schema:
            type: boolean
        - name: take
          in: query
          description: |
            Limit number of latest checks returned per page (default 10, max 100).
            For example, set `take=2` to retrieve only the last 2 checks if you have a lot of historic data.
          schema:
            type: integer
            minimum: 1
            maximum: 100
            default: 10
        - name: folder
          in: query
          description: |
            Filter by folder. By default returns pages from the main view.
            - `*` - Return all pages across all folders
            - `/app/pages/folders/myfolder` - Return pages in a specific folder (use the pathname from the UI)
          schema:
            type: string
        - name: workspace_id
          in: query
          description: Target a specific workspace (required when account has multiple workspaces).
          schema:
            type: integer
        - name: filters[tags][]
          in: query
          description: |
            Only return pages carrying these labels. By default a page must
            carry ALL listed labels; set `filters[tags_mode]=or` to match ANY.
          schema:
            type: array
            items:
              type: string
        - name: filters[tags_mode]
          in: query
          description: How multiple `filters[tags][]` values combine, `and` (default) or `or`.
          schema:
            type: string
            enum: [and, or]
        - name: filters[exclude_tags][]
          in: query
          description: |
            Hide pages carrying any of these labels. Combines with the other
            label filters, e.g. `filters[tags][]=promo&filters[exclude_tags][]=archived`
            returns pages labelled promo that are not also labelled archived.
          schema:
            type: array
            items:
              type: string
        - name: filters[tag_rules]
          in: query
          description: |
            Advanced label rules as a JSON array of `{op, labels}` rows that
            combine with AND. `op` is one of `any` (has at least one), `all`
            (has every one) or `none` (has none of them). Example, "(promo OR
            sale) AND NOT archived":
            `[{"op":"any","labels":["promo","sale"]},{"op":"none","labels":["archived"]}]`.
            Max 10 rules, 100 labels per rule. Invalid rules return a 422.
          schema:
            type: string
      responses:
        '200':
          description: List of pages
          content:
            application/json:
              schema:
                type: array
                items:
                  $ref: '#/components/schemas/Page'
        '401':
          $ref: '#/components/responses/Unauthorized'
        '422':
          $ref: '#/components/responses/ValidationError'
        '429':
          $ref: '#/components/responses/RateLimited'

    post:
      tags: [Pages]
      summary: Create a new page
      description: |
        Create a new monitored page with full configuration options.
        Returns the created Page object with HTTP 201. If `ignore_duplicates` is true and a matching page exists, returns HTTP 200 with the existing page.
      operationId: createPage
      requestBody:
        required: true
        content:
          application/json:
            schema:
              $ref: '#/components/schemas/PageCreateRequest'
      responses:
        '201':
          description: Page created
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/Page'
        '200':
          description: Duplicate page found (when ignore_duplicates is true)
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/Page'
        '422':
          $ref: '#/components/responses/ValidationError'

  /pages/{id}:
    get:
      tags: [Pages]
      summary: Get page details
      description: Retrieve full configuration and latest check data for a specific page.
      operationId: getPage
      parameters:
        - $ref: '#/components/parameters/pageId'
        - name: take
          in: query
          description: Limit number of latest checks returned.
          schema:
            type: integer
        - name: simple
          in: query
          description: Return simplified response without configuration options.
          schema:
            type: boolean
      responses:
        '200':
          description: Page details
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/Page'
        '404':
          $ref: '#/components/responses/NotFound'

    put:
      tags: [Pages]
      summary: Update a page
      description: Update the configuration of an existing monitored page.
      operationId: updatePage
      parameters:
        - $ref: '#/components/parameters/pageId'
      requestBody:
        required: true
        content:
          application/json:
            schema:
              $ref: '#/components/schemas/PageUpdateRequest'
      responses:
        '200':
          description: Page updated
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/Page'
        '422':
          $ref: '#/components/responses/ValidationError'

    delete:
      tags: [Pages]
      summary: Delete a page
      description: Delete a monitored page and all its history.
      operationId: deletePage
      parameters:
        - $ref: '#/components/parameters/pageId'
      responses:
        '200':
          description: Page deleted
        '404':
          $ref: '#/components/responses/NotFound'

  /pages/{id}/history:
    get:
      tags: [History]
      summary: Get check history
      description: Retrieve the full check history for a monitored page.
      operationId: getPageHistory
      parameters:
        - $ref: '#/components/parameters/pageId'
        - name: take
          in: query
          description: Limit number of checks returned.
          schema:
            type: integer
        - name: simple
          in: query
          description: Return simplified response.
          schema:
            type: boolean
      responses:
        '200':
          description: List of checks
          content:
            application/json:
              schema:
                type: array
                items:
                  $ref: '#/components/schemas/Check'
        '404':
          $ref: '#/components/responses/NotFound'

  /pages/{id}/checks/{checkId}/diff.png:
    get:
      tags: [History]
      summary: Get text diff as image
      description: |
        Retrieve a rendered diff card as a PNG, comparing a check with its
        previous check. The card includes the page name, monitored domain,
        check time (in the workspace timezone), AI summary, matched keywords,
        and the highlighted diff, rendered according to the workspace's
        notification format settings.

        The image is always 1440 px wide (a 720 px card at 2x retina scale).
        Height depends on the diff and the `height`/`compact` parameters below;
        when content exceeds the height budget the card ends with an
        "N more lines not shown" footer instead of clipping silently.

        Responses for a specific `checkId` are cached server-side; `latest` is
        rendered on demand.
      operationId: getDiffImage
      parameters:
        - $ref: '#/components/parameters/pageId'
        - $ref: '#/components/parameters/checkId'
        - name: height
          in: query
          required: false
          description: |
            Height budget for the card, in **logical pixels**. The returned
            PNG is 2x retina, so `height=600` produces an image at most
            1200 px tall. Out-of-range values are clamped to the documented
            range rather than rejected. Overrides `compact` when both are sent.
          schema:
            type: integer
            minimum: 400
            maximum: 2600
          example: 900
        - name: compact
          in: query
          required: false
          description: |
            Shorthand for the chat-embed height budget (920 logical px, the
            same card Telegram/Slack/Discord/Teams notifications embed). Use
            it when the image will sit in a preview box that downscales by its
            longest side. Ignored when `height` is present.
          schema:
            type: boolean
            default: false
      responses:
        '200':
          description: Diff card PNG (1440 px wide, height per the parameters above)
          content:
            image/png:
              schema:
                type: string
                format: binary
        '404':
          $ref: '#/components/responses/NotFound'

  /pages/{id}/checks/{checkId}/diff.html:
    get:
      tags: [History]
      summary: Get text diff as HTML
      description: Retrieve an HTML-formatted text diff comparing a check with its previous check.
      operationId: getDiffHtml
      parameters:
        - $ref: '#/components/parameters/pageId'
        - $ref: '#/components/parameters/checkId'
      responses:
        '200':
          description: HTML diff
          content:
            text/html:
              schema:
                type: string
        '404':
          $ref: '#/components/responses/NotFound'

  /pages/{id}/checks/{checkId}/diff.markdown:
    get:
      tags: [History]
      summary: Get text diff as Markdown
      description: Retrieve a Markdown-formatted text diff comparing a check with its previous check.
      operationId: getDiffMarkdown
      parameters:
        - $ref: '#/components/parameters/pageId'
        - $ref: '#/components/parameters/checkId'
      responses:
        '200':
          description: Markdown diff
          content:
            text/markdown:
              schema:
                type: string
        '404':
          $ref: '#/components/responses/NotFound'

  /pages/{id}/checks/latest/screenshot:
    get:
      tags: [Screenshots]
      summary: Get latest screenshot
      description: |
        Retrieve the full-page screenshot from the most recent check.

        You may also pass the token via query parameter for convenience when embedding images:
        ```
        /api/pages/:id/checks/latest/screenshot?api_token=your_token
        ```
        Do not expose your token publicly.
      operationId: getLatestScreenshot
      parameters:
        - $ref: '#/components/parameters/pageId'
      responses:
        '200':
          description: Screenshot image
          content:
            image/png:
              schema:
                type: string
                format: binary
        '404':
          $ref: '#/components/responses/NotFound'

  /pages/{id}/checks/latest/diff:
    get:
      tags: [Screenshots]
      summary: Get latest visual diff
      description: Retrieve a visual screenshot diff comparing the latest check with the previous check.
      operationId: getLatestVisualDiff
      parameters:
        - $ref: '#/components/parameters/pageId'
      responses:
        '200':
          description: Visual diff image
          content:
            image/png:
              schema:
                type: string
                format: binary
        '404':
          $ref: '#/components/responses/NotFound'

  /pages/{id}/checks/{checkId}/screenshot:
    get:
      tags: [Screenshots]
      summary: Get check screenshot
      description: Retrieve the full-page screenshot for a specific check. Use "latest" as checkId for the most recent check.
      operationId: getCheckScreenshot
      parameters:
        - $ref: '#/components/parameters/pageId'
        - $ref: '#/components/parameters/checkId'
      responses:
        '200':
          description: Screenshot image
          content:
            image/png:
              schema:
                type: string
                format: binary
        '404':
          $ref: '#/components/responses/NotFound'

  /pages/{id}/checks/{checkId}/diff:
    get:
      tags: [Screenshots]
      summary: Get visual diff for check
      description: Retrieve a visual screenshot diff comparing a check with its previous check.
      operationId: getCheckVisualDiff
      parameters:
        - $ref: '#/components/parameters/pageId'
        - $ref: '#/components/parameters/checkId'
      responses:
        '200':
          description: Visual diff image
          content:
            image/png:
              schema:
                type: string
                format: binary
        '404':
          $ref: '#/components/responses/NotFound'

  /pages/{id}/recordings:
    get:
      tags: [Recordings]
      summary: List recordings
      description: |
        List the recordings of checks that found no change, newest first.
        Available on the Ultimate plan.

        Results are paged with a cursor: pass a response's `next_cursor` as
        `cursor` to get the next page. `next_cursor` is null on the last page.
      operationId: listRecordings
      parameters:
        - $ref: '#/components/parameters/pageId'
        - name: since
          in: query
          description: Only recordings made at or after this time (ISO 8601).
          schema:
            type: string
            format: date-time
          example: '2026-09-01T00:00:00Z'
        - name: until
          in: query
          description: Only recordings made before this time (ISO 8601). Must be later than `since`.
          schema:
            type: string
            format: date-time
        - name: cursor
          in: query
          description: The `next_cursor` of the previous page.
          schema:
            type: string
        - name: limit
          in: query
          description: Recordings per page.
          schema:
            type: integer
            minimum: 1
            maximum: 100
            default: 50
      responses:
        '200':
          description: One page of recordings, newest first
          content:
            application/json:
              schema:
                type: object
                properties:
                  data:
                    type: array
                    items:
                      $ref: '#/components/schemas/Recording'
                  next_cursor:
                    type: string
                    nullable: true
                    description: Pass as `cursor` to get the next page. Null on the last page.
        '404':
          $ref: '#/components/responses/NotFound'
        '422':
          $ref: '#/components/responses/ValidationError'

  /pages/{id}/recordings/{recordingId}/screenshot:
    get:
      tags: [Recordings]
      summary: Get recording screenshot
      description: |
        The full-page screenshot of a recording. Available on the Ultimate plan.

        Returns a placeholder image when the recording has no screenshot or no
        longer exists, or when screenshots are turned off for your team.

        You may pass the token as a query parameter when embedding the image:
        ```
        /api/pages/:id/recordings/:recordingId/screenshot?api_token=your_token
        ```
        Do not expose your token publicly.
      operationId: getRecordingScreenshot
      parameters:
        - $ref: '#/components/parameters/pageId'
        - $ref: '#/components/parameters/recordingId'
      responses:
        '200':
          description: Screenshot image (JPEG), or a placeholder image (PNG)
          content:
            image/jpeg:
              schema:
                type: string
                format: binary
            image/png:
              schema:
                type: string
                format: binary
        '404':
          $ref: '#/components/responses/NotFound'

  /pages/{id}/recordings/{recordingId}/archive:
    get:
      tags: [Recordings]
      summary: Download recording web archive
      description: |
        Download a recording's web archive. Available on the Ultimate plan.

        A `wacz` archive opens in any WACZ viewer, such as ReplayWeb.page, where
        you can browse the page as it was at the time of the check. An `html`
        archive is a single HTML file. The recording's `archive_type` says which
        one you get.
      operationId: getRecordingArchive
      parameters:
        - $ref: '#/components/parameters/pageId'
        - $ref: '#/components/parameters/recordingId'
      responses:
        '200':
          description: The web archive
          content:
            application/wacz+zip:
              schema:
                type: string
                format: binary
            text/html:
              schema:
                type: string
        '402':
          $ref: '#/components/responses/PaymentRequired'
        '404':
          $ref: '#/components/responses/NotFound'

  /pages/{id}/recordings/{recordingId}/archive/warc:
    get:
      tags: [Recordings]
      summary: Download recording WARC
      description: |
        Download a recording's capture as a WARC file, the standard format of
        web archives. Available on the Ultimate plan, for recordings whose
        `archive_type` is `wacz`.
      operationId: getRecordingWarc
      parameters:
        - $ref: '#/components/parameters/pageId'
        - $ref: '#/components/parameters/recordingId'
      responses:
        '200':
          description: The WARC file
          content:
            application/warc:
              schema:
                type: string
                format: binary
        '402':
          $ref: '#/components/responses/PaymentRequired'
        '404':
          $ref: '#/components/responses/NotFound'

  /pages/move-to-board:
    post:
      tags: [Review Boards]
      summary: Move multiple pages to a review board
      description: |
        For each page in `page_ids`, moves its latest review card to `target_board` and
        marks every unseen update detected on that page as reviewed (zeroes the unseen
        counter and stamps `seen` on all outstanding checks).
        Idempotent: pages whose latest review is already on the target board are silently
        re-stamped (no error). Unresolved entries in `page_ids` (unknown IDs, unknown
        slugs, or items from another workspace) are silently skipped; `moved_count`
        reflects the actual number moved.
      operationId: bulkMoveReviews
      requestBody:
        required: true
        content:
          application/json:
            schema:
              type: object
              required: [page_ids, target_board]
              example:
                page_ids: [123, "my-page-slug"]
                target_board: "reviewed"
              properties:
                page_ids:
                  type: array
                  description: |
                    List of page identifiers. Each entry may be either:
                    - a numeric **page ID** (e.g. `123`) — the internal `id`
                      returned by `GET /pages`
                    - a string **page slug** (e.g. `"pagecrawl-pricing-page-a1b2"`)
                      — the URL-safe identifier from the page's in-app URL
                      `https://pagecrawl.io/app/pages/{slug}`, also returned as
                      `slug` by `GET /pages`

                    Numeric strings (e.g. `"123"`) are always treated as IDs, never slugs.
                  items:
                    oneOf:
                      - { type: integer, description: "Page ID (numeric, from GET /pages `id`)" }
                      - { type: string, description: "Page slug (URL-safe identifier, from GET /pages `slug`)" }
                target_board:
                  type: string
                  description: |
                    Target review board. Accepts:
                    - board label (e.g. `"Reviewed"`, `"To Review"`, or any custom
                      board name you've created in the app)
                    - system key (`"to_review"` or `"reviewed"`) — recommended for
                      the built-in boards because it survives label renames
      responses:
        '200':
          description: Move result
          content:
            application/json:
              schema:
                type: object
                properties:
                  success: { type: boolean, example: true }
                  moved_count:
                    type: integer
                    description: Number of pages actually moved (excludes unresolved entries).
        '400':
          description: Missing required fields or unresolvable target board.
          content:
            application/json:
              schema:
                type: object
                properties:
                  error: { type: string }
        '401':
          $ref: '#/components/responses/Unauthorized'

  /data-sources:
    post:
      tags: [Data Sources]
      summary: Create a data source
      description: |
        Create a push-based monitor that receives values instead of fetching a URL.

        Leave `fields` out to track a single value, or define named fields to
        track several values (for example price and stock) on one data source.
        The response is the monitor object, including `id`, `slug`, `elements`,
        `ingest_url` (the secret capability URL you can POST values to with no
        Authorization header; the URL is the credential, treat it like a
        password), and `ingest_email`, the data source's private inbound email
        address (see the Data Sources tag description for how email ingestion
        works). Pushed values get the same change comparison, history, charts,
        notifications, and AI summaries as monitored pages.

        **Example:**
        ```bash
        curl -X POST "https://pagecrawl.io/api/data-sources" \
          -H "Authorization: Bearer YOUR_API_TOKEN" \
          -H "Content-Type: application/json" \
          -d '{
            "name": "Competitor API price",
            "fields": [
              { "label": "price", "type": "number" },
              { "label": "stock", "type": "number" }
            ]
          }'
        ```
      operationId: createDataSource
      requestBody:
        required: true
        content:
          application/json:
            schema:
              $ref: '#/components/schemas/DataSourceCreateRequest'
      responses:
        '201':
          description: |
            Data source created. Returns the monitor object with `check_type`
            set to `push` and both `ingest_url` (the secret capability URL,
            no auth needed, treat it like a password) and `ingest_email`
            populated. The same two fields are returned when fetching the
            single page; the pages list does not include them.
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/Page'
        '401':
          $ref: '#/components/responses/Unauthorized'
        '403':
          description: Plan limit reached (total pages cap, or the Free plan API active-monitor cap).
        '422':
          $ref: '#/components/responses/ValidationError'

  /ingest/{token}:
    post:
      tags: [Data Sources]
      summary: Push a value
      description: |
        Push one or more values into a data source using its secret ingest
        URL. No Authorization header is needed: the long random token in the
        URL is the credential, so treat the whole URL like a password. Copy
        it from the `ingest_url` field returned when the data source is
        created (also shown in the data source's "Send data" panel in the
        app).

        The endpoint accepts JSON bodies and form-encoded bodies, so pushing
        a value is a true one-liner:

        ```bash
        curl -d value=42 https://pagecrawl.io/api/ingest/data-a1b2c3d4e5f6g7h8i9j0k1l2
        ```

        **More examples:**
        ```bash
        # Multiple named fields (form-encoded)
        curl -d 'values[price]=9.99' -d 'values[stock]=12' \
          https://pagecrawl.io/api/ingest/data-a1b2c3d4e5f6g7h8i9j0k1l2

        # JSON variant
        curl -H "Content-Type: application/json" \
          -d '{ "values": { "price": 9.99, "stock": 12 } }' \
          https://pagecrawl.io/api/ingest/data-a1b2c3d4e5f6g7h8i9j0k1l2
        ```

        Send `value` for a single-value data source, or `values` (a map of
        field name to value) for multi-field data sources. Field names in
        `values` that do not exist yet are created automatically:
        numeric-looking values become number fields, everything else becomes
        a text field.

        Pushing the same value again returns `"changed": false` and stores
        nothing new, so it is safe to push on a schedule without flooding your
        history. When a value does change, notifications go out on the data
        source's configured channels shortly after the value is received.

        The same secret is also the local part of the data source's inbound
        email address, so `echo "42" | mail data-a1b2c3d4e5f6g7h8i9j0k1l2@ingest.pagecrawl.io`
        records a value too (see the Data Sources tag description).

        For workspace-scoped scripts that already hold an API token, use the
        authenticated alternative,
        [Push a value (authenticated)](#operation/ingestValue).
      operationId: ingestValueByUrl
      security: []
      parameters:
        - name: token
          in: path
          required: true
          description: |
            The secret from the data source's `ingest_url` (everything after
            `/api/ingest/`). Always `data-` followed by 24 characters, e.g.
            `data-a1b2c3d4e5f6g7h8i9j0k1l2`. This token is the credential:
            anyone who has it can push values into the data source, so keep
            it out of client-side code and public repositories.
          schema:
            type: string
            pattern: '^data-[a-z0-9]{24}$'
          example: data-a1b2c3d4e5f6g7h8i9j0k1l2
      requestBody:
        required: true
        content:
          application/json:
            schema:
              $ref: '#/components/schemas/IngestRequest'
            examples:
              singleValue:
                summary: Push a single value
                value:
                  value: 42
              namedFields:
                summary: Push multiple named fields
                value:
                  values:
                    price: 9.99
                    stock: 12
          application/x-www-form-urlencoded:
            schema:
              $ref: '#/components/schemas/IngestRequest'
            examples:
              singleValue:
                summary: Push a single value (curl -d value=42)
                value:
                  value: 42
      responses:
        '200':
          description: Value received.
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/IngestResponse'
        '403':
          description: Data source is disabled.
        '404':
          description: Unknown ingest URL. The token does not match any data source.
        '422':
          $ref: '#/components/responses/ValidationError'
        '429':
          $ref: '#/components/responses/RateLimited'
    get:
      tags: [Data Sources]
      summary: Read current values
      description: |
        Read the current value of each field on a data source using the same
        secret ingest URL you push to. No Authorization header is needed. This
        is symmetric with the push endpoint, so you can round-trip data with a
        single URL and no client library:

        ```bash
        curl https://pagecrawl.io/api/ingest/data-a1b2c3d4e5f6g7h8i9j0k1l2
        ```

        Add `?history=<n>` (up to 100) to also return the most recent points,
        newest first. Reading does not count against your check allowance.
        Because the token both reads and writes, treat the URL as a full
        credential. Responses are returned with `Cache-Control: no-store`.
      operationId: readValuesByUrl
      security: []
      parameters:
        - name: token
          in: path
          required: true
          description: The secret from the data source's `ingest_url` (everything after `/api/ingest/`).
          schema:
            type: string
            pattern: '^data-[a-z0-9]{24}$'
          example: data-a1b2c3d4e5f6g7h8i9j0k1l2
        - name: history
          in: query
          required: false
          description: Number of recent points to include (newest first). 0 (default) returns current values only.
          schema:
            type: integer
            minimum: 0
            maximum: 100
          example: 20
      responses:
        '200':
          description: Current values for the data source.
          content:
            application/json:
              schema:
                type: object
                properties:
                  monitor_id:
                    type: integer
                  name:
                    type: string
                  last_received_at:
                    type: string
                    format: date-time
                    nullable: true
                    description: When the most recent value was recorded, or null before the first push.
                  fields:
                    type: array
                    items:
                      type: object
                      properties:
                        label: { type: string }
                        type: { type: string }
                        value: { type: string, nullable: true }
                        changed: { type: boolean }
                  values:
                    type: object
                    additionalProperties: { type: string, nullable: true }
                    description: Flat label to current value map.
                  history:
                    type: array
                    description: Present only when history is requested.
                    items:
                      type: object
                      properties:
                        recorded_at: { type: string, format: date-time }
                        values:
                          type: object
                          additionalProperties: { type: string, nullable: true }
              example:
                monitor_id: 55
                name: Competitor API price
                last_received_at: '2026-07-27T09:14:00+00:00'
                fields:
                  - { label: price, type: number, value: '9.99', changed: true }
                  - { label: stock, type: number, value: '12', changed: false }
                values: { price: '9.99', stock: '12' }
        '404':
          description: Unknown ingest URL. The token does not match any data source.
        '422':
          $ref: '#/components/responses/ValidationError'
        '429':
          $ref: '#/components/responses/RateLimited'

  /ingest/{id}:
    post:
      tags: [Data Sources]
      summary: Push a value (authenticated)
      description: |
        Authenticated alternative to the secret ingest URL, for
        workspace-scoped scripts that already hold an API token. Pass the
        data source's slug (or numeric ID) in the path and your token in the
        `Authorization` header. If you just want the simplest possible push,
        use [Push a value](#operation/ingestValueByUrl) with the data
        source's `ingest_url` instead.

        Send `value` for a
        single-value data source, or `values` (a map of field name to value)
        for multi-field data sources. Field names in `values` that do not
        exist yet are created automatically: numeric-looking values become
        number fields, everything else becomes a text field.

        Pushing the same value again returns `"changed": false` and stores
        nothing new, so it is safe to push on a schedule without flooding your
        history. When a value does change, notifications go out on the data
        source's configured channels shortly after the value is received.

        **Examples:**
        ```bash
        # Single value
        curl -X POST "https://pagecrawl.io/api/ingest/my-kpi-source" \
          -H "Authorization: Bearer YOUR_API_TOKEN" \
          -H "Content-Type: application/json" \
          -d '{ "value": 42 }'

        # Multiple named fields
        curl -X POST "https://pagecrawl.io/api/ingest/my-kpi-source" \
          -H "Authorization: Bearer YOUR_API_TOKEN" \
          -H "Content-Type: application/json" \
          -d '{ "values": { "price": 9.99, "stock": 12 } }'
        ```
      operationId: ingestValue
      parameters:
        - $ref: '#/components/parameters/dataSourceId'
      requestBody:
        required: true
        content:
          application/json:
            schema:
              $ref: '#/components/schemas/IngestRequest'
            examples:
              singleValue:
                summary: Push a single value
                value:
                  value: 42
              namedFields:
                summary: Push multiple named fields
                value:
                  values:
                    price: 9.99
                    stock: 12
      responses:
        '200':
          description: Value received.
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/IngestResponse'
        '401':
          $ref: '#/components/responses/Unauthorized'
        '403':
          description: Data source is disabled.
        '404':
          $ref: '#/components/responses/NotFound'
        '422':
          $ref: '#/components/responses/ValidationError'
        '429':
          $ref: '#/components/responses/RateLimited'

components:
  securitySchemes:
    bearerAuth:
      type: http
      scheme: bearer
      description: |
        API token from Settings > API. Pass as `Authorization: Bearer <token>`.
    apiTokenQuery:
      type: apiKey
      in: query
      name: api_token
      description: |
        API token passed as query parameter. Useful for screenshot URLs.
        Do not expose publicly.

  parameters:
    pageId:
      name: id
      in: path
      required: true
      description: |
        Page ID or slug.
        - **ID** — numeric, e.g. `123`. Returned as `id` by `GET /pages`.
        - **Slug** — URL-safe string, e.g. `pagecrawl-pricing-page-a1b2`. Returned
          as `slug` by `GET /pages`; matches the in-app URL `https://pagecrawl.io/app/pages/{slug}`.
      schema:
        oneOf:
          - type: integer
          - type: string
    dataSourceId:
      name: id
      in: path
      required: true
      description: |
        Data source slug or ID. Prefer the slug in scripts and examples.
        - **Slug** - URL-safe string, e.g. `my-kpi-source`. Returned
          as `slug` on the created data source; matches the in-app URL.
        - **ID** - numeric. Returned as `id` when the data source is created;
          also accepted.
      schema:
        oneOf:
          - type: integer
          - type: string
    checkId:
      name: checkId
      in: path
      required: true
      description: Check ID (integer) or "latest" for the most recent check.
      schema:
        oneOf:
          - type: integer
          - type: string
            enum: [latest]
    recordingId:
      name: recordingId
      in: path
      required: true
      description: Recording ID, the `id` returned by [List recordings](#operation/listRecordings).
      schema:
        type: string
        pattern: '^[A-Za-z0-9]{16}$'

  responses:
    Unauthorized:
      description: Authentication required or invalid token.
    NotFound:
      description: Resource not found.
    ValidationError:
      description: Validation failed.
      content:
        application/json:
          schema:
            type: object
            properties:
              message:
                type: string
              errors:
                type: object
                additionalProperties:
                  type: array
                  items:
                    type: string
    RateLimited:
      description: Too many requests. Retry after the period indicated in the Retry-After header.
    PaymentRequired:
      description: Not included in your plan. Web archives and recording every check are available on the Ultimate plan.

  schemas:
    Page:
      type: object
      properties:
        id:
          type: integer
          description: Internal identifier.
        name:
          type: string
          description: Page label or title.
        slug:
          type: string
          description: |
            Stable, URL-safe identifier for the page (lowercase letters, digits,
            and hyphens). Auto-generated from the page name on creation and unique
            within the team. The slug is what appears in the in-app URL, e.g.
            `https://pagecrawl.io/app/pages/{slug}`, and is accepted anywhere the
            API takes a `{id}` path parameter or a `page_ids` entry. Unlike the
            numeric `id`, the slug is exposed in shareable URLs, so prefer it for
            human-readable references; use `id` for machine-to-machine integrations.
            The slug stays the same when a page is renamed. If a page is deleted
            and re-added with the same name, it will typically receive the same slug.
          example: pagecrawl-pricing-page-a1b2
        url:
          type: string
          format: uri
          description: Monitored URL.
        url_tld:
          type: string
          description: Top-level domain of the URL.
        check_type:
          type: string
          enum: [crawl, push]
          description: |
            How this monitor gets its values.
            - `crawl` - PageCrawl fetches the configured URL on a schedule.
            - `push` - The monitor is a data source; values arrive via the
              ingest endpoint or its private inbound email address.
        ingest_url:
          type: string
          format: uri
          nullable: true
          description: |
            Secret capability URL for push data sources (e.g.
            `https://pagecrawl.io/api/ingest/data-a1b2c3d4e5f6g7h8i9j0k1l2`).
            POST values to it with no Authorization header; the URL itself is
            the credential, so treat it like a password. Returned on create
            and when fetching a single page; not included in the pages list.
            Only present on `push` data sources; null for crawl monitors.
        ingest_email:
          type: string
          nullable: true
          description: |
            Private inbound email address for push data sources (e.g.
            `data-a1b2c3d4e5f6g7h8i9j0k1l2@ingest.pagecrawl.io`). The local
            part is the same secret as `ingest_url`. Emailing a
            value to this address records it like an API push. The address
            itself is the credential, so keep it secret. Returned on create
            and when fetching a single page; not included in the pages list.
            Only present on
            `push` data sources; null for crawl monitors.
        created_at:
          type: string
          format: date-time
          nullable: true
          description: When the page was added (ISO 8601).
        last_checked_at:
          type: string
          format: date-time
          nullable: true
          description: When the last check was performed. Null if never checked or check in progress.
        latest:
          type: object
          nullable: true
          description: Data from the latest check.
          properties:
            numeric:
              type: boolean
              description: Whether the primary tracked element is a number.
            contents:
              type: string
              nullable: true
              description: Text content of the primary tracked element.
            changed_at:
              type: string
              format: date-time
              nullable: true
              description: When the check was performed.
            difference:
              type: integer
              nullable: true
              description: |
                Difference from the previous check in percent (0-100). For text elements, the share of the text that differs. For visual elements, the share of the captured area that changed; an Entire page or Top 2 screens capture taller than one screen is measured against one screen instead. Null when nothing was compared, such as a first capture.
            human_difference:
              type: string
              nullable: true
              description: Human-readable difference description (e.g. "The value has increased by 5%").
            three_month_difference:
              type: integer
              nullable: true
              description: Text difference percentage compared to a check 3 months ago.
            three_month_human_difference:
              type: string
              nullable: true
              description: Human-readable 3-month difference.
        status:
          type: string
          description: |
            Current page status. Common values:
            - `ok` - Checks running normally
            - `error` - Page failed to load
            - `disabled` - Monitoring paused
        failed:
          type: integer
          description: Number of consecutive failed checks.
        elements:
          type: array
          description: Configured tracked elements.
          items:
            $ref: '#/components/schemas/TrackedElement'
        tags:
          type: array
          description: Tags assigned to the page.
          items:
            type: object
            properties:
              id:
                type: integer
              label:
                type: string
              color:
                type: string

    TrackedElement:
      type: object
      description: |
        A configured tracked element. The fields shown here are ones `PUT /pages/{id}` accepts back,
        so an element read from a page can be resent on an update without losing its scope or
        threshold.
      properties:
        id:
          type: integer
          description: Element ID.
        selector:
          type: string
          description: |
            CSS or XPath selector (or other applicable value). For `visual` elements it is the
            page-absolute rectangle `x,y,width,height` in CSS pixels.
        type:
          type: string
          description: Type of tracked element (e.g. text, fullpage).
        label:
          type: string
          nullable: true
          description: Short label shown for the element.
        threshold:
          type: integer
          nullable: true
          description: Alert threshold in percent, as described on `PageCreateRequest`.
        visual_scope:
          type: string
          nullable: true
          enum: [fold, top_two, page]
          description: |
            What a visual element is configured to capture, null for the fixed rectangle in `selector`.
            Resend it with the element on update, since a `selector` submitted without it clears the
            stored scope.
        format_text:
          type: boolean
          description: |
            Whether "Capture formatted text" is on. Checks captured while it is on hold Markdown source
            in the element's `contents` rather than plain text, as described on `Check.elements[].contents`.
        value_type:
          type: string
          nullable: true
          readOnly: true
          enum: [auto, number, price, percentage, date, boolean, list, text]
          description: |
            `ai_extract` elements only: the kind of value the extraction holds, detected from its first
            usable answer. `auto` while it is still detecting; null for every other type and for an
            `ai_extract` element created before typed values (until its prompt is edited). Read-only:
            it is ignored when sent. For `number`, `price` and `percentage`, a check's `contents` is the
            plain number (for example `1299`) and its `original` holds the page's wording (`$1,299.00`),
            the same as a `price` element.

    Check:
      type: object
      properties:
        id:
          type: integer
          description: Internal identifier.
        status:
          type: string
          description: Check status.
        seen:
          type: string
          format: date-time
          nullable: true
          description: When the check was first viewed in the UI.
        created_at:
          type: string
          format: date-time
          description: When the check was performed.
        content_type:
          type: string
          nullable: true
          description: Content-Type of the page response.
        visual_diff:
          type: integer
          nullable: true
          description: Visual difference percentage compared to previous check (0-100). Null if not calculated.
        elements:
          type: array
          description: Changes detected for each tracked element.
          items:
            type: object
            properties:
              id:
                type: integer
                description: |
                  ID of this reading (the per-check record). Changes on every check;
                  do not use it as a stable key. Use `element_id` to match a value to
                  a specific tracked element across checks.
              element_id:
                type: integer
                description: |
                  Stable ID of the tracked element this value belongs to. Same across
                  every check. Matches an `id` in the page's `elements` array.
              contents:
                type: string
                nullable: true
                description: |
                  Text content of the tracked element.

                  When the element has "Capture formatted text" on (`format_text` on the element in
                  the page's `elements`), this is Markdown source rather than plain text: headings,
                  lists, links and tables are written in Markdown, and characters in the page's own
                  text that Markdown would read as formatting are escaped the CommonMark way, so a
                  page showing `AT&T #1` is returned as `AT&amp;T \#1`. Render it as Markdown to get
                  the text as the page shows it. In `latest.contents` and in webhook payloads the
                  escaping is already removed.

                  Note: except in PDF elements, whose Markdown has always been escaped, checks
                  captured before this escaping was introduced hold those characters unescaped,
                  and a page keeps capturing that way until it is next updated.
              difference:
                type: integer
                nullable: true
                description: |
                  Difference from the previous check in percent (0-100). For text elements, the share of the text that differs. For visual elements, the share of the captured area that changed; an Entire page or Top 2 screens capture taller than one screen is measured against one screen instead. Null when nothing was compared, such as a first capture.
              hash:
                type: string
                description: MD5 hash of the contents.
              changed:
                type: boolean
                description: Whether the element changed since the previous check.
              original:
                type: string
                nullable: true
                description: |
                  Original text before number extraction. Only populated for number, price, rating, and reviews types.
                  The "contents" field contains the extracted number only.
              elements:
                type: integer
                description: Number of DOM elements matching the selector. Note "text" tracked element captures only the first one.
        history:
          type: array
          description: Historical values for the primary tracked element.
          items:
            type: object

    Recording:
      type: object
      description: |
        A recording of a check that found no change, kept because the page
        records every check. Available on the Ultimate plan.
      properties:
        id:
          type: string
          description: Recording ID (16 characters).
          example: Xk3pQ9rT2vLm8nBw
        created_at:
          type: string
          format: date-time
          description: When the check was performed.
        has_screenshot:
          type: boolean
          description: Whether a full-page screenshot was recorded.
        has_archive:
          type: boolean
          description: Whether a web archive was recorded.
        archive_type:
          type: string
          nullable: true
          enum: [wacz, html]
          description: The web archive's format, or null when there is none.
        screenshot_url:
          type: string
          format: uri
          nullable: true
          description: API URL of the screenshot (authenticate as for any request). Null when there is none.
        archive_url:
          type: string
          format: uri
          nullable: true
          description: API URL of the web archive. Null when there is none.

    PageCreateRequest:
      type: object
      required: [url, name, elements, frequency]
      properties:
        url:
          type: string
          format: uri
          description: URL to monitor.
        name:
          type: string
          description: Label for the page.
        ignore_duplicates:
          type: boolean
          default: false
          description: If true and a page with the same URL/selector exists, returns the existing page (HTTP 200) instead of creating a duplicate.
        elements:
          type: array
          description: |
            Tracked elements to monitor. At least one element is required. The first element is the primary element.

            Example for full-page tracking:
            ```json
            [{ "type": "fullpage", "selector": "*", "label": "Page body" }]
            ```
          minItems: 1
          items:
            type: object
            required: [type, selector]
            properties:
              type:
                type: string
                description: |
                  Element type to track. Common values:
                  - `fullpage` - Track all visible text on the page
                  - `text` - Track text of a specific element (requires `selector`)
                  - `number` - Extract a numeric value from a specific element
                  - `price` - Track price + availability (use selector `*` for auto-detect)
                  - `rating`, `reviews` - Track aggregate rating / review count
                  - `availability` - Track in-stock / out-of-stock state
                  - `http_status` - Track HTTP response code only
                  - `html`, `html_multiple` - Track raw HTML
                  - `javascript` - Execute JS and track returned value
                  - `visual` - Track visual diff of an element region
                  - `json_path` - Track a value at a JSONPath
                  - `ai_extract` - Use AI to extract a value described by `prompt`
                  - `feed` - Treat URL as RSS/Atom/JSON feed
              selector:
                type: string
                description: |
                  CSS selector, XPath, or JSONPath. Use `*` for full-page or price auto-detect.

                  `visual` elements are the exception: the selector is a page-absolute rectangle
                  `x,y,width,height` in CSS pixels (e.g. `100,200,800,600`), and any other value is
                  rejected. A rectangle is required even when `visual_scope` is set, where it is the
                  seed the element falls back to; send `0,0,1920,1080` when you have no particular
                  area in mind.
              label:
                type: string
                description: Short label (e.g. "Price", "Quantity").
              prompt:
                type: string
                maxLength: 2000
                description: Required for `ai_extract` elements; the per-element extraction prompt.
              threshold:
                type: integer
                nullable: true
                description: |
                  Alert threshold in percent. Text: how much of the text must differ. Number and price: the percent increase or decrease. Visual: how much of the captured area must change, measured against one screen when the area is larger than one screen. Null means any change (1% for visual). -1 records changes without alerting.
              visual_scope:
                type: string
                nullable: true
                enum: [fold, top_two, page]
                description: |
                  For visual elements, what each check captures: `fold` (first screen), `top_two` (top two screens), `page` (entire page), or null for the fixed area in `selector`.

                  Sending an element's `selector` without `visual_scope` clears a stored scope, because a
                  freshly drawn rectangle wins over a scope declared earlier. Include `visual_scope` on
                  every update that resends the element, or an "Entire page" element quietly reverts to
                  the seed rectangle.
              currency:
                type: string
                pattern: '^[A-Za-z]{3}$'
                description: ISO 4217 currency code (e.g. `USD`, `EUR`). Applies to `price` elements.
        frequency:
          type: integer
          description: |
            Check frequency in minutes. Available values depend on your plan:
            - 1 (every minute)
            - 2, 3, 5 (every N minutes)
            - 15, 30, 45 (every N minutes)
            - 60 (hourly), 120, 180, 360 (every N hours)
            - 720 (twice daily), 1440 (daily)
            - 2880, 4320, 5760, 7200, 8640 (every N days)
            - 10080 (weekly), 14400, 20160 (every N weeks)
          enum: [1, 2, 3, 5, 15, 30, 45, 60, 120, 180, 360, 720, 1440, 2880, 4320, 5760, 7200, 8640, 10080, 14400, 20160]
        location:
          type: string
          default: random1
          description: |
            Server location for requests.
            - `random1` - Random Location (EU Datacenter proxies, default)
            - `residential1` - Random Location (US Premium proxies, Enterprise/Ultimate plans)
            - `residential_premium` - Residential Proxy with geo-targeting (extra fee, Enterprise/Ultimate plans). Requires `residential_country`.
            - `fixed1` - Fixed IP
            - `lon1` - London, UK
            - `fra1` - Frankfurt, DE
            - `tel1` - Tel Aviv, IL
            - `ny1` - New York, US
            - `sfo3` - San Francisco, US
            - `tor1` - Toronto, CA
          enum:
            - random1
            - residential1
            - residential_premium
            - fixed1
            - lon1
            - fra1
            - tel1
            - ny1
            - sfo3
            - tor1
        residential_country:
          type: string
          description: ISO 3166-1 alpha-2 country code (lowercase, e.g. `us`, `gb`, `de`). Required when `location` is `residential_premium`.
        folder_id:
          type: integer
          description: Save page in a specific folder.
        template_id:
          type: integer
          description: Use a specific template configuration.
        tags:
          type: array
          description: 'Tags to assign to the page (e.g. ["tag1", "another tag"]). Tags that don''t exist in the workspace are created automatically.'
          items:
            type: string
        notifications:
          type: array
          description: "Notification channels to enable (e.g. ['slack', 'mail', 'discord', 'telegram', 'teams'])."
          items:
            type: string
            enum: [mail, slack, discord, telegram, teams]
        notification_emails:
          description: |
            Additional Cc recipients for email notifications. Accepts either a
            comma-separated string of addresses (resolved into verified emails on
            the team) or an array of verified-email IDs. Subject to your plan's
            Cc limit per page.
          oneOf:
            - type: string
              description: Comma-separated email addresses.
            - type: array
              items:
                type: integer
                description: Verified-email ID.
        slack_channel:
          type: string
          format: uri
          description: Override the workspace Slack webhook for this page.
        discord_webhook:
          type: string
          format: uri
          description: Override the workspace Discord webhook for this page.
        ms_teams_webhook:
          type: string
          format: uri
          description: Microsoft Teams Workflow webhook URL (must contain `.logic.azure.com` or `.api.powerplatform.com` and `/workflows/`).
        telegram_id:
          type: integer
          description: Telegram chat ID to notify (validated by sending a test message at create time).
        skip_first_notification:
          type: boolean
          default: false
          description: Do not send notification if change is detected after a configuration update.
        fail_silently:
          type: integer
          minimum: 0
          description: |
            Controls error notification frequency:
            - `0` - Send an error notification on every failed check (default).
            - `1` - Never send error notifications.
            - `N` (N > 1) - Send an error notification only every Nth consecutive
              failure (e.g. `5` notifies on the 5th, 10th, 15th… failure in a row).
        rules_enabled:
          type: boolean
          default: false
          description: Enable notification rules.
        rules_and:
          type: boolean
          default: false
          description: When true, all rules must match to trigger a notification (AND logic). When false, any rule match triggers it (OR logic).
        rules:
          type: array
          description: Notification rules to control when alerts are sent.
          items:
            $ref: '#/components/schemas/NotificationRule'
        actions:
          type: array
          description: |
            Actions to perform before capturing the page.

            Example:
            ```json
            [{ "type": "scroll_to_bottom" }, { "type": "remove_cookies_v2" }]
            ```
          items:
            $ref: '#/components/schemas/Action'
        screenshots:
          type: boolean
          description: Capture a full-page screenshot on each check. Required for visual diffs.
        archive_enabled:
          type: boolean
          nullable: true
          description: |
            Store a web archive (WACZ) of each detected change (Ultimate plan).
            Needs a check frequency of 360 minutes (6 hours) or slower. `null`
            (the default) follows the page's template, then the workspace's web
            archiving setting. That setting Off overrides `true`.
        every_check_screenshots:
          type: boolean
          nullable: true
          description: |
            Record every check: keep the full-page screenshot of every check,
            including checks that found no change (Ultimate plan). Needs
            `screenshots` on. `null` (the default) follows the page's template,
            then the workspace setting. The workspace setting Off overrides
            `true`. See [Recordings](#tag/Recordings).
            Sending `true` without the Ultimate plan returns a validation error.
        every_check_archives:
          type: boolean
          nullable: true
          description: |
            Record every check: keep a web archive of every check, including
            checks that found no change (Ultimate plan). Needs a check frequency
            of 360 minutes (6 hours) or slower, and works even when
            `archive_enabled` is off. `null` (the default) follows the page's
            template, then the workspace setting. The workspace setting Off
            overrides `true`. See [Recordings](#tag/Recordings).
            Sending `true` without the Ultimate plan returns a validation error.
        ai_summaries_enabled:
          type: boolean
          description: |
            Generate AI summaries for detected changes (requires AI configured on
            the workspace).

            **Default behaviour** when this field is omitted:
            - `true` for content-style monitors (the first element's `type` is
              `text`, `fullpage`, or any AI/content element)
            - `false` for numeric/status monitors (`number`, `price`,
              `http_status`, `availability`, `boolean`, `rating`, `reviews`)
              where a prose summary adds little
            - **Always forced to `true`** when `ai_page_focus` is provided — the
              focus hint is a strong intent signal that you want AI summaries on.
        ai_page_focus:
          type: string
          maxLength: 1000
          description: |
            Free-text hint to the AI about what matters on this page (e.g.
            "Focus on pricing tier changes; ignore footer updates"). Influences
            change summaries and priority scoring. Providing this field
            automatically enables `ai_summaries_enabled`.
        disabled:
          type: boolean
          default: false
          description: Create the page with monitoring disabled. Enabling returns HTTP 406 if account page limit is exceeded.
        track_type:
          type: string
          default: one
          enum: [one, multiple, from_file, scan]
          description: |
            - `one` - Track a single URL (default)
            - `multiple` - Create multiple pages from a list of URLs (requires `urls`)
            - `from_file` - Create pages from an uploaded URL list (requires `urls`)
            - `scan` - Crawl a starting URL and create a page per discovered URL (requires `start_url`, `levels`)
        urls:
          description: |
            URLs to track when `track_type` is `multiple` or `from_file`.
            Supports arrays or newline-separated strings.
            Append `||` to customize page titles: `"example.com||My Page"`.
          oneOf:
            - type: array
              items:
                type: string
            - type: string
        start_url:
          type: string
          format: uri
          description: Starting URL for `track_type=scan`.
        levels:
          type: string
          enum: ['all', '1', '2', '3', '4']
          description: Crawl depth for `track_type=scan`.
        url_regex:
          type: string
          description: Optional regex filter applied to discovered URLs during `scan`.
        report_id:
          type: integer
          description: If provided, the new page is added to this Report.
        # ---- Advanced (rarely used) ----
        advanced:
          type: boolean
          default: false
          description: Toggle to surface advanced settings (auth_username/password, user_agent, proxies, headers) in the UI. Does not affect API behaviour.
        auth_id:
          type: integer
          description: Use a saved authentication configuration.
        auth_username:
          type: string
          description: HTTP Basic Authentication username. Cannot be used with custom proxies.
        auth_password:
          type: string
          description: HTTP Basic Authentication password.
        check_always:
          type: boolean
          default: false
          description: Continue checking even when the page returns 4xx/5xx errors (normally checking speed is reduced and eventually stopped).
        smart_retries:
          type: boolean
          default: false
          description: Retry transient failures (timeouts, proxy errors) before marking a check as failed.
        rate_limits:
          type: boolean
          default: false
          description: Apply per-check rate limiting to spread load.
        user_agent:
          type: string
          description: Custom browser User-Agent string.
        proxies:
          type: string
          description: Newline-separated list of proxies. A random proxy is selected per check.
        headers:
          type: string
          description: Custom HTTP headers as a JSON-encoded string (e.g. `{"X-Auth":"..."}`).
        device:
          type: string
          description: Device profile to emulate (see `/settings.devices`).
        timezone:
          type: string
          description: IANA timezone for the browser session (e.g. `Europe/London`).
        language:
          type: string
          description: Browser language code (see `/settings.languages`).

    PageUpdateRequest:
      description: Same fields as PageCreateRequest. All fields are optional for updates.
      type: object
      properties:
        url:
          type: string
          format: uri
        name:
          type: string
        elements:
          type: array
          description: |
            Replaces the monitor's elements, so resend every element you want to keep, with the fields
            you want kept. In particular a `visual` element sent with a `selector` but no `visual_scope`
            loses its scope and falls back to that rectangle.
          items:
            type: object
            properties:
              id:
                type: integer
                description: |
                  Element ID from the page response. Send it to update that element in place. An element
                  sent without an `id` is created as a new one, and an element missing from the array is
                  deleted.
              type:
                type: string
              selector:
                type: string
              label:
                type: string
              threshold:
                type: integer
                nullable: true
              visual_scope:
                type: string
                nullable: true
                enum: [fold, top_two, page]
        frequency:
          type: integer
        location:
          type: string
        disabled:
          type: boolean
        notifications:
          type: array
          items:
            type: string
        rules_enabled:
          type: boolean
        rules:
          type: array
          items:
            $ref: '#/components/schemas/NotificationRule'
        actions:
          type: array
          items:
            $ref: '#/components/schemas/Action'
        tags:
          type: array
          items:
            type: string
        archive_enabled:
          type: boolean
          nullable: true
          description: See `archive_enabled` on PageCreateRequest. Send `null` to follow the template or workspace again.
        every_check_screenshots:
          type: boolean
          nullable: true
          description: See `every_check_screenshots` on PageCreateRequest (Ultimate plan). Send `null` to follow the template or workspace again.
        every_check_archives:
          type: boolean
          nullable: true
          description: See `every_check_archives` on PageCreateRequest (Ultimate plan). Send `null` to follow the template or workspace again.

    TrackSimpleRequest:
      type: object
      required: [url]
      properties:
        url:
          type: string
          format: uri
          description: URL to monitor.
        name:
          type: string
          description: Optional human-readable label. Defaults to the URL.
        frequency:
          type: integer
          default: 1440
          description: |
            Check frequency in minutes. Available values depend on your plan
            (see `PageCreateRequest.frequency` for the full list). Defaults to
            `1440` (daily).
        tracking_mode:
          type: string
          description: |
            What to track on the page:
            - `fullpage` - All visible text on the page (default)
            - `content_only` - Page text excluding navigation, headers, and footers
            - `reader` - Reader-mode (article) content only, ideal for blog posts and articles
            - `feed` - Parse the URL as an RSS/Atom/JSON feed and track new items
            - `price` - Auto-detect prices and track availability
            - `seo` - Track on-page SEO signals (title, meta, headings)
            - `leaderboard` - Track ranking changes on leaderboards/listings
            - `pdf_extract` - Extract and track text from a PDF at the URL
            - `http_status` - Track only the HTTP response status code
            - `json_path` - Track a value at a JSON path (requires `selector` as the JSONPath)
            - `specific_text` - Track text of a specific element (requires `selector`)
            - `specific_number` - Track a numeric value from a specific element (requires `selector`)
            - `ai_extract` - Use AI to extract a value described in natural language (requires `prompt`)
          enum:
            - fullpage
            - content_only
            - reader
            - feed
            - price
            - seo
            - leaderboard
            - pdf_extract
            - http_status
            - json_path
            - specific_text
            - specific_number
            - ai_extract
        selector:
          type: string
          description: |
            CSS selector, XPath, or (for `json_path` mode) a JSONPath expression
            pointing at a specific element/value to track. Required for
            `specific_text`, `specific_number`, and `json_path` tracking modes.
            If omitted in other modes, the full page is tracked.
        element_type:
          type: string
          description: |
            Element-level type override. Most users should set `tracking_mode`
            instead; `element_type` is the lower-level value persisted on the
            tracked element (e.g. `fullpage`, `text`, `number`, `price`,
            `availability`, `boolean`, `http_status`). See the Pages API.
        prompt:
          type: string
          description: |
            Natural-language description of what to extract. Required when
            `tracking_mode` is `ai_extract`. Example: "Extract the current
            ETH/USD price as a number".
        elements:
          type: array
          description: |
            Optional explicit list of tracked elements. Use this when a single
            `tracking_mode` + `selector` isn't enough (e.g. an `ai_extract`
            monitor that should extract several distinct values, each with its
            own prompt). Overrides `tracking_mode`/`selector`/`element_type`
            when present.
          items:
            type: object
            required: [type]
            properties:
              type:
                type: string
                description: Element type (e.g. `text`, `number`, `fullpage`, `ai_extract`).
              selector:
                type: string
              label:
                type: string
                maxLength: 255
              prompt:
                type: string
                maxLength: 2000
                description: For AI element types; the per-element extraction prompt.
        folder_id:
          type: integer
          description: Save the page in a specific folder.
        template_id:
          type: integer
          description: Apply a saved template configuration.
        auth_id:
          type: integer
          description: Use a saved authentication configuration.
        location:
          type: string
          description: |
            Server/proxy location used for checks. See `PageCreateRequest.location`
            for the full enum (e.g. `random1`, `residential1`, `lon1`, `ny1`).
        proxies:
          type: string
          description: Newline-separated custom proxies (one per line). A random proxy is picked per check.
        headers:
          type: object
          description: Custom HTTP headers as key-value pairs.
          additionalProperties:
            type: string
        user_agent:
          type: string
          description: Custom browser User-Agent string.
        notifications:
          type: array
          description: Notification channels to enable for this page.
          items:
            type: string
            enum: [mail, slack, discord, telegram]
        notification_emails:
          description: |
            Additional Cc recipients for email notifications. Accepts either a
            comma-separated string of email addresses (the API will resolve them
            to verified-email IDs) or an array of verified-email IDs you've
            already added to the team. Subject to your plan's Cc limit per page.
          oneOf:
            - type: string
              description: Comma-separated list of email addresses.
            - type: array
              items:
                type: integer
                description: Verified-email ID.
        slack_channel:
          type: string
          format: uri
          description: Override the workspace Slack webhook for this page.
        discord_webhook:
          type: string
          format: uri
          description: Override the workspace Discord webhook for this page.
        telegram_id:
          type: integer
          description: Telegram chat ID to notify (validated by sending a test message at create time).
        check_always:
          type: boolean
          default: false
          description: Continue checking even when the page returns 4xx/5xx errors.
        smart_retries:
          type: boolean
          default: false
          description: Retry transient failures (timeouts, proxy errors) before marking a check as failed.
        fail_silently:
          type: integer
          minimum: 0
          description: |
            Error notification frequency:
            - `0` - Send a notification on every failed check (default).
            - `1` - Never send error notifications.
            - `N` (N > 1) - Send an error notification only every Nth consecutive failure.
        ignore_duplicates:
          type: boolean
          default: false
          description: |
            If a page with the same URL (and selector, when applicable) already
            exists, return that existing page with HTTP 200 instead of creating
            a duplicate. New pages return 201.
        rules_enabled:
          type: boolean
          default: false
          description: Enable notification rules for this page.
        rules_and:
          type: boolean
          default: false
          description: When true, all rules must match to trigger a notification (AND). When false, any rule match triggers it (OR).
        rules:
          type: array
          description: Notification rules controlling when alerts are sent.
          items:
            $ref: '#/components/schemas/NotificationRule'
        ai_summaries_enabled:
          type: boolean
          description: |
            Generate AI summaries for detected changes (requires AI configured on
            the workspace).

            **Default behaviour** when this field is omitted:
            - `true` for content-style tracking modes (`fullpage`, `content_only`,
              `reader`, `feed`, `leaderboard`, `seo`, `specific_text`, `ai_extract`)
            - `false` for numeric/status modes (`price`, `specific_number`,
              `http_status`) where a prose summary adds little
            - **Always forced to `true`** when `ai_page_focus` is provided — the
              focus hint is a strong intent signal that you want AI summaries on.
        ai_page_focus:
          type: string
          maxLength: 500
          description: |
            Free-text hint to the AI about what matters on this page (e.g.
            "Focus on pricing tier changes; ignore footer updates"). Influences
            change summaries and priority scoring. Providing this field
            automatically enables `ai_summaries_enabled`.
        tags:
          type: array
          description: |
            Tags to assign to this page (e.g. `["pricing", "competitor"]`).
            Tags that don't exist in the workspace are created automatically.
          items:
            type: string
            maxLength: 255
        track_availability:
          type: boolean
          default: true
          description: Also monitor product availability. Only applies to `price` tracking mode.
        actions:
          type: array
          description: Actions to perform before capturing the page, same shape and rules as `PageCreateRequest.actions`.
          items:
            $ref: '#/components/schemas/Action'
        engine:
          type: string
          enum: [stealth, fast]
          description: Pin the capture engine, `stealth` for bot-protected sites or `fast` for static pages and XML feeds. Omit for the default browser engine.
        fullpage_level:
          type: string
          enum: [all, content, reader]
          description: For `fullpage` tracking, how much of the page text to keep, `all` (default), `content` (no navigation, headers or footers) or `reader` (article body only).
        feed_selector:
          type: string
          description: For `feed` tracking, the CSS selector of one repeated item on an HTML page, or the dot-notation path to the items array in a JSON response. Takes precedence over `selector`. Omit for RSS/Atom feeds.
        feed_item_limit:
          type: integer
          minimum: 1
          description: For `feed` tracking, how many newest items to keep per check. Clamped to your plan's cap.

    DataSourceCreateRequest:
      type: object
      required: [name]
      properties:
        name:
          type: string
          description: Label for the data source.
        fields:
          type: array
          description: |
            Optional named fields to track. Leave this out to track a single
            value. Each field becomes a tracked element on the data source, so
            it gets its own history, charts, and notification rules. Fields can
            also be created implicitly later by pushing new names in the
            `values` map of the ingest endpoint.
          items:
            type: object
            required: [label]
            properties:
              label:
                type: string
                description: Field name (e.g. "price", "stock"). Used as the key in the `values` map when pushing.
              type:
                type: string
                enum: [number, price, text]
                description: |
                  Field type. Optional, defaults to `text` when omitted.
                  - `number` - Numeric value (charted, supports numeric notification rules)
                  - `price` - Monetary value
                  - `text` - Free text
              currency:
                type: string
                description: Optional currency for `price` fields (e.g. `USD`).
              threshold:
                type: number
                description: Optional numeric change threshold for the field.
        ai_summaries_enabled:
          type: boolean
          default: true
          description: |
            Whether AI summaries and priority scoring run on each recorded
            value. On by default for data sources. Set to `false` for
            high-frequency numeric metrics where you only want charts and
            thresholds. No effect on workspaces without AI access.
        ai_prompt:
          type: string
          description: |
            Optional guidance for AI summaries: what the value means and what
            to flag (e.g. "Competitor changelog entry. Summarize what changed
            and flag pricing or plan changes."). Stored as the data source's
            AI focus.
        email_sender_filter:
          type: string
          description: |
            Optional sender restriction for email ingestion. Accepts an exact
            email address (e.g. `reports@example.com`) or an exact domain
            (e.g. `example.com`). When set, emails from any other sender are
            ignored.
      example:
        name: Competitor API price
        fields:
          - label: price
            type: number
          - label: stock
            type: number
        ai_prompt: Flag pricing or plan changes.

    IngestRequest:
      description: |
        Send `value` for a single-value data source, or `values` for named
        fields. Field names in `values` that do not exist yet are created
        automatically (numeric-looking values become number fields, everything
        else text).

        Value size: a single value can be up to 500,000 characters (larger
        values return a 422). The first 65,535 characters are stored (10,000
        on the Free plan); longer values also keep a checksum of the full
        content, appended to the stored portion as a marker line, so a change
        anywhere in the value is still detected.
        Set `return_summary: true` to also get the AI `summary` and `score`
        for the change in the response (default off; only populated when a
        change was detected and AI summaries are enabled). The response always
        includes `changed`, so you know whether the value changed without it.
      oneOf:
        - type: object
          required: [value]
          properties:
            value:
              description: The value to record. Numbers and strings are accepted.
              oneOf:
                - type: number
                - type: string
            return_summary:
              type: boolean
              description: Also return the AI summary and score in the response. Default false.
          example:
            value: 42
        - type: object
          required: [values]
          properties:
            values:
              type: object
              description: Map of field name to value.
              additionalProperties:
                oneOf:
                  - type: number
                  - type: string
            return_summary:
              type: boolean
              description: Also return the AI summary and score in the response. Default false.
          example:
            values:
              price: 9.99
              stock: 12

    IngestResponse:
      type: object
      properties:
        changed:
          type: boolean
          description: |
            Whether at least one value differs from the previous reading.
            Pushing the same value again returns `false` and stores nothing new.
        check_id:
          type: integer
          nullable: true
          description: |
            ID of the check recorded for this push, usable with the History
            and diff endpoints. May be null when nothing new was stored.
        monitor_id:
          type: integer
          description: ID of the data source the values were recorded on.
        values:
          type: object
          description: The stored values, keyed by field name.
          additionalProperties:
            type: string
        created_fields:
          type: array
          description: Names of any fields created automatically by this push.
          items:
            type: string
        summary:
          type: string
          nullable: true
          description: |
            AI summary of the change. Present only when the request set
            `return_summary: true`. Null when no change was detected, AI
            summaries are off for the data source, or AI processing was
            deferred.
        score:
          type: number
          nullable: true
          description: |
            AI priority score (0-100) for the change. Present only when the
            request set `return_summary: true`. Null under the same conditions
            as `summary`.
      example:
        changed: true
        check_id: 123
        monitor_id: 55
        values:
          price: "9.99"
          stock: "12"
        created_fields: []

    NotificationRule:
      type: object
      required: [type]
      properties:
        type:
          type: string
          description: |
            Rule type:
            - `text_difference` - Text difference percent compared to the previous check
            - `added` - Text added
            - `removed` - Text removed
            - `eq` - Equals
            - `neq` - Not equals
            - `contains` - Contains
            - `ncontains` - Not contains
            - `increased` - Number increased by x% (number types only)
            - `decreased` - Number decreased by x% (number types only)
            - `is` - Number is (number types only)
            - `gt` - Greater than (number types only)
            - `lt` - Less than (number types only)
            - `gte` - Greater than or equals (number types only)
            - `lte` - Less than or equals (number types only)
            - `ignore_text` - Ignore lines of text when comparing (value: one line per ignored text)
            - `ignore_digits` - Ignore numbers when comparing (no value)
            - `ignore_dates` - Ignore dates and times when comparing (no value)
          enum:
            - text_difference
            - added
            - removed
            - eq
            - neq
            - contains
            - ncontains
            - increased
            - decreased
            - is
            - gt
            - lt
            - gte
            - lte
            - ignore_text
            - ignore_digits
            - ignore_dates
        value:
          description: |
            Rule value. String for text rules, number for numeric rules.
            Not used for ignore_digits or ignore_dates.
          oneOf:
            - type: string
            - type: number
        element:
          description: Tracked element ID or index (e.g. 0 for the first tracked element).
          oneOf:
            - type: integer
            - type: string

    Action:
      type: object
      required: [type]
      properties:
        type:
          type: string
          description: |
            Action type:
            - `scroll_to_bottom` - Scroll to bottom of the page
            - `remove_dates` - Deprecated and no longer applied: sending it turns on the `ignore_dates` rule for every eligible tracked element instead, which leaves the captured text intact. A monitor that already carries the action keeps it until you remove it.
            - `remove_cookies` - Remove known cookie windows (deprecated, use remove_cookies_v2)
            - `remove_cookies_v2` - Remove known cookie windows
            - `back` - Go back to previous page
            - `click` - Click on an element (requires selector)
            - `hover` - Hover on an element (requires selector)
            - `type` - Type text (requires selector and value)
            - `select` - Select dropdown option (requires selector and value)
            - `remove_element` - Remove an element (requires selector)
            - `wait_element` - Wait for element to appear (requires selector)
            - `wait` - Wait N seconds, up to 10 seconds, max 3 times (requires value)
            - `javascript` - Execute JavaScript code (requires value)
          enum:
            - scroll_to_bottom
            - remove_dates
            - remove_cookies
            - remove_cookies_v2
            - back
            - click
            - hover
            - type
            - select
            - remove_element
            - wait_element
            - wait
            - javascript
        selector:
          type: string
          description: CSS or XPath selector (for click, hover, type, select, remove_element, wait_element).
        value:
          type: string
          description: |
            Value for the action:
            - `type`: text to type
            - `select`: option value or label
            - `wait`: number of seconds
            - `javascript`: JavaScript code to execute
