# Website Change Monitor – Page & Price Change Alerts (`glidepath/website-change-monitor`) Actor

Website change monitor: watch pages or CSS selectors (prices, stock, policies, docs) with plain HTTP and get new, changed and removed lines since the last run. Input: page URLs. Output: changed pages + diffs. $1.00/1k checks.

- **URL**: https://apify.com/glidepath/website-change-monitor.md
- **Developed by:** [Glidepath](https://apify.com/glidepath) (community)
- **Categories:** Developer tools, Automation, Marketing
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.80 / 1,000 page checks

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Website Change Monitor – Page & Price Change Alerts

**Website change monitor for prices, stock, policies and docs:** watch any public web page, or just the part of it that matters (a price, a stock label, a changelog), and get only what changed since the last run: new, changed and removed pages with the added and removed lines. Plain HTTP, no browser, so checks are fast and cheap.

Schedule it hourly or daily, connect Slack or e-mail, and you have change alerts without writing code.

### Who uses it

- **E-commerce and pricing teams:** watch competitors' price and stock elements (`.price`, `#availability`) and get the old and new value.
- **Legal, compliance and procurement:** track terms of service, privacy policies, supplier pages and public notices.
- **Developers and product managers:** follow changelogs, release pages, API docs and status pages.
- **Marketers and SEO:** see when a competitor changes a landing page, a plan or a headline.

### How to use it

1. Put the pages in **Page URLs**, one per line. Optionally add a **CSS selector** to watch only part of each page (e.g. `.price`), or use **Pages with their own selector** for different selectors per page.
2. Run it once. The first run saves every page as `new` (your baseline) in a memory store in your Apify account.
3. Schedule it (hourly, daily, weekly). Every later run returns only pages that changed or disappeared, with the lines that were added and removed.

### Input example

```json
{
  "urls": ["https://www.python.org/downloads/", "https://example.com/terms"],
  "pages": [
    {"url": "https://shop.example.com/product/123", "selector": ".price", "label": "Competitor price"},
    {"url": "https://shop.example.com/product/123", "selector": "#stock", "label": "Competitor stock"}
  ],
  "ignorePatterns": ["Last updated .*", "\\d+ people are viewing"],
  "includeUnchanged": false,
  "stateStoreName": "my-competitor-watchlist"
}
```

Use one **Memory store name** per watchlist: each store remembers the last version of its pages, so two different schedules never mix their baselines.

### Output example

```json
{
  "url": "https://shop.example.com/product/123",
  "finalUrl": null,
  "label": "Competitor price",
  "selector": ".price",
  "changeType": "changed",
  "checkedAt": "2026-10-07T18:00:00Z",
  "previousCheckedAt": "2026-10-06T18:00:00Z",
  "httpStatus": 200,
  "title": "Acme Pro – Acme Store",
  "textLength": 11,
  "hash": "5f1c0d0e8a4b0c3f9e2d7a1b6c4e8f20",
  "previousHash": "a3b9e0f1c2d4e5f60718293a4b5c6d7e",
  "addedCount": 1,
  "removedCount": 1,
  "addedLines": ["$35 / month"],
  "removedLines": ["$29 / month"],
  "text": "$35 / month"
}
```

| Field | Type | Description |
|---|---|---|
| `url` | string | The page you asked to check |
| `finalUrl` | string / null | Where the page ended up after redirects (null when it is the same URL) |
| `label` | string / null | Your label from **Pages with their own selector** |
| `selector` | string / null | CSS selector used (null = whole page) |
| `changeType` | string | `new` (first check in this memory store), `changed`, `removed` (page now 404/410), `unchanged` (only with **Also output unchanged pages**) |
| `checkedAt`, `previousCheckedAt` | string | This check and the version it was compared with (ISO 8601 UTC) |
| `httpStatus` | integer | HTTP status of the page |
| `title` | string / null | Page `<title>` |
| `textLength` | integer | Characters of normalised text compared |
| `hash`, `previousHash` | string / null | Fingerprints of the compared text |
| `addedCount`, `removedCount` | integer | Lines added and removed since the last check |
| `addedLines`, `removedLines` | array | The lines themselves (up to **Max added/removed lines**) |
| `text` | string / null | Current text, first 5,000 characters, e-mail addresses and phone numbers removed |

Pages that could not be checked are listed in the run's key-value store as `ERRORS` (with the reason: `ROBOTS_DISALLOWED`, `BLOCKED`, `FETCH_FAILED`, `NOT_FOUND`, `SELECTOR_NOT_FOUND`, `EMPTY_PAGE`, `UNSUPPORTED_CONTENT`, `REFUSED`), and a run summary as `SUMMARY`.

### Use it with

- **Schedules:** save your input as a Task and run it every hour, day or week; each run compares with the previous one.
- **Alerts:** the Integrations tab sends results to Slack, e-mail, Google Sheets or a webhook. With the default settings a run only outputs changed pages, so an empty run means "nothing changed".
- **Make, Zapier and n8n:** start a workflow when a run finishes and read the changed pages.
- **API and code:** the Apify API, the JavaScript and Python clients, or CSV/JSON/Excel downloads.
- **AI agents:** the Apify MCP server exposes this Actor as a tool; the input is a list of URLs.

### Pricing

Pay per page check: **$1.00 per 1,000 page checks** on the Apify Free plan; on paid plans **$1.00** (Starter), **$0.90** (Scale) and **$0.80** (Business and higher) (the price for your plan is shown on the Pricing tab). The examples below use the Free-plan price. One check = one page (or one page + selector) fetched and compared, whether or not it changed. Pages that are blocked, disallowed by robots.txt, not found on the first check or fail are never charged.

- 20 competitor pages checked daily → 600 checks/month → about **$1.00/month** or less.
- 200 pages checked hourly → about 144,000 checks/month → about **$144.00/month**.
- 2,000 pages checked once a week → about 8,600 checks/month → about **$9.00/month**.

A small start fee of $0.00005 per run applies. Set a **maximum cost per run** in the run options and the Actor stops cleanly when it is reached; pages it could not get to keep their old version and are compared next run.

### Limits and notes

- **Plain HTTP only, no browser.** Pages that build their content with JavaScript may show little or no text (reported as `EMPTY_PAGE`, not charged). Many shops and docs render server-side and work fine.
- **robots.txt is respected.** Before checking a page the Actor reads the site's robots.txt (RFC 9309, including `*` and `$` rules) and skips pages it disallows for crawlers (`ROBOTS_DISALLOWED`, not charged). If robots.txt itself can't be read (server error), the site is skipped to be safe.
- **No bot-check bypass, no proxies, no logins.** A Cloudflare/DataDome-style challenge, a 401/403 or a persistent 429 is reported as `BLOCKED` and not charged. There are no cookie, login or header inputs, and private or internal addresses are refused.
- **Politeness:** at most one request per second per website, five websites in parallel, pages up to 5 MB.
- Up to 1,000 pages per run. Text is compared line by line after whitespace is normalised; use **Ignore text matching** for dates, counters or rotating banners so they don't trigger false changes.
- E-mail addresses and phone numbers are removed from all output text.

### FAQ

**Is it legal to monitor websites?** The Actor only reads public pages that the site's robots.txt allows for crawlers, at a polite rate, without logging in or getting around any protection, and it removes contact details from its output. Whether you may use a site's content depends on that site's terms and your purpose; check them for the pages you add. If in doubt, ask a lawyer.

**Why is a page `changed` when nothing visible changed?** Something in its text changed: often a date, a counter or a rotating promo. Add an ignore pattern for it, or watch only the relevant part with a CSS selector.

**Why did a page show `EMPTY_PAGE` or `SELECTOR_NOT_FOUND`?** The content is probably loaded by JavaScript after the page opens. Try a selector on server-rendered parts, or the site's RSS feed or sitemap.

**Can I monitor a page behind a login?** No. The Actor only checks public pages and never takes passwords or cookies.

**How do I get alerts only when something changes?** Keep **Also output unchanged pages** off and add a Slack or e-mail integration to your scheduled Task; runs with no changes produce no rows.

### Changelog

See [CHANGELOG.md](CHANGELOG.md).

### Support

Found a problem or need a feature? Open an issue on the Actor's Issues tab. We read every one.

# Changelog

This Actor's version history is a separate document: https://apify.com/glidepath/website-change-monitor/changelog.md

# Actor input Schema

## `urls` (type: `array`):

Public pages to check, one per line (pricing pages, product pages, policies, docs, changelogs). Each run compares every page with the previous run of the same memory store.

## `cssSelector` (type: `string`):

Watch only the part of each page that matches this selector, e.g. '.price', '#stock', 'main article'. Leave empty to watch the whole page's visible text.

## `pages` (type: `array`):

Advanced: a list of {"url": ..., "selector": ..., "label": ...} objects, for pages that need different selectors. Added to the URLs above.

## `includeUnchanged` (type: `boolean`):

Off: only new, changed and removed pages are saved. On: every checked page gets a row (changeType 'unchanged' when nothing changed). Every checked page is charged either way.

## `includeText` (type: `boolean`):

Add the page's (or selector's) current text, first 5,000 characters, e-mails and phone numbers removed.

## `ignorePatterns` (type: `array`):

Regular expressions removed before comparing, for parts that change on every visit (dates, counters, 'Last updated …'). Up to 20.

## `maxDiffLines` (type: `integer`):

How many added and removed lines to show per changed page (the counts are always complete).

## `stateStoreName` (type: `string`):

Named key-value store in your account that keeps the last version of each page. Use one name per watchlist (letters, digits and '-').

## Actor input object example

```json
{
  "urls": [
    "https://www.python.org/downloads/",
    "https://www.rfc-editor.org/rfc/rfc9309.html"
  ],
  "includeUnchanged": false,
  "includeText": true,
  "maxDiffLines": 50,
  "stateStoreName": "glidepath-change-monitor"
}
```

# Actor output Schema

## `changes` (type: `string`):

No description

## `summary` (type: `string`):

No description

## `errors` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        "https://www.python.org/downloads/",
        "https://www.rfc-editor.org/rfc/rfc9309.html"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("glidepath/website-change-monitor").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "urls": [
        "https://www.python.org/downloads/",
        "https://www.rfc-editor.org/rfc/rfc9309.html",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("glidepath/website-change-monitor").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    "https://www.python.org/downloads/",
    "https://www.rfc-editor.org/rfc/rfc9309.html"
  ]
}' |
apify call glidepath/website-change-monitor --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,glidepath/website-change-monitor"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/gwYJrVAMO7cndj5Rn/builds/XVUvhYCsFYwcMrMJr/openapi.json
