# Website Change Monitor – Text Diff per Page (`pontio/website-change-monitor`) Actor

Watch web pages, or one part of each with a CSS selector, and get the lines added and removed since the last run. Schedule it and pay per page checked.

- **URL**: https://apify.com/pontio/website-change-monitor.md
- **Developed by:** [Gabor Molnar](https://apify.com/pontio) (community)
- **Categories:** Automation, Developer tools, Marketing
- **Stats:** 2 total users, 1 monthly users, 0.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$1.00 / 1,000 page checkeds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Website Change Monitor – Text Diff per Page

Give it a list of pages and it tells you, for each, whether its text changed since the last run, and which lines were added and removed. Watch a whole page or just one part of it with a CSS selector, such as a pricing table, a terms-of-service section or a job list. Put it on an Apify schedule and it becomes a monitor. You pay per page checked, and a check that couldn't see the page is free.

### What you get

One dataset item per page. Here is an example: a pricing page whose price moved between two scheduled runs.

```json
{
  "url": "https://acme.com/pricing",
  "selector": "#plans",
  "reasonCode": "changed",
  "httpStatus": 200,
  "contentHash": "5f0c…e21a",
  "previousCheckedAt": "2026-09-26T06:00:04.118Z",
  "added": ["$12 per month", "New: SSO on every plan"],
  "removed": ["$10 per month"],
  "addedCount": 2,
  "removedCount": 1,
  "checkedAt": "2026-09-27T06:00:03.902Z"
}
```

The first run for a page stores its baseline and reports `first_seen`. Every run after that reports `unchanged`, `changed` or `gone`.

### Use it when

- You track competitors' pricing pages, plans or feature lists.
- You need to know when terms of service, a privacy policy or a regulation page is edited.
- You watch a careers page, a changelog or a status page for new entries.
- An agent needs "what's new on this page since yesterday" as data, not a screenshot.

### Input

| Field | Type | Required | What it does |
| --- | --- | --- | --- |
| `urls` | array of strings | yes | Up to 1,000 `http` or `https` URLs per run. A URL is checked once per run, however it is spelled: case in the scheme and host, a default port and a `#fragment` make no difference. The path and query string do. |
| `selector` | string | no | A CSS selector (`#plans`, `main .pricing-table`, `article`) for the part of each page to watch. Defaults to the whole page body. Applies to every URL in the run: to watch different parts of different pages, use one run (or one scheduled task) per selector. |
| `monitorName` | string | no | Keeps this monitor's snapshots apart from any other monitor watching the same URLs. Use one per scheduled task. Change it to start over from a fresh baseline. Lowercase letters, digits and hyphens, up to 40. |

### How a change is detected

The page's visible text is read one line per paragraph, heading, list item, table cell or other block, with whitespace collapsed. Scripts, styles, markup and attributes are not watched, so a rotated security token or a new asset hash in the HTML is not a change. A `text/plain` or other non-HTML response is watched line by line as it is.

`changed` means the watched text differs from the previous snapshot. `added` and `removed` list the lines that appear in only one of the two versions, up to 100 each, and `addedCount` and `removedCount` give the full numbers. A change that only reorders lines, or changes how often the same line repeats, is `changed` with both lists empty. The same goes for a change past the first 10,000 lines of a very long page: it is detected, but only the first 10,000 lines are stored and diffed.

Snapshots are kept in a key-value store named `website-change-monitor` in your own Apify account, one per page, selector and `monitorName`. Deleting that store resets every baseline.

### Pricing

Pay per event, $1.00 per 1,000 pages checked ($0.001 each), charged on the `page-checked` event.

A check is charged when it compared the page against its baseline, or took the baseline. `unchanged` is an answer too, and on a schedule it is the one you get most often.

| `reasonCode` | Charged |
| --- | --- |
| `first_seen` (no earlier snapshot for this page, selector and `monitorName`: the baseline is taken) | yes |
| `unchanged` (the watched text is the same as last time) | yes |
| `changed` (the watched text differs; see `added` and `removed`) | yes |
| `gone` (the page now answers 404 or 410; `removed` lists what it held) | yes |
| `not_found` (404 or 410 on a page never seen before, most likely a wrong URL) | no |
| `blocked` (401, 403 or 429: a bot wall or login) | no |
| `unreachable` (no answer, a timeout, a 5xx or any other error status) | no |
| `selector_not_found` (the page loaded but the selector matched nothing, or is not valid CSS) | no |
| `no_content` (the page, or the selected part, has no text, typically a page that renders with JavaScript) | no |
| `invalid_url` (not an `http(s)` URL on a domain name: an IP address, `localhost` and credentials in the URL are refused before any request) | no |

A free check leaves the stored baseline alone, so the next check that can see the page still compares against the last one that did.

If you set a maximum total charge for a run, the Actor stops as soon as that limit is reached instead of working for free. Items after that point are left out of the dataset and not charged, and the run log says how many; submit them in a new run.

### Limits

- Pages are read as served, without running JavaScript. A page that renders its content client-side shows up as `no_content`, or its selector as `selector_not_found`.
- Only text is watched. An image, a style or a link target that changes without its text changing is not reported.
- Text that differs on every load (a live counter, a "posted 3 minutes ago" line) reports a change every run. Point `selector` at the part you care about to leave it out.
- If a run fails after a snapshot is saved but before its row is written, that one change is not reported; the next run compares against the new snapshot.
- Pages are fetched one after another. If the run reaches its timeout it stops there; the rows already written stay in the dataset.

### Data sources

The pages themselves, fetched over HTTP(S). No third-party service, and nothing leaves your Apify account.

# Changelog

This Actor's version history is a separate document: https://apify.com/pontio/website-change-monitor/changelog.md

# Actor input Schema

## `urls` (type: `array`):

Pages to watch, as http(s) URLs (max 1000 per run). A repeated URL is checked and charged once.

## `selector` (type: `string`):

CSS selector for the part of each page to watch, e.g. #pricing or main .content. Defaults to the whole page body. Applies to every URL in the run.

## `monitorName` (type: `string`):

Keeps this monitor's snapshots apart from other monitors on the same input (lowercase letters, digits, hyphens). Change it to start over from a fresh baseline. Defaults to one shared set of snapshots.

## Actor input object example

```json
{
  "urls": [
    "https://example.com"
  ]
}
```

# Actor output Schema

## `results` (type: `string`):

One item per input, in this run's default dataset.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        "https://example.com"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("pontio/website-change-monitor").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "urls": ["https://example.com"] }

# Run the Actor and wait for it to finish
run = client.actor("pontio/website-change-monitor").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    "https://example.com"
  ]
}' |
apify call pontio/website-change-monitor --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,pontio/website-change-monitor"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/q4zDJjmPEfjYSfYph/builds/r0BmibPeReAkFnZXM/openapi.json
