# Website Change Monitor — Hash & Text Diff (No Browser, No LLM) (`ingenious_quip_bxq/website-change-monitor`) Actor

Monitor URL lists for content changes: SHA-256 hash, difflib text diff, optional CSS selector. Chain from sitemap / url-status. HTTP only — no browser, no LLM. Failed checks free. 256 MB.

- **URL**: https://apify.com/ingenious_quip_bxq/website-change-monitor.md
- **Developed by:** [新世紀書僮](https://apify.com/ingenious_quip_bxq) (community)
- **Categories:** SEO tools, Developer tools, Automation
- **Stats:** 2 total users, 1 monthly users, 0.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.50 / 1,000 page checkeds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Website Change Monitor — Hash & Text Diff

**Watch a list of URLs for content changes — SHA-256 hash, optional CSS selector, and a `difflib` text diff. No browser. No paid LLM.**
Paste URLs, or chain from Sitemap URL Discovery / URL Status Checker datasets. Snapshots persist in a named (or ID) key-value store so scheduled runs can compare against the last check. Default memory: **256 MB**. Failed network checks are free by default.

### What you get

- 🔐 **Content hash** — SHA-256 of normalized page text (script/style stripped)
- 🧩 **Optional CSS selector** — hash only `main` / `article` / `#content`
- 📝 **Text diff** — stdlib `difflib` unified diff + lines added/removed (no LLM “explanation”)
- 🧺 **Persistent snapshots** — named KV store (or explicit store ID) across runs
- 🔌 **Chain-friendly** — same `{ "url": ... }` shape as Sitemap / URL-status outputs
- 💸 **PPE** — low start + per successful check + extra only when hash changes; failures free by default
- 💾 **Light** — HTTP GET only, target ≤512 MB

### Measured results

Local + private cloud benches (2026-10-02 Asia/Taipei). PPE locked after cloud margin benches — see `docs/PRICING.md`.

| Test | Result |
|---|---|
| Local smoke (2 URLs, first / second / forced-change) | first×2 → unchanged×2 → changed×1 + `page-changed` charge; peak **78 MB** |
| Local errors (1 ok + 1 bad DNS) | charged **1** check only (fail free); peak **78 MB** |
| Cloud smoke `rK92EHmHttVUrjPhY` / `dncJ8mOaNHB5OgzUh` (2 URLs, 256 MB, build **0.1.2**) | **SUCCEEDED**; peak **86 MB**; settled **~$0.00037** |
| Cloud fail-free `ZOyhRn1x65VdFG73d` | check×1 charged; settled **$0.000120** |
| Cloud bench `P89Kp73Zbw8wYUv6e` (5 URLs) | peak **88 MB**; settled **$0.000492** (~$0.000098 / URL) |
| Cloud changed `h7LQ6fUdvbnNWOTl0` | check×2 + changed×1; peak **84 MB**; settled **$0.000341** |

### Use cases

- **Scheduled content watch** on a sitemap-derived URL list
- **Competitor / pricing page** change alerts (hash + diff, not marketing AI)
- **Re-crawl trigger** — only re-run doc/enrichment Actors when `changed=true`
- **Post-hygiene monitor** after URL Status Checker

### How to use

1. Add URLs, and/or a **Source dataset ID** from Sitemap / URL-status.
2. Set a **Snapshot store name** (or ID) you will reuse on every schedule.
3. Optional: CSS selector, ignore regexes, output filter `changed`.
4. Click **Start**. Dataset + `SUMMARY` / `OUTPUT` in the run KV; snapshots in the snapshot store.

#### Input example

```json
{
  "urls": [
    { "url": "https://example.com/" },
    { "url": "https://example.org/" }
  ],
  "snapshotStoreName": "website-change-snapshots",
  "storePreviousText": true,
  "maxUrls": 100,
  "outputFilter": "all"
}
```

Chain from Sitemap / URL-status:

```json
{
  "datasetId": "<upstream-run-default-dataset-id>",
  "snapshotStoreName": "my-site-watch",
  "cssSelector": "main",
  "outputFilter": "changed"
}
```

#### Output example (one dataset item)

```json
{
  "url": "https://example.com/",
  "finalUrl": "https://example.com/",
  "httpStatus": 200,
  "ok": true,
  "contentHash": "b94d27b9…",
  "previousHash": "a3f1…",
  "changed": true,
  "isFirstSeen": false,
  "diffSummary": {
    "isFirstSeen": false,
    "linesAdded": 2,
    "linesRemoved": 1,
    "unifiedDiff": "--- previous\n+++ current\n@@ …",
    "truncated": false
  },
  "checkedAt": "2026-10-02T01:00:00+00:00",
  "errorClass": null
}
```

#### Key-value store records (run)

| Key | Content |
|---|---|
| `SUMMARY` | Counts: totalChecked, changed, firstSeen, failed, charged, peakMemoryMb |
| `OUTPUT` | Same summary |

Snapshots are written to the **snapshot** store (not the run default KV), keyed as `snap:<sha256(url)>`.

### Pricing

Pay per event (private lock; see `docs/PRICING.md`):

| Event | Price |
|---|---|
| Actor start (Apify synthetic) | **$0.001** / GB (1 GB min) |
| Page checked (primary) | **$0.0015** per successful fetch+hash |
| Page changed | **$0.008** when hash differs from previous snapshot |

Failed checks (DNS / timeout / network) are **not charged** unless you enable `chargeFailedChecks`. First-seen pages charge `page-checked` only (not `page-changed`).

### Chaining

```
Sitemap URL Discovery → URL Status Checker → Website Change Monitor
                                              └─ on changed → re-run doc/enrichment
```

### Limits

- HTTP only — no headless browser / JS rendering (keeps memory and CU low).
- Dynamic pages that inject clocks/CSRF into the main text may need `ignorePatterns` or a tighter `cssSelector`.
- Very large HTML is truncated at `maxBodyBytes` (default 2 MB).

### License

Actor source: AGPL-3.0. Dependencies: `httpx` (BSD-3), `selectolax` (MIT), `apify` (Apache-2.0). See `docs/LICENSE_REVIEW.md`.

# Changelog

This Actor's version history is a separate document: https://apify.com/ingenious_quip_bxq/website-change-monitor/changelog.md

# Actor input Schema

## `urls` (type: `array`):

One or more URLs. Accepts {"url": "..."} objects (same shape as Sitemap / URL-status dataset rows). Plain strings and remote lists (requestsFromUrl) also work.

## `datasetId` (type: `string`):

Optional. Apify dataset ID from Sitemap URL Discovery or URL Status Checker. Reads the `url` field from each item.

## `keyValueStoreId` (type: `string`):

Optional. Load URLs from a KV record such as DOC_TO_MARKDOWN_INPUT.

## `keyValueRecordKey` (type: `string`):

Key inside the source key-value store. Ignored unless a store ID is set.

## `snapshotStoreId` (type: `string`):

Optional. Existing Apify key-value store ID that holds previous hashes/text. Prefer this for stable cross-run state. If empty, Snapshot store name is used.

## `snapshotStoreName` (type: `string`):

Named KV store used when Snapshot store ID is empty. Created if missing. Use one name per monitoring project.

## `storePreviousText` (type: `boolean`):

Save normalized text in the snapshot so the next run can emit a unified diff. Turn off to store hashes only (cheaper storage; hash change still detected).

## `maxStoredTextChars` (type: `integer`):

Cap for text kept in each snapshot (characters). Diffs use this stored text.

## `cssSelector` (type: `string`):

Optional. Restrict hashing/diff to matching nodes (e.g. main, article, #content). Empty = whole body after stripping script/style.

## `ignorePatterns` (type: `array`):

Regex patterns removed from text before hashing (timestamps, CSRF tokens, counters). Applied with MULTILINE.

## `maxUrls` (type: `integer`):

Stop after checking this many unique URLs (0 = no limit).

## `maxBodyBytes` (type: `integer`):

Truncate each response body after this many bytes before extraction.

## `outputFilter` (type: `string`):

Which rows to save. SUMMARY always covers every check. Successful checks are still charged when filtered out.

## `maxDiffLines` (type: `integer`):

Cap lines in diffSummary.unifiedDiff.

## `chargeFailedChecks` (type: `boolean`):

Default off: DNS/timeout/network failures are free. Turn on only if you want every attempted URL billed as page-checked.

## `maxConcurrency` (type: `integer`):

Parallel HTTP fetches.

## `requestTimeoutSecs` (type: `integer`):

Timeout for each HTTP GET.

## `userAgent` (type: `string`):

Optional custom User-Agent.

## `proxyConfiguration` (type: `object`):

Optional Apify Proxy if a site blocks data-center IPs.

## Actor input object example

```json
{
  "urls": [
    {
      "url": "https://example.com/"
    },
    {
      "url": "https://example.org/"
    }
  ],
  "keyValueRecordKey": "DOC_TO_MARKDOWN_INPUT",
  "snapshotStoreName": "website-change-snapshots",
  "storePreviousText": true,
  "maxStoredTextChars": 100000,
  "ignorePatterns": [],
  "maxUrls": 100,
  "maxBodyBytes": 2000000,
  "outputFilter": "all",
  "maxDiffLines": 40,
  "chargeFailedChecks": false,
  "maxConcurrency": 5,
  "requestTimeoutSecs": 30,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `results` (type: `string`):

Dataset items: url, httpStatus, contentHash, previousHash, changed, isFirstSeen, diffSummary, checkedAt, errorClass.

## `summary` (type: `string`):

KV SUMMARY: totalChecked, changed, firstSeen, failed, charged, durationSecs, peakMemoryMb.

## `output` (type: `string`):

KV OUTPUT: same as SUMMARY for sibling Actor consistency.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        {
            "url": "https://example.com/"
        },
        {
            "url": "https://example.org/"
        }
    ],
    "maxUrls": 100,
    "outputFilter": "all",
    "proxyConfiguration": {
        "useApifyProxy": false
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("ingenious_quip_bxq/website-change-monitor").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "urls": [
        { "url": "https://example.com/" },
        { "url": "https://example.org/" },
    ],
    "maxUrls": 100,
    "outputFilter": "all",
    "proxyConfiguration": { "useApifyProxy": False },
}

# Run the Actor and wait for it to finish
run = client.actor("ingenious_quip_bxq/website-change-monitor").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    {
      "url": "https://example.com/"
    },
    {
      "url": "https://example.org/"
    }
  ],
  "maxUrls": 100,
  "outputFilter": "all",
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}' |
apify call ingenious_quip_bxq/website-change-monitor --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,ingenious_quip_bxq/website-change-monitor"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/CdpRF02biK1VfZ25o/builds/gahKEsoZHxBZuuVIU/openapi.json
