Website Change Monitor — Page Diffs, Alerts & Webhooks avatar

Website Change Monitor — Page Diffs, Alerts & Webhooks

Pricing

from $1.20 / 1,000 change detecteds

Go to Apify Store
Website Change Monitor — Page Diffs, Alerts & Webhooks

Website Change Monitor — Page Diffs, Alerts & Webhooks

Watch any list of web pages on a schedule. Every run returns one row per page and, when a page changed, a row with the added and removed lines, a similarity score and your keyword hits. Title, links and status too. Optional webhook. No browser, no API key, no per-site setup.

Pricing

from $1.20 / 1,000 change detecteds

Rating

0.0

(0)

Developer

Insight Solutions

Insight Solutions

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

9 hours ago

Last modified

Share

Give it a list of web pages, put it on a schedule, and get back a row per page per run — plus a change row with the added and removed lines whenever a page's text changed. Every change row carries a similarity score, the share of the page that changed, and which of your keywords appeared or disappeared. Title, meta description, links and HTTP status can be watched too, and an optional webhook gets every change in one POST.

It reads the page the way the Website to Markdown Actor does — navigation, footers and cookie banners stripped — so a new menu item or a rotated banner is not a "change". No browser, no API key, no per-site setup: paste URLs and go. Pages that cannot be read come back as free, typed diagnostic rows, never as invoices.

At a glance

Input — this is the Store prefill; paste it and run:

{ "urls": ["https://docs.python.org/3/whatsnew/index.html", "https://www.federalreserve.gov/newsevents/pressreleases.htm",
"https://news.ycombinator.com/", "https://www.gov.uk/government/organisations/hm-treasury"],
"mode": "monitor", "firstRunBehavior": "baseline-only", "watch": ["text", "title"], "ignore": [],
"keywords": [], "minChangePercent": 0, "includeText": false, "maxTextChars": 20000, "webhookUrl": "",
"maxUrls": 100, "stateStoreName": "website-change-monitor-state", "maxConcurrency": 5,
"maxRunSecs": 240, "proxyConfiguration": { "useApifyProxy": true } }

Output — one page-check row per page per run, with changed, changedFacets, status, title, wordCount and the text hash; and one change row per change, with addedLines, removedLines, changePercent, similarity, keywordsAdded and keywordsRemoved (full list under Output reference). A page that could not be read comes back as a free diagnostic row (ok: false, errorType, error) instead of a charge.

Price — $0.40 per 1,000 page checks + $2.00 per 1,000 changes detected (+ $0.001 per run); diagnostics and empty runs free; no API key, no browser, limited permissions, works over the Apify MCP server (mcp.apify.com) and with x402 agentic payments.

From code — client.actor("insight.solutions/website-change-monitor").call(run_input={"urls": ["https://news.ycombinator.com/"]}) with apify-client, or POST https://api.apify.com/v2/acts/insight.solutions~website-change-monitor/run-sync-get-dataset-items.


What you get

A page changed since the last run: one change row (abridged — null columns left out; every row carries the same 49 columns):

{
"ok": true,
"rowType": "change",
"changeType": "text",
"url": "https://www.gov.uk/government/organisations/hm-treasury",
"status": 200,
"title": "HM Treasury",
"addedLines": ["- Spending Review 2026"],
"removedLines": ["- Spending Review 2025"],
"moreAdded": 0,
"moreRemoved": 0,
"addedChars": 22,
"removedChars": 22,
"similarity": 0.994,
"changePercent": 0.6,
"keywordsAdded": ["spending review"],
"keywordsRemoved": ["spending review"],
"previousHash": "1884a2f00dd1e7294fb0f6d52cc66693d30c69b66250a71a5be86f20bdfea45b",
"currentHash": "4852d9e2c94d9ee3d74e5658dfb61e2886befa54f824e7a40f8ba472926e63d7",
"previousCheckedAt": "2026-09-30T06:00:00.000Z",
"lastChangedAt": "2026-10-01T06:00:00.000Z",
"changeCount": 1,
"scrapedAt": "2026-10-01T06:00:00.000Z",
"source": "www.gov.uk",
"sourceUrl": "https://www.gov.uk/government/organisations/hm-treasury"
}

Next to it, the page's page-check row for the same run says "changed": true, "changedFacets": ["text"], "wordCount": 702, "linksCount": 76 and the new textHash. On a quiet day every page gets its page-check row with "changed": false and there are no change rows — so the dataset is never empty and you can see that every page was really checked.

(This row comes from the test suite: the real gov.uk capture, with the "Spending Review 2025" link edited to 2026, run with keywords: ["spending review", "budget"].)


Quick start

Watch a pricing page daily. Create a task with your URL, then Schedule → daily:

{ "urls": ["https://www.example.com/pricing"], "watch": ["text", "title"] }

The first run stores the page (one page-check row, isBaseline: true). Every run after that returns changed: false, or a text change row with the lines that moved.

Alert only when a keyword moves. Report a change only if one of your words is in what was added or removed:

{ "urls": ["https://www.example.com/terms", "https://www.example.com/pricing"],
"keywords": ["price", "fee", "discontinued"], "ignore": ["^Last updated"], "minChangePercent": 1 }

A change without any keyword in it is not written as a change row (and not billed as one), but the page-check row still says changed: true, so nothing is hidden.

Send changes to a webhook. Set webhookUrl and each run that finds at least one change POSTs one JSON document:

{
"actorRunId": "HG7ML7M8z78YcAPEB",
"runAt": "2026-10-01T06:00:00.000Z",
"urlsChecked": 4,
"changeCount": 2,
"changesIncluded": 2,
"changesDropped": 0,
"textDropped": false,
"datasetUrl": "https://api.apify.com/v2/datasets/<id>/items?clean=true&format=json",
"changes": [ { "rowType": "change", "changeType": "text", "url": "…", "addedLines": ["…"], "…": "…" } ]
}

changes holds the change rows exactly as they are in the dataset. The body is capped at 256 KB: past that the text column is dropped first, then rows from the end, and changesDropped says how many (they are all in the dataset). Zapier's and Make's catch-hook triggers take this JSON as it is. Slack's own incoming webhooks expect a text field this payload does not have, so route it through one of those, or use the Slack integration on the Actor's Integrations tab.


How the first run works

A monitor needs something to compare with, and on the first run there is nothing. So the first time a page is seen, the Actor stores it and writes one page-check row with isBaseline: true — billed as a page check, because the page was fetched and read like any other. No change rows are written unless you set firstRunBehavior: "emit-all", which also writes a new change row carrying the page's whole text (billed as a change).

From the second run on, each page is compared with its stored snapshot. Snapshots live in a named key-value store (stateStoreName, default website-change-monitor-state) that the Actor creates on its first run and re-opens on every run after it; an Actor's default store is new on every run, so it could not remember anything. With the prefill, the Hacker News front page changes on nearly every run; the Python "What's New" index and the gov.uk page change when those sites publish.

A page added to the list later gets its own baseline on its first run; the others carry on comparing.


Input reference

FieldTypeDefaultWhat it does
urlsstring[]—Pages to watch, http:// or https://. A missing scheme becomes https://, a #fragment is dropped, a ?query is kept. Duplicates are watched once.
modemonitor | snapshotmonitormonitor compares with the stored snapshot and stores the new one. snapshot fetches and extracts only: page-check rows, no comparison, the state store is never opened.
firstRunBehaviorbaseline-only | emit-allbaseline-onlyFirst sight of a page: store it and write a baseline page-check row; emit-all also writes a new change row with the full text.
watchstring[]["text"]What to compare: text, title, metaDescription, links, status, html. See What counts as a change.
ignorestring[][]Regular expressions; a text line matching any of them is left out before comparing. Plain patterns are case-insensitive; /pattern/flags sets flags. An invalid one is a free invalid-input row and is skipped.
keywordsstring[][]Case-insensitive. When set, a change row is written only if a keyword is in the added or removed lines.
minChangePercentnumber, 0–1000Skip text changes smaller than this share of the page (changePercent). Text facet only.
includeTextbooleanfalsePut the full compared text in the text column of page-check and change rows.
maxTextCharsinteger, 1,000–200,00020000How much of each page is kept and compared, cut at a whole line. Longer pages say truncated: true.
webhookUrlstring""POST every change row of the run to this URL, once, when there is at least one.
maxUrlsinteger, 1–1,000100Pages past this are not checked or charged; one free row says how many.
stateStoreNamestringwebsite-change-monitor-stateThe named key-value store for snapshots. One per watchlist.
maxConcurrencyinteger, 1–205Pages read at once. Requests to one host are at least one second apart regardless.
maxRunSecsinteger, 30–3,600240Wall-clock budget. Pages not reached in time get a free deadline row and keep their snapshots.
proxyConfigurationobject{ "useApifyProxy": true }Apify datacenter proxy by default. Switch to RESIDENTIAL for sites that refuse datacenter addresses (slower, and residential traffic costs more).

Output reference

Every row has the same 49 columns, in the same order, so a CSV export has one header whatever a run produced. A column that does not apply to a row is null. Three row types:

page-check — one per page per run that was fetched and compared (paid: page-check). url, finalUrl, redirectedTo, status, contentType, lastModified, etag, title, metaDescription, wordCount, textChars, truncated, textHash, titleHash, linksCount, changed, changedFacets, isBaseline, previousCheckedAt, lastChangedAt, changeCount (changes detected since the baseline), fetchMs, text (with includeText), outageGuard.

change — one per change (paid: change). changeType is one of:

changeTypeWhenWhat the row carries
textthe readable text changedaddedLines, removedLines (line-based longest-common-subsequence diff; at most 200 each, lines cut to 500 characters, moreAdded/moreRemoved count the rest), addedChars, removedChars, similarity (2 × common lines ÷ (lines before + lines after)), changePercent (100 × (1 − similarity)), keywordsAdded, keywordsRemoved, previousHash, currentHash
titlethe title changedpreviousTitle, title; the old and new titles as removedLines/addedLines
metathe meta description changedthe old and new description as removedLines/addedLines
linkssame-site links in the content were added or removedlinksAdded, linksRemoved, similarity over the two link sets
statusthe status or redirect target changedpreviousStatus, status, redirectedTo; e.g. HTTP 200 → HTTP 200 → https://…/new
htmlthe raw HTML's SHA-256 changedpreviousHash, currentHash
newfirst sight, with emit-allevery line as addedLines, and the whole text
unreachablea page that was readable last run could not be readstatus (the failing one, or null), previousStatus; the reason is on the diagnostic row beside it
restoreda page is back after being unreachablepreviousStatus, status

Every change row also has url, finalUrl, title, status, previousCheckedAt, lastChangedAt and changeCount, and text when includeText is on.

diagnostic — free, ok: false. errorType is one of:

errorTypeMeaning
invalid-inputan entry that is not an http(s) page URL, an invalid ignore pattern, or pages past maxUrls
robots-disallowedthe site's robots.txt disallows the URL for every crawler; it was not fetched
blocked401/403/429, an anti-bot interstitial, or a proxy refusal — after one retry from a fresh exit IP
httpany other non-2xx answer (status says which), a host that did not answer, or more than 3 redirects
timeoutno answer within 20 seconds
not-htmla 2xx that is not a web page — a PDF, an image, JSON; the error names the content type
too-largea body over 5 MB
deadlinethe run's maxRunSecs ran out before the page was read
budgetyour maximum run cost was reached before the page's rows could be delivered
state-lockedanother run is using the same state store; this run stopped before fetching anything
state-write-failedthe page's new snapshot could not be saved (the check itself was delivered and billed)
webhookthe webhook POST failed (the error names the host and the status), or webhookUrl is not usable
upstream-formata 2xx page whose HTML could not be read
run-failedthe Actor itself hit an unexpected error

Views in the Console: Changes, Checks and Diagnostics.


What counts as a change

Facets. watch picks which parts of a page are compared:

  • text — the readable content, normalised: whitespace collapsed per line, blank lines dropped. A line over 400 characters is split into sentences (and a run-on without sentences into word chunks whose boundaries depend only on the words themselves) so that one edit in a long paragraph is one changed line, not a rewritten paragraph. GitHub's release page, for one, arrives from the extractor as a single 10,000-character line; split, an edit to one release note is one line out of about 80.
  • title — the page title. metaDescription — the meta description (change type meta).
  • links — the set of same-site links in the readable content. Order does not matter; navigation links are not in it, for the same reason navigation text is not in text.
  • status — the HTTP status and, if the page redirects, where to.
  • html — the SHA-256 of the raw HTML. Every byte counts — a rotated nonce, a build ID, a timestamp in a comment — so this is noisy and off by default.

Every facet is stored on every run whether it is watched or not, so switching one on later compares with a real earlier reading instead of reporting every page as changed.

ignore patterns drop whole lines before hashing and comparing — the standard fix for "Last updated …", view counters and "posted 3 hours ago". They are applied to the stored text too, at compare time, so adding a pattern never reports the lines it now hides as removed, and removing one never reports them as added.

minChangePercent skips text changes smaller than that share of the page. The share is changePercent = 100 × (1 − similarity), where similarity counts lines: changing one line of a 100-line page is about 1 %. It applies to the text facet only — a new title or a new redirect target is news at any size.

keywords gate change rows: with keywords set, a row is written only if a keyword occurs in its added or removed lines (for links rows, in the link URLs). unreachable and restored rows are never keyword-gated, and an html row has no lines, so with keywords set it is never written.

A change that keywords or minChangePercent filtered out is still a change to the page: the page-check row says changed: true, the snapshot moves on, and changeCount counts it. It is not reported again on the next run.

Pages that go down. A 4xx/5xx, a timeout or a block is a free diagnostic row. If the page was readable on the last run, the run also writes one unreachable change; while it stays down, further failures are free. When it answers again, a restored row says so — and if its text changed while it was down, the text row comes with it. The stored text is never overwritten by a failure.

The outage guard. If more than half of the pages that were readable last run fail in the same run (with at least two such pages), that looks like a proxy or network fault rather than that many sites going down together. Those failures are held back: free diagnostic rows with outageGuard: true, no unreachable changes, snapshots kept. If the next run sees the same pages fail, it believes it and reports them.


No browser: what this Actor sees and what it does not

This Actor reads the HTML the server sends. It does not run JavaScript. If the text you care about is not in View Source, this Actor will not see it.

That is a real limit, and two pages in the test set show it. The OpenAI business pricing page renders its prices in the browser; the server sends three short lines, and that is what gets compared. The Federal Reserve press-release page in the prefill draws its list of releases with JavaScript; what the server sends — and what is compared — is the index of years and FOMC pages beside it. Both are checked correctly; they are just checked for what their servers publish. For pages like that, look for a server-rendered alternative (an archive page, a plain-HTML list, a print view) and watch that instead.


Pricing

Pay per event:

EventFREEBRONZESILVERGOLD
Run started (actor-start)$0.001$0.001$0.001$0.001
Page checked (page-check)$0.0004$0.0004$0.00032$0.00024
Change detected (change)$0.002$0.002$0.0016$0.0012

page-check is charged for every page that was fetched and compared, including its first (baseline) check. change is charged per change row, on top of the page check.

RunBills
The prefill, first run: 4 baselines$0.001 + 4 × $0.0004 = $0.0026
The prefill, a later run with 2 changes$0.001 + $0.0016 + 2 × $0.002 = $0.0066
50 pages hourly for a day, 4 changes1,200 × $0.0004 + 4 × $0.002 + 24 × $0.001 = $0.512 / day
50 pages daily for a month, 20 changes1,500 × $0.0004 + 20 × $0.002 + 30 × $0.001 = $0.67 / month

A page whose title and text both changed is two change rows and two charges; watch fewer facets if you only want one.

What you are never charged for

SituationBilled?
A page that answers 4xx/5xx, times out, is blocked, is not HTML or is over 5 MBNo. If it was readable last run, one unreachable change is billed, once
A page disallowed by robots.txtNo — it is not fetched
A page the run never reached (maxRunSecs, or your cost limit)No
Failures held back by the outage guardNo
Invalid URLs, invalid ignore patterns, the webhook POST, a failed webhookNo
A change filtered out by keywords or minChangePercentNo change charge (the page check is billed as usual)
A run whose URLs were usable but none could be checkedNothing at all, start fee included. It finishes SUCCEEDED with zero results, a status message that says so ("0 results. 3 diagnostic rows explain why — blocked (2), http (1). Nothing was charged.") and a free row per page saying why. It finishes FAILED only when there was nothing usable to attempt (no URL, or every entry invalid), the state store is locked by another run, or the Actor itself hit an error
A change already deliveredNo — a page's snapshot moves on only once its rows are delivered and paid for, so a change is billed exactly once

The start fee is charged only after a page has answered and a paid row has been written.


FAQ

How do I schedule it? Save your input as a task, then open Schedules and run the task hourly, daily or on any cron expression. Keep the same stateStoreName on every run of the schedule: that store is the monitor's memory.

Should each project have its own state store? Yes — one stateStoreName per watchlist. Two tasks with different URL lists can share a store (each page has its own record), but two schedules that run at the same time cannot, see below. A distinct name per project also keeps their snapshots' 90-day expiry independent.

Why was my second run refused with state-locked? Two runs sharing one state store would both compare with the same snapshots and then overwrite each other, reporting every change twice. So the run that finds the store locked stops at once with one free row, and finishes FAILED so that a scheduler notices. The lock expires with the holding run's maxRunSecs plus five minutes, so a crashed run cannot hold it forever. Give overlapping schedules different stateStoreName values.

How long are snapshots kept? As long as the page is in your list. A page that has not been in any run's input for 90 days has its snapshot deleted. Each snapshot also keeps the dates of its last 90 changes.

Does it follow robots.txt? Yes. It fetches each site's robots.txt once per run and obeys the User-agent: * group; a disallowed page is a free robots-disallowed row and is never requested. A robots.txt that is missing or cannot be read disallows nothing. Crawl-delay is not applied; the Actor's own one-second spacing per host is.

Why did my first run not report any changes? Because it had nothing to compare with — see How the first run works. Set firstRunBehavior: "emit-all" if you want the full text of every page up front.

Can I diff pages I only need once, without keeping state? Use mode: "snapshot": the text, title, hashes and links of each page, no comparison, and the state store is never opened.

Does it work from an AI agent? Yes. It runs with limited permissions and pay-per-event pricing and never enters Standby, so it works over the Apify MCP server and with x402 agentic payments.


Limitations

  • No browser. Text drawn by JavaScript is not seen (see above). There is no login, and no cookies are sent, so pages behind a sign-in are not monitored.
  • Blocks happen. Sites behind a bot filter can refuse Apify's datacenter addresses. The Actor rotates the exit IP once and then reports a free blocked row rather than hammering the site. The residential proxy group helps with some of those sites and not others, and costs more per page.
  • One request per second per host. Fifty pages on one site take at least fifty seconds, whatever maxConcurrency says; raise maxRunSecs for long single-site lists.
  • Long pages are compared on their first maxTextChars characters. A change further down is not seen, and the row says truncated: true. Raising the cap does not report the newly visible part as a change.
  • Changing maxTextChars, watch or the extraction changes what is compared. The Actor avoids reporting a false change for the cap and the ignore patterns; a site redesign is a real, large text change.
  • Very different versions are diffed approximately. Past 2,000 line edits (a complete rewrite of a long page) the diff matches lines as a multiset instead of computing the exact sequence, so a line that moved counts as unchanged.
  • At most 3 redirects are followed and bodies over 5 MB are not read.
  • The upstream format may change. Sites change their markup all the time — that is what this Actor is for — and a change of layout is reported as the text change it produces. Extraction is written against real captured pages and is tolerant of odd HTML, but an extraction that stops making sense comes back as a typed diagnostic row, free, saying what happened.

Use it from an AI agent, or from code

One JSON object in, one flat array out. The Integrations tab can also push results to Slack, a webhook, Zapier, Make, Google Sheets, Snowflake or BigQuery.

curl -X POST "https://api.apify.com/v2/acts/insight.solutions~website-change-monitor/run-sync-get-dataset-items?token=$APIFY_TOKEN" \
-H "Content-Type: application/json" \
-d '{"urls":["https://www.example.com/pricing"],"watch":["text","title"],"stateStoreName":"pricing-watch"}'
# pip install apify-client
from apify_client import ApifyClient
client = ApifyClient("<APIFY_TOKEN>")
run = client.actor("insight.solutions/website-change-monitor").call(run_input={
"urls": ["https://www.example.com/pricing", "https://www.example.com/terms"],
"watch": ["text", "title"],
"ignore": ["^Last updated"],
"keywords": ["price", "fee"],
"stateStoreName": "competitor-watch",
})
for row in client.dataset(run["defaultDatasetId"]).iterate_items():
if row["rowType"] == "change":
print(row["url"], row["changeType"], f'{row["changePercent"]}%', row["keywordsAdded"])
for line in row["addedLines"] or []:
print(" +", line)
for line in row["removedLines"] or []:
print(" -", line)
elif row["rowType"] == "diagnostic":
print("could not check", row["input"], row["errorType"], row["error"])

  • Public pages only. No login, no session, no cookies of anyone's. The Actor requests each page as an anonymous visitor would.
  • robots.txt is honoured for every monitored URL, and requests to one host are at least one second apart.
  • No personal data is extracted. A row carries a page's readable text, the lines that changed, and metadata about the page (status, title, hashes, links). Nothing is parsed out of the text: no names, emails or phone numbers are picked out or enriched. If a page you monitor publishes personal data, it is in that page's text as published, and you are responsible for whether you may keep it.
  • Webhook URLs are treated as secrets. Only the host is ever written to the log or to a row.
  • You are responsible for how you use the pages you monitor, including any terms of use that apply to them.

Our other Actors

Every Insight Solutions Actor is pay-per-result with no browser, no login and no API key, and every one of them returns free diagnostic rows instead of billing for failures. Prices are per 1,000 results.

Video, audio & social

News, documents & the web

Business, finance & jobs

Apps & games