# Website Content Change Checks (`jj_toolworks/website-content-change-checks`) Actor

Compare current static page content with explicit customer-supplied snapshots. Get stable hashes, bounded change evidence and a new snapshot without treating fetch failures as deletions.

- **URL**: https://apify.com/jj_toolworks/website-content-change-checks.md
- **Developed by:** [JJ Toolworks](https://apify.com/jj_toolworks) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$4.00 / 1,000 completed change checks

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Website Content Change Checks

Check a supplied list of static public pages against explicit previous snapshots. Get current hashes, changed/unchanged results, bounded text evidence and a fresh snapshot. Failed fetches never become deletion claims.

### First run

```json
{
  "pages": [{"url": "https://example.com/", "recordId": "example-domain"}],
  "baseline": {"records": []},
  "preset": "general"
}
```

The first successful check returns `initialized`, with current content for your next run. Copy the complete `SNAPSHOT` JSON object into the next input's `baseline`. An external workflow can retrieve that file and supply it on a later scheduled run. This Actor does not silently read or overwrite a shared cross-customer store, send messages, or set up a schedule.

You can also use a compatible `SNAPSHOT` from JJ Toolworks Bulk Website Content Export. Keep `preset`, `contentSelector`, `excludeSelectors` and `minTextCharacters` identical. The extraction profile includes these settings and the extractor version.

### Baseline contract

`baseline.records` contains up to 100 records, keyed by normalized `sourceUrl`. A valid record needs `complete: true`, lowercase SHA-256 `contentHash`, and the original `extractionProfile`. Include `contentText` for before/after line evidence and `fetchedAt` for the evidence time. The hash must match the exact supplied text. Hash-only records can identify a change but cannot provide removed text. Conflicting records for one source URL are flagged instead of selecting one silently.

A baseline can contain 1 MiB text per record and 10 MiB text total. The tool validates these limits before fetching. Different original URLs remain distinct snapshot keys even when they redirect to the same final URL.

### Results

Filter `recordType = change_check` in the dataset.

| Status | Meaning |
|---|---|
| initialized | Current extraction completed; no prior baseline for this URL |
| unchanged | Complete current content and compatible prior hashes match |
| changed | Complete current content and compatible prior hashes differ |
| baseline_incompatible | Current extraction completed but the prior extraction profile differs; no change conclusion |
| baseline_invalid | Current extraction completed but prior evidence fails validation; no change conclusion |
| current_unavailable | Fetching or extraction failed, hit a cap, or produced insufficient content; no change/deletion conclusion |
| invalid_input | The supplied URL is unsupported |

Successful unique checks include `currentContentText`, current/previous hashes, source/final URL, evidence times and extraction profile. Changed compatible text baselines can include added/removed line excerpts. Excerpts are capped at 200 lines of 500 characters per direction. Very large comparisons omit line evidence while retaining full hash comparison and the available current content; this is visible in `diffMeaning`.

Every input is represented in the `SNAPSHOT` file. When a current check fails and valid prior evidence exists, that prior evidence is retained with `carriedForward: true`, its original fetched time, a new `lastCheckedAt`, and the failed `currentStatus`. It is not freshly fetched content. Duplicate snapshot records omit repeated text but preserve their hashes/profiles for a later hash comparison.

No page is considered deleted because it is absent from the input or unavailable. This is content evidence within your selected page section, not an interpretation of pricing, contracts, opportunities or business significance.

### Pricing and limits

Configured price: **$4 per 1,000 completed page checks** ($0.004 each). A first successful run without a baseline is a billable `initialized` snapshot. With a supplied valid compatible baseline, completed `unchanged` or `changed` checks are billable. Invalid, corrupt, conflicting or profile-incompatible supplied baselines do not create a success event, even if current content was fetched and returned. Failed/empty/truncated current pages and duplicate final URL/profile checks are also unbilled. `complete` describes the comparison or initialization; `currentComplete` separately describes the current extraction.

Up to 100 pages/run; 2 MiB fetched HTML/page; 1 MiB extracted text+Markdown/page; 30 seconds per request. A 12 MiB run budget for current content and bounded evidence produces visible `run_content_limit` statuses rather than silently dropping later pages. HTML-only; no browser, screenshots, OCR, automatic email or implicit monitoring state.

### Useful variations

Use an exact vendor-terms selector for purchasing review, a client article/main selector for agency copy, a notice section for human review, or a docs-content selector for software documentation. Exclude known rotating time widgets to reduce irrelevant changes. Exclusions are part of the comparison profile, so changing them requires a fresh compatible baseline.

Existing category reference: https://apify.com/jakubbalada/content-checker. See `MARKET.md` for the narrower buyer hypothesis and the limits of observed marketplace demand.

Text hashes normalize line whitespace and indentation. They can ignore a spacing-only source change; they are not HTML-byte, visual-layout or code-semantics hashes. Markdown retains preformatted code where available.

### Price and run controls

The launch price is **$4 per 1,000 completed units** ($0.004 per event), as defined above. Check the current Pricing tab before running. Status and duplicate rows do not add this custom event. Set a maximum charge appropriate for your batch; a run stops when its event budget is exhausted. `RUN-SUMMARY` records actual accepted events and execution totals.

Automatic replay of an interrupted run is disabled to prevent duplicate charges. Keep the available output, then start a new run for a new execution. A later run is a new billable execution. Completed dataset rows and individually saved files remain available when a budget stops a run. Combined exports, manifests, and snapshots assembled at the end may be absent after a budget stop or interruption; use a sufficient run budget when you need those aggregate files. The initial release uses documented input limits and public source access; external source changes can require maintenance.

# Actor input Schema

## `pages` (type: `array`):

1–100 public page URLs or {url, recordId} objects. An omitted page is never inferred to be deleted.

## `baseline` (type: `object`):

Optional SNAPSHOT object containing a records array. Records need sourceUrl, complete:true, contentHash and extractionProfile; contentText enables line evidence. Maximum 100 records and 10 MiB total baseline text.

## `preset` (type: `string`):

Prefers relevant main-content containers when no explicit selector is supplied. Preset is part of the extraction profile.

## `contentSelector` (type: `string`):

Optional exact section selector; maximum 500 characters. A missing match produces a free selector_not_found status.

## `excludeSelectors` (type: `array`):

Up to 20 CSS selectors to remove, such as .last-updated or .cookie-banner. These settings change the comparison profile.

## `minTextCharacters` (type: `integer`):

Minimum letters/digits after cleaning. Pages below this criterion are not success events.

## Actor input object example

```json
{
  "pages": [
    {
      "url": "https://example.com/",
      "recordId": "example-domain"
    }
  ],
  "baseline": {
    "records": []
  },
  "preset": "general",
  "excludeSelectors": [],
  "minTextCharacters": 40
}
```

# Actor output Schema

## `results` (type: `string`):

No description

## `summary` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "pages": [
        {
            "url": "https://example.com/",
            "recordId": "example-domain"
        }
    ],
    "baseline": {
        "records": []
    },
    "preset": "general"
};

// Run the Actor and wait for it to finish
const run = await client.actor("jj_toolworks/website-content-change-checks").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "pages": [{
            "url": "https://example.com/",
            "recordId": "example-domain",
        }],
    "baseline": { "records": [] },
    "preset": "general",
}

# Run the Actor and wait for it to finish
run = client.actor("jj_toolworks/website-content-change-checks").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "pages": [
    {
      "url": "https://example.com/",
      "recordId": "example-domain"
    }
  ],
  "baseline": {
    "records": []
  },
  "preset": "general"
}' |
apify call jj_toolworks/website-content-change-checks --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,jj_toolworks/website-content-change-checks"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/Y9hIgluVe8qZWyidT/builds/RArvReBPvzKtWtSpp/openapi.json
