Website Content Change Checks avatar

Website Content Change Checks

Pricing

$4.00 / 1,000 completed change checks

Go to Apify Store
Website Content Change Checks

Website Content Change Checks

Compare current static page content with explicit customer-supplied snapshots. Get stable hashes, bounded change evidence and a new snapshot without treating fetch failures as deletions.

Pricing

$4.00 / 1,000 completed change checks

Rating

0.0

(0)

Developer

JJ Toolworks

JJ Toolworks

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

a day ago

Last modified

Categories

Share

Check a supplied list of static public pages against explicit previous snapshots. Get current hashes, changed/unchanged results, bounded text evidence and a fresh snapshot. Failed fetches never become deletion claims.

First run

{
"pages": [{"url": "https://example.com/", "recordId": "example-domain"}],
"baseline": {"records": []},
"preset": "general"
}

The first successful check returns initialized, with current content for your next run. Copy the complete SNAPSHOT JSON object into the next input's baseline. An external workflow can retrieve that file and supply it on a later scheduled run. This Actor does not silently read or overwrite a shared cross-customer store, send messages, or set up a schedule.

You can also use a compatible SNAPSHOT from JJ Toolworks Bulk Website Content Export. Keep preset, contentSelector, excludeSelectors and minTextCharacters identical. The extraction profile includes these settings and the extractor version.

Baseline contract

baseline.records contains up to 100 records, keyed by normalized sourceUrl. A valid record needs complete: true, lowercase SHA-256 contentHash, and the original extractionProfile. Include contentText for before/after line evidence and fetchedAt for the evidence time. The hash must match the exact supplied text. Hash-only records can identify a change but cannot provide removed text. Conflicting records for one source URL are flagged instead of selecting one silently.

A baseline can contain 1 MiB text per record and 10 MiB text total. The tool validates these limits before fetching. Different original URLs remain distinct snapshot keys even when they redirect to the same final URL.

Results

Filter recordType = change_check in the dataset.

StatusMeaning
initializedCurrent extraction completed; no prior baseline for this URL
unchangedComplete current content and compatible prior hashes match
changedComplete current content and compatible prior hashes differ
baseline_incompatibleCurrent extraction completed but the prior extraction profile differs; no change conclusion
baseline_invalidCurrent extraction completed but prior evidence fails validation; no change conclusion
current_unavailableFetching or extraction failed, hit a cap, or produced insufficient content; no change/deletion conclusion
invalid_inputThe supplied URL is unsupported

Successful unique checks include currentContentText, current/previous hashes, source/final URL, evidence times and extraction profile. Changed compatible text baselines can include added/removed line excerpts. Excerpts are capped at 200 lines of 500 characters per direction. Very large comparisons omit line evidence while retaining full hash comparison and the available current content; this is visible in diffMeaning.

Every input is represented in the SNAPSHOT file. When a current check fails and valid prior evidence exists, that prior evidence is retained with carriedForward: true, its original fetched time, a new lastCheckedAt, and the failed currentStatus. It is not freshly fetched content. Duplicate snapshot records omit repeated text but preserve their hashes/profiles for a later hash comparison.

No page is considered deleted because it is absent from the input or unavailable. This is content evidence within your selected page section, not an interpretation of pricing, contracts, opportunities or business significance.

Pricing and limits

Configured price: $4 per 1,000 completed page checks ($0.004 each). A first successful run without a baseline is a billable initialized snapshot. With a supplied valid compatible baseline, completed unchanged or changed checks are billable. Invalid, corrupt, conflicting or profile-incompatible supplied baselines do not create a success event, even if current content was fetched and returned. Failed/empty/truncated current pages and duplicate final URL/profile checks are also unbilled. complete describes the comparison or initialization; currentComplete separately describes the current extraction.

Up to 100 pages/run; 2 MiB fetched HTML/page; 1 MiB extracted text+Markdown/page; 30 seconds per request. A 12 MiB run budget for current content and bounded evidence produces visible run_content_limit statuses rather than silently dropping later pages. HTML-only; no browser, screenshots, OCR, automatic email or implicit monitoring state.

Useful variations

Use an exact vendor-terms selector for purchasing review, a client article/main selector for agency copy, a notice section for human review, or a docs-content selector for software documentation. Exclude known rotating time widgets to reduce irrelevant changes. Exclusions are part of the comparison profile, so changing them requires a fresh compatible baseline.

Existing category reference: https://apify.com/jakubbalada/content-checker. See MARKET.md for the narrower buyer hypothesis and the limits of observed marketplace demand.

Text hashes normalize line whitespace and indentation. They can ignore a spacing-only source change; they are not HTML-byte, visual-layout or code-semantics hashes. Markdown retains preformatted code where available.

Price and run controls

The launch price is $4 per 1,000 completed units ($0.004 per event), as defined above. Check the current Pricing tab before running. Status and duplicate rows do not add this custom event. Set a maximum charge appropriate for your batch; a run stops when its event budget is exhausted. RUN-SUMMARY records actual accepted events and execution totals.

Automatic replay of an interrupted run is disabled to prevent duplicate charges. Keep the available output, then start a new run for a new execution. A later run is a new billable execution. Completed dataset rows and individually saved files remain available when a budget stops a run. Combined exports, manifests, and snapshots assembled at the end may be absent after a budget stop or interruption; use a sufficient run budget when you need those aggregate files. The initial release uses documented input limits and public source access; external source changes can require maintenance.