PDF Change Monitor & Linked Document Tracker avatar

PDF Change Monitor & Linked Document Tracker

Pricing

from $5.00 / 1,000 document checkeds

Go to Apify Store
PDF Change Monitor & Linked Document Tracker

PDF Change Monitor & Linked Document Tracker

Monitor public PDFs even when their download URL changes. Give a stable webpage or a direct PDF URL and get typed link, content, page and metadata changes across runs, with last-good state that failures never overwrite.

Pricing

from $5.00 / 1,000 document checkeds

Rating

0.0

(0)

Developer

Vadim Bezrukov

Vadim Bezrukov

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

14 hours ago

Last modified

Share

Monitor public PDFs even when their download URL changes. Give the Actor a stable webpage or a direct PDF URL and receive structured link, content, page and metadata changes across runs. Every target keeps its own last-good state, so a temporary error never becomes a false "document removed" alert.

Who needs it

  • Compliance and regulatory teams watching guidance, forms, policies and standards that are republished under new file names.
  • Procurement and operations tracking supplier price lists, manuals, specifications and tender documents.
  • Legal and risk monitoring terms, privacy policies and public reports published as PDF.
  • Data and RAG pipelines that must re-ingest a document only when its text actually changed.

The stable page, changing PDF problem

The IRS "About Form 1040" page always links to the current Form 1040 PDF. When a new revision is published the page stays put while the file, and often the file URL, changes. A monitor keyed on the PDF URL loses the document; a monitor keyed on the page cannot tell you what changed inside the PDF.

This Actor resolves the intended PDF from the page on every run using your selection criteria (linkText, hrefRegex, cssSelector), fetches and parses it, and compares it with the last successful observation of the same logical document (targetId). A new URL with identical content is exactly one event: PDF_LINK_CHANGED.

Change types

changeTypes valueMeaning
BASELINEFirst successful observation of this targetId (stored; emitted unless baselineMode=silentBaseline).
PDF_ADDEDThe page previously had no matching link (deterministic LINK_NOT_FOUND) and now has one.
PDF_LINK_CHANGEDThe resolved PDF URL differs from the last good observation. Not a content claim.
FILE_CHANGEDThe PDF bytes differ (SHA-256). Emitted alone when the text layer is identical.
CONTENT_CHANGEDThe normalized text differs; changedPages and textDiff carry evidence.
PAGE_ADDED / PAGE_REMOVEDPage count increased / decreased.
METADATA_CHANGEDTitle, author, subject, keywords, creator, producer or dates differ.

Several change types can appear on one row. A URL-only change never produces CONTENT_CHANGED; a byte-only change (re-saved file, new producer stamp) never fabricates a semantic change.

30-second example

{
"monitorKey": "policy-watch",
"targets": [
{
"targetId": "irs-form-1040",
"label": "Current IRS Form 1040",
"pageUrl": "https://www.irs.gov/forms-pubs/about-form-1040",
"linkText": "Form 1040 PDF"
},
{
"targetId": "supplier-price-list",
"directPdfUrl": "https://supplier.example.com/files/price-list-2026.pdf"
}
],
"mode": "changesOnly"
}

First run: one BASELINE row per target. Later runs with the default mode: "all": one row per target, including NO_CHANGE, so every scheduled run documents what was verified. mode: "changesOnly" emits rows only for changed or unresolved targets (best for alert webhooks). Keep monitorKey and targetId stable; that pair is the document's identity. Last-good state lives in the named key-value store pdf-link-change-monitor-state in your account, not in the run's temporary store, so scheduled runs and Tasks share it.

Criteria intersect. With pageUrl and no criteria, the page must contain exactly one .pdf link.

  • linkText: exact visible anchor text, case and whitespace insensitive.
  • hrefRegex: regular expression against the absolute link URL, e.g. "f1040\\.pdf$".
  • cssSelector: descendant selector subset (tag, #id, .class, [attr=value], [attr$=value]) that must contain or match the anchor, e.g. "div.current-products a". Combinators >, +, ~, , and pseudo-classes are rejected explicitly.

Zero matches is LINK_NOT_FOUND; more than one distinct URL is AMBIGUOUS_LINK. Both rows list the candidates seen on the page so you can fix the criteria. The Actor never guesses.

Dataset output

One row per target per run. Key fields (see examples/sample_output.json):

FieldNotes
recordTypeOBSERVATION (verified document) or UNRESOLVED
statusBASELINE, NO_CHANGE, CHANGED, LINK_NOT_FOUND, AMBIGUOUS_LINK, FAILED, INVALID_INPUT
targetId, label, monitorKeyYour identity for the document
parentPageUrl, previousPdfUrl, currentPdfUrl, finalPdfUrl, matchedLinkTextLink resolution provenance; currentPdfUrl is the compared link, finalPdfUrl the redirect target that served the bytes
pageCount, previousPageCount, fileSizeBytes, contentType, httpEtag, httpLastModified, pdfVersionDocument facts
metadataPDF info dictionary fields when present
fileHash, contentHash, metadataHash, fingerprintSHA-256 fingerprints; fingerprint covers URL + file + content + metadata
textLayerAvailablefalse for scanned/image-only PDFs (file-level monitoring continues, no OCR)
changeTypes, changes, changedPages, textDiffTyped changes, field-level before/after, up to 50 changed pages with excerpts, bounded unified diff
candidates, error, warningsWhy a target is unresolved; non-fatal notes such as MIME_MISMATCH
observedAt, previousObservedAt, source, sourceId, sourceUrl, schemaVersionHistory-ready provenance

Dataset views: Changes, Current observations, Errors / unresolved targets. RUN_SUMMARY and CHECKS in the run's key-value store carry the per-run counts and per-target receipts.

Scheduling and webhooks

Create a Task from your input, schedule it (daily for regulatory sources, weekly for manuals) and add a webhook on ACTOR.RUN.SUCCEEDED that reads the default Dataset. In changesOnly mode an empty Dataset means "verified, nothing changed". Filter on status == "CHANGED" for alerts and on recordType == "UNRESOLVED" for configuration or source problems.

curl -X POST "https://api.apify.com/v2/acts/automa-flow~pdf-link-change-monitor/runs?token=$APIFY_TOKEN" \
-H "Content-Type: application/json" -d @input.json

Pricing

Pay per event, charged only for verified work:

EventPriceWhen
apify-actor-start$0.005Platform start fee, once per GB of run memory (1 event at the default 1024 MB)
document-checked$0.005One logical document fetched, validated, parsed and compared, including an unchanged check
pdf-page-processed$0.0001Each page whose text was extracted and fingerprinted (text extraction is the real compute cost)
pdf-mib-processed$0.0003Each started MiB downloaded for a successfully checked document (minimum 1)

Examples at 1024 MB: one 2-page tax form is $0.0105 per run; a 492-page, 6 MiB standard is $0.0611; 10 policy PDFs of 20 pages checked daily cost about $0.078 per run; 100 supplier documents averaging 20 pages and 3 MiB cost about $0.80 per run. Estimate before running: $0.005 x GB + documents x $0.005 + pages x $0.0001 + MiB x $0.0003, and set maxTotalChargeUsd accordingly. A run that cannot pay for one document check is rejected before any request as BUDGET_EXCEEDED.

Never charged: LINK_NOT_FOUND, AMBIGUOUS_LINK, failed fetches, invalid PDFs, oversized files, invalid input, retries and restarted runs. When the platform accepts fewer events than delivered, the run ends BILLING_LIMIT_REACHED with results and state retained and no rebilling.

Failure semantics

SituationRow statusState
Page fetched, deterministic zero matchLINK_NOT_FOUNDuntouched (remembered as absent only before any document was ever seen)
Several distinct links matchAMBIGUOUS_LINKuntouched
PDF URL returns 404/410FAILED (PDF_HTTP_404)untouched; not a removal
Timeout, 429/5xx after 3 attempts, network errorFAILED (HTTP_503, NETWORK_ERROR, ...)untouched
HTML or challenge page served as PDFFAILED (NOT_A_PDF)untouched
pageUrl itself serves a PDFFAILED (PAGE_IS_PDF, use directPdfUrl)untouched
Malformed, encrypted, oversized or too-long PDFFAILED (PDF_PARSE_FAILED, PDF_ENCRYPTED, PDF_TOO_LARGE, PDF_PAGE_LIMIT)untouched
Unchanged documentNO_CHANGEuntouched, still a billed check

One failed target never affects the others. The run fails only when every target failed (SOURCE_FAILED) or was invalid (INVALID_INPUT), when state could not be committed after rows were delivered (STATE_PERSISTENCE_FAILED, nothing billed, next run re-emits) or on a billing problem (BILLING_UNCERTAIN, BILLING_LIMIT_REACHED). A restarted run never repeats Dataset rows or charges.

  • HTTP only: no browser, no proxy, no CAPTCHA or login handling. Sources that require them are reported as FAILED (HTTP_403 and similar), never bypassed.
  • No OCR: scanned PDFs are monitored by bytes, page count and metadata with textLayerAvailable=false.
  • Retries: at most 3 attempts for 408/429/5xx and network errors with exponential backoff, Retry-After honoured, 200 retries per run.
  • Bounds: 100 targets, maxPdfSizeMb 1-100, maxPagesPerPdf 1-2000, 8 MiB parent pages, 5 redirects, diff of 400 lines / 60 kB, 50 changed pages. Text above 2 MiB per document is compared by hash only (no unified diff).
  • Compute: text extraction is CPU-bound and Apify grants about one CPU core per 4096 MB. A 492-page standard took 152 s at the default 1024 MB (measured 2026-09-14). For baskets of long documents run with 4096 MB (same cost per page, four times faster) or raise the run timeout; parallelism is bounded automatically by memory and maxPdfSizeMb.
  • SSRF protection: only public http(s) URLs on ports 80/443; loopback, private, link-local, multicast, reserved and cloud metadata addresses are rejected at input, after DNS resolution and on every redirect.
  • You are responsible for having the right to access and process each source; public availability does not override a publisher's terms. The Actor stores normalized text only to compute diffs for you and does not redistribute documents as a catalog. It collects no personal data beyond what the PDF metadata already exposes.

API and MCP

Run it from the Apify API, the JavaScript or Python clients, or from an AI agent through MCP: https://mcp.apify.com?tools=automa-flow/pdf-link-change-monitor. Example agent request: "Watch the PDF linked as 'Form 1040 PDF' on https://www.irs.gov/forms-pubs/about-form-1040 under targetId irs-form-1040 and tell me when its content or link changes." The agent passes that as one targets entry, reads status and changeTypes from the Dataset and RUN_SUMMARY.status from the key-value store; the input schema, Dataset schema and this README use the same vocabulary, so no extra mapping is needed. Runs require the caller's own Apify authentication and are billed to that account.