Go to example tasks
Watch a PDF and get it only when it changes
One row per PDF: the text, the page count, the title, the author and the dates the file carries, plus a hash of the text so you can tell when it changes. The first run returns everything and remembers it; later runs return only what moved, with the previous hash and the change in length beside it. On a day when nothing changed, one row says so and nothing is charged.
PDF Text Extractor - Watch PDFs and Get Only What Changedneverempty/pdf-text-extractor-monitor
url
title
numPages
pagesExtracted
+26 fieldsTextNumberBooleanListObject
Input
PDF URLs:https://www.irs.gov/pub/irs-pdf/fw9.pdf+1
Maximum pages per PDF:0
Maximum PDFs per run:25
Maximum PDF size (MB):50
Timeout per PDF (seconds):30
Full text:true
Per-page text:false
Remove repeated headers and footers:true
Minimum characters:0
Keywords
Keyword match:any
Exclude keywords
Monitoring mode: return a PDF only when it has changed:true
Forget what was already returned:false
Max concurrency:5
Output fields
url
title
numPages
pagesExtracted
charCount
wordCount
contentHash
previousContentHash
charCountDelta
author
createdAt
modifiedAt
status
source
finalUrl
httpStatus
sizeBytes
subject
keywords
creator
producer
pdfVersion
isEncrypted
text
pages
previousCharCount
isFirstCheck
scrapedAt
pdfKey
reason
Sign up on Apify01
Create your Apify account to access the PDF Text Extractor - Watch PDFs and Get Only What Changed.
Start the run02
The Actor will start running based on the input automatically.
Receive the output03
Monitor the progress in real-time. You will be notified as soon as your dataset is complete and ready for review.
Integrate into your workflow04
The final output is delivered in JSON, CSV, or Excel format, ready to be plugged into your workflow.
