Go to example tasks
Keep only PDFs that mention a word
One row per PDF: the text, the page count, the title, the author and the dates the file carries, plus a hash of the text so you can tell when it changes. The keywords are matched against the text, the title, the subject and the author. Nothing extra is downloaded to apply them: the filter runs on the text that was already extracted.
PDF Text Extractor - Watch PDFs and Get Only What Changedneverempty/pdf-text-extractor-monitor
url
title
numPages
pagesExtracted
+26 fieldsTextNumberBooleanListObject
Input
PDF URLs:https://www.irs.gov/pub/irs-pdf/fw9.pdf+2
Maximum pages per PDF:0
Maximum PDFs per run:25
Maximum PDF size (MB):50
Timeout per PDF (seconds):30
Full text:true
Per-page text:false
Remove repeated headers and footers:true
Minimum characters:0
Keywords:taxpayer
Keyword match:any
Exclude keywords
Monitoring mode: return a PDF only when it has changed:false
Forget what was already returned:false
Max concurrency:5
Output fields
url
title
numPages
pagesExtracted
charCount
wordCount
contentHash
previousContentHash
charCountDelta
author
createdAt
modifiedAt
status
source
finalUrl
httpStatus
sizeBytes
subject
keywords
creator
producer
pdfVersion
isEncrypted
text
pages
previousCharCount
isFirstCheck
scrapedAt
pdfKey
reason
Sign up on Apify01
Create your Apify account to access the PDF Text Extractor - Watch PDFs and Get Only What Changed.
Start the run02
The Actor will start running based on the input automatically.
Receive the output03
Monitor the progress in real-time. You will be notified as soon as your dataset is complete and ready for review.
Integrate into your workflow04
The final output is delivered in JSON, CSV, or Excel format, ready to be plugged into your workflow.
