Website Document Inventory & PDF Finder
Pricing
$3.00 / 1,000 unique document link discovereds
Website Document Inventory & PDF Finder
Crawl public websites for PDF, Word, Excel, CSV, PowerPoint, RTF, and OpenDocument links with source context and deterministic document categories.
Pricing
$3.00 / 1,000 unique document link discovereds
Rating
0.0
(0)
Developer
月 明
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
3 days ago
Last modified
Categories
Share
Build a clean inventory of public document links published by websites. The Actor crawls a bounded set of same-site HTML pages and returns one source-attributed record for every unique document it finds.
Why use it
Document links are often scattered across investor-relations pages, resource centers, support sections, public meeting pages, policy pages, and download libraries. This Actor turns those links into structured rows without downloading document contents.
Useful for investor-relations/public-report inventories, product manuals and datasheets, public forms and policies, RFP/tender discovery, research pipelines, and migration checks.
Supported document types
PDF, DOC/DOCX, XLS/XLSX, CSV, PPT/PPTX, RTF, ODT, ODS, and ODP.
Links can be recognized from the URL extension, the anchor's download filename, or a declared document MIME type.
Output
One dataset row is saved per unique document with:
- site
- documentUrl
- extension
- fileName
- anchorText
- category
- sourcePage
- sourceContext
- sameSite
- discoveryDepth
Category labels are deterministic keyword hints, using URL, anchor text, and nearby text. Categories include financial reports, manuals/datasheets, policy/legal files, RFP/tender files, agendas/minutes, forms/applications, brochures/catalogs, menus/price lists, sustainability/ESG documents, research/whitepapers, or other.
The default key-value store contains OUTPUT, a run summary with pages scanned/failed, documents found/saved, per-site errors, and stop reasons.
Example input
{"startUrls": [{"url": "https://www.w3.org/WAI/WCAG21/Techniques/pdf/PDF3"}],"maxPagesPerSite": 5,"maxDepth": 1,"maxDocumentsPerSite": 100,"extensions": ["pdf", "doc", "docx", "xls", "xlsx", "csv", "ppt", "pptx"],"respectRobotsTxt": true}
Bounded and safe by default
- Public HTTP(S) URLs and standard ports only.
- Credentials in URLs are rejected.
- DNS/IP checks reject non-public targets before requests and redirects.
- Redirects are capped and every destination is revalidated.
- robots.txt is respected by default.
- Same-site crawling only; linked documents are not fetched.
- Maximum 20 input sites, 50 HTML page attempts per site, crawl depth 2, and 1,000 unique documents per site.
- HTML response bodies are capped at 2 MB.
- No login, form submission, CAPTCHA bypass, document download, OCR, or private-network access.
Pricing
$0.003 per unique document saved across the entire run ($3 per 1,000 documents). Pages with no matching document do not create a paid dataset record, and failed page requests are not charged as document results.
Always check the live Pricing tab because Store pricing can change.
Limitations
The Actor reads server-rendered HTML only. It finds links rather than document contents, so JavaScript-only links, OCR, content validation, and authenticated/private documents are outside scope. Category labels are convenience filters, not legal or accounting conclusions.
Only process public websites you are permitted to access and follow applicable site terms, copyright, privacy, and data-use rules.
Start with HTML pages containing document links. Direct file URLs are rejected so supplied filenames are never counted as newly discovered results. Duplicate document URLs across input sites are saved and charged once per run. An all-page-failure run fails explicitly after saving its diagnostics.