Get started
Product
Back
Start here!
Ready-to-run tools for your AI agents and apps. Just pick one and go.
Browse 59,444 Actors
Apify platform
Apify Store
Actors for any job on the web
Actors
Build and run serverless programs
Integrations
Connect with apps and services
MCP
Give your AI access to Actors
Anti-blocking
Scrape without getting blocked
Proxy
Rotate scraper IP addresses
Open source
Crawlee
Web scraping and crawling library
Solutions
MCP server configuration
Configure your Apify MCP server with Actors and tools for seamless integration with MCP clients.
Start building
Apify for
Enterprise
Startups
Universities
Nonprofits
Use cases
Data for generative AI
Data for AI agents
Lead generation
Market research
View more →
Consulting
Apify Professional Services
Apify Partners
Developers
Documentation
Full reference for the Apify platform
Actor templates
Python, JavaScript, and TypeScript
Web scraping academy
Courses for beginners and experts
Monetize your code
Publish your Actors and get paid
Learn
API reference
CLI
SDK
Earn from your code
$1.4M paid out last month. Many developers earn over $3k.
Start earning now
Resources
Help and support
Advice and answers about Apify
Actor ideas
Get inspired to build Actors
Changelog
See what’s new on Apify
Customer stories
Find out how others use Apify
Company
About Apify
Contact us
Blog
Live events
Partners
Jobs
We're hiring!
Join our Discord
Talk to other builders
Pricing
Contact sales
Document Text Extractor
Pay per event
automation-lab/layout-aware-text-extractor
Extract ordered plain text from public PDF, image, and web page URLs while preserving readable headings, lists, tables, page breaks, metadata, and quality warnings.
Rating
0.0
(0)
Developer
Stas Persiianenko
Actor stats
0
Bookmarked
3
Total users
2
Monthly active users
2 days ago
Last modified
Categories
Developer tools
AI
Share
fetchbase/document-to-markdown
Convert PDF and Word (DOCX) documents into clean Markdown, text, or JSON. Smart PDF paragraph reflow, page markers for RAG citations, full DOCX structure (headings, lists, tables), custom auth headers. No browser — parses in seconds. Charged per page processed — no startup fee.
Fetchbase
vivianferreira/pdf-page-splitter
Split any PDF into individual pages instantly. Extract all pages, specific pages (1,3,5), or ranges (1-5). Handles up to 50,000 pages. Flat $0.005 per run. Perfect first step for document processing pipelines — chain with OCR, table extraction, and text analysis actors.
Vivian Ferreira
automation-lab/pdf-text-extractor
Extract text, metadata, and page-by-page content from PDF files. Provide PDF URLs and get structured JSON with full text, per-page text, page count, author, title, creation date, and more. Export as JSON, CSV, or Excel. No browser or proxy needed.
261
jungle_synthesizer/dea-arcos-prescriber-crawler
Crawl DEA administrative actions, registration revocations, suspensions, and controlled-substance applications from the Federal Register API. Extracts DEA registration numbers, action types, registrant names, and states for opioid litigation, pharma sales, and compliance.
BowTiedRaccoon
andok/pdf-text-converter
Convert bulk PDF documents via URL into clean, raw text. The perfect document scraper for LLMs, vector databases, and RAG pipelines.
Andok
26
ely_source/pdf-ocr
Read scanned PDFs and get clean text per page, with a confidence score for each one. Pages that already carry a text layer are read directly and never billed as OCR, so mixed batches come back faster, more accurately and cheaper.
Alexandre Leclerc
lukaskrivka/rust-input-function-example
Dynamically compile and run input-provided page function. Like Cheerio Scraper but in Rust.
Lukáš Křivka
5
automation-lab/pdf-structured-table-extractor
Extract structured tables from public PDF URLs into JSON rows and coordinate-aware cells with page provenance, confidence scores, and parsing warnings.
lexis-solutions/german-insolvency-announcements-scraper
Scrape official German insolvency announcements from neu.insolvenzbekanntmachungen.de. Extract full texts, dates, courts & companies with filters for states, subjects & dates. Ideal for credit risk, BI & legal research.
Lexis Solutions
13
studio-amba/staatsblad-scraper
Scrape publications from the Belgian Official Gazette (Belgisch Staatsblad / Moniteur Belge). Laws, royal decrees, ministerial orders with NUMAC IDs and PDF links.
Studio Amba
silentflow/linkedin-ads-scraper
Scrape LinkedIn Ads Library without cookies or login. Extract ad creatives, targeting data, impressions, and advertiser info. Search by keywords, scrape specific ads, or get all ads from any company.
SilentFlow
19