Get started
Product
Back
Start here!
Ready-to-run tools for your AI agents and apps. Just pick one and go.
Browse 81,977 Actors
Apify platform
Apify Store
Actors for any job on the web
Actors
Build and run serverless programs
Integrations
Connect with apps and services
MCP
Give your AI access to Actors
Anti-blocking
Scrape without getting blocked
Proxy
Rotate scraper IP addresses
Open source
Crawlee
Web scraping and crawling library
Solutions
MCP server configuration
Configure your Apify MCP server with Actors and tools for seamless integration with MCP clients.
Start building
Apify for
Enterprise
Startups
Universities
Nonprofits
Use cases
Data for generative AI
Data for AI agents
Lead generation
Market research
View more →
Consulting
Professional services
Apify Partners
Developers
Documentation
Full reference for the Apify platform
Actor templates
Python, JavaScript, and TypeScript
Web scraping academy
Courses for beginners and experts
Monetize your code
Publish your Actors and get paid
Learn
API reference
CLI
SDK
Earn from your code
$1.6M paid out last month. Many developers earn over $3k.
Start earning now
Resources
Help and support
Advice and answers about Apify
Actor ideas
Get inspired to build Actors
Changelog
See what’s new on Apify
Customer stories
Find out how others use Apify
Company
About Apify
Contact us
Blog
Live events
Doing Good
Jobs
We're hiring!
Join our Discord
Talk to other builders
Pricing
Contact sales
Web Text Extractor
from $10.00 / 1,000 results
rl1987/web-text-extractor
Rating
0.0
(0)
Developer
R.L.
Actor stats
0
Bookmarked
42
Total users
Monthly active users
a month ago
Last modified
Categories
Developer tools
Share
scrapesage/ocr-text-extractor
Extract text from images (PNG, JPG, TIFF, WEBP, BMP, GIF) and scanned PDFs with Tesseract OCR in 28 languages. Per-page text, confidence scores, optional word boxes, automatic text-layer detection for born-digital PDFs. No API key, no browser. JSON, CSV, Excel.
Scrape Sage
5
simple.actor/page-text-reader
Scrape any web page as clean text: the full visible prose, with scripts, navigation, headers, footers, booking widgets and cookie banners stripped out. Discover mode also returns every same-site link with its anchor text. JavaScript is rendered only when a page needs it.
Simple Actor
3
automa-flow/ai-schema-web-extractor
Extract schema-validated JSON from URLs with your own LLM key. Starts on fast HTTP and renders JavaScript only when the page needs it, with source evidence for the fields.
Vadim Bezrukov
2
software_mechanics/page-aware-pdf-text
Extract existing text from batches of public HTTPS PDFs. Get page-by-page text, page numbers, title and author, plus clear results for files with no text or download errors. Up to 25 PDFs per run; 10 MiB and 100 pages per PDF. No OCR.
Software Mechanics
subimpact/rag-web-browser-lite
Web search and fetch tool for AI agents and RAG pipelines. Queries Google Search, scrapes the top N pages over raw HTTP, and returns clean Markdown. Fixed $0.002 per page — no surprise compute bills.
subimpact
alleserojje/web-page-to-markdown
Convert any web page into clean, readable Markdown or plain text — ads, navigation and boilerplate stripped out. Built for feeding AI agents, LLMs and RAG pipelines. Pay only per page.
Pedro Resende
unbrowseai/web-page-to-markdown
Turn any list of URLs into clean Markdown, text, HTML and links for LLMs and RAG. Renders JavaScript when needed, returns title, description and metadata. Pay per page, failed pages are free.
Unbrowse AI
drain54/web-to-markdown
Extract clean, LLM-ready Markdown, titles, and metadata from any public webpage. Strips ads, navigation, and boilerplate.
Indra Darmawan
rl1987/idealista-api-scraper
Scrape Idealista real-estate listings (ES/IT/PT/FR) directly from the Idealista mobile API — price, size, rooms, location, coordinates, property type, URL and the seller/agency CONTACT (commercial name + phone number).
accountable_eel/pdf-text-extractor
Extracts text or Markdown from PDFs by URL, for a list of PDF URLs from a crawl, a Sheet, or a CRM export. Rebuilds paragraphs and headings from PDF fonts, with best-effort tables, then feeds a RAG pipeline, an LLM prompt, or a row to Google Sheets. No OCR: scanned PDFs are flagged, not charged.
Adrian Voss
Description
JSON example
Start URLs
start_urls
Required
URLs to start with
Mode
mode
Optional
Text extraction mode