Universal Web Scraper
Pricing
from $1.00 / 1,000 page-scrapeds
Pricing
from $1.00 / 1,000 page-scrapeds
Rating
0.0
(0)
Developer
Nalintha Alwis
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
5 days ago
Last modified
Categories
Share
🌐 Universal Web Scraper (Fast, Parallel & AI-Ready)
The Universal Web Scraper is a high-speed, enterprise-grade scraper built on Crawlee and Apify. It extracts clean, structured data from any website on the internet with blazing-fast parallel concurrency, dynamic JavaScript rendering, and AI-ready Markdown formatting.
Powered by a transparent Pay-Per-Event (PPE) pricing model, you only pay for the exact pages and features you scrape—no subscription lock-in.
⚡ Key Features
- 🚀 High Concurrency & Parallelism: Configure parallel workers (1 to 100) to scrape hundreds of pages in seconds without bottlenecks.
- 🤖 AI & LLM Ready Markdown: Strips clutter, headers, footers, cookie notices, and ads to produce clean GitHub-Flavored Markdown (GFM) ideal for RAG, vector embeddings, ChatGPT, and Claude.
- ⚙️ Dual-Engine Architecture:
- Cheerio Engine: Lightning-fast HTTP-based scraper for static, SSR, blog, e-commerce, and documentation sites.
- Playwright Engine: Full headless Chromium browser for JavaScript-heavy Single Page Applications (React, Vue, Angular, Next.js).
- 📊 Deep Structured Data Extraction:
- Schema.org JSON-LD: Automatically parses all embedded microdata and
application/ld+jsonblocks. - HTML Tables to JSON: Detects tables and transforms them into structured arrays and key-value objects.
- Schema.org JSON-LD: Automatically parses all embedded microdata and
- 📬 Automated Contact & Lead Discovery: Regex-powered extraction of business email addresses, telephone numbers, and 10+ social profile links (Twitter/X, LinkedIn, Facebook, Instagram, YouTube, GitHub, TikTok, Discord, Telegram, WhatsApp).
- 📸 Screenshots & PDF Generation: Capture high-resolution full-page screenshots (PNG) and printable PDF documents in Playwright mode, automatically uploaded to Apify Key-Value Store.
- ⚡ Resource Optimization: Built-in request blocking in Playwright mode (blocks ads, analytics, fonts, and heavy media) for up to 10x faster execution and minimal memory usage.
- 🛡️ Anti-Block & Apify Proxy: Built-in support for Apify Residential and Datacenter proxies to bypass geo-restrictions, Cloudflare, and rate limits.
💰 Pay-Per-Event (PPE) Pricing
This Actor utilizes Apify's Pay-Per-Event (PPE) pricing model. You are charged strictly based on the operations performed:
| Event Name | Event Title | Price (USD) | Description |
|---|---|---|---|
page-scraped | Standard Web Page Scraped | $0.001 / page | Base charge per scraped web page ($1.00 per 1,000 pages). Includes title, text, status, and metadata. |
markdown-extracted | AI-Ready Markdown Extraction | $0.002 / page | Charged when content is converted into clean GitHub-Flavored Markdown for LLMs and vector search. |
structured-data | Deep Structured Data | $0.003 / page | Charged when Schema.org JSON-LD, parsed HTML data tables, or contact details are extracted. |
browser-rendered | Headless Browser JS Execution | $0.005 / page | Charged for dynamic SPA rendering with headless Chromium (Playwright engine only). |
screenshot-captured | High-Res Screenshot Captured | $0.005 / shot | Charged per full-page screenshot saved to Apify Key-Value Store. |
pdf-generated | Full Page PDF Document | $0.005 / PDF | Charged per full-page printable PDF generated and saved to Key-Value Store. |
💡 Cost Examples:
- 1,000 Standard Pages (Cheerio + Markdown): 1,000 × ($0.001 + $0.002) = $3.00 USD
- 100 Dynamic E-Commerce Product Pages (Playwright + Tables + JSON-LD): 100 × ($0.001 + $0.005 + $0.003) = $0.90 USD
📄 License
ISC © Nal1ntha