Universal Web Scraper
Pricing
from $1.00 / 1,000 page-scrapeds
Select 'Cheerio' for ultra-fast parallel HTTP scraping of static/SSR pages, or 'Playwright' for full headless Chromium browser execution with dynamic JavaScript rendering.
Maximum number of parallel requests/pages running simultaneously. Higher values speed up crawling significantly.
Maximum total number of pages to process before the actor finishes.
If enabled, crawler will discover and follow internal links on each scraped page up to Max Crawl Depth.
Maximum link depth to follow when Crawl Links is enabled (0 = start URLs only, 1 = direct links, etc.).
Convert page content into clean, clutter-free GitHub-Flavored Markdown for LLMs, RAG, and AI vector search.
Extract page title, description, OpenGraph tags (og:*), Twitter cards, canonical URL, language, and favicon.
Parse all Schema.org structured data scripts (application/ld+json) embedded in the page.
Extract clean main body text, word count, character count, and estimated reading time.
Auto-detect and convert HTML
elements into structured JSON objects and rows.Automatically detect email addresses, telephone numbers, and social media profile URLs (Twitter, LinkedIn, Facebook, Instagram, YouTube, GitHub).
Extract list of all images (src, alt, dimensions), videos, audio, and downloadable document links.
Include a list of all internal and external hyperlinks found on the page with anchor text.
Include the full unparsed raw HTML source code in each record.
Capture high-resolution full-page screenshot and save to the default Key-Value Store.
Generate clean printable PDF document of the page and save to Key-Value Store.
In Playwright mode, blocks images, videos, web fonts, and advertising scripts for ultra-fast, smooth page loads.
Optional CSS selector to wait for before extracting data in Playwright mode (e.g. '.product-list', '#main-content').
Configure Apify Proxy or custom proxies to bypass rate limits, geo-restrictions, and anti-scraping protections.
{ "useApifyProxy": false}