# Universal Web Scraper (`apificasion/universal-web-scraper`) Actor

Universal Web Scraping solution for your marketing and seo needs

- **URL**: https://apify.com/apificasion/universal-web-scraper.md
- **Developed by:** [Nalintha Alwis](https://apify.com/apificasion) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.00 / 1,000 page-scrapeds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

<p align="center">
  <img src="assets/logo.jpg" alt="Universal Web Scraper Logo" width="220" style="border-radius: 20px;" />
</p>

<h1 align="center">🌐 Universal Web Scraper (Fast, Parallel & AI-Ready)</h1>

<p align="center">
  <a href="https://apify.com"><img src="https://img.shields.io/badge/Apify-Actor-blue?logo=apify" alt="Apify Actor" /></a>
  <a href="https://apify.com"><img src="https://img.shields.io/badge/Pricing-Pay--Per--Event-green" alt="Pricing" /></a>
  <a href="https://crawlee.dev"><img src="https://img.shields.io/badge/Engine-Cheerio%20%7C%20Playwright-orange" alt="Engine" /></a>
  <a href="#-key-features"><img src="https://img.shields.io/badge/Output-Markdown%20%7C%20JSON--LD%20%7C%20Tables-purple" alt="AI Ready" /></a>
</p>

The **Universal Web Scraper** is a high-speed, enterprise-grade scraper built on [Crawlee](https://crawlee.dev) and [Apify](https://apify.com). It extracts clean, structured data from **any website on the internet** with blazing-fast parallel concurrency, dynamic JavaScript rendering, and AI-ready Markdown formatting.

Powered by a transparent **Pay-Per-Event (PPE)** pricing model, you only pay for the exact pages and features you scrape—no subscription lock-in.

***

### ⚡ Key Features

- 🚀 **High Concurrency & Parallelism**: Configure parallel workers (1 to 100) to scrape hundreds of pages in seconds without bottlenecks.
- 🤖 **AI & LLM Ready Markdown**: Strips clutter, headers, footers, cookie notices, and ads to produce clean GitHub-Flavored Markdown (GFM) ideal for RAG, vector embeddings, ChatGPT, and Claude.
- ⚙️ **Dual-Engine Architecture**:
  - **Cheerio Engine**: Lightning-fast HTTP-based scraper for static, SSR, blog, e-commerce, and documentation sites.
  - **Playwright Engine**: Full headless Chromium browser for JavaScript-heavy Single Page Applications (React, Vue, Angular, Next.js).
- 📊 **Deep Structured Data Extraction**:
  - **Schema.org JSON-LD**: Automatically parses all embedded microdata and `application/ld+json` blocks.
  - **HTML Tables to JSON**: Detects tables and transforms them into structured arrays and key-value objects.
- 📬 **Automated Contact & Lead Discovery**: Regex-powered extraction of business email addresses, telephone numbers, and 10+ social profile links (Twitter/X, LinkedIn, Facebook, Instagram, YouTube, GitHub, TikTok, Discord, Telegram, WhatsApp).
- 📸 **Screenshots & PDF Generation**: Capture high-resolution full-page screenshots (PNG) and printable PDF documents in Playwright mode, automatically uploaded to Apify Key-Value Store.
- ⚡ **Resource Optimization**: Built-in request blocking in Playwright mode (blocks ads, analytics, fonts, and heavy media) for up to 10x faster execution and minimal memory usage.
- 🛡️ **Anti-Block & Apify Proxy**: Built-in support for Apify Residential and Datacenter proxies to bypass geo-restrictions, Cloudflare, and rate limits.

***

### 💰 Pay-Per-Event (PPE) Pricing

This Actor utilizes Apify's **Pay-Per-Event (PPE)** pricing model. You are charged strictly based on the operations performed:

| Event Name | Event Title | Price (USD) | Description |
| :--- | :--- | :--- | :--- |
| `page-scraped` | Standard Web Page Scraped | **$0.001** / page | Base charge per scraped web page ($1.00 per 1,000 pages). Includes title, text, status, and metadata. |
| `markdown-extracted` | AI-Ready Markdown Extraction | **$0.002** / page | Charged when content is converted into clean GitHub-Flavored Markdown for LLMs and vector search. |
| `structured-data` | Deep Structured Data | **$0.003** / page | Charged when Schema.org JSON-LD, parsed HTML data tables, or contact details are extracted. |
| `browser-rendered` | Headless Browser JS Execution | **$0.005** / page | Charged for dynamic SPA rendering with headless Chromium (Playwright engine only). |
| `screenshot-captured` | High-Res Screenshot Captured | **$0.005** / shot | Charged per full-page screenshot saved to Apify Key-Value Store. |
| `pdf-generated` | Full Page PDF Document | **$0.005** / PDF | Charged per full-page printable PDF generated and saved to Key-Value Store. |

#### 💡 Cost Examples:

- **1,000 Standard Pages (Cheerio + Markdown)**: 1,000 × ($0.001 + $0.002) = **$3.00 USD**
- **100 Dynamic E-Commerce Product Pages (Playwright + Tables + JSON-LD)**: 100 × ($0.001 + $0.005 + $0.003) = **$0.90 USD**

***

### 📄 License

ISC © [Nal1ntha](https://github.com/Nal1ntha)

# Actor input Schema

## `startUrls` (type: `array`):

List of URLs to scrape. You can provide any website URL.

## `crawlerType` (type: `string`):

Select 'Cheerio' for ultra-fast parallel HTTP scraping of static/SSR pages, or 'Playwright' for full headless Chromium browser execution with dynamic JavaScript rendering.

## `maxConcurrency` (type: `integer`):

Maximum number of parallel requests/pages running simultaneously. Higher values speed up crawling significantly.

## `maxRequestsPerCrawl` (type: `integer`):

Maximum total number of pages to process before the actor finishes.

## `crawlLinks` (type: `boolean`):

If enabled, crawler will discover and follow internal links on each scraped page up to Max Crawl Depth.

## `maxCrawlDepth` (type: `integer`):

Maximum link depth to follow when Crawl Links is enabled (0 = start URLs only, 1 = direct links, etc.).

## `extractMarkdown` (type: `boolean`):

Convert page content into clean, clutter-free GitHub-Flavored Markdown for LLMs, RAG, and AI vector search.

## `extractMetadata` (type: `boolean`):

Extract page title, description, OpenGraph tags (og:\*), Twitter cards, canonical URL, language, and favicon.

## `extractSchema` (type: `boolean`):

Parse all Schema.org structured data scripts (application/ld+json) embedded in the page.

## `extractArticles` (type: `boolean`):

Extract clean main body text, word count, character count, and estimated reading time.

## `extractTables` (type: `boolean`):

Auto-detect and convert HTML <table> elements into structured JSON objects and rows.

## `extractContacts` (type: `boolean`):

Automatically detect email addresses, telephone numbers, and social media profile URLs (Twitter, LinkedIn, Facebook, Instagram, YouTube, GitHub).

## `extractMedia` (type: `boolean`):

Extract list of all images (src, alt, dimensions), videos, audio, and downloadable document links.

## `extractLinks` (type: `boolean`):

Include a list of all internal and external hyperlinks found on the page with anchor text.

## `rawHtml` (type: `boolean`):

Include the full unparsed raw HTML source code in each record.

## `captureScreenshot` (type: `boolean`):

Capture high-resolution full-page screenshot and save to the default Key-Value Store.

## `generatePdf` (type: `boolean`):

Generate clean printable PDF document of the page and save to Key-Value Store.

## `blockMedia` (type: `boolean`):

In Playwright mode, blocks images, videos, web fonts, and advertising scripts for ultra-fast, smooth page loads.

## `waitForSelector` (type: `string`):

Optional CSS selector to wait for before extracting data in Playwright mode (e.g. '.product-list', '#main-content').

## `proxyConfiguration` (type: `object`):

Configure Apify Proxy or custom proxies to bypass rate limits, geo-restrictions, and anti-scraping protections.

## `debugLog` (type: `boolean`):

Turn on detailed debug logs in the run console for troubleshooting.

## Actor input object example

```json
{
  "startUrls": [
    {
      "url": "https://news.ycombinator.com"
    }
  ],
  "crawlerType": "cheerio",
  "maxConcurrency": 10,
  "maxRequestsPerCrawl": 20,
  "crawlLinks": false,
  "maxCrawlDepth": 1,
  "extractMarkdown": true,
  "extractMetadata": true,
  "extractSchema": true,
  "extractArticles": true,
  "extractTables": true,
  "extractContacts": true,
  "extractMedia": false,
  "extractLinks": false,
  "rawHtml": false,
  "captureScreenshot": false,
  "generatePdf": false,
  "blockMedia": true,
  "proxyConfiguration": {
    "useApifyProxy": false
  },
  "debugLog": false
}
```

# Actor output Schema

## `scrapedPages` (type: `string`):

Complete dataset of all scraped web pages including Markdown, JSON-LD, tables, and contacts

## `overview` (type: `string`):

Quick overview of URLs, titles, HTTP status, and response times

## `aiContent` (type: `string`):

Clean GitHub-Flavored Markdown and articles formatted for LLMs and RAG

## `contacts` (type: `string`):

Extracted email addresses, telephone numbers, and social links

## `structuredData` (type: `string`):

Schema.org JSON-LD structured data and parsed HTML tables

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "startUrls": [
        {
            "url": "https://news.ycombinator.com"
        }
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("apificasion/universal-web-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "startUrls": [{ "url": "https://news.ycombinator.com" }] }

# Run the Actor and wait for it to finish
run = client.actor("apificasion/universal-web-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "startUrls": [
    {
      "url": "https://news.ycombinator.com"
    }
  ]
}' |
apify call apificasion/universal-web-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,apificasion/universal-web-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/wmgWse8EFWLV26fGO/builds/oi02dExhu1wkd8Nyw/openapi.json
