# Web Content Scraper (`stealth_mode/web-content-scraper`) Actor

Scrape web pages effortlessly with undetected browser technology. Extract HTML, markdown, JavaScript results, and cookies from any URL. Features configurable timeouts, proxy support, and retry logic — ideal for content aggregation, data extraction, and automated workflows without detection.

- **URL**: https://apify.com/stealth\_mode/web-content-scraper.md
- **Developed by:** [Stealth mode](https://apify.com/stealth_mode) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $2.00 / 1,000 results

This Actor is paid per event and usage. You are charged both the fixed price for specific events and for Apify platform usage.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Web Content Scraper: Extract HTML, Text & Structured Data from Any Website

***

### What Is Web Scraping?

Web scraping automates the extraction of information from websites. Instead of manually copying and pasting content, a scraper retrieves structured data at scale. This is essential for market research, content aggregation, competitive analysis, and data-driven decision-making. The **Web Content Scraper** removes the friction from this process, offering granular control over what data you collect and how the page loads.

***

### Overview

The **Web Content Scraper** is a versatile tool for extracting structured and unstructured content from web pages. It supports:

- Running custom JavaScript to extract dynamic data
- Waiting for specific page elements to load before collection
- Exporting multiple output formats (HTML, markdown, JSON)
- Collecting cookies for authenticated sessions
- Automatic retry logic for failed requests
- **Undetected browser technology** that bypasses anti-bot protections (Cloudflare, Akamai)
- Proxy support to avoid detection and IP blocking

Ideal users include:

- **Content aggregators** building automated publishing pipelines
- **Data analysts** extracting market intelligence from web sources
- **Developers** automating web data collection from protected websites
- **Researchers** gathering structured datasets from websites
- **Quality assurance teams** testing website rendering across scenarios
- **Enterprise teams** scraping heavily-protected sites requiring anti-bot circumvention

The scraper uses an undetected browser engine that mimics human behavior, making it difficult for security systems to identify it as automated traffic. If you encounter blocking or rate limits, enable proxy rotation in the Scrape options to distribute requests across different IP addresses.

***

### Input Configuration

The scraper accepts a JSON configuration object controlling how pages are loaded and what data is extracted:

```json
{
    "urls": [
        "https://www.markepear.dev/blog/developer-marketing-guide"
    ],
    "init_js_script": null,
    "cookies": null,
    "wait_element": true,
    "element_selector": "#id",
    "js_script": "async () => { return document.title; }",
    "js_timeout": 10,
    "get_js_result": true,
    "get_html": true,
    "get_content": true,
    "get_cookies": true,
    "ignore_url_failures": true,
    "proxy": null
}
```

#### Key Input Parameters

| Parameter | Type | Purpose |
|-----------|------|---------|
| `urls` | array | List of web page URLs to scrape. Supports bulk input. |
| `cookies` | array | Session cookies to send with requests (optional). Format: `[{"name": "cna", "value": "abc123", "domain": ".example.com", "path": "/", "expires": null}]` |
| `wait_element` | boolean | Pause execution until a specific element appears on the page (Timeout in 5s) (requires `element_selector`). Useful for waiting on dynamic content. |
| `element_selector` | string | CSS selector for the element to wait for (e.g., `#main-content`, `.article-body`). |
| `init_js_script` | string | JavaScript code to run *before* navigating to the URL. Use for injecting utilities or overriding browser properties. |
| `js_script` | string | JavaScript code to run *after* the page loads. Example: `async () => { return $('title').text(); }` |
| `js_timeout` | integer | Maximum seconds to wait for the JS script to complete. Default: `10`. |
| `max_retries_per_url` | integer | How many times to retry a failed URL. Default: `2`. |
| `proxy` | object | Proxy configuration (optional). Leave null to use direct connection or set `useApifyProxy: true` for residential/datacenter proxies. |
| `ignore_url_failures` | `boolean` | If `true`, the scraper continues running when a URL fails instead of stopping the entire run. Recommended for bulk jobs. Default: `true`. |

> **Tip:** Use `wait_element` when targeting pages with lazy-loaded content or single-page applications where elements render after initial HTML load.

***

### Output Format

Each URL produces a record with up to seven fields, depending on your configuration:

```json
{
  "url": "https://www.markepear.dev/blog/developer-marketing-guide",
  "js_result": "Developer marketing guide (by a dev tool startup CMO)",
  "html": "<!DOCTYPE html><!-- Last Published: Fri Jun 05 2026 12:33:40 GMT+0000 (Coordinated Universal Time) --> ...",
  "markdown": "Contents\n[What is developer marketing?](#what-is-developer-marketing) [How is marketing to developers different than just marketing?](#how-is-marketing-to-developers-different-than-just-marketing) [Best practices of marketing to developers](#best-practices-of-marketing-to-developers) [How to market to software developers with a plan?] ...",
  "content": "Contents\nDeveloper marketing guide (by a dev tool startup CMO)\nThis is a 6000-word guide to developer marketing written for practitioners by a practitioner.\nI helped grow a machine learning dev tool startup from 0 to a Series A and learned a lot about marketing to devs along the way...",
  "json_content": {
    "title": "Developer marketing guide (by a dev tool startup CMO)",
    "author": null,
    "hostname": "markepear.dev",
    "date": "2026-09-01",
    "categories": "",
    "tags": "",
    "fingerprint": "",
    "id": null,
    "license": null,
    "comments": "",
    "raw_text": "Contents Developer marketing guide (by a dev tool startup CMO) This is a 6000-word ...",
    "text": "Contents\nDeveloper marketing guide (by a dev tool startup CMO)\nThis is a 6000-word guide to developer marketing ...",
    "language": null,
    "image": "https://cdn.prod.website-files.com/6161939cdc6297e03f7803e0/65158b7336e68a13c8461f3d_Developer%20marketing%20guide%20(by%20a%20dev%20tool%20startup%20CMO).png",
    "pagetype": "website",
    "source": "https://www.markepear.dev/blog/developer-marketing-guide",
    "source-hostname": "markepear.dev",
    "excerpt": "Sep 01, 2026 - In 6000 words, I share what I learned about developer marketing from talking to hundreds of practitioners and years of growing a dev tool startup."
  },
  "cookies": ""
}
```

#### Core Output Fields

| Field | Description | Use Case |
|-------|-------------|----------|
| `URL` | The source URL that was scraped. | Reference and tracking. |
| `JS Result` | The return value from your custom `js_script`. | Custom data extraction via JavaScript. |
| `HTML` | Full page HTML after rendering. | Archiving, further parsing, or DOM analysis. |
| `Markdown` | Clean, plain-text markdown extracted from the main content area. | Content aggregation and publishing workflows. |
| `Content` | Extracted main content text with navigation, ads, and boilerplate removed. | Reading-focused applications and content extraction. |
| `JSON Content` | Structured JSON representation of the page content. | Direct integration into databases or APIs. |
| `Cookies` | Browser cookies present after page load. | Session management and authentication flows. |

> **Note:** Only fields matching your output configuration (`get_html: true`, `get_content: true`, etc.) are included in results.

***

### How to Use

#### Step 1: Prepare Your URLs

Compile a list of web pages to scrape. URLs can be:

- Individual article pages
- Search results
- Product listings
- Directory pages

#### Step 2: Configure Output Options

Decide which data you need:

- `get_html`: Set to `true` if you need the raw HTML for further parsing.
- `get_content`: Set to `true` to extract clean, readable text without ads or navigation.
- `get_js_result`: Set to `true` if you have a custom JavaScript query.
- `get_cookies`: Set to `true` for session tracking or authentication.

#### Step 3: Handle Dynamic Content

If the page uses JavaScript to load content:

- Set `wait_element: true` and provide a CSS selector (e.g., `.article-body`) for an element that appears after rendering.
- Optionally write a custom `js_script` to extract specific data: `async () => { return document.querySelector('.price').innerText; }`

#### Step 4: Run the Scraper

Start the process. The scraper will:

1. Load each URL in a headless browser
2. Wait for elements (if configured)
3. Execute JavaScript (if provided)
4. Collect the requested output
5. Retry failed URLs up to `max_retries_per_url` times

#### Step 5: Export Results

Download output as JSON, CSV, or Excel for downstream processing.

**Best practices:**

- Use cookies for authenticated scraping to access paywalled or member-only content.
- Set realistic `js_timeout` values (10-30 seconds depending on page complexity).
- Test with a single URL before scaling to bulk jobs.
- Monitor retry counts to identify consistently problematic pages.

***

### Use Cases & Business Value

**Content Aggregation:** Automatically collect articles, blog posts, and news items from multiple sources for republishing or analysis.

**Market Intelligence:** Extract pricing, product descriptions, and competitor data from e-commerce sites to track market trends.

**Research & Academic:** Gather structured datasets from public websites for analysis without manual data entry.

**SEO Monitoring:** Scrape page titles, meta descriptions, and headings to audit site structure and compliance.

**Lead Generation:** Collect contact information, company details, and job postings from business directories.

**Automation:** Feed extracted content into downstream workflows, databases, or machine learning pipelines.

The Web Content Scraper eliminates hours of manual work, enabling data-driven workflows at scale while maintaining flexibility through custom JavaScript and configurable output options.

***

### Conclusion

The **Web Content Scraper** is a powerful, flexible solution for anyone needing to extract data from websites programmatically. With support for cookies, JavaScript execution, element waiting, and multiple output formats, it adapts to virtually any scraping scenario — from simple content extraction to complex data aggregation pipelines. Start with a single URL today and scale to thousands of pages with confidence.

> **Legal note:** Always respect website Terms of Service, robots.txt, and applicable laws (e.g., CFAA, GDPR) when scraping. Some sites prohibit automated access; verify terms before proceeding.

# Actor input Schema

## `urls` (type: `array`):

Add the URLs of the web pages you want to scrape. You can paste URLs one by one, or use the Bulk edit section to add a prepared list.

## `wait_element` (type: `boolean`):

Wait until the element matching the selector below is displayed before running the JS script. Requires 'Element selector'.

## `element_selector` (type: `string`):

CSS selector of the element to wait for: an id (#main-content) or a class (.article-body). Required when 'Wait for an element' is enabled.

## `init_js_script` (type: `string`):

JS script that runs before the browser navigates to the URL, and before any of the page's own scripts. Useful for overriding browser properties, injecting helpers, or setting up hooks. It does not have access to the page DOM yet, and its return value is not collected. To run code after the page has loaded and get a result, use 'JS script you want to run' instead.

## `js_script` (type: `string`):

Add the JS script you want to run and get the result.

## `js_timeout` (type: `integer`):

Timeout on JS script run, in seconds.

## `cookies` (type: `array`):

Cookies to set before loading the page. A list of cookie objects with fields: name, value, domain, path, expires (null for a session cookie). Example: \[{"name": "cna", "value": "abc123", "domain": ".example.com", "path": "/", "expires": null}]

## `get_js_result` (type: `boolean`):

Return the result of the JS script in the output.

## `get_html` (type: `boolean`):

Return the page HTML after it has loaded.

## `get_content` (type: `boolean`):

Return the main content extracted from the page (text/markdown), without navigation, ads and other boilerplate.

## `get_cookies` (type: `boolean`):

Return the cookies of the page after it has loaded.

## `max_retries_per_url` (type: `integer`):

Limit the number of retries for each URL if an error occurs during the data scraping process.

## `ignore_url_failures` (type: `boolean`):

If true, the scraper will continue running even if some URLs fail to be scraped.

## `proxy` (type: `object`):

Select proxies to be used by your scraper.

## Actor input object example

```json
{
  "urls": [
    "https://www.markepear.dev/blog/developer-marketing-guide"
  ],
  "wait_element": false,
  "js_script": "async () => { return document.title; }",
  "js_timeout": 10,
  "cookies": [
    {
      "name": "cna",
      "value": "abc123",
      "domain": ".example.com",
      "path": "/",
      "expires": null
    }
  ],
  "get_js_result": true,
  "get_html": true,
  "get_content": false,
  "get_cookies": false,
  "max_retries_per_url": 2,
  "ignore_url_failures": true,
  "proxy": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        "https://www.markepear.dev/blog/developer-marketing-guide"
    ],
    "element_selector": "",
    "js_script": async () => { return document.title; },
    "js_timeout": 10,
    "cookies": [
        {
            "name": "cna",
            "value": "abc123",
            "domain": ".example.com",
            "path": "/",
            "expires": null
        }
    ],
    "max_retries_per_url": 2,
    "ignore_url_failures": true,
    "proxy": {
        "useApifyProxy": false
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("stealth_mode/web-content-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "urls": ["https://www.markepear.dev/blog/developer-marketing-guide"],
    "element_selector": "",
    "js_script": "async () => { return document.title; }",
    "js_timeout": 10,
    "cookies": [{
            "name": "cna",
            "value": "abc123",
            "domain": ".example.com",
            "path": "/",
            "expires": None,
        }],
    "max_retries_per_url": 2,
    "ignore_url_failures": True,
    "proxy": { "useApifyProxy": False },
}

# Run the Actor and wait for it to finish
run = client.actor("stealth_mode/web-content-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    "https://www.markepear.dev/blog/developer-marketing-guide"
  ],
  "element_selector": "",
  "js_script": "async () => { return document.title; }",
  "js_timeout": 10,
  "cookies": [
    {
      "name": "cna",
      "value": "abc123",
      "domain": ".example.com",
      "path": "/",
      "expires": null
    }
  ],
  "max_retries_per_url": 2,
  "ignore_url_failures": true,
  "proxy": {
    "useApifyProxy": false
  }
}' |
apify call stealth_mode/web-content-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,stealth_mode/web-content-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/r29UlWi2W7ubFdedZ/builds/gSbSy5sFcPwrq3YMU/openapi.json
