Web Content Scraper avatar

Web Content Scraper

Pricing

from $2.00 / 1,000 results

Go to Apify Store
Web Content Scraper

Web Content Scraper

Scrape web pages effortlessly with undetected browser technology. Extract HTML, markdown, JavaScript results, and cookies from any URL. Features configurable timeouts, proxy support, and retry logic — ideal for content aggregation, data extraction, and automated workflows without detection.

Pricing

from $2.00 / 1,000 results

Rating

0.0

(0)

Developer

Stealth mode

Stealth mode

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

a day ago

Last modified

Categories

Share

Web Content Scraper: Extract HTML, Text & Structured Data from Any Website


What Is Web Scraping?

Web scraping automates the extraction of information from websites. Instead of manually copying and pasting content, a scraper retrieves structured data at scale. This is essential for market research, content aggregation, competitive analysis, and data-driven decision-making. The Web Content Scraper removes the friction from this process, offering granular control over what data you collect and how the page loads.


Overview

The Web Content Scraper is a versatile tool for extracting structured and unstructured content from web pages. It supports:

  • Running custom JavaScript to extract dynamic data
  • Waiting for specific page elements to load before collection
  • Exporting multiple output formats (HTML, markdown, JSON)
  • Collecting cookies for authenticated sessions
  • Automatic retry logic for failed requests
  • Undetected browser technology that bypasses anti-bot protections (Cloudflare, Akamai)
  • Proxy support to avoid detection and IP blocking

Ideal users include:

  • Content aggregators building automated publishing pipelines
  • Data analysts extracting market intelligence from web sources
  • Developers automating web data collection from protected websites
  • Researchers gathering structured datasets from websites
  • Quality assurance teams testing website rendering across scenarios
  • Enterprise teams scraping heavily-protected sites requiring anti-bot circumvention

The scraper uses an undetected browser engine that mimics human behavior, making it difficult for security systems to identify it as automated traffic. If you encounter blocking or rate limits, enable proxy rotation in the Scrape options to distribute requests across different IP addresses.


Input Configuration

The scraper accepts a JSON configuration object controlling how pages are loaded and what data is extracted:

{
"urls": [
"https://www.markepear.dev/blog/developer-marketing-guide"
],
"init_js_script": null,
"cookies": null,
"wait_element": true,
"element_selector": "#id",
"js_script": "async () => { return document.title; }",
"js_timeout": 10,
"get_js_result": true,
"get_html": true,
"get_content": true,
"get_cookies": true,
"ignore_url_failures": true,
"proxy": null
}

Key Input Parameters

ParameterTypePurpose
urlsarrayList of web page URLs to scrape. Supports bulk input.
cookiesarraySession cookies to send with requests (optional). Format: [{"name": "cna", "value": "abc123", "domain": ".example.com", "path": "/", "expires": null}]
wait_elementbooleanPause execution until a specific element appears on the page (Timeout in 5s) (requires element_selector). Useful for waiting on dynamic content.
element_selectorstringCSS selector for the element to wait for (e.g., #main-content, .article-body).
init_js_scriptstringJavaScript code to run before navigating to the URL. Use for injecting utilities or overriding browser properties.
js_scriptstringJavaScript code to run after the page loads. Example: async () => { return $('title').text(); }
js_timeoutintegerMaximum seconds to wait for the JS script to complete. Default: 10.
max_retries_per_urlintegerHow many times to retry a failed URL. Default: 2.
proxyobjectProxy configuration (optional). Leave null to use direct connection or set useApifyProxy: true for residential/datacenter proxies.
ignore_url_failuresbooleanIf true, the scraper continues running when a URL fails instead of stopping the entire run. Recommended for bulk jobs. Default: true.

Tip: Use wait_element when targeting pages with lazy-loaded content or single-page applications where elements render after initial HTML load.


Output Format

Each URL produces a record with up to seven fields, depending on your configuration:

{
"url": "https://www.markepear.dev/blog/developer-marketing-guide",
"js_result": "Developer marketing guide (by a dev tool startup CMO)",
"html": "<!DOCTYPE html><!-- Last Published: Fri Jun 05 2026 12:33:40 GMT+0000 (Coordinated Universal Time) --> ...",
"markdown": "Contents\n[What is developer marketing?](#what-is-developer-marketing) [How is marketing to developers different than just marketing?](#how-is-marketing-to-developers-different-than-just-marketing) [Best practices of marketing to developers](#best-practices-of-marketing-to-developers) [How to market to software developers with a plan?] ...",
"content": "Contents\nDeveloper marketing guide (by a dev tool startup CMO)\nThis is a 6000-word guide to developer marketing written for practitioners by a practitioner.\nI helped grow a machine learning dev tool startup from 0 to a Series A and learned a lot about marketing to devs along the way...",
"json_content": {
"title": "Developer marketing guide (by a dev tool startup CMO)",
"author": null,
"hostname": "markepear.dev",
"date": "2026-09-01",
"categories": "",
"tags": "",
"fingerprint": "",
"id": null,
"license": null,
"comments": "",
"raw_text": "Contents Developer marketing guide (by a dev tool startup CMO) This is a 6000-word ...",
"text": "Contents\nDeveloper marketing guide (by a dev tool startup CMO)\nThis is a 6000-word guide to developer marketing ...",
"language": null,
"image": "https://cdn.prod.website-files.com/6161939cdc6297e03f7803e0/65158b7336e68a13c8461f3d_Developer%20marketing%20guide%20(by%20a%20dev%20tool%20startup%20CMO).png",
"pagetype": "website",
"source": "https://www.markepear.dev/blog/developer-marketing-guide",
"source-hostname": "markepear.dev",
"excerpt": "Sep 01, 2026 - In 6000 words, I share what I learned about developer marketing from talking to hundreds of practitioners and years of growing a dev tool startup."
},
"cookies": ""
}

Core Output Fields

FieldDescriptionUse Case
URLThe source URL that was scraped.Reference and tracking.
JS ResultThe return value from your custom js_script.Custom data extraction via JavaScript.
HTMLFull page HTML after rendering.Archiving, further parsing, or DOM analysis.
MarkdownClean, plain-text markdown extracted from the main content area.Content aggregation and publishing workflows.
ContentExtracted main content text with navigation, ads, and boilerplate removed.Reading-focused applications and content extraction.
JSON ContentStructured JSON representation of the page content.Direct integration into databases or APIs.
CookiesBrowser cookies present after page load.Session management and authentication flows.

Note: Only fields matching your output configuration (get_html: true, get_content: true, etc.) are included in results.


How to Use

Step 1: Prepare Your URLs

Compile a list of web pages to scrape. URLs can be:

  • Individual article pages
  • Search results
  • Product listings
  • Directory pages

Step 2: Configure Output Options

Decide which data you need:

  • get_html: Set to true if you need the raw HTML for further parsing.
  • get_content: Set to true to extract clean, readable text without ads or navigation.
  • get_js_result: Set to true if you have a custom JavaScript query.
  • get_cookies: Set to true for session tracking or authentication.

Step 3: Handle Dynamic Content

If the page uses JavaScript to load content:

  • Set wait_element: true and provide a CSS selector (e.g., .article-body) for an element that appears after rendering.
  • Optionally write a custom js_script to extract specific data: async () => { return document.querySelector('.price').innerText; }

Step 4: Run the Scraper

Start the process. The scraper will:

  1. Load each URL in a headless browser
  2. Wait for elements (if configured)
  3. Execute JavaScript (if provided)
  4. Collect the requested output
  5. Retry failed URLs up to max_retries_per_url times

Step 5: Export Results

Download output as JSON, CSV, or Excel for downstream processing.

Best practices:

  • Use cookies for authenticated scraping to access paywalled or member-only content.
  • Set realistic js_timeout values (10-30 seconds depending on page complexity).
  • Test with a single URL before scaling to bulk jobs.
  • Monitor retry counts to identify consistently problematic pages.

Use Cases & Business Value

Content Aggregation: Automatically collect articles, blog posts, and news items from multiple sources for republishing or analysis.

Market Intelligence: Extract pricing, product descriptions, and competitor data from e-commerce sites to track market trends.

Research & Academic: Gather structured datasets from public websites for analysis without manual data entry.

SEO Monitoring: Scrape page titles, meta descriptions, and headings to audit site structure and compliance.

Lead Generation: Collect contact information, company details, and job postings from business directories.

Automation: Feed extracted content into downstream workflows, databases, or machine learning pipelines.

The Web Content Scraper eliminates hours of manual work, enabling data-driven workflows at scale while maintaining flexibility through custom JavaScript and configurable output options.


Conclusion

The Web Content Scraper is a powerful, flexible solution for anyone needing to extract data from websites programmatically. With support for cookies, JavaScript execution, element waiting, and multiple output formats, it adapts to virtually any scraping scenario — from simple content extraction to complex data aggregation pipelines. Start with a single URL today and scale to thousands of pages with confidence.

Legal note: Always respect website Terms of Service, robots.txt, and applicable laws (e.g., CFAA, GDPR) when scraping. Some sites prohibit automated access; verify terms before proceeding.