# Bulk Website Content Export (`jj_toolworks/bulk-website-content-export`) Actor

Export bounded static website pages as clean text and Markdown with source URLs, content hashes, extraction profiles and complete per-input coverage statuses.

- **URL**: https://apify.com/jj_toolworks/bulk-website-content-export.md
- **Developed by:** [JJ Toolworks](https://apify.com/jj_toolworks) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$3.00 / 1,000 completed results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Bulk Website Content Export

Export static public HTML as cleaned text and Markdown with source URLs, timestamps, content hashes and clear extraction coverage. This is useful for known knowledge sources, vendor research, content inventories and public notice review.

### Start with exact pages

```json
{
  "urls": [{"url": "https://example.com/", "recordId": "example-domain"}],
  "mode": "pages",
  "preset": "general"
}
```

For a bounded site section, choose `mode: "crawl"`, `maxDepth: 1`, `maxPagesPerRoot: 10`, and a run-wide `maxPages`. Discovery follows the final root host. Root depth is zero. Cross-host redirects from discovered pages are reported as out of scope.

### Extraction controls

`contentSelector` selects an exact CSS section. A selector that matches nothing returns a free `selector_not_found` result. Without an explicit selector, the Actor prefers main/article regions, then falls back to the body. `preset` adjusts preferred containers for general pages, knowledge bases, vendor research, content inventories, public notices or education.

`excludeSelectors` removes up to 20 unwanted sections, such as a cookie banner or timestamp widget. Without an explicit content selector, common navigation/header/footer/aside areas are removed. Output retains ordinary headings, lists, tables, code and links where represented in the supplied HTML. Row/column-spanning tables are flagged and are not expanded into a spreadsheet model.

The minimum extraction criterion is 40 letters/digits after cleaning by default; `minTextCharacters` can be set from 20–1,000. This is a transparent minimum-content check, not a guarantee that every important section was captured.

### Output and snapshots

Filter `recordType = page` for page results. Successful records contain `contentText`, `markdown`, `contentHash`, `extractionProfile`, `extractionVersion`, `extractionSettings`, `sourceUrl`, `finalUrl`, `title`, `fetchedAt` and warnings. Failed or limited records retain the URL and reason; no success is fabricated from HTTP 200.

Each supplied URL/root also gets `recordType = coverage`: pages processed, complete/failed/duplicate pages, pages omitted by the page cap, and links outside the requested depth. `complete_within_scope` means that the selected scope finished; it does not mean the entire website was inventoried. A complete page can be charged even if other pages cause its root crawl to be incomplete.

`recordType = run_summary` includes a `snapshotUrl` to the `SNAPSHOT` JSON object containing successful unique page evidence. That object can be supplied as an explicit baseline to JJ Toolworks Website Content Change Checks using identical extraction settings. Failed or omitted pages are excluded from this content-export snapshot, with their statuses retained in the dataset. The change-check Actor's snapshots separately preserve prior successful evidence after a failed current fetch.

Input duplicates and distinct inputs that redirect to the same final URL/profile retain mapping/status rows without a second fee. Duplicate rows reference the original final URL and omit repeated content text/Markdown.

### Pricing and limits

Configured price: **$3 per 1,000 complete extracted pages** ($0.003 each), deduplicated by final URL and extraction profile per run. Empty, too-thin, blocked, unsupported, timed-out or truncated pages have no success event. A complete page is defined by the chosen extraction scope and published content criteria.

| Limit | Maximum |
|---|---:|
| Supplied URLs/roots | 100 |
| Unique page fetches/run | 200; default 100 |
| Pages/root in crawl mode | 50; default 10 |
| Depth | 3 |
| Fetched HTML/page | 2 MiB |
| Text + Markdown/page | 1 MiB |
| Run content output | 20 MiB; later pages get a visible run_content_limit status |
| Request time | 30 seconds |
| Stored discovery links/page | 500; omitted links are counted |

The Actor uses static HTTP only. JavaScript shells, unsupported content types and access challenges produce explicit results. It does not run a browser, access logins, perform OCR, create embeddings or interpret legal/commercial meaning.

### Efficient industry usage

A help-center selector improves knowledge-base extraction; a product-description selector supports supplier research; a notice-section selector supports public notice review. These are settings in one maintained tool. Keep the extraction profile unchanged between snapshots to compare content reliably.

Implementation interface reference: https://www.crummy.com/software/BeautifulSoup/bs4/doc/. The source bundle's `MARKET.md` records competitor usage observations and economics hypotheses; listed market usage is not evidence of revenue.

Text hashes normalize line whitespace and indentation. They can ignore a spacing-only source change; they are not HTML-byte, visual-layout or code-semantics hashes. Markdown retains preformatted code where available.

### Price and run controls

The launch price is **$3 per 1,000 completed units** ($0.003 per event), as defined above. Check the current Pricing tab before running. Status and duplicate rows do not add this custom event. Set a maximum charge appropriate for your batch; a run stops when its event budget is exhausted. `RUN-SUMMARY` records actual accepted events and execution totals.

Automatic replay of an interrupted run is disabled to prevent duplicate charges. Keep the available output, then start a new run for a new execution. A later run is a new billable execution. Completed dataset rows and individually saved files remain available when a budget stops a run. Combined exports, manifests, and snapshots assembled at the end may be absent after a budget stop or interruption; use a sufficient run budget when you need those aggregate files. The initial release uses documented input limits and public source access; external source changes can require maintenance.

# Actor input Schema

## `urls` (type: `array`):

1–100 public URLs, or {url, recordId} objects. Your record ID appears on page and coverage rows.

## `mode` (type: `string`):

pages fetches only supplied pages. crawl follows same-host links within the explicit limits.

## `maxDepth` (type: `integer`):

Only used in crawl mode. Root is depth 0; links beyond depth are counted separately.

## `maxPagesPerRoot` (type: `integer`):

Only used in crawl mode. Pending pages omitted by this cap appear in root coverage.

## `maxPages` (type: `integer`):

Duplicate requested URLs share a fetched result; aliases may still require one redirect request before deduplication.

## `preset` (type: `string`):

Prefers relevant main-content containers when no explicit selector is supplied. Preset is part of the extraction profile.

## `contentSelector` (type: `string`):

Optional exact section selector; maximum 500 characters. A missing match produces a free selector_not_found status.

## `excludeSelectors` (type: `array`):

Up to 20 CSS selectors to remove, such as .last-updated or .cookie-banner. These settings change the comparison profile.

## `minTextCharacters` (type: `integer`):

Minimum letters/digits after cleaning. Pages below this criterion are not success events.

## Actor input object example

```json
{
  "urls": [
    {
      "url": "https://example.com/",
      "recordId": "example-domain"
    }
  ],
  "mode": "pages",
  "maxDepth": 1,
  "maxPagesPerRoot": 10,
  "maxPages": 100,
  "preset": "general",
  "excludeSelectors": [],
  "minTextCharacters": 40
}
```

# Actor output Schema

## `results` (type: `string`):

No description

## `summary` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        {
            "url": "https://example.com/",
            "recordId": "example-domain"
        }
    ],
    "mode": "pages",
    "preset": "general"
};

// Run the Actor and wait for it to finish
const run = await client.actor("jj_toolworks/bulk-website-content-export").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "urls": [{
            "url": "https://example.com/",
            "recordId": "example-domain",
        }],
    "mode": "pages",
    "preset": "general",
}

# Run the Actor and wait for it to finish
run = client.actor("jj_toolworks/bulk-website-content-export").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    {
      "url": "https://example.com/",
      "recordId": "example-domain"
    }
  ],
  "mode": "pages",
  "preset": "general"
}' |
apify call jj_toolworks/bulk-website-content-export --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,jj_toolworks/bulk-website-content-export"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/rb3iqxQ6roOB8oCWz/builds/ru0yh8QjbYsdTEGwa/openapi.json
