Bulk Website Content Export
Pricing
$3.00 / 1,000 completed results
Bulk Website Content Export
Export bounded static website pages as clean text and Markdown with source URLs, content hashes, extraction profiles and complete per-input coverage statuses.
Pricing
$3.00 / 1,000 completed results
Rating
0.0
(0)
Developer
JJ Toolworks
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
a day ago
Last modified
Categories
Share
Export static public HTML as cleaned text and Markdown with source URLs, timestamps, content hashes and clear extraction coverage. This is useful for known knowledge sources, vendor research, content inventories and public notice review.
Start with exact pages
{"urls": [{"url": "https://example.com/", "recordId": "example-domain"}],"mode": "pages","preset": "general"}
For a bounded site section, choose mode: "crawl", maxDepth: 1, maxPagesPerRoot: 10, and a run-wide maxPages. Discovery follows the final root host. Root depth is zero. Cross-host redirects from discovered pages are reported as out of scope.
Extraction controls
contentSelector selects an exact CSS section. A selector that matches nothing returns a free selector_not_found result. Without an explicit selector, the Actor prefers main/article regions, then falls back to the body. preset adjusts preferred containers for general pages, knowledge bases, vendor research, content inventories, public notices or education.
excludeSelectors removes up to 20 unwanted sections, such as a cookie banner or timestamp widget. Without an explicit content selector, common navigation/header/footer/aside areas are removed. Output retains ordinary headings, lists, tables, code and links where represented in the supplied HTML. Row/column-spanning tables are flagged and are not expanded into a spreadsheet model.
The minimum extraction criterion is 40 letters/digits after cleaning by default; minTextCharacters can be set from 20–1,000. This is a transparent minimum-content check, not a guarantee that every important section was captured.
Output and snapshots
Filter recordType = page for page results. Successful records contain contentText, markdown, contentHash, extractionProfile, extractionVersion, extractionSettings, sourceUrl, finalUrl, title, fetchedAt and warnings. Failed or limited records retain the URL and reason; no success is fabricated from HTTP 200.
Each supplied URL/root also gets recordType = coverage: pages processed, complete/failed/duplicate pages, pages omitted by the page cap, and links outside the requested depth. complete_within_scope means that the selected scope finished; it does not mean the entire website was inventoried. A complete page can be charged even if other pages cause its root crawl to be incomplete.
recordType = run_summary includes a snapshotUrl to the SNAPSHOT JSON object containing successful unique page evidence. That object can be supplied as an explicit baseline to JJ Toolworks Website Content Change Checks using identical extraction settings. Failed or omitted pages are excluded from this content-export snapshot, with their statuses retained in the dataset. The change-check Actor's snapshots separately preserve prior successful evidence after a failed current fetch.
Input duplicates and distinct inputs that redirect to the same final URL/profile retain mapping/status rows without a second fee. Duplicate rows reference the original final URL and omit repeated content text/Markdown.
Pricing and limits
Configured price: $3 per 1,000 complete extracted pages ($0.003 each), deduplicated by final URL and extraction profile per run. Empty, too-thin, blocked, unsupported, timed-out or truncated pages have no success event. A complete page is defined by the chosen extraction scope and published content criteria.
| Limit | Maximum |
|---|---|
| Supplied URLs/roots | 100 |
| Unique page fetches/run | 200; default 100 |
| Pages/root in crawl mode | 50; default 10 |
| Depth | 3 |
| Fetched HTML/page | 2 MiB |
| Text + Markdown/page | 1 MiB |
| Run content output | 20 MiB; later pages get a visible run_content_limit status |
| Request time | 30 seconds |
| Stored discovery links/page | 500; omitted links are counted |
The Actor uses static HTTP only. JavaScript shells, unsupported content types and access challenges produce explicit results. It does not run a browser, access logins, perform OCR, create embeddings or interpret legal/commercial meaning.
Efficient industry usage
A help-center selector improves knowledge-base extraction; a product-description selector supports supplier research; a notice-section selector supports public notice review. These are settings in one maintained tool. Keep the extraction profile unchanged between snapshots to compare content reliably.
Implementation interface reference: https://www.crummy.com/software/BeautifulSoup/bs4/doc/. The source bundle's MARKET.md records competitor usage observations and economics hypotheses; listed market usage is not evidence of revenue.
Text hashes normalize line whitespace and indentation. They can ignore a spacing-only source change; they are not HTML-byte, visual-layout or code-semantics hashes. Markdown retains preformatted code where available.
Price and run controls
The launch price is $3 per 1,000 completed units ($0.003 per event), as defined above. Check the current Pricing tab before running. Status and duplicate rows do not add this custom event. Set a maximum charge appropriate for your batch; a run stops when its event budget is exhausted. RUN-SUMMARY records actual accepted events and execution totals.
Automatic replay of an interrupted run is disabled to prevent duplicate charges. Keep the available output, then start a new run for a new execution. A later run is a new billable execution. Completed dataset rows and individually saved files remain available when a budget stops a run. Combined exports, manifests, and snapshots assembled at the end may be absent after a budget stop or interruption; use a sufficient run budget when you need those aggregate files. The initial release uses documented input limits and public source access; external source changes can require maintenance.