Website Content Extractor
Pricing
from $9.00 / 1,000 results
Website Content Extractor
Extract clean text and markdown from docs, pricing, product, policy, and help-center URLs for RAG datasets and content operations.
Pricing
from $9.00 / 1,000 results
Rating
0.0
(0)
Developer
naoki anzai
Maintained by CommunityActor stats
1
Bookmarked
26
Total users
3
Monthly active users
8 days ago
Last modified
Categories
Share
Run the next report
Turn extracted public pages into capped Site QA report/export runs:
- Run a content QA report when you need title, metadata, thin-content, and page-quality issues grouped into a reviewable report.
- Run an indexability and AI crawler readiness report when you need robots, sitemap, canonical, noindex, schema, and AI crawler access checks.
These follow-on actors use user-supplied public URLs only. They do not provide ranking guarantees, legal advice, or automated messaging workflows.
AI builders, content ops, SEO teams, and documentation teams use this actor to turn Public website pages supplied by the user into a clean dataset for Site QA & Content Intelligence Pack. Provide focused source inputs, keep the first run small, and expand only after the output shape is useful. Each emitted row includes source context, timestamps, and fields designed for monitoring, QA, research, or workflow handoff.
Store Quickstart
Start with 5 to 20 URLs from one domain, review extracted markdown quality, then expand to sitemap or scheduled checks.
Recommended first run:
{"urls": ["https://example.com/docs"],"outputFormat": "markdown","limit": 10,"delivery": "dataset","dryRun": false}
Input examples
Docs pages
{"urls": ["https://example.com/docs"],"outputFormat": "markdown","limit": 10,"delivery": "dataset","dryRun": false}
Pricing and product pages
{"urls": ["https://example.com/pricing","https://example.com/product"],"outputFormat": "text","limit": 20,"delivery": "dataset","dryRun": false}
Webhook handoff
{"urls": ["https://example.com/help"],"outputFormat": "markdown","delivery": "webhook","webhookUrl": "https://example.com/webhook","dryRun": false}
Sample output
{"meta": {"actorName": "website-content-extractor","actorTitle": "Website Content Extractor","bundle": "Site QA & Content Intelligence Pack","fetchedAt": "2026-05-06T00:00:00.000Z","totalRows": 1},"rows": [{"actorName": "website-content-extractor","rowType": "web_content","url": "https://example.com/docs","title": "Example Docs","markdown": "# Example Docs\nUseful content.","wordCount": 240,"sourceUrl": "https://example.com/docs","fetchedAt": "2026-05-06T00:00:00.000Z"}],"warnings": []}
Output fields
rowTypeurltitlemarkdowntextwordCountmetadatasourceUrlfetchedAt
Rows also include source URLs, fetch timestamps, warnings when a source is partial, and stable IDs when the workflow supports recurring change detection.
See also (Content extraction cluster)
- Article Content Extractor & Reader Scraper — Article-specific extraction (byline, publish date, hero image) for news/blog/press URLs.
Pricing and no-change runs
$0.001 actor start and $0.009 per useful content row. Failed/no-content rows should stay out of the default dataset.
The default dataset is the billable surface. Dry runs, validation-only runs, missing-key warnings, and unchanged recurring polls should not write payable default-dataset rows.
Compliance guardrails
- Fetch public pages supplied by the user.
- Respect site policies, rate limits, and robots guidance where applicable.
- Use output for content operations, QA, and RAG workflows.
- Do not use provider emblems or wording that implies approval by an upstream data provider.
See also
Related report Actors
Use these follow-on Actors when you want a capped, decision-ready report instead of more raw rows. They use public or user-provided inputs, respect maxChargeUsd, and do not promise rankings, revenue, conversion lifts, or sales outcomes.
- Website Lighthouse & WCAG Regression Report - turn public pages into Lighthouse, axe-core, and regression evidence.
- SaaS Pricing Page Monitor - turn public pricing pages into competitor pricing decision reports.
- Ad Landing Page Offer Intelligence - turn public landing pages into CRO offer and proof checklists.
Related paid report workflows
If this Actor gave you raw rows or source context, these follow-on report Actors are designed for a small capped paid run. They help make a decision, not just collect more data.
- Website Lighthouse & WCAG Regression Report - audit performance, accessibility, SEO, and regression signals with a capped report run.
- SaaS Pricing Page Monitor & Competitor Price Change Alerts - decide whether a public competitor pricing page changed in a way that affects packaging or sales messaging. Entry $3 /
pricing_snapshot_report; premium $15 /competitor_pricing_report. - Ad Landing Page Offer Intelligence & CRO Gap Report - decide which public landing-page offer gaps to fix before increasing ad spend. Entry $3 /
landing_offer_report; premium $15 /cro_gap_report_pack.
Keep maxChargeUsd equal to the selected tier. Internal links are traffic aids only; real proof requires accounted paid usage.
💾 Save it for later: click the bookmark icon at the top of the Apify Store page if you'd like to come back to it. Bookmarks help other engineers find this actor via Apify's discovery surfaces.
⭐ Was Website Content Extractor useful for your page content extraction?
If this actor saved you time, please leave a 5★ rating on Apify Store — it takes 10 seconds, helps other engineers and analysts discover it, and keeps updates free.
Have a feature request, bug, or sample workflow you'd like to share? Open an issue — we read every one and use them to prioritise the next release.