Article Scraper & Text Extractor – Clean Markdown for LLM/RAG
Pricing
$1.00 / 1,000 article extracteds
Article Scraper & Text Extractor – Clean Markdown for LLM/RAG
Article scraper & web page text extraction: turn any list of URLs into clean article text and Markdown (HTML to Markdown) with title, author, date, site, language and word count. Boilerplate removed; content extraction for LLM/RAG. No browser, fast.
Pricing
$1.00 / 1,000 article extracteds
Rating
0.0
(0)
Developer
Forever Tools
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
an hour ago
Last modified
Categories
Share
Article Scraper & Clean Text Extractor to Markdown for LLM/RAG
Paste a list of URLs and get the main content of every page as clean plain text and Markdown, with title, author, published date, site name, language, description and word count. Menus, sidebars, footers, cookie notices, scripts and styles are removed (using Mozilla Readability, the engine behind Firefox Reader View), so what you get is the article itself, ready to chunk and embed.
Use it to: feed web pages into an LLM or RAG pipeline, build a knowledge base from docs or blog posts, collect articles for summarisation or classification, archive readable copies of pages, or get word counts and metadata for a list of URLs.
It does not use a browser, so it's fast and cheap. The trade-off: it reads the HTML the server sends and does not run JavaScript (see Limitations).
Features
- Bulk: any number of URLs, 5 in parallel by default (up to 20). If one URL fails, the error goes in its row and the run keeps going.
- Markdown for LLMs: headings, lists, links (made absolute), bold/italic, code blocks and simple tables become
GitHub-flavoured Markdown. Complex tables (merged cells, no header row, e.g. Wikipedia infoboxes) become
Label: valuelines instead of raw HTML. - Plain text with paragraph breaks kept, plus a
wordCount. - Metadata from JSON-LD (schema.org Article/NewsArticle/BlogPosting), OpenGraph and standard meta tags, with Readability's guesses as a fallback. Dates are normalised to ISO 8601 when they parse.
- Retries with backoff for timeouts, network errors and HTTP 408/425/429/5xx.
- Encoding aware: the charset from the HTTP header or
<meta charset>is used to decode the page. - Optional cleaned HTML of the main content (
includeHtml).
Input
{"urls": ["https://en.wikipedia.org/wiki/Sourdough", "paulgraham.com/greatwork.html"],"includeHtml": false,"maxConcurrency": 5,"maxRetries": 2,"timeoutSecs": 30}
Output (one dataset row per URL)
{"url": "https://paulgraham.com/greatwork.html","finalUrl": "https://paulgraham.com/greatwork.html","statusCode": 200,"title": "How to Do Great Work","author": null,"publishedDate": null,"siteName": null,"language": null,"description": "If you collected lists of techniques for doing great work in a lot of different fields, what would the intersection look like? ...","text": "July 2023\nIf you collected lists of techniques for doing great work ...","markdown": "July 2023\n\nIf you collected lists of techniques ...","wordCount": 11801,"error": null}
- Fields the page doesn't provide are
null(the example page has no author or language tags). finalUrlis the URL after redirects.statusCodeis the final HTTP status.errorisnullon success. On failure (HTTP 4xx/5xx, DNS error, timeout, non-HTML content such as a PDF, invalid URL) the row still appears witherrorset, e.g."HTTP 404"or"fetch failed: ENOTFOUND", and empty content.- With
includeHtml: truean extrahtmlfield holds the cleaned main-content HTML.
The Console has two views: Articles (metadata table) and Content (Markdown). Export as JSON, CSV or Excel, or fetch via the API.
Pricing
Pay per event: $1 per 1,000 articles ($0.001 per successfully extracted URL). Failed URLs (error rows) are not charged. Apify platform usage is billed separately as usual. If you set a maximum charge for the run, the actor stops when it is reached.
FAQ
Does it work on any website? It works on pages whose content is in the HTML the server returns: most news sites, blogs, documentation, Wikipedia and similar. It won't see content that only appears after JavaScript runs.
What about home pages or listing pages? They are extracted, but a home page has no single "article", so the result may be a list of headlines, a section of the page, or some navigation text. The actor is built for article and documentation pages.
Can it get past paywalls, logins or CAPTCHAs? No. It sees what a logged-out visitor sees. Bot-protected sites may return an error or a challenge page.
Is the Markdown good for chunking? Yes, that's the goal: headings are #-style, and there are no scripts, styles or
layout tables. Split on headings or paragraphs.
Why is title sometimes "Page - Site"? The OpenGraph/<title> value is returned as the site publishes it.
Limitations
- No JavaScript rendering. Single-page apps and pages that load content client-side may come back empty or partial.
- Only HTML (and plain text) pages. PDFs, images and other files get an error row.
- Pages over 15 MB are skipped with an error.
- Which part of a page counts as "main content" is decided by Readability's heuristics. They are good on articles and can be wrong on unusual layouts.
- Respect the terms of the sites you extract from.
Support
Open an issue on the actor's Issues tab and include the URL and your input. Issues are answered asynchronously.
Not affiliated with Mozilla or any website you extract from. Built and maintained with AI assistance.
Related tools
Other actors by the same developer (same flat pay-per-result pricing, no subscription):
- Apple App Store Reviews Scraper (Multi-Country)
- Company Jobs Scraper: Workday, Greenhouse, Lever, Ashby
- Bulk Domain Checker — WHOIS/RDAP, DNS, SPF/DMARC, SSL Expiry
- Bulk PageSpeed Insights & Core Web Vitals Checker
- PDF to Text Extractor (Bulk, with Metadata)
- Website SEO Audit Crawler
- Sitemap Extractor & Bulk URL Status Checker
- Website Tech Stack Detector (CMS, Framework, Analytics)
- Website Screenshot – Bulk Full Page PNG, JPEG & PDF
Integrations
Run it from the Apify API, a schedule, or no-code tools: the Apify apps for Zapier, Make and n8n can start any public actor ("Run Actor") and read its dataset. AI agents can call it through the Apify MCP server.