Article Text Extractor: Clean Text and Markdown from URLs
Pricing
from $1.00 / 1,000 article extracteds
Article Text Extractor: Clean Text and Markdown from URLs
Main text of the article and blog URLs you list, as plain text and Markdown, with title, author, date and language. Article-body F1 0.959 on a held-out half of a public benchmark. One request per URL, robots.txt obeyed; blocked pages reported, not bypassed.
Pricing
from $1.00 / 1,000 article extracteds
Rating
0.0
(0)
Developer
Jack Valmadre
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
32 minutes ago
Last modified
Categories
Share
Article Text Extractor
Give it the URLs of articles or blog posts; get back the main text as plain text and Markdown, with the title, author, publication date and language the page states. The extractor aims to leave out menus, sidebars, comments and footers (see the accuracy figures below for how well that works on a public test set).
What it does
- One HTTP GET per URL you name. Links on the page are not followed; nothing else is crawled.
- robots.txt is always obeyed (read once per site, also for redirect targets; Crawl-delay honoured up to 10 seconds); at most one request at a time per site.
- Pages behind a bot check, a login or an access-denied answer are reported as
blockedand not retried. No proxies, no browser, no paywall bypass. - Main-text extraction uses the open-source trafilatura library (Apache-2.0) with settings we chose on a public benchmark (below).
What it does not do
- It does not run JavaScript, so pages that only load their text with JavaScript come back as
no-article. - It does not log in, solve captchas or get around paywalls.
- Author, date and language are taken as the page states them and are not checked; the date can be the page's first publication date rather than its last update.
- You are responsible for having the right to use the text of the pages you choose.
Accuracy (measured by us, September 2026)
On the public scrapinghub article-extraction benchmark (181 news and blog pages saved in 2019, each with a reference copy of its article text), we chose the extraction settings on one half of the pages and then measured the other half (93 pages): article-body F1 0.959 (± 0.012; precision 0.950, recall 0.968). F1 here is the word overlap between our text and the reference text, averaged over pages, computed with the benchmark's own scoring script (trafilatura 2.2.0, run 2026-09-26).
What this number does and does not tell you:
- It is the benchmark's own measure on its own pages. Your pages will differ: the pages are from 2019 and mostly English-language news, so modern layouts, other languages and JavaScript-heavy sites are not covered.
- Only the article body is scored; title, author, date and language are not.
- On the same 93 pages, trafilatura's default settings also score 0.959; our settings trade a little recall for precision (less stray text), they do not raise the average. The number describes the open-source extractor this actor runs, not something unique to it.
Input
| Field | Default | Meaning |
|---|---|---|
urls | - | Article URLs, one per line (up to 1,000 per run). |
startUrls, articleUrls, links, url | - | Other names for the same thing, for API callers and AI agents. All are merged and duplicates removed. |
includeMarkdown | true | Also return Markdown. |
includeText | true | Return plain text. |
timeoutSecs | 20 | Give up on a page after this many seconds. |
maxConcurrency | 5 | Different sites fetched in parallel. |
Use with AI agents
Any of the URL field names works ({"urls": ["https://example.com/post"]} or {"url": "..."}). Each row carries status, so an agent can tell an extracted article (ok) from no-article, blocked, disallowed-by-robots, error and skipped.
Output (one dataset row per URL)
url, finalUrl (after redirects), httpStatus, status, error, title, author, date (YYYY-MM-DD), language (as declared by the page, e.g. en-GB), siteName, text, markdown, wordCount, fetchedAt. The run's OUTPUT record counts URLs by status.
Pricing
Pay per event: US$0.001 per article extracted (article-extracted, rows with status ok), plus a start fee of US$0.00005 per GB of run memory, charged once per run (at most US$0.00005 at the default 512 MB). Rows with any other status (blocked, disallowed by robots.txt, no article found, failed) are not charged. There is no extra charge for Apify platform usage. Apify shows the price before you run, and you can set a maximum cost per run: once no further article fits, the remaining URLs get a free skipped row.
Privacy
The actor stores nothing beyond the run's own dataset. It reads only the pages you name and returns what those pages show; it does not look up, enrich or combine information about people.
Support
Questions and bug reports: open an issue in the Issues tab on this actor's page. We aim to respond within 14 days. Replies are written with AI assistance; a human owner can be reached on request.
About
Published by Madrasco and built and maintained with AI assistance.