Healthline Article Scraper - Text, Authors & Medical Review
Pricing
from $1.00 / 1,000 results
Healthline Article Scraper - Text, Authors & Medical Review
Scrape Healthline articles: headline, full body text, authors, publication and modification dates, keywords and section — plus the medical reviewer and last-reviewed date, which is what makes health content auditable. HTTP-only, no account.
Pricing
from $1.00 / 1,000 results
Rating
0.0
(0)
Developer
Farhan Febrian Nauval
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
2 days ago
Last modified
Categories
Share
Healthline Article Scraper — Text, Authors & Medical Review
Scrape Healthline articles: headline, full body text, authors, dates, keywords — plus the medical reviewer and last-reviewed date, which is what makes health content auditable. HTTP-only, no account, no browser.
Input
| Field | Type | Default | Description |
|---|---|---|---|
articleUrls | array | required | Healthline article URLs |
maxRounds | integer | 2 | TLS ladder repeats on a refused page |
maxConcurrency | integer | 4 | Articles in parallel |
proxyConfiguration | object | Apify proxy | Optional — a bare IP answered in testing |
Output
{"_input": "https://www.healthline.com/health/type-2-diabetes","_source": "S1-jsonld+article","_scrapedAt": "2026-09-08T…Z","url": "https://www.healthline.com/health/type-2-diabetes","headline": "Understanding Type 2 Diabetes","description": "Everything you've wanted to know about type 2 diabetes…","authors": ["Ann Pietrangelo"],"reviewedBy": ["Alana Biggers, M.D., MPH"],"lastReviewed": "2025-06-27T10:17:00Z","datePublished": "2020-06-17T09:49:00Z","dateModified": "…","articleSection": "Uncategorized","keywords": "…","bodyText": "Heart disease is the leading cause of death…","wordCount": 644,"paragraphCount": 22,"headings": ["Symptoms", "Causes", …],"otherSchemaTypes": ["VideoObject", "MedicalCondition"]}
Failure rows carry _error (not_found, bad_input, blocked, unexpected_shape) and no payload.
How it works
Metadata comes from the page's MedicalWebPage JSON-LD node; the body is read from the rendered
<article> element, because Healthline does not put articleBody in the JSON-LD.
Body text is joined with a separator, deliberately. Reading paragraph text without one
concatenates inline elements: According to the <a>Centers for Disease Control</a> comes out as
According to theCenters. The result still reads as prose, which is exactly why that bug survives
a casual look at the output — so paragraphs are extracted with an explicit space separator and the
result was checked for glued-word artefacts (zero remaining).
A 404 here is not small. A missing article still returns ~148 KB of site chrome, so neither a
200 nor a large body proves anything. The actor requires both a JSON-LD block and an <article>
element before it will parse, and reports not_found otherwise.
Known limits
| Limit | Detail |
|---|---|
| Not every article is medically reviewed | reviewedBy and lastReviewed were present on 2 of 3 test articles — that reflects the source, not a parse failure |
articleSection is often "Uncategorized" | Passed through as Healthline publishes it |
| No comments | Healthline does not publish reader comments |
| URLs only | There is no search or sitemap crawl here — supply the article URLs you want |