Healthline Article Scraper - Text, Authors & Medical Review avatar

Healthline Article Scraper - Text, Authors & Medical Review

Pricing

from $1.00 / 1,000 results

Go to Apify Store
Healthline Article Scraper - Text, Authors & Medical Review

Healthline Article Scraper - Text, Authors & Medical Review

Scrape Healthline articles: headline, full body text, authors, publication and modification dates, keywords and section — plus the medical reviewer and last-reviewed date, which is what makes health content auditable. HTTP-only, no account.

Pricing

from $1.00 / 1,000 results

Rating

0.0

(0)

Developer

Farhan Febrian Nauval

Farhan Febrian Nauval

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

2 days ago

Last modified

Categories

Share

Healthline Article Scraper — Text, Authors & Medical Review

Scrape Healthline articles: headline, full body text, authors, dates, keywords — plus the medical reviewer and last-reviewed date, which is what makes health content auditable. HTTP-only, no account, no browser.

Input

FieldTypeDefaultDescription
articleUrlsarrayrequiredHealthline article URLs
maxRoundsinteger2TLS ladder repeats on a refused page
maxConcurrencyinteger4Articles in parallel
proxyConfigurationobjectApify proxyOptional — a bare IP answered in testing

Output

{
"_input": "https://www.healthline.com/health/type-2-diabetes",
"_source": "S1-jsonld+article",
"_scrapedAt": "2026-09-08T…Z",
"url": "https://www.healthline.com/health/type-2-diabetes",
"headline": "Understanding Type 2 Diabetes",
"description": "Everything you've wanted to know about type 2 diabetes…",
"authors": ["Ann Pietrangelo"],
"reviewedBy": ["Alana Biggers, M.D., MPH"],
"lastReviewed": "2025-06-27T10:17:00Z",
"datePublished": "2020-06-17T09:49:00Z",
"dateModified": "…",
"articleSection": "Uncategorized",
"keywords": "…",
"bodyText": "Heart disease is the leading cause of death…",
"wordCount": 644,
"paragraphCount": 22,
"headings": ["Symptoms", "Causes", …],
"otherSchemaTypes": ["VideoObject", "MedicalCondition"]
}

Failure rows carry _error (not_found, bad_input, blocked, unexpected_shape) and no payload.

How it works

Metadata comes from the page's MedicalWebPage JSON-LD node; the body is read from the rendered <article> element, because Healthline does not put articleBody in the JSON-LD.

Body text is joined with a separator, deliberately. Reading paragraph text without one concatenates inline elements: According to the <a>Centers for Disease Control</a> comes out as According to theCenters. The result still reads as prose, which is exactly why that bug survives a casual look at the output — so paragraphs are extracted with an explicit space separator and the result was checked for glued-word artefacts (zero remaining).

A 404 here is not small. A missing article still returns ~148 KB of site chrome, so neither a 200 nor a large body proves anything. The actor requires both a JSON-LD block and an <article> element before it will parse, and reports not_found otherwise.

Known limits

LimitDetail
Not every article is medically reviewedreviewedBy and lastReviewed were present on 2 of 3 test articles — that reflects the source, not a parse failure
articleSection is often "Uncategorized"Passed through as Healthline publishes it
No commentsHealthline does not publish reader comments
URLs onlyThere is no search or sitemap crawl here — supply the article URLs you want