Article body
articleBody
Optional
Full article text, joined from
tags inside the #Content element (stable id across multiple CMS template generations). Only page 1 is captured for paginated photo-essay articles.
China Daily Scraper
Pricing
from $3.50 / 1,000 results
Extract full article text or the newest headlines from China Daily (chinadaily.com.cn), China's official English-language state newspaper -- no account or API key needed.
Input
_input
Optional
The exact input value (URL, or 'latest') that produced this record.
Source strategy
_source
Optional
Which extraction strategy emitted this record (e.g. S1-html, S2-html-listing).
Scraped at
_scrapedAt
Optional
UTC ISO 8601 timestamp when this record was scraped.
Error
_error
Optional
Present only on failed extractions -- the error kind.
Error detail
_errorDetail
Optional
Free-form diagnostic detail, present only alongside _error.
Headline
headline
Optional
Article headline (from ld+json NewsArticle node).
Description
description
Optional
Short article description -- absent on wire-service/video-type content with no named byline.
Article body
articleBody
Optional
Full article text, joined from
tags inside the #Content element (stable id across multiple CMS template generations). Only page 1 is captured for paginated photo-essay articles.
Publisher
publisher
Optional
Publisher name, extracted from the ld+json publisher object's 'name' field -- inconsistently either 'China Daily' or 'chinadaily.com.cn' depending on which CMS template rendered the page, passed through raw (not normalized).
Date published
datePublished
Optional
ISO 8601 publish date, from ld+json.
Date modified
dateModified
Optional
ISO 8601 last-modified date, from ld+json.
Image
image
Optional
Lead image value as emitted by upstream ld+json (raw passthrough, shape varies).
Title
title
Optional
Headline text scraped from the homepage link (latest mode) -- the longest non-empty anchor text seen for that URL.
Link
link
Optional
Article URL (latest mode).
URL date
urlDate
Optional
Day-precision date (YYYY-MM-DD) derived from the article URL path, not a true publish timestamp (latest mode only). Fetch the article for the exact datePublished.