CNN Galleries & Live Stories Scraper avatar

CNN Galleries & Live Stories Scraper

Pricing

from $2.10 / 1,000 results

Go to Apify Store
CNN Galleries & Live Stories Scraper

CNN Galleries & Live Stories Scraper

Collects CNN photo galleries with every image, caption and credit, and live blogs with their full timeline of updates, from CNN US, International and en Espanol -- including an archive of 1,587 gallery and 546 live-story monthly partitions back to 2015.

Pricing

from $2.10 / 1,000 results

Rating

0.0

(0)

Developer

Ibnu Adzim

Ibnu Adzim

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

4 days ago

Last modified

Categories

Share

Two of CNN's richest formats, from the three editions that publish them.

EditionsCNN US, CNN International, CNN en Español
Galleriesevery photo with its caption and credit — verified 33, 39 and 30 images on real galleries
Live storiesthe full timeline of updates — verified 21, 42 and 38 updates on real live blogs
Archive1,587 gallery and 546 live-story monthly partitions, back to 2015

Example input

{
"editions": ["us", "espanol"],
"contentKinds": ["gallery", "liveStory"],
"maxItemsPerEdition": 15
}

Output

GALLERY rows carry galleryImages — an array of {imageUrl, imageAlt, imageCaption, imageCredit} — plus galleryImageCount.

LIVE_STORY rows carry liveStoryUpdates — an array of {updateHeadline, updateBody, updatePublishedAt, updateUrl} — plus liveStoryCoverageStart, liveStoryCoverageEnd and liveStoryStatus, so you can tell an active live blog from a concluded one.

Both share the dataset with SEARCH_SUMMARY and ERROR rows.

Limits

  • Only three editions publish these formats. Every licensee edition and CNN Arabic return HTTP 404 on /sitemap/gallery.xml and /sitemap/live-story.xml. Asking for one returns an ERROR row naming the reason rather than an empty result.
  • Galleries have no gallery-specific structured data. CNN's gallery pages carry a plain NewsArticle JSON-LD block — no ImageGallery node, no per-image entries. The photographs exist only in the HTML, so this actor reads them from the page. A JSON-LD-only extractor returns a gallery with zero pictures.
  • A live story's articleBody is empty. The NewsArticle node on a live-story page has articleBody: ""; the real content is in LiveBlogPosting.liveBlogUpdate. Reading articleBody returns an empty story.
  • Live stories are large. 40+ updates each with a body is normal, so maxItemsPerEdition defaults to 15.