CNN Videos Scraper avatar

CNN Videos Scraper

Pricing

from $2.10 / 1,000 results

Go to Apify Store
CNN Videos Scraper

CNN Videos Scraper

Collects video metadata from the six CNN editions that publish a video feed -- US, International, Espanol, Arabic, Indonesia and Portugal -- returning title, description, duration, thumbnail, upload date and embed URL, with archive access back to 2015. Never downloads the media itself.

Pricing

from $2.10 / 1,000 results

Rating

0.0

(0)

Developer

Ibnu Adzim

Ibnu Adzim

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

4 days ago

Last modified

Categories

Share

CNN Videos Scraper (6 Editions)

Collects video metadata — never the media files — from the six CNN editions that publish a video feed.

Editions with videoUS, International, en Español, Arabic, Indonesia, Portugal
Editions withoutBrasil, Chile, Czechia, Greece, Japan, Türkiye — each returns an ERROR row explaining why
Returnstitle, description, duration (normalised to seconds), thumbnail, upload date, content and embed URLs, transcript when published
Archiveback to 2015 on US / International / en Español

Example input

{
"editions": ["us", "indonesia", "portugal"],
"maxItemsPerEdition": 25,
"includeVideoDetail": true
}

Two metadata sources

Which one an edition offers changes what a cheap run can return:

  • CNN Indonesia and CNN Portugal embed Google Video sitemap tags (<video:video>) directly in the feed — title, description, thumbnail and duration arrive without any detail fetch. feedCarriesVideoTags is true on their summary rows.
  • US, International, en Español and Arabic publish bare URL sitemaps. Their metadata exists only in a VideoObject JSON-LD block on each video's page, so includeVideoDetail must stay on for those.

Durations are normalised to videoDurationSeconds from both forms — sitemaps give plain seconds, JSON-LD gives ISO-8601 PT1M34S — so one comparable number exists regardless of which edition a row came from.

Limits

  • Six editions have no video feed at all. They are not silently dropped: each returns an ERROR row naming the reason. CNN Greece is a notable case — cnn.gr/sitemap/videos answers 200 with a valid <sitemapindex>, but its children are the site's article, pages and tags sitemaps. It is the generic index served for any /sitemap/<anything> path, not a video feed.
  • CNN Indonesia's video sitemap carries no date field. No <video:publication_date>, no <lastmod>. The publish date is derived from the timestamp in each video's URL and flagged videoUploadDateSource: "url-derived" so a consumer can tell it apart from an upstream-supplied date.
  • CNN Arabic's video archive is stale and ordered oldest-first. Its ten pages run from 2014; the newest content is on the last page, which this actor reads first.
  • Media files are never downloaded. videoContentUrl and videoEmbedUrl are returned as published; fetching them is out of scope.