CNN Videos Scraper
Pricing
from $2.10 / 1,000 results
CNN Videos Scraper
Collects video metadata from the six CNN editions that publish a video feed -- US, International, Espanol, Arabic, Indonesia and Portugal -- returning title, description, duration, thumbnail, upload date and embed URL, with archive access back to 2015. Never downloads the media itself.
Pricing
from $2.10 / 1,000 results
Rating
0.0
(0)
Developer
Ibnu Adzim
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
4 days ago
Last modified
Categories
Share
CNN Videos Scraper (6 Editions)
Collects video metadata — never the media files — from the six CNN editions that publish a video feed.
| Editions with video | US, International, en Español, Arabic, Indonesia, Portugal |
| Editions without | Brasil, Chile, Czechia, Greece, Japan, Türkiye — each returns an ERROR row explaining why |
| Returns | title, description, duration (normalised to seconds), thumbnail, upload date, content and embed URLs, transcript when published |
| Archive | back to 2015 on US / International / en Español |
Example input
{"editions": ["us", "indonesia", "portugal"],"maxItemsPerEdition": 25,"includeVideoDetail": true}
Two metadata sources
Which one an edition offers changes what a cheap run can return:
- CNN Indonesia and CNN Portugal embed Google Video sitemap tags
(
<video:video>) directly in the feed — title, description, thumbnail and duration arrive without any detail fetch.feedCarriesVideoTagsistrueon their summary rows. - US, International, en Español and Arabic publish bare URL sitemaps. Their
metadata exists only in a
VideoObjectJSON-LD block on each video's page, soincludeVideoDetailmust stay on for those.
Durations are normalised to videoDurationSeconds from both forms — sitemaps
give plain seconds, JSON-LD gives ISO-8601 PT1M34S — so one comparable number
exists regardless of which edition a row came from.
Limits
- Six editions have no video feed at all. They are not silently dropped:
each returns an ERROR row naming the reason. CNN Greece is a notable case —
cnn.gr/sitemap/videosanswers 200 with a valid<sitemapindex>, but its children are the site's article, pages and tags sitemaps. It is the generic index served for any/sitemap/<anything>path, not a video feed. - CNN Indonesia's video sitemap carries no date field. No
<video:publication_date>, no<lastmod>. The publish date is derived from the timestamp in each video's URL and flaggedvideoUploadDateSource: "url-derived"so a consumer can tell it apart from an upstream-supplied date. - CNN Arabic's video archive is stale and ordered oldest-first. Its ten pages run from 2014; the newest content is on the last page, which this actor reads first.
- Media files are never downloaded.
videoContentUrlandvideoEmbedUrlare returned as published; fetching them is out of scope.