CNN Sections & Tags Scraper
Pricing
from $2.10 / 1,000 results
CNN Sections & Tags Scraper
Enumerates the section and tag taxonomy of six CNN editions -- slug, URL, hierarchy depth and last-modified date -- so you know which section values the other CNN scrapers accept. Covers 198 US sections, 1,137 Espanol sections and 2,236 Arabic tags.
Pricing
from $2.10 / 1,000 results
Rating
0.0
(0)
Developer
Ibnu Adzim
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
4 days ago
Last modified
Categories
Share
The discovery actor for the CNN family. It enumerates the section and tag
taxonomy each edition publishes, so you know exactly which values the other CNN
scrapers' sections input will accept instead of guessing slugs against a 404.
| Edition | Sections | Tags |
|---|---|---|
| CNN US / International | 198 | — |
| CNN en Español | 1,137 | — |
| CNN Arabic | 53 | 2,236 |
| CNN Greece | 53 | 106 child sitemaps |
| CNN Indonesia | 9 | — |
| Brasil, Chile, Czechia, Portugal, Japan, Türkiye | none published | — |
Example input
{"editions": ["us", "arabic", "indonesia"],"taxonomyKinds": ["section", "tag"]}
Output
SECTION and TAG rows carry taxonomySlug (the value to pass to a sibling
actor), taxonomyUrl, taxonomyDepth (1 for /health, 2 for /world/africa)
and the sitemap's own lastmod, changefreq and priority. Turn on
includeTaxonomyDetail to also fetch each page's title, heading and description.
Defaults to metadata-only — one request per edition — because the slug and URL alone answer "what sections exist", and CNN Arabic's 2,236 tags would otherwise mean 2,236 extra requests.
Limits
- Only CNN Arabic and CNN Greece publish a tag sitemap. Asking for tags on another edition returns an ERROR row saying so rather than an empty result.
- CNN Greece splits its tags across 106 child sitemaps.
maxIndexChildrencaps how many are read (default 5) and the summary row records when more were available, so a partial listing is never silent. - CNN Indonesia has no taxonomy sitemap of its own. Its sections are derived
from the child URLs of its main sitemap index (
/{section}/sitemap_news.xml). Because there are three sitemap files per section, rows are de-duplicated on the canonical section URL andtaxonomyDiscoveredViarecords the sitemap file the section was found through. - CNN en Español counts author and topic hubs as sections, which is why it
reports 1,137 against the flagship's 198. Filter on
taxonomyDepthto get just the top level.