CNN Sections & Tags Scraper avatar

CNN Sections & Tags Scraper

Pricing

from $2.10 / 1,000 results

Go to Apify Store
CNN Sections & Tags Scraper

CNN Sections & Tags Scraper

Enumerates the section and tag taxonomy of six CNN editions -- slug, URL, hierarchy depth and last-modified date -- so you know which section values the other CNN scrapers accept. Covers 198 US sections, 1,137 Espanol sections and 2,236 Arabic tags.

Pricing

from $2.10 / 1,000 results

Rating

0.0

(0)

Developer

Ibnu Adzim

Ibnu Adzim

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

4 days ago

Last modified

Categories

Share

The discovery actor for the CNN family. It enumerates the section and tag taxonomy each edition publishes, so you know exactly which values the other CNN scrapers' sections input will accept instead of guessing slugs against a 404.

EditionSectionsTags
CNN US / International198
CNN en Español1,137
CNN Arabic532,236
CNN Greece53106 child sitemaps
CNN Indonesia9
Brasil, Chile, Czechia, Portugal, Japan, Türkiyenone published

Example input

{
"editions": ["us", "arabic", "indonesia"],
"taxonomyKinds": ["section", "tag"]
}

Output

SECTION and TAG rows carry taxonomySlug (the value to pass to a sibling actor), taxonomyUrl, taxonomyDepth (1 for /health, 2 for /world/africa) and the sitemap's own lastmod, changefreq and priority. Turn on includeTaxonomyDetail to also fetch each page's title, heading and description.

Defaults to metadata-only — one request per edition — because the slug and URL alone answer "what sections exist", and CNN Arabic's 2,236 tags would otherwise mean 2,236 extra requests.

Limits

  • Only CNN Arabic and CNN Greece publish a tag sitemap. Asking for tags on another edition returns an ERROR row saying so rather than an empty result.
  • CNN Greece splits its tags across 106 child sitemaps. maxIndexChildren caps how many are read (default 5) and the summary row records when more were available, so a partial listing is never silent.
  • CNN Indonesia has no taxonomy sitemap of its own. Its sections are derived from the child URLs of its main sitemap index (/{section}/sitemap_news.xml). Because there are three sitemap files per section, rows are de-duplicated on the canonical section URL and taxonomyDiscoveredVia records the sitemap file the section was found through.
  • CNN en Español counts author and topic hubs as sections, which is why it reports 1,137 against the flagship's 198. Filter on taxonomyDepth to get just the top level.