Sitemap Analyzer - URL Structure, Freshness & SEO Issues avatar

Sitemap Analyzer - URL Structure, Freshness & SEO Issues

Pricing

from $50.00 / 1,000 analysed sites

Go to Apify Store
Sitemap Analyzer - URL Structure, Freshness & SEO Issues

Sitemap Analyzer - URL Structure, Freshness & SEO Issues

Analyse any site's XML sitemaps: find every sitemap, count URLs, map site structure and depth, measure lastmod freshness, and get a ranked list of sitemap problems.

Pricing

from $50.00 / 1,000 analysed sites

Rating

0.0

(0)

Developer

Apisight

Apisight

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

2 days ago

Last modified

Share

Sitemap Analyzer — URL Structure, Freshness & SEO Issues

Point it at a domain. It finds every XML sitemap the site publishes, reads them all (including gzipped ones and nested sitemap indexes), and returns a structural analysis of the site plus a ranked list of problems worth fixing.

No sitemap URL needed — discovery happens automatically via robots.txt and well-known paths. You can also pass a sitemap URL directly if you already know it.

What you get per site

Inventory — every sitemap file found, its type (index or urlset), URL count, byte size, whether it was gzipped, and how it was discovered.

Scale and structure

  • Total and unique URL counts, and how many URLs are duplicated
  • URL count per top-level section — the actual shape of the site
  • Depth distribution, max depth and average depth

Content freshness

  • lastmod coverage as a percentage
  • URLs modified in the last 7 / 30 / 90 / 365 days, older than a year, or dated in the future
  • Oldest and newest lastmod dates

Media and internationalisation — image, video, Google News and hreflang alternate counts.

Ranked issues — errors first, then warnings, then informational notes. Each has a stable code so you can filter or track them over time:

CodeSeverityMeaning
sitemap-unreadableerrorA sitemap returned an error or malformed XML
too-many-urlserrorOver the sitemaps.org limit of 50,000 URLs in one file
file-too-largeerrorOver the 50 MB uncompressed limit
no-robots-txtwarningrobots.txt unreachable, so crawlers can't discover the sitemap there
sitemap-not-in-robotswarningrobots.txt exists but has no Sitemap: line
cross-host-urlswarningSitemap lists URLs on other hosts
mixed-schemeswarningMixes http:// and https:// URLs
no-lastmodwarningNo <lastmod> anywhere
duplicate-urlswarningThe same URL appears more than once
future-lastmodwarning<lastmod> dated in the future
uniform-priorityinfoEvery URL has identical <priority>, making it meaningless
inconsistent-trailing-slashinfoMixed trailing-slash conventions risk duplicate content
truncatedinfoThe maxUrls limit stopped the analysis early

Input

{
"startUrls": ["ghost.org", "wordpress.org"],
"maxUrls": 50000,
"maxSitemaps": 50,
"includeUrlList": false,
"proxyConfiguration": { "useApifyProxy": true }
}
FieldTypeDefaultNotes
startUrlsarray—Domains, or a direct sitemap URL
maxUrlsinteger50000Per site. Results flag whether analysis was truncated
maxSitemapsinteger50Caps files fetched when a sitemap index is used
includeUrlListbooleanfalseAdds every URL with its lastmod, changefreq, priority
maxConcurrencyinteger5Sites analysed in parallel

Output (abridged)

{
"domain": "ghost.org",
"status": "ok",
"discoveryMethod": "robots.txt",
"sitemapCount": 8,
"sitemapIndexCount": 1,
"totalUrls": 1003,
"uniqueUrls": 1003,
"duplicateUrlCount": 0,
"lastmodCoveragePct": 100.0,
"freshness": {
"last7Days": 41, "last30Days": 88, "last90Days": 122,
"last365Days": 402, "olderThan1Year": 350, "inTheFuture": 0
},
"newestLastmod": "2026-09-30",
"oldestLastmod": "2019-02-22",
"maxDepth": 3,
"avgDepth": 1.98,
"topSections": { "resources": 428, "themes": 260, "help": 140 },
"mediaCounts": { "images": 684, "videos": 0, "newsItems": 0, "hreflangAlternates": 0 },
"errorCount": 0,
"warningCount": 0,
"issues": []
}

Three output views are provided: Overview (one row per site), Issues (one row per problem), and Sitemap files (one row per sitemap document).

Pricing and what you are charged for

EventPriceWhen
site-analysed$0.05Once per site where at least one sitemap was read and analysed
apify-actor-start$0.00005Once per run

You are only charged when a sitemap was actually analysed. Three outcomes cost you nothing:

  • status: "no-sitemap" — the site genuinely publishes no sitemap. That is a real finding, and you still get it, but we don't think you should pay for an empty result when you're screening a list of domains.
  • status: "blocked" — bot protection served a challenge instead of the sitemap.
  • status: "error" — the host never responded.

Honest limitations

  • Bot protection can block sitemap access. Some sites put robots.txt and sitemaps behind a challenge. These are reported as blocked, with the reason, and never billed.
  • A site with no discoverable sitemap may still have one at an unconventional path that is neither in robots.txt nor one of the well-known locations. Pass the URL directly if you know it.
  • lastmod is self-reported. Some CMSs stamp every URL with the build date, which makes freshness look better than it is. Treat 100% coverage with identical dates sceptically — the freshness buckets make that pattern easy to spot.
  • Nested indexes are followed three levels deep, which covers essentially every real site while bounding runtime.
  • No personal data. This Actor reads robots.txt and XML sitemaps only — public technical files. It collects no personal data and is not a contact-scraping tool.

Common uses

  • SEO audits — inventory a site's indexable URLs and find sitemap problems before they cost you crawl budget.
  • Migration checks — compare structure and URL counts before and after a replatform.
  • Competitive research — see how large a competitor's site is and which sections dominate.
  • Content freshness monitoring — track how much of a site is genuinely being updated.
  • Crawl seeding — enable includeUrlList to get a clean, deduplicated URL list to feed another tool.