Sitemap Analyzer - URL Structure, Freshness & SEO Issues
Pricing
from $50.00 / 1,000 analysed sites
Sitemap Analyzer - URL Structure, Freshness & SEO Issues
Analyse any site's XML sitemaps: find every sitemap, count URLs, map site structure and depth, measure lastmod freshness, and get a ranked list of sitemap problems.
Pricing
from $50.00 / 1,000 analysed sites
Rating
0.0
(0)
Developer
Apisight
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
2 days ago
Last modified
Categories
Share
Sitemap Analyzer — URL Structure, Freshness & SEO Issues
Point it at a domain. It finds every XML sitemap the site publishes, reads them all (including gzipped ones and nested sitemap indexes), and returns a structural analysis of the site plus a ranked list of problems worth fixing.
No sitemap URL needed — discovery happens automatically via robots.txt and well-known
paths. You can also pass a sitemap URL directly if you already know it.
What you get per site
Inventory — every sitemap file found, its type (index or urlset), URL count, byte size, whether it was gzipped, and how it was discovered.
Scale and structure
- Total and unique URL counts, and how many URLs are duplicated
- URL count per top-level section — the actual shape of the site
- Depth distribution, max depth and average depth
Content freshness
lastmodcoverage as a percentage- URLs modified in the last 7 / 30 / 90 / 365 days, older than a year, or dated in the future
- Oldest and newest
lastmoddates
Media and internationalisation — image, video, Google News and hreflang alternate counts.
Ranked issues — errors first, then warnings, then informational notes. Each has a stable
code so you can filter or track them over time:
| Code | Severity | Meaning |
|---|---|---|
sitemap-unreadable | error | A sitemap returned an error or malformed XML |
too-many-urls | error | Over the sitemaps.org limit of 50,000 URLs in one file |
file-too-large | error | Over the 50 MB uncompressed limit |
no-robots-txt | warning | robots.txt unreachable, so crawlers can't discover the sitemap there |
sitemap-not-in-robots | warning | robots.txt exists but has no Sitemap: line |
cross-host-urls | warning | Sitemap lists URLs on other hosts |
mixed-schemes | warning | Mixes http:// and https:// URLs |
no-lastmod | warning | No <lastmod> anywhere |
duplicate-urls | warning | The same URL appears more than once |
future-lastmod | warning | <lastmod> dated in the future |
uniform-priority | info | Every URL has identical <priority>, making it meaningless |
inconsistent-trailing-slash | info | Mixed trailing-slash conventions risk duplicate content |
truncated | info | The maxUrls limit stopped the analysis early |
Input
{"startUrls": ["ghost.org", "wordpress.org"],"maxUrls": 50000,"maxSitemaps": 50,"includeUrlList": false,"proxyConfiguration": { "useApifyProxy": true }}
| Field | Type | Default | Notes |
|---|---|---|---|
startUrls | array | — | Domains, or a direct sitemap URL |
maxUrls | integer | 50000 | Per site. Results flag whether analysis was truncated |
maxSitemaps | integer | 50 | Caps files fetched when a sitemap index is used |
includeUrlList | boolean | false | Adds every URL with its lastmod, changefreq, priority |
maxConcurrency | integer | 5 | Sites analysed in parallel |
Output (abridged)
{"domain": "ghost.org","status": "ok","discoveryMethod": "robots.txt","sitemapCount": 8,"sitemapIndexCount": 1,"totalUrls": 1003,"uniqueUrls": 1003,"duplicateUrlCount": 0,"lastmodCoveragePct": 100.0,"freshness": {"last7Days": 41, "last30Days": 88, "last90Days": 122,"last365Days": 402, "olderThan1Year": 350, "inTheFuture": 0},"newestLastmod": "2026-09-30","oldestLastmod": "2019-02-22","maxDepth": 3,"avgDepth": 1.98,"topSections": { "resources": 428, "themes": 260, "help": 140 },"mediaCounts": { "images": 684, "videos": 0, "newsItems": 0, "hreflangAlternates": 0 },"errorCount": 0,"warningCount": 0,"issues": []}
Three output views are provided: Overview (one row per site), Issues (one row per problem), and Sitemap files (one row per sitemap document).
Pricing and what you are charged for
| Event | Price | When |
|---|---|---|
site-analysed | $0.05 | Once per site where at least one sitemap was read and analysed |
apify-actor-start | $0.00005 | Once per run |
You are only charged when a sitemap was actually analysed. Three outcomes cost you nothing:
status: "no-sitemap"— the site genuinely publishes no sitemap. That is a real finding, and you still get it, but we don't think you should pay for an empty result when you're screening a list of domains.status: "blocked"— bot protection served a challenge instead of the sitemap.status: "error"— the host never responded.
Honest limitations
- Bot protection can block sitemap access. Some sites put
robots.txtand sitemaps behind a challenge. These are reported asblocked, with the reason, and never billed. - A site with no discoverable sitemap may still have one at an unconventional path that
is neither in
robots.txtnor one of the well-known locations. Pass the URL directly if you know it. lastmodis self-reported. Some CMSs stamp every URL with the build date, which makes freshness look better than it is. Treat 100% coverage with identical dates sceptically — the freshness buckets make that pattern easy to spot.- Nested indexes are followed three levels deep, which covers essentially every real site while bounding runtime.
- No personal data. This Actor reads
robots.txtand XML sitemaps only — public technical files. It collects no personal data and is not a contact-scraping tool.
Common uses
- SEO audits — inventory a site's indexable URLs and find sitemap problems before they cost you crawl budget.
- Migration checks — compare structure and URL counts before and after a replatform.
- Competitive research — see how large a competitor's site is and which sections dominate.
- Content freshness monitoring — track how much of a site is genuinely being updated.
- Crawl seeding — enable
includeUrlListto get a clean, deduplicated URL list to feed another tool.