Sitemap SEO Audit & Monitor - Broken Links & Hreflang
Pricing
from $3.00 / 1,000 page audits
Sitemap SEO Audit & Monitor - Broken Links & Hreflang
Audit every page in your sitemap: broken links, redirects, noindex, canonicals, titles, meta descriptions, hreflang, AI-crawler robots.txt. Alerts on changes.
Pricing
from $3.00 / 1,000 page audits
Rating
0.0
(0)
Developer
Dave West
Maintained by CommunityActor stats
1
Bookmarked
2
Total users
1
Monthly active users
21 hours ago
Last modified
Categories
Share
Sitemap SEO Audit & Monitor: broken links, redirects, meta tags, hreflang and AI crawlers
Run a technical SEO audit on every page in your sitemap and get alerts when something breaks. Give it a website, a sitemap URL or a list of pages. It finds the sitemap, checks each page for the issues that hurt indexing and rankings, and scores the site from 0 to 100. It also checks which AI crawlers robots.txt allows and whether /llms.txt exists. Schedule it to see on every run which issues are new, which you fixed, and which pages started returning 404 or became noindex.
- One dataset row per page, listing its issues with a severity and a plain-English message.
- A site summary (JSON) and a self-contained HTML report.
- A change log between runs, so you can use it as an SEO monitoring tool.
- Pay only for pages audited: $0.003 per page plus $0.02 per run for the site report.
Use cases
- Technical SEO audit of a client site in minutes, with no desktop crawler to install.
- Sitemap hygiene: find sitemap URLs that return 404, redirect, are noindex or canonicalised elsewhere. These waste crawl budget and confuse search engines.
- Website monitoring after releases and migrations: schedule a daily or weekly run and catch new 404s, accidental
noindextags, broken canonicals and redirect chains before Google does. - Broken link checker: deduplicated broken internal links (and external links if you turn that on), with the pages that link to them.
- International SEO: hreflang code validation,
x-default, return links (reciprocity) and hreflang targets that don't return 200. - AI search readiness: see whether GPTBot, OAI-SearchBot, ChatGPT-User, PerplexityBot, ClaudeBot, Google-Extended, CCBot and Applebot-Extended may crawl the site, and whether
/llms.txtexists. - Agencies and freelancers: one run per client site, with an HTML report you can share.
What is checked
| Area | Checks (issue codes) |
|---|---|
| Status & speed | 4xx / 5xx (HTTP_4XX, HTTP_5XX), network errors (FETCH_FAILED), RATE_LIMITED, slow responses over 3 s (SLOW_RESPONSE), HTML over 1 MB (PAGE_TOO_LARGE), non-HTML URLs (NOT_HTML) |
| Redirects | REDIRECTED, chains of 2+ hops (REDIRECT_CHAIN), REDIRECT_LOOP, TOO_MANY_REDIRECTS, HTTPS to HTTP (HTTPS_TO_HTTP_REDIRECT) |
| Indexability | noindex from meta robots or X-Robots-Tag (NOINDEX), blocked for Googlebot by robots.txt (BLOCKED_BY_ROBOTS_TXT) |
| Canonical | missing, multiple, pointing elsewhere, canonical target not 200 (CANONICAL_*) |
| Title & meta description | missing, too short (<30 / <70), too long (>60 / >160), duplicated across the site (TITLE_*, META_DESCRIPTION_*) |
| Headings & language | H1_MISSING, H1_MULTIPLE, HTML_LANG_MISSING, HTML_LANG_INVALID |
| Hreflang | invalid codes (e.g. en-UK, jp), missing x-default, missing self-reference, conflicting codes, missing return links between audited pages, targets that don't return 200 (HREFLANG_*) |
| Links | broken internal links, broken external links (optional), internal links to redirecting URLs |
| Content | images without alt, mixed content (http:// resources on https pages) |
| Sitemap hygiene | sitemap URLs that aren't 200, redirect, are noindex, are canonicalised elsewhere or are blocked by robots.txt (SITEMAP_URL_*) |
| Site level | robots.txt (found, unreachable, Crawl-delay), AI crawler access table, /llms.txt, sitemap discovery (robots.txt Sitemap: lines, common locations, sitemap indexes, gzip) |
Every issue has a fixed severity:
- error: breaks indexing or the user experience. Fix these first.
- warning: likely hurts rankings, snippets or crawl efficiency.
- notice: worth a look, but often intentional.
Health score (0-100)
pageScore = 100 - 15 x errors - 5 x warnings - 1 x notices (distinct issue codes on the page, floored at 0)healthScore = round( average pageScore of all fetched pages - site penalties ), clamped to 0..100site penalties: -5 robots.txt unreachable, -5 no sitemap found, -5 homepage blocked for Googlebot
Pages that weren't fetched because robots.txt blocks this crawler are left out of the average. The formula is also stored in every SUMMARY (healthScoreFormula).
Quick start
- Enter your website (e.g.
https://www.example.com) in Website or sitemap URL. - Set Max pages. Large sitemaps are sampled evenly across their sitemap files.
- Click Start. A 30-page audit usually finishes in 10-40 seconds.
- Open the HTML report in the Output tab, or the dataset for per-page details.
- To monitor a site, schedule the Actor (daily or weekly). From the second run on, every row has a
changeflag and the summary lists new and fixed issues.
Leave the input empty for a free preview on a small example site. Preview runs are never charged.
Want just the list of URLs in a sitemap, or alerts when pages are added or removed? Use Sitemap Extractor & Change Monitor.
Input example
{"startUrl": "https://crawlee.dev","maxPages": 30,"checkExternalLinks": false,"maxLinkChecks": 1000,"compareWithPreviousRun": true}
| Field | Default | Notes |
|---|---|---|
startUrl | - | A homepage or a sitemap / sitemap index URL. One site per run. |
urlList | - | Audit exactly these pages instead of the sitemap. |
maxPages | 500 | 1-5,000. Each audited page is one billable event. |
checkExternalLinks | false | Also check links to other sites. |
maxLinkChecks | 1000 | Distinct link targets fetched for the broken-link check. Pages that were already audited are free. |
compareWithPreviousRun | true | Store state per site in a named key-value store and report changes. |
stateStoreName | sitemap-seo-audit-state | Use different names to keep separate histories. |
followLinksIfNoSitemap | true | Crawl internal links from the homepage when there is no sitemap. |
respectRobotsTxt | true | Skip disallowed URLs (reported for free) and honour Crawl-delay up to 10 s. |
maxConcurrency / minDelayBetweenRequestsMs | 4 / 150 | Polite defaults. |
Output
Dataset: one row per page. Views: Overview, Indexability, Content, Links & performance, Changes. Below is a real row from crawlee.dev, shortened:
{"url": "https://crawlee.dev/blog/crawlee-for-python-v1","statusCode": 200,"responseTimeMs": 569,"pageSizeBytes": 156820,"indexable": true,"canonicalStatus": "self","title": "Crawlee for Python v1 | Crawlee for JavaScript · Build reliable crawlers. Fast.","titleLength": 79,"metaDescriptionLength": 47,"h1Count": 1,"htmlLang": "en","hreflang": [{ "lang": "en", "href": "https://crawlee.dev/blog/crawlee-for-python-v1", "validCode": true, "targetStatus": 200, "reciprocal": true }],"brokenInternalLinkCount": 1,"brokenInternalLinks": [{ "url": "https://www.crawlee.dev/python/docs/examples/playwright-crawler-with-fingeprint-generator", "status": 404 }],"imagesMissingAlt": 0,"issues": [{ "code": "BROKEN_INTERNAL_LINKS", "severity": "error", "message": "1 internal link(s) are broken, e.g. https://www.crawlee.dev/python/docs/examples/playwright-crawler-with-fingeprint-generator (404)." },{ "code": "TITLE_TOO_LONG", "severity": "warning", "message": "Title is 79 characters and may be cut off in results (aim for at most 60)." },{ "code": "META_DESCRIPTION_TOO_SHORT", "severity": "notice", "message": "Meta description is 47 characters (aim for 70-160)." }],"pageScore": 78,"change": { "status": "existing", "changed": false, "newIssues": [], "fixedIssues": [], "new404": false, "newlyNoindexed": false, "previousStatusCode": 200 }}
Key-value store:
SUMMARY: health score and formula, page counts and status codes, issue counts by code and severity, top issues, duplicate titles and descriptions, the AI crawler table, llms.txt, robots.txt, sitemap files, link-check stats, andchanges(new and fixed issues, new 404s, newly noindexed pages, new and removed pages, change in score).OUTPUT.html: a self-contained report you can open in a browser or send to a client.
Excerpt from a real SUMMARY (crawlee.dev, 30 pages, second run):
{"healthScore": 92,"pages": { "audited": 30, "indexable": 30, "statusCodes": { "200": 30 }, "avgResponseTimeMs": 549 },"issuesBySeverity": { "error": 1, "warning": 39, "notice": 17 },"topIssues": [{ "code": "BROKEN_INTERNAL_LINKS", "severity": "error", "pages": 1 }, { "code": "TITLE_TOO_LONG", "severity": "warning", "pages": 30 }],"aiCrawlers": [{ "bot": "GPTBot", "owner": "OpenAI", "access": "allowed", "hasSpecificRules": false }],"llmsTxt": { "exists": true, "httpStatus": 200 },"changes": { "comparedWithPreviousRun": true, "newIssues": [], "fixedIssues": [], "new404s": [], "newlyNoindexed": [] }}
Pricing
Pay per event, with no subscription:
| Event | Price | When |
|---|---|---|
page-audited | $0.003 ($3 per 1,000 pages) | each audited page saved to the dataset |
site-report | $0.02 | once per run, after the summary and report are built. Only charged if at least one page was billed |
- Examples: 30 pages = $0.11. 500 pages = $1.52. 5,000 pages = $15.02.
- Free rows: pages that robots.txt blocks (not fetched) are listed but not charged. Preview runs (empty input) cost nothing.
- Max cost per run: the Actor respects the limit you set. It audits only as many pages as fit, always keeps room for the site report, and stops gracefully.
FAQ
How is the sitemap found? Through the Sitemap: lines in robots.txt, then common locations (/sitemap.xml, /sitemap_index.xml, /wp-sitemap.xml, and more). Sitemap indexes, gzip sitemaps, text sitemaps and RSS/Atom feeds are supported. If no sitemap exists, the Actor follows internal links from the homepage.
What happens with very large sitemaps? Up to maxPages URLs are sampled evenly across all sitemap files, so every section of the site is represented. The summary shows how many URLs the sitemaps contain.
Does it render JavaScript? No. It audits the HTML your server sends, which is what search engines index first. Pages that build titles or links only in the browser may show missing titles or few links.
Will it overload my server? No. Defaults are 4 parallel requests and at least 150 ms between requests. It honours robots.txt (including Crawl-delay) and backs off on 429 and 503 responses.
How does change monitoring work? After each run, a small state record per site is stored in a named key-value store. The next run compares each page's status code, noindex and issue codes with it. Use a different stateStoreName for separate histories.
Why is a page "blocked by robots.txt" but not charged? When robots.txt disallows this crawler, the page isn't fetched. It is still listed so you can see it, but you don't pay for it.
What if I abort a run? Pages audited before the abort are still saved and appear in the summary and report. Link checks may be incomplete. The site report isn't charged and the monitoring state isn't updated.
Can I audit several sites? Yes, with one run per site. Create a task per site and schedule them.
Support
Found a false positive or missing a check? Open an issue on the Actor's Issues tab and include the URL and the issue code.