Sitemap and Broken Link Checker with On-Page SEO Audit
Pricing
Pay per usage
Sitemap and Broken Link Checker with On-Page SEO Audit
Check your sitemap URLs, any list of URLs, or every link on your pages: status, redirect chain, noindex, canonical, robots.txt, and broken links with the page and anchor text they sit on. On-page SEO rules (title, meta description, H1, canonical, alt text, schema, hreflang) with fix hints.
Pricing
Pay per usage
Rating
0.0
(0)
Developer
Jack Valmadre
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
a day ago
Last modified
Categories
Share
Find the URLs and links on your site that search engines and visitors trip over, as data you can sort and filter:
- Broken links on your pages, each with the page it sits on and its anchor text, so you know exactly where to fix it.
- Sitemap problems: listed URLs that are broken, redirected,
noindex, canonicalised to another URL, or blocked by your own robots.txt. - Status and redirect chains for any list of URLs, for example before and after a site migration.
- On-page SEO issues on every HTML page it reads: missing or duplicate titles, meta descriptions, H1s and canonicals, missing
lang, viewport, imagealttext and structured data, invalid JSON-LD and hreflang values. Each issue comes with a severity and a one-line fix hint.
No login to your site, no browser extension. Give it a sitemap, a list of URLs, or the pages whose links you want checked.
Pick a mode
| You want to… | Set mode to | Give it | You get |
|---|---|---|---|
| Check the URLs in your sitemap | sitemap (default) | a sitemap URL or just your site's address; it finds the sitemaps itself (robots.txt Sitemap: lines, /sitemap.xml, sitemap index files, .gz sitemaps) | one row per URL listed in the sitemaps |
| Check a list of URLs (redirect maps, migrations, old links) | urlList | the URLs in urls (or one per line in url) | one row per URL |
| Find broken links on pages | pageLinks | page URLs, or a sitemap URL to check the links on every page it lists | one row per unique link or image found, and one row per page with its broken and redirected links |
pageLinks does not follow links to discover more pages of your site. To check the links on your whole site, give it your sitemap URL; the pages it lists are read and every unique link on them is checked once.
What it flags
URLs and pages (problems on sitemap and urlList rows, and on page rows in pageLinks):
| Code | Meaning |
|---|---|
broken | the URL ends in a 4xx or 5xx status |
redirected | the URL redirects; every hop is listed in redirectChain with its status |
redirect-loop, too-many-redirects | the redirects go round in a circle, or run past 10 hops |
noindex | noindex or none in <meta name="robots"> (or googlebot / bingbot), or in an X-Robots-Tag header |
canonical-elsewhere | the page's <link rel="canonical">, or a Link: rel="canonical" header (used on PDFs and feeds), points to a different URL |
blocked-by-robots, redirect-target-blocked-by-robots | the site's own robots.txt disallows the URL, or a redirect leads into a disallowed path; the URL is reported but not requested |
fetch-error | no usable answer: the name doesn't resolve, the connection fails or times out, the site's security certificate fails (expired, self-signed, incomplete chain), or the address has an invalid port (a URL that cannot be read at all is invalid-url) |
has-broken-links | pageLinks page rows: at least one link on the page is broken (see brokenLinks) |
Also possible: robots-unreachable (the site's robots.txt answered with a server error, so its URLs were skipped), other-host-not-checked, invalid-url, unexpected-status.
A sitemap should list only final, indexable, canonical URLs, so every flagged sitemap row is a URL you probably want to fix or take out of the sitemap.
Links (problems on pageLinks link rows):
| Code | Meaning |
|---|---|
broken | 4xx (404, 410 and so on) |
server-error | 5xx |
access-denied | 401, 403 or 429: the linked site refused our checker. Many large sites do this to all automated clients, so check these by hand before removing them |
fetch-error | no usable answer, as above (a site that no longer exists usually shows up here) |
redirected, redirect-loop, too-many-redirects | as above; finalUrl is where the link ends up |
blocked-by-robots | the linked site's robots.txt disallows our checker, so the link was not requested |
mailto:, tel:, javascript: and same-page # links are skipped, and so is a link whose address cannot be read as a URL at all. On a page row, links.broken counts links that are broken, server-error, fetch-error, redirect-loop or too-many-redirects; access-denied links are counted separately in links.accessDenied.
On-page rules
On every HTML page it reads (in all modes), seoIssues lists the rules the page fails, each with a severity, a detail (such as the length found) and a fix hint, and seo holds what was found: title, meta description, H1s, lang, JSON-LD / microdata types, image and missing-alt counts, hreflang codes.
| Rule | Severity | Fails when |
|---|---|---|
title-missing | error | no <title>, or it is empty (a <title> inside an inline SVG doesn't count) |
title-multiple | warning | more than one <title> |
title-too-long / title-too-short | notice | over 60 / under 30 characters |
meta-description-missing | warning | no <meta name="description">, or it is empty |
meta-description-multiple | warning | more than one |
meta-description-too-long / meta-description-too-short | notice | over 155 / under 70 characters |
h1-missing / h1-multiple | warning / notice | no <h1> / more than one |
canonical-missing / canonical-multiple | notice / warning | no <link rel="canonical"> / several different ones |
html-lang-missing | notice | no lang on <html> |
viewport-missing | notice | no <meta name="viewport"> |
img-alt-missing | notice | an <img> has no alt attribute (alt="" is fine) |
structured-data-missing | notice | no JSON-LD, microdata or RDFa |
json-ld-invalid | error | a JSON-LD block is not valid JSON |
hreflang-invalid | warning | an hreflang value is not a language code or x-default |
The length limits are the ones common audit tools use; search engines publish no fixed limits. Notices are worth a look, not necessarily a fix: several <h1> elements, for example, are valid HTML. Turn the rules off with checkOnPage: false.
Input
| Field | What it does | Default |
|---|---|---|
mode | sitemap, urlList or pageLinks (see above) | sitemap |
url | A sitemap, site or page URL; several allowed, separated by spaces or new lines. Also accepted: urls, startUrls, sitemapUrls, sitemapUrl | required |
maxUrls | Stop after this many unique URLs (in pageLinks: pages) | 5,000 |
maxLinks | pageLinks: stop after this many unique links | 2,000 |
checkExternalLinks | pageLinks: also check links to other sites (their robots.txt is obeyed) | true |
checkImages | pageLinks: also check <img src> | true |
onlyProblems | Leave healthy rows out of the dataset | false |
checkPageTags | Read each HTML page (first 1 MB) for meta robots and canonical | true |
checkOnPage | Run the on-page rules on every HTML page read | true |
maxSitemaps | Limit on sitemap files read (indexes included) | 100 |
respectRobotsTxt | Don't request URLs robots.txt disallows (they are still reported) | true |
sameHostOnly | sitemap mode: don't request URLs on other hosts than the site | true |
delaySecs | Spacing between requests to one host; a larger robots.txt Crawl-delay, up to 10 s, wins | 1 |
timeoutSecs | Give up on a request that takes longer than this | 20 |
maxConcurrencyPerHost | 1 or 2 requests at once per host | 2 |
The prefilled example checks a small demo sitemap we host in Apify storage (a sitemap index, a plain and a gzip sitemap, 6 listed URLs): one page marked noindex, one whose canonical points elsewhere, one deleted page (404), one http:// URL that redirects, and one duplicate. The storage host sends X-Robots-Tag: none on every file, so every demo page is also flagged noindex. Replace it with your own sitemap or site.
Examples:
{ "url": "https://example.com/sitemap_index.xml", "onlyProblems": true }
{ "mode": "urlList", "urls": ["http://example.com/old-page", "https://example.com/pricing"] }
{ "mode": "pageLinks", "url": "https://example.com/sitemap.xml", "maxLinks": 5000, "onlyProblems": true }
Output
A broken link found in pageLinks mode (a real row from one of our test runs on a public blog post):
{"type": "link","url": "http://stackoverflow.com/jobs","ok": false,"problems": ["redirected", "broken"],"finalStatus": 404,"finalUrl": "https://stackoverflow.com/jobs","redirectCount": 1,"redirectChain": [{"url": "http://stackoverflow.com/jobs", "status": 301, "location": "https://stackoverflow.com/jobs"}],"kind": "link","internal": false,"foundOn": [{"page": "https://www.joelonsoftware.com/2000/08/09/the-joel-test-12-steps-to-better-code/", "anchorText": "Stack Overflow Jobs", "nofollow": false}],"foundOnCount": 1,"error": null}
The page row for that post lists links (found, checked, broken, redirected, accessDenied, notChecked), brokenLinks and redirectedLinks (each with URL, anchor text and final status), plus its own status, canonical, noindex and seoIssues.
A sitemap or urlList row:
{"type": "url","url": "http://jekyllrb.com/docs","ok": false,"problems": ["redirected", "canonical-elsewhere"],"finalStatus": 200,"finalUrl": "http://jekyllrb.com/docs/","redirectCount": 1,"redirectChain": [{"url": "http://jekyllrb.com/docs", "status": 301, "location": "http://jekyllrb.com/docs/"}],"noindex": false,"canonicalUrl": "https://jekyllrb.com/docs/","canonicalElsewhere": true,"blockedByRobots": false,"seoIssues": [{"rule": "meta-description-too-long", "severity": "notice", "detail": "222 characters", "fix": "Descriptions over 155 characters are usually cut off."},{"rule": "h1-multiple", "severity": "notice", "detail": "2 <h1> elements", "fix": "Several <h1> elements: fine in HTML5, but one clear main heading is the usual advice."}]}
Rows also carry sitemap and lastmod (sitemap mode), contentType, method, error and checkedAt. The last row ("type": "summary", also saved as the OUTPUT record) gives the totals: sitemaps read (status, URL count, gzip, parse errors), URLs and links checked, duplicates, caps reached, and counts per problem, per on-page rule and per final status.
How it checks, and how we test it
- Each URL or link gets a
HEADrequest, or aGETif the server refusesHEAD. Redirects are followed one hop at a time (up to 10) so every hop is recorded. HTML pages that answer 200 are read (first 1 MB) for meta robots, canonical, the on-page rules and, inpageLinks, their links. - Each unique link is checked once, however many pages it appears on;
foundOnlists up to 20 of those pages with the anchor text used on each. - Broken sitemap XML is reported, and the URLs read before the error are still checked.
- Polite by design: robots.txt is read once per site and obeyed, also for links to other sites. At most 2 requests at a time per host, spaced by
delaySecsor the site'sCrawl-delay. User agentMadrascoSitemapHealth, so site owners can allow or block it by name. - We test it on public sites we didn't build and compare each result with a second tool:
curlfor statuses and redirects, and a separate implementation of each on-page rule on a different HTML parser; every difference is checked by hand. Latest check (27 Sep 2026, 7 sites on WordPress, Shopify, Ghost, Jekyll, Hugo and Drupal): the on-page rules agreed on 85 of 86 pages (on the other, the site sent a different page to each tool); link statuses and redirect chains agreed with curl on 91 of 97 links (the other 6: two where curl skipped a redirect hop that we recorded, three on a site that rate-limits automated clients and answered each tool differently, one we did not request because robots.txt disallowed it).
Limits
- It checks what the server returns to an automated HTTP client. It does not run JavaScript, so titles, canonicals,
noindexor links added by scripts are not seen. - Some sites answer automated clients differently from browsers: they refuse them (
access-denied) or send them somewhere else. Checkaccess-deniedlinks by hand before removing them. - A site whose security certificate is incomplete may still open in a browser but is reported as
fetch-error; theerrorfield says which certificate check failed. - It reports the signals a search engine reads; it can't tell you how a search engine will treat a page, and it gives no ranking advice or SEO score.
pageLinkschecks the pages you give it (or the pages your sitemap lists); it does not crawl your site to find pages.- Pages behind logins or bot walls come back as errors (such as 403); it does not try to get past them.
Use with AI agents
Call it with {"url": "<sitemap or site>"}, {"mode": "urlList", "urls": [...]} or {"mode": "pageLinks", "url": "<page or sitemap>"}. Read problems on each row; rows with "ok": true need nothing. The type: "summary" row gives the totals.
Pricing
No charge from us for now: you pay only Apify's platform usage of your run, which is small at the default 256 MB memory. We may add a per-URL price in a later release; any price is shown on this page and by Apify before you start a run.
Support
Use the Issues tab on this actor's page. Built and maintained by Madrasco, with AI assistance.