Sitemap and Broken Link Checker with On-Page SEO Audit avatar

Sitemap and Broken Link Checker with On-Page SEO Audit

Pricing

Pay per usage

Go to Apify Store
Sitemap and Broken Link Checker with On-Page SEO Audit

Sitemap and Broken Link Checker with On-Page SEO Audit

Check your sitemap URLs, any list of URLs, or every link on your pages: status, redirect chain, noindex, canonical, robots.txt, and broken links with the page and anchor text they sit on. On-page SEO rules (title, meta description, H1, canonical, alt text, schema, hreflang) with fix hints.

Pricing

Pay per usage

Rating

0.0

(0)

Developer

Jack Valmadre

Jack Valmadre

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

a day ago

Last modified

Share

Find the URLs and links on your site that search engines and visitors trip over, as data you can sort and filter:

  • Broken links on your pages, each with the page it sits on and its anchor text, so you know exactly where to fix it.
  • Sitemap problems: listed URLs that are broken, redirected, noindex, canonicalised to another URL, or blocked by your own robots.txt.
  • Status and redirect chains for any list of URLs, for example before and after a site migration.
  • On-page SEO issues on every HTML page it reads: missing or duplicate titles, meta descriptions, H1s and canonicals, missing lang, viewport, image alt text and structured data, invalid JSON-LD and hreflang values. Each issue comes with a severity and a one-line fix hint.

No login to your site, no browser extension. Give it a sitemap, a list of URLs, or the pages whose links you want checked.

Pick a mode

You want to…Set mode toGive itYou get
Check the URLs in your sitemapsitemap (default)a sitemap URL or just your site's address; it finds the sitemaps itself (robots.txt Sitemap: lines, /sitemap.xml, sitemap index files, .gz sitemaps)one row per URL listed in the sitemaps
Check a list of URLs (redirect maps, migrations, old links)urlListthe URLs in urls (or one per line in url)one row per URL
Find broken links on pagespageLinkspage URLs, or a sitemap URL to check the links on every page it listsone row per unique link or image found, and one row per page with its broken and redirected links

pageLinks does not follow links to discover more pages of your site. To check the links on your whole site, give it your sitemap URL; the pages it lists are read and every unique link on them is checked once.

What it flags

URLs and pages (problems on sitemap and urlList rows, and on page rows in pageLinks):

CodeMeaning
brokenthe URL ends in a 4xx or 5xx status
redirectedthe URL redirects; every hop is listed in redirectChain with its status
redirect-loop, too-many-redirectsthe redirects go round in a circle, or run past 10 hops
noindexnoindex or none in <meta name="robots"> (or googlebot / bingbot), or in an X-Robots-Tag header
canonical-elsewherethe page's <link rel="canonical">, or a Link: rel="canonical" header (used on PDFs and feeds), points to a different URL
blocked-by-robots, redirect-target-blocked-by-robotsthe site's own robots.txt disallows the URL, or a redirect leads into a disallowed path; the URL is reported but not requested
fetch-errorno usable answer: the name doesn't resolve, the connection fails or times out, the site's security certificate fails (expired, self-signed, incomplete chain), or the address has an invalid port (a URL that cannot be read at all is invalid-url)
has-broken-linkspageLinks page rows: at least one link on the page is broken (see brokenLinks)

Also possible: robots-unreachable (the site's robots.txt answered with a server error, so its URLs were skipped), other-host-not-checked, invalid-url, unexpected-status.

A sitemap should list only final, indexable, canonical URLs, so every flagged sitemap row is a URL you probably want to fix or take out of the sitemap.

Links (problems on pageLinks link rows):

CodeMeaning
broken4xx (404, 410 and so on)
server-error5xx
access-denied401, 403 or 429: the linked site refused our checker. Many large sites do this to all automated clients, so check these by hand before removing them
fetch-errorno usable answer, as above (a site that no longer exists usually shows up here)
redirected, redirect-loop, too-many-redirectsas above; finalUrl is where the link ends up
blocked-by-robotsthe linked site's robots.txt disallows our checker, so the link was not requested

mailto:, tel:, javascript: and same-page # links are skipped, and so is a link whose address cannot be read as a URL at all. On a page row, links.broken counts links that are broken, server-error, fetch-error, redirect-loop or too-many-redirects; access-denied links are counted separately in links.accessDenied.

On-page rules

On every HTML page it reads (in all modes), seoIssues lists the rules the page fails, each with a severity, a detail (such as the length found) and a fix hint, and seo holds what was found: title, meta description, H1s, lang, JSON-LD / microdata types, image and missing-alt counts, hreflang codes.

RuleSeverityFails when
title-missingerrorno <title>, or it is empty (a <title> inside an inline SVG doesn't count)
title-multiplewarningmore than one <title>
title-too-long / title-too-shortnoticeover 60 / under 30 characters
meta-description-missingwarningno <meta name="description">, or it is empty
meta-description-multiplewarningmore than one
meta-description-too-long / meta-description-too-shortnoticeover 155 / under 70 characters
h1-missing / h1-multiplewarning / noticeno <h1> / more than one
canonical-missing / canonical-multiplenotice / warningno <link rel="canonical"> / several different ones
html-lang-missingnoticeno lang on <html>
viewport-missingnoticeno <meta name="viewport">
img-alt-missingnoticean <img> has no alt attribute (alt="" is fine)
structured-data-missingnoticeno JSON-LD, microdata or RDFa
json-ld-invaliderrora JSON-LD block is not valid JSON
hreflang-invalidwarningan hreflang value is not a language code or x-default

The length limits are the ones common audit tools use; search engines publish no fixed limits. Notices are worth a look, not necessarily a fix: several <h1> elements, for example, are valid HTML. Turn the rules off with checkOnPage: false.

Input

FieldWhat it doesDefault
modesitemap, urlList or pageLinks (see above)sitemap
urlA sitemap, site or page URL; several allowed, separated by spaces or new lines. Also accepted: urls, startUrls, sitemapUrls, sitemapUrlrequired
maxUrlsStop after this many unique URLs (in pageLinks: pages)5,000
maxLinkspageLinks: stop after this many unique links2,000
checkExternalLinkspageLinks: also check links to other sites (their robots.txt is obeyed)true
checkImagespageLinks: also check <img src>true
onlyProblemsLeave healthy rows out of the datasetfalse
checkPageTagsRead each HTML page (first 1 MB) for meta robots and canonicaltrue
checkOnPageRun the on-page rules on every HTML page readtrue
maxSitemapsLimit on sitemap files read (indexes included)100
respectRobotsTxtDon't request URLs robots.txt disallows (they are still reported)true
sameHostOnlysitemap mode: don't request URLs on other hosts than the sitetrue
delaySecsSpacing between requests to one host; a larger robots.txt Crawl-delay, up to 10 s, wins1
timeoutSecsGive up on a request that takes longer than this20
maxConcurrencyPerHost1 or 2 requests at once per host2

The prefilled example checks a small demo sitemap we host in Apify storage (a sitemap index, a plain and a gzip sitemap, 6 listed URLs): one page marked noindex, one whose canonical points elsewhere, one deleted page (404), one http:// URL that redirects, and one duplicate. The storage host sends X-Robots-Tag: none on every file, so every demo page is also flagged noindex. Replace it with your own sitemap or site.

Examples:

{ "url": "https://example.com/sitemap_index.xml", "onlyProblems": true }
{ "mode": "urlList", "urls": ["http://example.com/old-page", "https://example.com/pricing"] }
{ "mode": "pageLinks", "url": "https://example.com/sitemap.xml", "maxLinks": 5000, "onlyProblems": true }

Output

A broken link found in pageLinks mode (a real row from one of our test runs on a public blog post):

{
"type": "link",
"url": "http://stackoverflow.com/jobs",
"ok": false,
"problems": ["redirected", "broken"],
"finalStatus": 404,
"finalUrl": "https://stackoverflow.com/jobs",
"redirectCount": 1,
"redirectChain": [{"url": "http://stackoverflow.com/jobs", "status": 301, "location": "https://stackoverflow.com/jobs"}],
"kind": "link",
"internal": false,
"foundOn": [{"page": "https://www.joelonsoftware.com/2000/08/09/the-joel-test-12-steps-to-better-code/", "anchorText": "Stack Overflow Jobs", "nofollow": false}],
"foundOnCount": 1,
"error": null
}

The page row for that post lists links (found, checked, broken, redirected, accessDenied, notChecked), brokenLinks and redirectedLinks (each with URL, anchor text and final status), plus its own status, canonical, noindex and seoIssues.

A sitemap or urlList row:

{
"type": "url",
"url": "http://jekyllrb.com/docs",
"ok": false,
"problems": ["redirected", "canonical-elsewhere"],
"finalStatus": 200,
"finalUrl": "http://jekyllrb.com/docs/",
"redirectCount": 1,
"redirectChain": [{"url": "http://jekyllrb.com/docs", "status": 301, "location": "http://jekyllrb.com/docs/"}],
"noindex": false,
"canonicalUrl": "https://jekyllrb.com/docs/",
"canonicalElsewhere": true,
"blockedByRobots": false,
"seoIssues": [
{"rule": "meta-description-too-long", "severity": "notice", "detail": "222 characters", "fix": "Descriptions over 155 characters are usually cut off."},
{"rule": "h1-multiple", "severity": "notice", "detail": "2 <h1> elements", "fix": "Several <h1> elements: fine in HTML5, but one clear main heading is the usual advice."}
]
}

Rows also carry sitemap and lastmod (sitemap mode), contentType, method, error and checkedAt. The last row ("type": "summary", also saved as the OUTPUT record) gives the totals: sitemaps read (status, URL count, gzip, parse errors), URLs and links checked, duplicates, caps reached, and counts per problem, per on-page rule and per final status.

How it checks, and how we test it

  • Each URL or link gets a HEAD request, or a GET if the server refuses HEAD. Redirects are followed one hop at a time (up to 10) so every hop is recorded. HTML pages that answer 200 are read (first 1 MB) for meta robots, canonical, the on-page rules and, in pageLinks, their links.
  • Each unique link is checked once, however many pages it appears on; foundOn lists up to 20 of those pages with the anchor text used on each.
  • Broken sitemap XML is reported, and the URLs read before the error are still checked.
  • Polite by design: robots.txt is read once per site and obeyed, also for links to other sites. At most 2 requests at a time per host, spaced by delaySecs or the site's Crawl-delay. User agent MadrascoSitemapHealth, so site owners can allow or block it by name.
  • We test it on public sites we didn't build and compare each result with a second tool: curl for statuses and redirects, and a separate implementation of each on-page rule on a different HTML parser; every difference is checked by hand. Latest check (27 Sep 2026, 7 sites on WordPress, Shopify, Ghost, Jekyll, Hugo and Drupal): the on-page rules agreed on 85 of 86 pages (on the other, the site sent a different page to each tool); link statuses and redirect chains agreed with curl on 91 of 97 links (the other 6: two where curl skipped a redirect hop that we recorded, three on a site that rate-limits automated clients and answered each tool differently, one we did not request because robots.txt disallowed it).

Limits

  • It checks what the server returns to an automated HTTP client. It does not run JavaScript, so titles, canonicals, noindex or links added by scripts are not seen.
  • Some sites answer automated clients differently from browsers: they refuse them (access-denied) or send them somewhere else. Check access-denied links by hand before removing them.
  • A site whose security certificate is incomplete may still open in a browser but is reported as fetch-error; the error field says which certificate check failed.
  • It reports the signals a search engine reads; it can't tell you how a search engine will treat a page, and it gives no ranking advice or SEO score.
  • pageLinks checks the pages you give it (or the pages your sitemap lists); it does not crawl your site to find pages.
  • Pages behind logins or bot walls come back as errors (such as 403); it does not try to get past them.

Use with AI agents

Call it with {"url": "<sitemap or site>"}, {"mode": "urlList", "urls": [...]} or {"mode": "pageLinks", "url": "<page or sitemap>"}. Read problems on each row; rows with "ok": true need nothing. The type: "summary" row gives the totals.

Pricing

No charge from us for now: you pay only Apify's platform usage of your run, which is small at the default 256 MB memory. We may add a per-URL price in a later release; any price is shown on this page and by Apify before you start a run.

Support

Use the Issues tab on this actor's page. Built and maintained by Madrasco, with AI assistance.