SEO Audit Crawler: Broken Links, llms.txt & AI Bot Check avatar

SEO Audit Crawler: Broken Links, llms.txt & AI Bot Check

Pricing

from $8.00 / 1,000 page auditeds

Go to Apify Store
SEO Audit Crawler: Broken Links, llms.txt & AI Bot Check

SEO Audit Crawler: Broken Links, llms.txt & AI Bot Check

SEO audit crawler for any website: a full site audit of titles, meta descriptions, headings, canonical, noindex, hreflang, schema and broken links, plus duplicate titles, AI crawlers blocked by robots.txt and llms.txt. 0-100 SEO score and fix hints. No proxies, no login.

Pricing

from $8.00 / 1,000 page auditeds

Rating

0.0

(0)

Developer

Lindenwerk Data

Lindenwerk Data

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

12 hours ago

Last modified

Share

What is SEO Audit Crawler?

A bulk website SEO audit crawler and on-page SEO checker. Crawl whole websites and audit every page in one run: title, meta description, headings, canonical, noindex / indexability, hreflang, Open Graph, structured data (JSON-LD and microdata), images without alt text, mixed content, response time, and broken internal and external links. Each website also gets a site summary: which AI crawlers (GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot, Google-Extended, CCBot, Applebot-Extended, Bytespider ...) robots.txt allows or blocks, llms.txt, XML sitemap, HTTP to HTTPS redirect, duplicate titles and descriptions, average score and top issues.

Every finding comes with a severity (error / warning / notice) and a plain-English fix, so the output is ready for a client report, a spreadsheet or an AI agent. No browser, no login, no proxies, no API key. USD 8 per 1,000 pages.

Who it's for

  • SEO agencies and freelancers. Audit 10 client sites overnight, export the issues view to Excel, and send each client a fix list. Schedule it monthly to catch new 404s and noindex accidents.
  • In-house marketing and web teams. Check a site after a relaunch or CMS migration: redirect chains, lost canonicals, pages that turned noindex, broken internal links, missing meta descriptions.
  • Lead-gen and sales teams selling web services. Score prospects' websites and open the conversation with concrete problems ("23 broken links, no meta descriptions on 40 pages").
  • Publishers and site owners deciding on AI crawlers. See at a glance which AI training bots and AI search bots your robots.txt blocks, whether that is what you intended, and whether you publish an llms.txt.
  • AI agents and LLM workflows. Flat, small JSON rows with issue ids and fix hints, callable over the Apify MCP server ("audit example.com and tell me the top 5 fixes").

What it checks

Per page (one row per URL, 40 rule types):

AreaChecks
Status & redirectsHTTP 4xx / 5xx, redirect chains (2+ hops), redirected internal URLs, HTTP instead of HTTPS, slow response (> 1.5 s), very large HTML
Title & descriptionMissing, too long / too short, multiple tags, duplicates across the site
HeadingsMissing / empty / multiple H1, skipped heading levels
Indexabilitynoindex in meta robots or X-Robots-Tag, canonical missing / pointing elsewhere / cross-domain / multiple, robots.txt blocking Googlebot, not in XML sitemap
Internationallang attribute, invalid hreflang codes, hreflang without self-reference
Social & schemaOpen Graph tags, schema.org types from JSON-LD and microdata, invalid JSON-LD
Content & mediaThin content (< 200 words on indexable pages), images without alt, mixed content (http:// resources on https pages), viewport meta
LinksInternal / external / nofollow counts, broken links with URL, status and anchor text

Each page gets a score from 0 to 100 (100 minus 10 per error type, 4 per warning type, 1 per notice type) and a sorted issueIds list for easy filtering.

Per website (one summary row):

  • robots.txt: found, sitemaps declared, crawl-delay, whether it disallows us.
  • AI crawler matrix: for 16 AI bots plus Googlebot and Bingbot, whether the homepage is allowed, whether the bot is named explicitly, and a one-line policy such as blocks AI training, allows AI search or
    allows all AI crawlers (no AI-specific rules)
    . Cloudflare-style Content-Signal lines (e.g. search=yes, ai-train=no) are reported too. We read the rules; we never pretend to be those bots.
  • llms.txt / llms-full.txt: present, size, title, summary and link count. Reported as information only: Google says it doesn't use llms.txt for Search.
  • XML sitemap: found (robots.txt or /sitemap.xml), sitemap indexes followed, URL count.
  • HTTPS: does http:// redirect to https://, and is it permanent?
  • Duplicates: groups of pages sharing a title or meta description (indexable pages only).
  • Roll-up: pages audited, status code counts, indexable pages, average score, issues by severity, top issues, broken link samples, what was skipped and why the crawl stopped.

How to run an SEO site audit

  1. Click Start with the prefilled example (10 pages of crawlee.dev) to see real output in about 15 seconds.
  2. Paste your own websites into Websites and set Max pages per site.
  3. Choose how pages are found: crawl follows internal links, sitemap reads the XML sitemap, list audits only the URLs you paste.
  4. Look at the results in the Output tab (views such as "Overview", "Issues & fixes", "Broken links" and "Site summary & AI crawlers") or download them as JSON, CSV or Excel for a client report.
  5. To re-audit regularly, save the input as a task and add a schedule (for example weekly).
  6. Send results on with Apify integrations (Slack, Google Drive, Make, Zapier, n8n or a webhook), or call the Actor through the Apify API or from AI agents via MCP.

Input

Only startUrls is required.

{
"startUrls": ["https://crawlee.dev"],
"maxPagesPerSite": 10
}

Agency batch: 5 client sites, 500 pages each, external links too:

{
"startUrls": ["client-one.de", "client-two.com", "client-three.at", "client-four.ch", "client-five.fr"],
"maxPagesPerSite": 500,
"checkLinks": "all"
}

Audit only the blog, from the sitemap:

{
"startUrls": ["example.com"],
"mode": "sitemap",
"includeUrlPatterns": ["/blog/"],
"maxPagesPerSite": 1000
}

Exact URL list (e.g. your top landing pages), no site row:

{
"startUrls": ["https://example.com/pricing", "https://example.com/features", "https://example.com/"],
"mode": "list",
"includeSiteSummary": false
}
FieldDefaultWhat it does
startUrlsnoneWebsites or URLs, one per line. Aliases: urls, websites, domains, url, so inputs from other SEO Actors work unchanged.
modecrawlcrawl follows internal links breadth-first; sitemap audits the URLs in the XML sitemap; list audits exactly the URLs given
maxPagesPerSite100Audited pages per website (max 10,000)
maxDepth5Crawl mode: clicks from the start URL
checkLinksinternalinternal, all (adds external links) or none
maxLinkChecksPerSite300Cap on extra HEAD/GET link checks per site (link checks are free)
includeSiteSummarytrueAdd the site row (AI crawlers, llms.txt, sitemap, HTTPS, duplicates, top issues)
includeSubdomainsfalseAlso crawl blog.example.com etc.
includeUrlPatterns / excludeUrlPatternsnoneText or * wildcard filters, e.g. /blog/ or */tag/*
skipUrlsWithQuerytrueIgnore ?filter= style URL variants while crawling
maxConcurrencyPerSite3Parallel requests per website (max 8; a robots.txt Crawl-delay forces 1)
maxConcurrentSites3Websites audited in parallel
requestTimeoutSecs20Per-request timeout

Output

Page rows (rowType: "page") and one site row per website (rowType: "site"). Shortened sample from a real run:

{
"rowType": "page",
"site": "crawlee.dev",
"url": "https://crawlee.dev/blog",
"httpStatus": 200,
"title": "Crawlee Blog - learn how to build better scrapers | Crawlee for JavaScript · Build reliable crawlers. Fast.",
"titleLength": 107,
"metaDescriptionLength": 151,
"h1Count": 0,
"canonicalStatus": "self",
"indexable": true,
"structuredDataTypes": ["Blog"],
"wordCount": 1492,
"inSitemap": true,
"brokenLinks": [],
"issues": [
{"id": "title-too-long", "severity": "warning", "message": "Title is 107 characters (over 60).", "fix": "Shorten the <title> to about 60 characters so it is not cut off in results."},
{"id": "h1-missing", "severity": "warning", "message": "No <h1> heading.", "fix": "Add one visible <h1> that states the page topic."}
],
"score": 92,
"issueIds": ["h1-missing", "title-too-long"],
"chargedEvent": "page-audited"
}
{
"rowType": "site",
"site": "crawlee.dev",
"title": "Site summary: 10 pages audited, average score 95.5, AI crawlers: allows all AI crawlers (no AI-specific rules)",
"siteSummary": {
"aiCrawlers": {"policy": "allows all AI crawlers (no AI-specific rules)", "trainingBlocked": [], "searchBlocked": [],
"googlebotAllowed": true, "contentSignals": []},
"llms": {"llmsTxt": {"found": true, "bytes": 30542, "hasTitle": true, "linkCount": 361}},
"sitemap": {"found": true, "urlCount": 4846},
"averageScore": 95.5,
"topIssues": [{"id": "title-too-long", "severity": "warning", "pages": 7}],
"duplicateDescriptions": [{"count": 6, "urls": ["https://crawlee.dev/", "https://crawlee.dev/js"]}],
"brokenLinks": {"total": 0, "linksChecked": 285},
"crawlStoppedReason": "page limit reached"
},
"chargedEvent": "site-summary"
}

Dataset views: Overview, On-page, Issues & fixes, Broken links and Site summary & AI crawlers. Export as CSV, Excel, JSON or XML, or read through the API. RUN_SUMMARY in the key-value store holds charged events, free rows, HTTP requests, bytes downloaded and runtime.

fetchStatus values: ok, http_error (a 4xx/5xx page: a finding, charged), and free rows not_html, timeout, unreachable, refused (non-public address) and error (e.g. redirect loop). URLs that robots.txt disallows for us are not requested and produce no row; they are counted in the site summary.

Pricing (pay per event)

EventPriceWhen
page-auditedUSD 0.008 (USD 8 per 1,000 pages)A URL answered with an HTML page (including 404/500 pages, which are findings) and was audited
site-summaryUSD 0.01One summary row per website (optional, includeSiteSummary)
Free0Timeouts, unreachable hosts, non-HTML files, robots.txt-blocked URLs, all link checks
Actor startUSD 0.00005Apify platform minimum per run

A 100-page site costs USD 0.81. A 1,000-page site costs USD 8.01. Set Maximum cost per run in Apify Console: the Actor reserves budget before each page and stops cleanly at the limit, so it never charges more (parallel sites share the limit).

Use with AI agents (MCP)

{
"mcpServers": {
"apify": {
"url": "https://mcp.apify.com/?tools=lindenwerk/seo-audit-crawler",
"headers": {"Authorization": "Bearer <APIFY_TOKEN>"}
}
}
}

Example prompts:

  • "Audit the first 50 pages of example.com and list the five fixes with the biggest impact."
  • "Which AI crawlers does example.com block, and does it have an llms.txt?"
  • "Find all broken internal links on example.de and the pages they are on."

For short agent answers, use maxPagesPerSite: 20 and read the site row first. Over HTTP: POST https://api.apify.com/v2/acts/lindenwerk~seo-audit-crawler/run-sync-get-dataset-items?token=<APIFY_TOKEN> with the input JSON as the body.

How it works and limits

  • Plain HTTP requests (no headless browser), so it's fast and cheap: in our test, 300 pages on 3 sites took 37 s at 512 MB. Content injected only by client-side JavaScript isn't seen. Server-rendered and statically generated sites (WordPress, TYPO3, Shopify, Next.js/Docusaurus SSG ...) are audited fully.
  • Link checks use HEAD, confirmed with GET when HEAD is refused or answers 4xx (many servers mishandle HEAD). Pages already crawled aren't requested twice. 401, 403 and 429 answers are "restricted", not broken, because many sites refuse data-center requests.
  • The AI crawler matrix evaluates robots.txt rules for each bot's token on the homepage. It doesn't test firewall or CDN bot blocking (e.g. Cloudflare's AI bot block), which can deny a bot even when robots.txt allows it.
  • Some sites block all cloud-server traffic. Their pages come back as http_error 403 or as free timeouts. After 10 timeouts or connection failures in a row the Actor stops crawling that site.

Data, privacy and responsible use

  • Honours robots.txt (RFC 9309) for the user agent token LindenwerkSEOAudit, including Crawl-delay (up to 10 s). If robots.txt can't be fetched because of a server error, the site isn't crawled.
  • Polite by default: 3 parallel requests per site, a clear user agent, no proxies, no login, no captcha solving, no bot-protection workarounds.
  • No personal data. It reads technical page metadata only: tags, headings, link targets, word counts. It does not extract names, email addresses or phone numbers. mailto: and tel: links are ignored.
  • Audit sites you own or are allowed to audit (clients, prospects' public pages at a polite rate). You are responsible for how you use the results.

FAQ

Is this a Screaming Frog alternative? For the core HTTP audit, yes: crawl, status codes, titles, descriptions, headings, canonicals, noindex, hreflang, structured data, broken links, duplicates. It runs in the cloud, on a schedule, and through an API or AI agent. It doesn't render JavaScript or connect to Search Console.

Why are some pages missing? They are disallowed by robots.txt for our bot, outside the start URL's site, behind a ?query (see skipUrlsWithQuery), beyond maxDepth, or the page limit was reached. The site row tells you which.

Does it change anything on my site? No. It only sends GET and HEAD requests.

Is llms.txt important for SEO? No. We report it because many teams ask, but robots.txt is what controls AI crawler access.

Examples

Other Lindenwerk Data Actors

Deutsch (Kurzfassung)

Dieser Actor crawlt ganze Websites und prüft jede Seite: Title, Meta-Description, Überschriften, Canonical, noindex, hreflang, Open Graph, strukturierte Daten, Bilder ohne Alt-Text, Mixed Content und defekte interne und externe Links. Jeder Fund kommt mit Schweregrad und konkretem Lösungshinweis. Pro Website gibt es eine Zusammenfassung: welche KI-Crawler (GPTBot, ClaudeBot, PerplexityBot, Google-Extended ...) die robots.txt erlaubt oder sperrt, llms.txt, XML-Sitemap, HTTP-zu-HTTPS-Weiterleitung, doppelte Titles und Descriptions und die häufigsten Probleme. Ideal für SEO-Agenturen, Relaunch-Checks und Website-Leadlisten. Preis: USD 0,008 pro geprüfter Seite und USD 0,01 pro Website-Zusammenfassung; nicht erreichbare Seiten und Link-Checks sind kostenlos. Der Actor beachtet robots.txt, nutzt keine Proxys und erhebt keine personenbezogenen Daten.

Changelog

See the Changelog tab. Version 0.1 is the first public release.

Made by Lindenwerk Data.