SEO Audit Crawler: Broken Links, llms.txt & AI Bot Check
Pricing
from $8.00 / 1,000 page auditeds
SEO Audit Crawler: Broken Links, llms.txt & AI Bot Check
SEO audit crawler for any website: a full site audit of titles, meta descriptions, headings, canonical, noindex, hreflang, schema and broken links, plus duplicate titles, AI crawlers blocked by robots.txt and llms.txt. 0-100 SEO score and fix hints. No proxies, no login.
Pricing
from $8.00 / 1,000 page auditeds
Rating
0.0
(0)
Developer
Lindenwerk Data
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
12 hours ago
Last modified
Categories
Share
What is SEO Audit Crawler?
A bulk website SEO audit crawler and on-page SEO checker. Crawl whole websites and audit every page in one run: title, meta description, headings, canonical, noindex / indexability, hreflang, Open Graph, structured data (JSON-LD and microdata), images without alt text, mixed content, response time, and broken internal and external links. Each website also gets a site summary: which AI crawlers (GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot, Google-Extended, CCBot, Applebot-Extended, Bytespider ...) robots.txt allows or blocks, llms.txt, XML sitemap, HTTP to HTTPS redirect, duplicate titles and descriptions, average score and top issues.
Every finding comes with a severity (error / warning / notice) and a plain-English fix, so the output is ready for a client report, a spreadsheet or an AI agent. No browser, no login, no proxies, no API key. USD 8 per 1,000 pages.
Who it's for
- SEO agencies and freelancers. Audit 10 client sites overnight, export the issues view to Excel, and send each client a fix list. Schedule it monthly to catch new 404s and noindex accidents.
- In-house marketing and web teams. Check a site after a relaunch or CMS migration: redirect chains, lost
canonicals, pages that turned
noindex, broken internal links, missing meta descriptions. - Lead-gen and sales teams selling web services. Score prospects' websites and open the conversation with concrete problems ("23 broken links, no meta descriptions on 40 pages").
- Publishers and site owners deciding on AI crawlers. See at a glance which AI training bots and AI search bots your robots.txt blocks, whether that is what you intended, and whether you publish an llms.txt.
- AI agents and LLM workflows. Flat, small JSON rows with issue ids and fix hints, callable over the Apify MCP server ("audit example.com and tell me the top 5 fixes").
What it checks
Per page (one row per URL, 40 rule types):
| Area | Checks |
|---|---|
| Status & redirects | HTTP 4xx / 5xx, redirect chains (2+ hops), redirected internal URLs, HTTP instead of HTTPS, slow response (> 1.5 s), very large HTML |
| Title & description | Missing, too long / too short, multiple tags, duplicates across the site |
| Headings | Missing / empty / multiple H1, skipped heading levels |
| Indexability | noindex in meta robots or X-Robots-Tag, canonical missing / pointing elsewhere / cross-domain / multiple, robots.txt blocking Googlebot, not in XML sitemap |
| International | lang attribute, invalid hreflang codes, hreflang without self-reference |
| Social & schema | Open Graph tags, schema.org types from JSON-LD and microdata, invalid JSON-LD |
| Content & media | Thin content (< 200 words on indexable pages), images without alt, mixed content (http:// resources on https pages), viewport meta |
| Links | Internal / external / nofollow counts, broken links with URL, status and anchor text |
Each page gets a score from 0 to 100 (100 minus 10 per error type, 4 per warning type, 1 per notice type) and a
sorted issueIds list for easy filtering.
Per website (one summary row):
- robots.txt: found, sitemaps declared, crawl-delay, whether it disallows us.
- AI crawler matrix: for 16 AI bots plus Googlebot and Bingbot, whether the homepage is allowed, whether the bot is
named explicitly, and a one-line policy such as
blocks AI training, allows AI searchor. Cloudflare-styleallows all AI crawlers (no AI-specific rules)Content-Signallines (e.g.search=yes, ai-train=no) are reported too. We read the rules; we never pretend to be those bots. - llms.txt / llms-full.txt: present, size, title, summary and link count. Reported as information only: Google says it doesn't use llms.txt for Search.
- XML sitemap: found (robots.txt or /sitemap.xml), sitemap indexes followed, URL count.
- HTTPS: does
http://redirect tohttps://, and is it permanent? - Duplicates: groups of pages sharing a title or meta description (indexable pages only).
- Roll-up: pages audited, status code counts, indexable pages, average score, issues by severity, top issues, broken link samples, what was skipped and why the crawl stopped.
How to run an SEO site audit
- Click Start with the prefilled example (10 pages of crawlee.dev) to see real output in about 15 seconds.
- Paste your own websites into Websites and set Max pages per site.
- Choose how pages are found:
crawlfollows internal links,sitemapreads the XML sitemap,listaudits only the URLs you paste. - Look at the results in the Output tab (views such as "Overview", "Issues & fixes", "Broken links" and "Site summary & AI crawlers") or download them as JSON, CSV or Excel for a client report.
- To re-audit regularly, save the input as a task and add a schedule (for example weekly).
- Send results on with Apify integrations (Slack, Google Drive, Make, Zapier, n8n or a webhook), or call the Actor through the Apify API or from AI agents via MCP.
Input
Only startUrls is required.
{"startUrls": ["https://crawlee.dev"],"maxPagesPerSite": 10}
Agency batch: 5 client sites, 500 pages each, external links too:
{"startUrls": ["client-one.de", "client-two.com", "client-three.at", "client-four.ch", "client-five.fr"],"maxPagesPerSite": 500,"checkLinks": "all"}
Audit only the blog, from the sitemap:
{"startUrls": ["example.com"],"mode": "sitemap","includeUrlPatterns": ["/blog/"],"maxPagesPerSite": 1000}
Exact URL list (e.g. your top landing pages), no site row:
{"startUrls": ["https://example.com/pricing", "https://example.com/features", "https://example.com/"],"mode": "list","includeSiteSummary": false}
| Field | Default | What it does |
|---|---|---|
startUrls | none | Websites or URLs, one per line. Aliases: urls, websites, domains, url, so inputs from other SEO Actors work unchanged. |
mode | crawl | crawl follows internal links breadth-first; sitemap audits the URLs in the XML sitemap; list audits exactly the URLs given |
maxPagesPerSite | 100 | Audited pages per website (max 10,000) |
maxDepth | 5 | Crawl mode: clicks from the start URL |
checkLinks | internal | internal, all (adds external links) or none |
maxLinkChecksPerSite | 300 | Cap on extra HEAD/GET link checks per site (link checks are free) |
includeSiteSummary | true | Add the site row (AI crawlers, llms.txt, sitemap, HTTPS, duplicates, top issues) |
includeSubdomains | false | Also crawl blog.example.com etc. |
includeUrlPatterns / excludeUrlPatterns | none | Text or * wildcard filters, e.g. /blog/ or */tag/* |
skipUrlsWithQuery | true | Ignore ?filter= style URL variants while crawling |
maxConcurrencyPerSite | 3 | Parallel requests per website (max 8; a robots.txt Crawl-delay forces 1) |
maxConcurrentSites | 3 | Websites audited in parallel |
requestTimeoutSecs | 20 | Per-request timeout |
Output
Page rows (rowType: "page") and one site row per website (rowType: "site"). Shortened sample from a real run:
{"rowType": "page","site": "crawlee.dev","url": "https://crawlee.dev/blog","httpStatus": 200,"title": "Crawlee Blog - learn how to build better scrapers | Crawlee for JavaScript · Build reliable crawlers. Fast.","titleLength": 107,"metaDescriptionLength": 151,"h1Count": 0,"canonicalStatus": "self","indexable": true,"structuredDataTypes": ["Blog"],"wordCount": 1492,"inSitemap": true,"brokenLinks": [],"issues": [{"id": "title-too-long", "severity": "warning", "message": "Title is 107 characters (over 60).", "fix": "Shorten the <title> to about 60 characters so it is not cut off in results."},{"id": "h1-missing", "severity": "warning", "message": "No <h1> heading.", "fix": "Add one visible <h1> that states the page topic."}],"score": 92,"issueIds": ["h1-missing", "title-too-long"],"chargedEvent": "page-audited"}
{"rowType": "site","site": "crawlee.dev","title": "Site summary: 10 pages audited, average score 95.5, AI crawlers: allows all AI crawlers (no AI-specific rules)","siteSummary": {"aiCrawlers": {"policy": "allows all AI crawlers (no AI-specific rules)", "trainingBlocked": [], "searchBlocked": [],"googlebotAllowed": true, "contentSignals": []},"llms": {"llmsTxt": {"found": true, "bytes": 30542, "hasTitle": true, "linkCount": 361}},"sitemap": {"found": true, "urlCount": 4846},"averageScore": 95.5,"topIssues": [{"id": "title-too-long", "severity": "warning", "pages": 7}],"duplicateDescriptions": [{"count": 6, "urls": ["https://crawlee.dev/", "https://crawlee.dev/js"]}],"brokenLinks": {"total": 0, "linksChecked": 285},"crawlStoppedReason": "page limit reached"},"chargedEvent": "site-summary"}
Dataset views: Overview, On-page, Issues & fixes, Broken links and Site summary & AI crawlers.
Export as CSV, Excel, JSON or XML, or read through the API. RUN_SUMMARY in the key-value store holds charged events,
free rows, HTTP requests, bytes downloaded and runtime.
fetchStatus values: ok, http_error (a 4xx/5xx page: a finding, charged), and free rows not_html, timeout,
unreachable, refused (non-public address) and error (e.g. redirect loop). URLs that robots.txt disallows for
us are not requested and produce no row; they are counted in the site summary.
Pricing (pay per event)
| Event | Price | When |
|---|---|---|
page-audited | USD 0.008 (USD 8 per 1,000 pages) | A URL answered with an HTML page (including 404/500 pages, which are findings) and was audited |
site-summary | USD 0.01 | One summary row per website (optional, includeSiteSummary) |
| Free | 0 | Timeouts, unreachable hosts, non-HTML files, robots.txt-blocked URLs, all link checks |
| Actor start | USD 0.00005 | Apify platform minimum per run |
A 100-page site costs USD 0.81. A 1,000-page site costs USD 8.01. Set Maximum cost per run in Apify Console: the Actor reserves budget before each page and stops cleanly at the limit, so it never charges more (parallel sites share the limit).
Use with AI agents (MCP)
{"mcpServers": {"apify": {"url": "https://mcp.apify.com/?tools=lindenwerk/seo-audit-crawler","headers": {"Authorization": "Bearer <APIFY_TOKEN>"}}}}
Example prompts:
- "Audit the first 50 pages of example.com and list the five fixes with the biggest impact."
- "Which AI crawlers does example.com block, and does it have an llms.txt?"
- "Find all broken internal links on example.de and the pages they are on."
For short agent answers, use maxPagesPerSite: 20 and read the site row first. Over HTTP:
POST https://api.apify.com/v2/acts/lindenwerk~seo-audit-crawler/run-sync-get-dataset-items?token=<APIFY_TOKEN>
with the input JSON as the body.
How it works and limits
- Plain HTTP requests (no headless browser), so it's fast and cheap: in our test, 300 pages on 3 sites took 37 s at 512 MB. Content injected only by client-side JavaScript isn't seen. Server-rendered and statically generated sites (WordPress, TYPO3, Shopify, Next.js/Docusaurus SSG ...) are audited fully.
- Link checks use HEAD, confirmed with GET when HEAD is refused or answers 4xx (many servers mishandle HEAD). Pages already crawled aren't requested twice. 401, 403 and 429 answers are "restricted", not broken, because many sites refuse data-center requests.
- The AI crawler matrix evaluates robots.txt rules for each bot's token on the homepage. It doesn't test firewall or CDN bot blocking (e.g. Cloudflare's AI bot block), which can deny a bot even when robots.txt allows it.
- Some sites block all cloud-server traffic. Their pages come back as
http_error403 or as free timeouts. After 10 timeouts or connection failures in a row the Actor stops crawling that site.
Data, privacy and responsible use
- Honours robots.txt (RFC 9309) for the user agent token
LindenwerkSEOAudit, including Crawl-delay (up to 10 s). If robots.txt can't be fetched because of a server error, the site isn't crawled. - Polite by default: 3 parallel requests per site, a clear user agent, no proxies, no login, no captcha solving, no bot-protection workarounds.
- No personal data. It reads technical page metadata only: tags, headings, link targets, word counts. It does not
extract names, email addresses or phone numbers.
mailto:andtel:links are ignored. - Audit sites you own or are allowed to audit (clients, prospects' public pages at a polite rate). You are responsible for how you use the results.
FAQ
Is this a Screaming Frog alternative? For the core HTTP audit, yes: crawl, status codes, titles, descriptions, headings, canonicals, noindex, hreflang, structured data, broken links, duplicates. It runs in the cloud, on a schedule, and through an API or AI agent. It doesn't render JavaScript or connect to Search Console.
Why are some pages missing? They are disallowed by robots.txt for our bot, outside the start URL's site, behind
a ?query (see skipUrlsWithQuery), beyond maxDepth, or the page limit was reached. The site row tells you which.
Does it change anything on my site? No. It only sends GET and HEAD requests.
Is llms.txt important for SEO? No. We report it because many teams ask, but robots.txt is what controls AI crawler access.
Examples
- Audit a website for SEO issues and broken links: a published example task with a ready-made input. Open it, adjust the input and run it in your own Apify account.
Other Lindenwerk Data Actors
- German & EU Tenders: Public Procurement Monitor: Ausschreibungen from oeffentlichevergabe.de and TED, deduplicated and scored to your CPV codes and keywords.
- France Tenders Monitor: BOAMP, TED & Marchés Publics: marchés publics from BOAMP and TED in one deduplicated, scored list.
- UK Tenders Monitor: Government Contracts & Find a Tender Alerts: Find a Tender notices (above and below threshold) as daily tender alerts, scored to your profile.
- SAM.gov Government Contract Opportunities Monitor: US federal bids and RFPs from SAM.gov, scored to your NAICS and set-asides.
- Website Technology & Tech Stack Detector: BuiltWith Alternative: CMS, shop system, analytics, consent manager and email provider of any website, in bulk.
- PDF to Markdown & RAG Chunks: Document Parser (DOCX/PPTX/XLSX): PDF, Word, PowerPoint and Excel to clean Markdown with inline tables and page-cited RAG chunks.
Deutsch (Kurzfassung)
Dieser Actor crawlt ganze Websites und prüft jede Seite: Title, Meta-Description, Überschriften, Canonical, noindex, hreflang, Open Graph, strukturierte Daten, Bilder ohne Alt-Text, Mixed Content und defekte interne und externe Links. Jeder Fund kommt mit Schweregrad und konkretem Lösungshinweis. Pro Website gibt es eine Zusammenfassung: welche KI-Crawler (GPTBot, ClaudeBot, PerplexityBot, Google-Extended ...) die robots.txt erlaubt oder sperrt, llms.txt, XML-Sitemap, HTTP-zu-HTTPS-Weiterleitung, doppelte Titles und Descriptions und die häufigsten Probleme. Ideal für SEO-Agenturen, Relaunch-Checks und Website-Leadlisten. Preis: USD 0,008 pro geprüfter Seite und USD 0,01 pro Website-Zusammenfassung; nicht erreichbare Seiten und Link-Checks sind kostenlos. Der Actor beachtet robots.txt, nutzt keine Proxys und erhebt keine personenbezogenen Daten.
Changelog
See the Changelog tab. Version 0.1 is the first public release.
Made by Lindenwerk Data.