Website URL Crawler & Link Extractor avatar

Website URL Crawler & Link Extractor

Pricing

from $0.15 / 1,000 website urls

Go to Apify Store
Website URL Crawler & Link Extractor

Website URL Crawler & Link Extractor

Crawl public websites and collect URLs from rendered navigation and sitemaps. Export a link map with source pages, depth, anchor text, HTTP facts, crawl status, and available sitemap metadata.

Pricing

from $0.15 / 1,000 website urls

Rating

0.0

(0)

Developer

Maxime Dupré

Maxime Dupré

Maintained by Community

Actor stats

0

Bookmarked

42

Total users

7

Monthly active users

5 days ago

Last modified

Share

Website URL Crawler is a web crawler for SEO specialists, developers, and site owners who need a structured map of a public website. It collects accepted URLs from rendered navigation and public sitemaps, with hierarchy, classification, crawl state, HTTP facts for loaded pages, and available sitemap metadata.

Each saved row is one accepted URL. Its source fields describe the discovery evidence attached to that saved match. Fields include:

  • startUrl and normalized url
  • discoverySources: rendered, sitemap, or both
  • parentUrl, depth, and readable anchorText when rendered navigation provides them. Anchor text contains visible link text without literal HTML element markup.
  • relationship and linkType
  • crawlStatus: crawled or discovered
  • httpStatusCode, finalUrl, and contentType for loaded pages
  • sitemapUrl, lastmod, priority, and changefreq when supplied by a sitemap

Unavailable source metadata is null.

🚀 Map a public website

  1. Add one or more public website URLs, page URLs, or bare domains.
  2. Choose the crawl scope.
  3. Set row, page, depth, and per-page link limits.
  4. Run the Actor and open its dataset.

Bare domains such as example.com are normalized to HTTPS. The Actor crawls public pages only; it does not log in, submit forms, scrape full page content, or check search-engine index status.

🧾 Input

Website URLs is required. Add at least one public website, domain, or page URL. The other public fields are:

Input fields

FieldTypeWhat it does
urlsarray of objectsRequired. Adds one or more public websites, domains, or page URLs to crawl.
urls[].urlstringRequired URL or domain for one crawl start.
keywordsarray of stringsKeeps URLs that contain all listed words or path parts. Leave empty to keep every matching URL.
crawlScopestringChooses same-host, same-domain, or include-external-discovered. Same host and same domain save only internal URLs. The external option also saves external URLs as discovered but never crawls them.
assetPolicystringChooses pages-only, include-documents, or include-all-links for the document, media, and other link types saved with page URLs.
ignoredExtensionsarray of stringsSkips links ending with these extensions unless assetPolicy is include-all-links.
maxResultsintegerOptional positive limit for accepted URL rows across the run. Leave empty to return all available URLs until the source is exhausted within the other crawl limits.
maxPagesPerStartUrlintegerPositive limit up to 10,000 for rendered pages opened per website. Defaults to 5.
maxDepthintegerNumber of link levels to follow from each website, from 0 to 30. Defaults to 2.
maxLinksPerPageintegerPositive limit up to 1,000 for links considered from each loaded page. Defaults to 120.

Input example

This example is copied from a successful current-beta run:

{
"urls": [
{
"url": "https://crawlee.dev"
}
],
"keywords": [],
"crawlScope": "same-domain",
"assetPolicy": "include-documents",
"ignoredExtensions": [
"jpg",
"jpeg",
"png",
"gif",
"webp",
"svg",
"zip",
"mp4",
"mp3"
],
"maxResults": 250,
"maxPagesPerStartUrl": 5,
"maxDepth": 2,
"maxLinksPerPage": 120
}

🧾 Output

Each accepted URL uses the same output shape. Source or HTTP metadata is null when it is not available.

Output fields

FieldTypeWhat it does
startUrlstringWebsite or domain that led to this URL.
urlstringAccepted URL.
discoverySourcesarray of stringsDiscovery values attached when the URL was accepted: rendered, sitemap, or both.
parentUrlstring or nullPage where the link was found, when known.
depthinteger or nullNumber of link steps from the start URL, when known.
anchorTextstring or nullReadable visible text for the rendered link when known, or null when unavailable. It never includes literal HTML element markup.
relationshipstringinternal or external relationship to the submitted website.
linkTypestringpage, document, media, or other.
crawlStatusstringcrawled when the page was loaded, or discovered when it was only found as a link.
httpStatusCodeinteger or nullHTTP status for a loaded page, when known.
finalUrlstring or nullURL reached after redirects, when known.
contentTypestring or nullHTTP content type for a loaded page, when known.
sitemapUrlstring or nullPublic sitemap that listed this URL, when known.
lastmodstring or nullLast-modified value from the sitemap, when known.
prioritynumber or nullSitemap priority value, when known.
changefreqstring or nullSitemap change frequency, when known.

Output example

This complete row is copied from a successful current-beta run:

{
"startUrl": "https://crawlee.dev/",
"url": "https://crawlee.dev/js",
"discoverySources": [
"sitemap",
"rendered"
],
"parentUrl": "https://crawlee.dev/",
"depth": 1,
"anchorText": "Get started with JS",
"relationship": "internal",
"linkType": "page",
"crawlStatus": "crawled",
"httpStatusCode": 200,
"finalUrl": "https://crawlee.dev/js",
"contentType": "text/html; charset=utf-8",
"sitemapUrl": "https://crawlee.dev/sitemap.xml",
"lastmod": null,
"priority": 0.5,
"changefreq": "weekly"
}

Rows are available through the Apify dataset and its standard export formats.

💳 Pricing

This Actor uses pay-per-event pricing. The Website URL event is charged once for each accepted URL saved to the dataset. Discovered URLs that are not accepted or saved do not create this event. Start with a small Max URL rows value, inspect the output, and then broaden the limits if needed.

🔌 Integrations

Use Apify datasets, the API, schedules, webhooks, and platform integrations to process or deliver the URL inventory. This walkthrough shows how to connect an Actor to other services:

❓ FAQ

Why can a URL be marked discovered without HTTP fields?

A URL can come from a sitemap or an unvisited link. Only pages that the crawl loads receive crawled state and available HTTP response facts.

When are external URLs saved?

Same host and Same domain save only internal URLs. Include external URLs as discovered also saves external URLs, but the Actor does not crawl them.

Can I crawl only one submitted page?

Set Max crawl depth to 0 and use a low Max pages per website. Public sitemap discovery can still add sitemap-backed URLs.

Does it parse sitemap indexes?

Yes. It reads public XML URL sitemaps and nested sitemap indexes. Sitemap metadata is included only when the source supplies it.

No. It provides HTTP facts for pages actually loaded during the crawl, which can support a downstream broken-link workflow. It does not promise an HTTP check for every discovered URL.

Does it scrape page text or private pages?

No. It extracts URL and link evidence from public pages and sitemaps. It does not return full page content, log in, or access private pages.

📝 Changelog

v1.0

  • Added sitemap discovery, sitemap metadata, URL keyword filtering, total row limits, and the merged URL inventory output.

v0.0

  • Initial release.

🆘 Support

For issues, questions, or feature requests, file a ticket and I'll fix or implement it in less than 24h 🫡

Made with ❤️ by Maxime Dupré