Website URL Crawler & Link Extractor
Pricing
from $0.15 / 1,000 website urls
Website URL Crawler & Link Extractor
Crawl public websites and collect URLs from rendered navigation and sitemaps. Export a link map with source pages, depth, anchor text, HTTP facts, crawl status, and available sitemap metadata.
Pricing
from $0.15 / 1,000 website urls
Rating
0.0
(0)
Developer
Maxime Dupré
Maintained by CommunityActor stats
0
Bookmarked
42
Total users
7
Monthly active users
5 days ago
Last modified
Categories
Share
🔗 Build a website URL inventory from links and sitemaps
Website URL Crawler is a web crawler for SEO specialists, developers, and site owners who need a structured map of a public website. It collects accepted URLs from rendered navigation and public sitemaps, with hierarchy, classification, crawl state, HTTP facts for loaded pages, and available sitemap metadata.
- Map available page URLs from a public site with Crawl Website for All URLs.
- Explore a public site and review accepted URL rows with Web Crawler.
- Build a clean URL list from a public website with URL Extractor Online.
- Follow link paths and keep source-page context with Link Crawler.
- Review parent pages, visible labels, and crawl depth with Internal Linking Review.
- Collect URLs and source metadata from XML sitemaps with Sitemap URL Extractor.
- Gather paths before redirects or a site move with Website Migration URL Inventory.
- Export loaded-page HTTP facts for a downstream check with Broken Link Workflow Crawler.
📦 URL inventory and link context
Each saved row is one accepted URL. Its source fields describe the discovery evidence attached to that saved match. Fields include:
startUrland normalizedurldiscoverySources:rendered,sitemap, or bothparentUrl,depth, and readableanchorTextwhen rendered navigation provides them. Anchor text contains visible link text without literal HTML element markup.relationshipandlinkTypecrawlStatus:crawledordiscoveredhttpStatusCode,finalUrl, andcontentTypefor loaded pagessitemapUrl,lastmod,priority, andchangefreqwhen supplied by a sitemap
Unavailable source metadata is null.
🚀 Map a public website
- Add one or more public website URLs, page URLs, or bare domains.
- Choose the crawl scope.
- Set row, page, depth, and per-page link limits.
- Run the Actor and open its dataset.
Bare domains such as example.com are normalized to HTTPS. The Actor crawls public pages only; it does not log in, submit forms, scrape full page content, or check search-engine index status.
🧾 Input
Website URLs is required. Add at least one public website, domain, or page URL. The other public fields are:
Input fields
| Field | Type | What it does |
|---|---|---|
urls | array of objects | Required. Adds one or more public websites, domains, or page URLs to crawl. |
urls[].url | string | Required URL or domain for one crawl start. |
keywords | array of strings | Keeps URLs that contain all listed words or path parts. Leave empty to keep every matching URL. |
crawlScope | string | Chooses same-host, same-domain, or include-external-discovered. Same host and same domain save only internal URLs. The external option also saves external URLs as discovered but never crawls them. |
assetPolicy | string | Chooses pages-only, include-documents, or include-all-links for the document, media, and other link types saved with page URLs. |
ignoredExtensions | array of strings | Skips links ending with these extensions unless assetPolicy is include-all-links. |
maxResults | integer | Optional positive limit for accepted URL rows across the run. Leave empty to return all available URLs until the source is exhausted within the other crawl limits. |
maxPagesPerStartUrl | integer | Positive limit up to 10,000 for rendered pages opened per website. Defaults to 5. |
maxDepth | integer | Number of link levels to follow from each website, from 0 to 30. Defaults to 2. |
maxLinksPerPage | integer | Positive limit up to 1,000 for links considered from each loaded page. Defaults to 120. |
Input example
This example is copied from a successful current-beta run:
{"urls": [{"url": "https://crawlee.dev"}],"keywords": [],"crawlScope": "same-domain","assetPolicy": "include-documents","ignoredExtensions": ["jpg","jpeg","png","gif","webp","svg","zip","mp4","mp3"],"maxResults": 250,"maxPagesPerStartUrl": 5,"maxDepth": 2,"maxLinksPerPage": 120}
🧾 Output
Each accepted URL uses the same output shape. Source or HTTP metadata is null when it is not available.
Output fields
| Field | Type | What it does |
|---|---|---|
startUrl | string | Website or domain that led to this URL. |
url | string | Accepted URL. |
discoverySources | array of strings | Discovery values attached when the URL was accepted: rendered, sitemap, or both. |
parentUrl | string or null | Page where the link was found, when known. |
depth | integer or null | Number of link steps from the start URL, when known. |
anchorText | string or null | Readable visible text for the rendered link when known, or null when unavailable. It never includes literal HTML element markup. |
relationship | string | internal or external relationship to the submitted website. |
linkType | string | page, document, media, or other. |
crawlStatus | string | crawled when the page was loaded, or discovered when it was only found as a link. |
httpStatusCode | integer or null | HTTP status for a loaded page, when known. |
finalUrl | string or null | URL reached after redirects, when known. |
contentType | string or null | HTTP content type for a loaded page, when known. |
sitemapUrl | string or null | Public sitemap that listed this URL, when known. |
lastmod | string or null | Last-modified value from the sitemap, when known. |
priority | number or null | Sitemap priority value, when known. |
changefreq | string or null | Sitemap change frequency, when known. |
Output example
This complete row is copied from a successful current-beta run:
{"startUrl": "https://crawlee.dev/","url": "https://crawlee.dev/js","discoverySources": ["sitemap","rendered"],"parentUrl": "https://crawlee.dev/","depth": 1,"anchorText": "Get started with JS","relationship": "internal","linkType": "page","crawlStatus": "crawled","httpStatusCode": 200,"finalUrl": "https://crawlee.dev/js","contentType": "text/html; charset=utf-8","sitemapUrl": "https://crawlee.dev/sitemap.xml","lastmod": null,"priority": 0.5,"changefreq": "weekly"}
Rows are available through the Apify dataset and its standard export formats.
💳 Pricing
This Actor uses pay-per-event pricing. The Website URL event is charged once for each accepted URL saved to the dataset. Discovered URLs that are not accepted or saved do not create this event. Start with a small Max URL rows value, inspect the output, and then broaden the limits if needed.
🔌 Integrations
Use Apify datasets, the API, schedules, webhooks, and platform integrations to process or deliver the URL inventory. This walkthrough shows how to connect an Actor to other services:
❓ FAQ
Why can a URL be marked discovered without HTTP fields?
A URL can come from a sitemap or an unvisited link. Only pages that the crawl loads receive crawled state and available HTTP response facts.
When are external URLs saved?
Same host and Same domain save only internal URLs. Include external URLs as discovered also saves external URLs, but the Actor does not crawl them.
Can I crawl only one submitted page?
Set Max crawl depth to 0 and use a low Max pages per website. Public sitemap discovery can still add sitemap-backed URLs.
Does it parse sitemap indexes?
Yes. It reads public XML URL sitemaps and nested sitemap indexes. Sitemap metadata is included only when the source supplies it.
Is this a full broken-link checker?
No. It provides HTTP facts for pages actually loaded during the crawl, which can support a downstream broken-link workflow. It does not promise an HTTP check for every discovered URL.
Does it scrape page text or private pages?
No. It extracts URL and link evidence from public pages and sitemaps. It does not return full page content, log in, or access private pages.
📝 Changelog
v1.0
- Added sitemap discovery, sitemap metadata, URL keyword filtering, total row limits, and the merged URL inventory output.
v0.0
- Initial release.
🆘 Support
For issues, questions, or feature requests, file a ticket and I'll fix or implement it in less than 24h 🫡
🔗 Related Actors
- Sitemap Sniffer: find public sitemap files and export their URL inventory before crawling page links.
- XML Sitemap Health Validator: check sitemap-listed URLs for HTTP status, redirects, response time, and sitemap metadata.
- Redirect Chain Checker: review redirect paths and terminal URLs for site migrations and link QA.
- Seobility SEO Checker: run public page SEO checks on selected URLs from the crawl.
- Webpage Diff Checker: monitor selected public pages and link changes after collecting their URLs.
Made with ❤️ by Maxime Dupré