Website URL Crawler & Link Extractor
Pricing
from $9.00 / 1,000 results
Website URL Crawler & Link Extractor
Crawl any website and extract every URL with its depth, parent page and anchor text. Build a sitemap or audit internal links. Works on static and JavaScript sites.
Pricing
from $9.00 / 1,000 results
Rating
0.0
(0)
Developer
Maged
Maintained by CommunityActor stats
0
Bookmarked
206
Total users
29
Monthly active users
17 hours ago
Last modified
Categories
Share
Crawl any website and extract all URLs in a hierarchical structure. Visualize your site architecture, audit internal links, discover broken paths, and build complete sitemaps — with configurable depth, domain filtering, and proxy support.
What does Website URL Crawler do?
Start from any URL and this Actor recursively follows links to map the full structure of a website. Each result includes the URL, its depth level, parent URL, and link text — giving you a complete picture of how pages connect.
It handles both static sites (fast mode) and JavaScript-heavy single-page apps (JavaScript rendering mode).
Why use this Actor?
- Full site mapping — crawl unlimited depth to discover every URL
- Hierarchy preserved — each URL includes its parent URL, depth, and anchor text
- Two crawl modes — fast HTTP for static sites, JS rendering for React/Vue/Angular apps
- Domain filtering — optionally restrict crawling to the same domain
- Extension filtering — skip PDFs, images, ZIPs, and other non-page assets
- Duplicate prevention — configurable deduplication to keep results clean
How to use Website URL Crawler
- Open the Actor and click Try for free
- Enter a
startUrl - Set
maxDepthandmaxChildrenPerLink - Run — the full URL tree appears in the Output tab
- Download as JSON or CSV, or connect via the Apify API
Input
{"startUrl": "https://example.com","maxDepth": 3,"maxChildrenPerLink": 20,"sameDomainOnly": true,"renderJavaScript": false,"allowDuplicates": false,"ignoredExtensions": ["pdf", "jpg", "png", "zip"]}
| Field | Type | Description | Default |
|---|---|---|---|
startUrl | string | The URL to start crawling from | required |
maxDepth | integer | Maximum link recursion depth (1–30) | 3 |
maxChildrenPerLink | integer | Max child links per page (1–100) | 20 |
sameDomainOnly | boolean | Only crawl URLs on the same domain | true |
renderJavaScript | boolean | Use JS rendering for dynamic pages | false |
allowDuplicates | boolean | Allow duplicate URLs in output | false |
ignoredExtensions | array | File extensions to skip | [] |
Output
[{"url": "https://example.com","name": null,"depth": 0,"parentUrl": null},{"url": "https://example.com/about","name": "About Us","depth": 1,"parentUrl": "https://example.com"},{"url": "https://example.com/about/team","name": "Our Team","depth": 2,"parentUrl": "https://example.com/about"}]
Output data fields
| Field | Type | Description |
|---|---|---|
url | string | The full URL |
name | string | Anchor text of the link (if available) |
depth | number | Depth level from the start URL |
parentUrl | string | The URL this link was found on |
Use cases
- Site audits — find orphaned pages, broken internal link paths, or redirect chains
- SEO analysis — map your site architecture to identify crawl depth issues
- Sitemap generation — build sitemaps for sites that don't have one
- Content migration — extract all URLs before moving to a new CMS
- Competitive research — map a competitor's full site structure
- QA testing — verify all pages are reachable from the homepage
How many results will I get?
You're charged per result, so the cost follows the number of rows below. Your plan's rates are on the Pricing tab.
| Input | Results |
|---|---|
| Small site | Under 100 rows |
| Medium site | About 1,000 rows |
| Large site | 10,000+ rows |
Each URL found is one row. Use Max depth and Max children per link to cap large sites, and keep JavaScript rendering off unless the site needs it — it is much slower.
FAQ
What is the difference between fast mode and JavaScript rendering? Fast mode (default) is about 10x faster and works for most static HTML sites. JavaScript rendering loads each page like a real visitor — use it for React, Vue, and Angular apps.
Can I crawl multiple sites in one run? This Actor starts from a single URL. Trigger multiple runs via the Apify API to crawl several sites in parallel.
Is this Actor maintained? Yes. For bugs or feature requests, open an issue in the Issues tab.
Found this Actor useful?
If this Actor saved you time, please leave a review on the Actor page. Reviews help other users discover it and take 30 seconds — every one genuinely matters.
For bugs, feature requests, or questions, open an issue in the Issues tab above.
⭐ Found this Actor useful? A quick review on the Store helps other users find it and keeps it maintained.