Website URL Crawler & Link Extractor avatar

Website URL Crawler & Link Extractor

Pricing

from $9.00 / 1,000 results

Go to Apify Store
Website URL Crawler & Link Extractor

Website URL Crawler & Link Extractor

Crawl any website and extract every URL with its depth, parent page and anchor text. Build a sitemap or audit internal links. Works on static and JavaScript sites.

Pricing

from $9.00 / 1,000 results

Rating

0.0

(0)

Developer

Maged

Maged

Maintained by Community

Actor stats

0

Bookmarked

206

Total users

29

Monthly active users

17 hours ago

Last modified

Share

Crawl any website and extract all URLs in a hierarchical structure. Visualize your site architecture, audit internal links, discover broken paths, and build complete sitemaps — with configurable depth, domain filtering, and proxy support.

What does Website URL Crawler do?

Start from any URL and this Actor recursively follows links to map the full structure of a website. Each result includes the URL, its depth level, parent URL, and link text — giving you a complete picture of how pages connect.

It handles both static sites (fast mode) and JavaScript-heavy single-page apps (JavaScript rendering mode).

Why use this Actor?

  • Full site mapping — crawl unlimited depth to discover every URL
  • Hierarchy preserved — each URL includes its parent URL, depth, and anchor text
  • Two crawl modes — fast HTTP for static sites, JS rendering for React/Vue/Angular apps
  • Domain filtering — optionally restrict crawling to the same domain
  • Extension filtering — skip PDFs, images, ZIPs, and other non-page assets
  • Duplicate prevention — configurable deduplication to keep results clean

How to use Website URL Crawler

  1. Open the Actor and click Try for free
  2. Enter a startUrl
  3. Set maxDepth and maxChildrenPerLink
  4. Run — the full URL tree appears in the Output tab
  5. Download as JSON or CSV, or connect via the Apify API

Input

{
"startUrl": "https://example.com",
"maxDepth": 3,
"maxChildrenPerLink": 20,
"sameDomainOnly": true,
"renderJavaScript": false,
"allowDuplicates": false,
"ignoredExtensions": ["pdf", "jpg", "png", "zip"]
}
FieldTypeDescriptionDefault
startUrlstringThe URL to start crawling fromrequired
maxDepthintegerMaximum link recursion depth (1–30)3
maxChildrenPerLinkintegerMax child links per page (1–100)20
sameDomainOnlybooleanOnly crawl URLs on the same domaintrue
renderJavaScriptbooleanUse JS rendering for dynamic pagesfalse
allowDuplicatesbooleanAllow duplicate URLs in outputfalse
ignoredExtensionsarrayFile extensions to skip[]

Output

[
{
"url": "https://example.com",
"name": null,
"depth": 0,
"parentUrl": null
},
{
"url": "https://example.com/about",
"name": "About Us",
"depth": 1,
"parentUrl": "https://example.com"
},
{
"url": "https://example.com/about/team",
"name": "Our Team",
"depth": 2,
"parentUrl": "https://example.com/about"
}
]

Output data fields

FieldTypeDescription
urlstringThe full URL
namestringAnchor text of the link (if available)
depthnumberDepth level from the start URL
parentUrlstringThe URL this link was found on

Use cases

  • Site audits — find orphaned pages, broken internal link paths, or redirect chains
  • SEO analysis — map your site architecture to identify crawl depth issues
  • Sitemap generation — build sitemaps for sites that don't have one
  • Content migration — extract all URLs before moving to a new CMS
  • Competitive research — map a competitor's full site structure
  • QA testing — verify all pages are reachable from the homepage

How many results will I get?

You're charged per result, so the cost follows the number of rows below. Your plan's rates are on the Pricing tab.

InputResults
Small siteUnder 100 rows
Medium siteAbout 1,000 rows
Large site10,000+ rows

Each URL found is one row. Use Max depth and Max children per link to cap large sites, and keep JavaScript rendering off unless the site needs it — it is much slower.

FAQ

What is the difference between fast mode and JavaScript rendering? Fast mode (default) is about 10x faster and works for most static HTML sites. JavaScript rendering loads each page like a real visitor — use it for React, Vue, and Angular apps.

Can I crawl multiple sites in one run? This Actor starts from a single URL. Trigger multiple runs via the Apify API to crawl several sites in parallel.

Is this Actor maintained? Yes. For bugs or feature requests, open an issue in the Issues tab.

Found this Actor useful?

If this Actor saved you time, please leave a review on the Actor page. Reviews help other users discover it and take 30 seconds — every one genuinely matters.

For bugs, feature requests, or questions, open an issue in the Issues tab above.


⭐ Found this Actor useful? A quick review on the Store helps other users find it and keeps it maintained.