Go to example tasks
Follow a sitemap index down to every child file
An index that points at other indexes is followed all the way down, gzip unpacked on the way, and every file that was read is reported together with the parent that led to it. On a large site that means a dozen or more files across two levels of nesting.
Sitemap URL Extractor — All Page URLs from a Websitepower_on/sitemap-url-list
Website
Record type
Sitemap
Found via
+4 fieldsTextNumberBooleanListObject
Input
Websites or sitemap URLs(required):https://wordpress.org
Max URLs per website:300
Max sitemap files per website:60
Include the sitemap report:true
Output fields
Website
Record type
Sitemap
Found via
Status
URLs
Child sitemaps
Note
Sign up on Apify01
Create your Apify account to access the Sitemap URL Extractor — All Page URLs from a Website.
Start the run02
The Actor will start running based on the input automatically.
Receive the output03
Monitor the progress in real-time. You will be notified as soon as your dataset is complete and ready for review.
Integrate into your workflow04
The final output is delivered in JSON, CSV, or Excel format, ready to be plugged into your workflow.
