# Sitemap Extractor & Bulk URL Status Checker (`forevertools/sitemap-url-status-checker`) Actor

Sitemap validation & URL status checker: auto-discover a site's sitemaps (robots.txt, index & .gz), extract every URL with lastmod, then check link status (HTTP status codes), redirect chains, noindex headers and response time. Or paste your own URL list.

- **URL**: https://apify.com/forevertools/sitemap-url-status-checker.md
- **Developed by:** [Forever Tools](https://apify.com/forevertools) (community)
- **Categories:** SEO tools, Developer tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$1.00 / 1,000 url processeds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Sitemap Extractor & Bulk URL Status Checker

Give it a domain and it **finds every sitemap** (robots.txt `Sitemap:` lines, common paths such as `/sitemap.xml`, sitemap indexes, gzipped sitemaps), **extracts every URL** with `lastmod`, `changefreq` and `priority`, and **checks each URL's health**: HTTP status, redirect chain, final URL, `X-Robots-Tag` header, content type and response time. You can also paste any list of URLs to check in bulk.

### Extract all URLs from a website's sitemap

Enter a domain or homepage in **Sites or sitemap URLs** and the actor discovers the sitemaps itself. You can also pass a direct `.xml` or `.xml.gz` sitemap URL. Sitemap indexes are followed, so you get the full URL inventory in one dataset, each row noting which sitemap file it came from. Turn off **Check HTTP status & redirects** to extract URLs only, which is faster and priced the same.

### Find broken links in a sitemap

With status checking on (the default), every URL is requested and rows are tagged with issue codes: `unreachable` (DNS, connection or timeout problems), `http-404` / `http-5xx` style codes for error responses, `redirects`, `redirect-chain` (more than one hop) and `noindex-header`. A `SUMMARY` record in the key-value store counts each issue, so you can see at a glance how many 404s your sitemap contains.

### Check redirect chains and noindex URLs in bulk

Sitemaps should list final, indexable URLs. Rows include `redirects` (hop count), `redirectChain` (each hop with its status code), `finalUrl` and `xRobotsTag`, so you can spot sitemap entries that redirect or carry a noindex header.

### Use cases

- **Weekly SEO health check:** schedule the actor and alert on new 404s or redirect chains.
- **Site-migration QA:** confirm every old URL redirects once to the right place.
- **Crawl-budget clean-up:** find sitemap URLs that are noindex or redirecting.
- **URL inventory:** get a site's complete URL list to feed a crawler or LLM pipeline.
- **Bulk URL check:** paste any list into **Extra URLs to check** without needing a sitemap.

### Input

```json
{
  "sites": ["example.com"],
  "urls": ["https://example.com/some-extra-page"],
  "checkStatus": true,
  "maxUrls": 1000,
  "maxConcurrency": 10
}
```

- **sites:** domains, homepages or direct sitemap URLs.
- **urls:** optional extra URLs to check directly.
- **checkStatus** (default on), **maxUrls** (default 10,000), **maxConcurrency** (default 10, max 50), and an optional **userAgent**.

### Output

One row per URL. Example:

```json
{
  "url": "http://example.com/old-page",
  "sitemap": "https://example.com/sitemap.xml",
  "lastmod": "2026-08-01",
  "statusCode": 200,
  "finalUrl": "https://example.com/new-page",
  "redirects": 2,
  "redirectChain": [
    { "url": "http://example.com/old-page", "statusCode": 301 },
    { "url": "https://example.com/old-page", "statusCode": 301 }
  ],
  "contentType": "text/html; charset=utf-8",
  "responseTimeMs": 231,
  "issues": ["redirects", "redirect-chain"]
}
```

Other fields: `changefreq`, `priority`, `xRobotsTag`, `error`. If a site is invalid or no sitemap is found, you get an error row with `site` and `error` (`invalid-site` or `no-sitemap-found`).

### Pricing

Pay per event: **$0.001 per URL** ($1 per 1,000 URLs), no subscription.

- 500-URL sitemap: about $0.50
- 10,000 URLs: about $10
- 100,000 URLs: about $100

**Max URLs** caps the run, and the actor stops when the charge limit is reached. Platform usage is billed by Apify per your plan.

### Limitations

- It reports status, redirects and headers only; it does not read page content, so it cannot detect soft 404s.
- Only sitemaps are used to discover URLs; it is not a link crawler.
- Response times are measured from the actor's location and vary.
- Use it on sites you own or may check.

### FAQ

#### How do I find all URLs in a sitemap?

Enter the domain in **Sites** and switch off status checking; you get one row per URL with `lastmod`, `changefreq` and `priority`.

#### How do I check a sitemap for 404 errors?

Leave status checking on and filter rows whose `issues` include `http-404`, or read the counts in `SUMMARY`.

#### Does it handle sitemap index files and .gz sitemaps?

Yes, nested indexes and gzipped sitemaps are followed.

#### What if no sitemap is found?

You get a row with `error: "no-sitemap-found"`; paste the URLs into **Extra URLs to check** instead.

#### Does it download the full pages?

No. Requests are plain GETs and only status and headers are read.

#### How much does it cost to check a 1,000-URL sitemap?

$1 at $0.001 per URL.

Built and maintained with AI assistance. Problems or requests: use the Issues tab.

### Related tools

Other actors by the same developer (same flat pay-per-result pricing, no subscription):

- [Apple App Store Reviews Scraper (Multi-Country)](https://apify.com/forevertools/apple-app-store-reviews)
- [Article Extractor – Clean Text & Markdown for LLM/RAG](https://apify.com/forevertools/article-extractor)
- [Company Jobs Scraper: Workday, Greenhouse, Lever, Ashby](https://apify.com/forevertools/ats-company-jobs)
- [Bulk Domain Checker — WHOIS/RDAP, DNS, SPF/DMARC, SSL Expiry](https://apify.com/forevertools/domain-whois-dns-ssl)
- [Bulk PageSpeed Insights & Core Web Vitals Checker](https://apify.com/forevertools/pagespeed-core-web-vitals)
- [PDF to Text Extractor (Bulk, with Metadata)](https://apify.com/forevertools/pdf-to-text-extractor)
- [Website SEO Audit Crawler](https://apify.com/forevertools/website-seo-audit)
- [Website Tech Stack Detector (CMS, Framework, Analytics)](https://apify.com/forevertools/website-tech-stack-detector)
- [Website Screenshot – Bulk Full Page PNG, JPEG & PDF](https://apify.com/forevertools/website-screenshot)

### Integrations

Run it from the Apify API, a schedule, or no-code tools: the Apify apps for **Zapier**, **Make** and **n8n** can start any public actor ("Run Actor") and read its dataset. AI agents can call it through the **Apify MCP server**.

# Actor input Schema

## `sites` (type: `array`):

Domains/homepages (sitemaps are found via robots.txt and common paths) or direct sitemap URLs (.xml or .xml.gz). Sitemap indexes are followed.

## `urls` (type: `array`):

Optional list of URLs to check directly (no sitemap needed).

## `checkStatus` (type: `boolean`):

Request every URL and record status code, redirect chain, final URL, content type, X-Robots-Tag and response time. Turn off to only extract sitemap URLs.

## `maxUrls` (type: `integer`):

Stop after this many URLs. You are charged per URL output.

## `maxConcurrency` (type: `integer`):

Parallel status checks. Lower for small servers.

## `userAgent` (type: `string`):

User-Agent header for requests.

## Actor input object example

```json
{
  "sites": [
    "crawlee.dev"
  ],
  "checkStatus": true,
  "maxUrls": 10000,
  "maxConcurrency": 10
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "sites": [
        "crawlee.dev"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("forevertools/sitemap-url-status-checker").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "sites": ["crawlee.dev"] }

# Run the Actor and wait for it to finish
run = client.actor("forevertools/sitemap-url-status-checker").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "sites": [
    "crawlee.dev"
  ]
}' |
apify call forevertools/sitemap-url-status-checker --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,forevertools/sitemap-url-status-checker"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/0bTZtBtbpThOQSkYd/builds/Ohqguw2cWbS81VRRQ/openapi.json
