# 网站全站链接提取 · Sitemap URL Extractor (`ordeal/sitemap-url-extractor`) Actor

输入域名，自动读取 sitemap.xml 并递归展开 sitemap 索引，支持 .gz 压缩 sitemap，输出全站 URL、lastmod、changefreq、priority。适合 SEO 排查、站点迁移盘点、爬虫种子准备。Sitemap parser, website URL list, SEO audit tool.

- **URL**: https://apify.com/ordeal/sitemap-url-extractor.md
- **Developed by:** [Li Sun](https://apify.com/ordeal) (community)
- **Categories:** SEO tools, Automation, Developer tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$5.00 / 1,000 解析一个 sitemaps

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## 网站全站链接提取（Sitemap & URL Extractor）

给一个网站地址，自动找到它的 sitemap，输出**全站 URL 清单**（含最后修改时间、更新频率、优先级）。

### 它会做什么

1. 先读 `robots.txt`，取里面声明的 sitemap（这是最可靠的入口）
2. 再试常见位置：`/sitemap.xml`、`/sitemap_index.xml`、`/sitemap-index.xml`、`/sitemap/sitemap.xml`
3. 遇到 **sitemap 索引**（一个 sitemap 里套多个 sitemap）会**递归展开**
4. 支持 **.gz 压缩的 sitemap** 与 gzip 传输（很多大站都是这样，未处理会拿到乱码）

### 输出字段

| 字段 | 说明 |
|---|---|
| `url` | 页面地址 |
| `lastmod` | 最后修改时间（若站点提供） |
| `changefreq` | 更新频率（若提供） |
| `priority` | 优先级（若提供） |
| `sourceSitemap` | 这条 URL 是从哪个 sitemap 文件里读到的 |

在 Apify 控制台 **Export** 可一键导出 CSV / Excel / JSON。

### 典型用途

- **SEO 排查**：拿到全站 URL 清单，检查收录、找失效页、比对新旧版本
- **迁移前盘点**：换域名/改版前，先把现有 URL 全量导出
- **内容审计**：统计站点规模、按 lastmod 找长期未更新的页面
- **给爬虫准备种子**：把 sitemap 变成抓取任务清单

### 输入参数

| 参数 | 默认 | 说明 |
|---|---|---|
| `startUrls` | — | 网站地址（首页即可），支持多个 |
| `maxSitemaps` | 20 | 最多解析几个 sitemap（超大站点会分成上千个，用它控额度） |
| `maxUrls` | 5000 | 最多输出多少条 URL |
| `respectRobots` | 开 | 读 robots.txt（同时也从里面取 Sitemap 声明） |

### 计费

按**成功解析的 sitemap 文件数**计费（`sitemap-parsed` 事件）。取不到或解析失败的 sitemap 不计费。

### 使用边界

- 只读取**公开的** sitemap 与 robots.txt，不做任何登录或绕过。
- 内置轻微限速（每个 sitemap 间隔 0.2 秒），不做并发轰炸。
- 请遵守目标网站的服务条款。

### 已知限制

- 站点**没有 sitemap** 时无法工作（这不是抓取器，不会去爬遍全站找链接）。若对方 robots.txt 与常见路径都没有声明，结果就是空——这是正常的，不是故障。
- 部分站点会**限制 sitemap 抓取**（例如维基百科对非浏览器 UA 返回 `403 Sitemap Restricted`），这类站点拿不到数据。
- `lastmod` / `priority` 只有站点自己写了才有；很多站点不写。

# Actor input Schema

## `startUrls` (type: `array`):

一个或多个网站地址（首页即可）。工具会自动找 /sitemap.xml，也会读取 robots.txt 里的 Sitemap 声明。

## `maxSitemaps` (type: `integer`):

防止超大站点（几十万 URL、上千个 sitemap）跑爆额度。

## `maxUrls` (type: `integer`):

输出上限，避免超大结果集。

## `respectRobots` (type: `boolean`):

开启后先读 robots.txt（同时也从里面取 Sitemap 声明）。

## Actor input object example

```json
{
  "startUrls": [
    "https://www.wikipedia.org"
  ],
  "maxSitemaps": 20,
  "maxUrls": 5000,
  "respectRobots": true
}
```

# Actor output Schema

## `overview` (type: `string`):

本次运行的全部结果，可在 Apify 控制台的 Export 里导出为 CSV / Excel / JSON。

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "startUrls": [
        "https://www.wikipedia.org"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("ordeal/sitemap-url-extractor").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "startUrls": ["https://www.wikipedia.org"] }

# Run the Actor and wait for it to finish
run = client.actor("ordeal/sitemap-url-extractor").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "startUrls": [
    "https://www.wikipedia.org"
  ]
}' |
apify call ordeal/sitemap-url-extractor --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,ordeal/sitemap-url-extractor"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/vRgmdrT3W0hfEqUzl/builds/RyrsIalK3GnTpR6Dr/openapi.json
