# 网页正文转 Markdown · Web Page to Clean Markdown (`ordeal/webpage-to-markdown`) Actor

输入网址，自动剥离导航、广告、脚本、侧栏，输出干净 Markdown 正文，保留标题层级、列表、代码块、引用、链接。喂给 AI 做摘要、RAG 知识库、语料清洗之前先洗一遍。HTML to Markdown, article extractor, LLM ready text.

- **URL**: https://apify.com/ordeal/webpage-to-markdown.md
- **Developed by:** [Li Sun](https://apify.com/ordeal) (community)
- **Categories:** AI, SEO tools, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$8.00 / 1,000 转换一个页面

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## 网页正文转 Markdown

给一个网址，输出干净的 **Markdown 正文**：自动剥离导航、广告、侧边栏、页脚、脚本与样式。

### 解决什么问题

把网页喂给 AI 之前，你拿到的是**一整页 HTML**：菜单、广告、cookie 提示、相关推荐全都混在正文里。
直接丢给模型，既浪费 token，又污染结果。这个工具只留正文。

### 输出字段

| 字段 | 说明 |
|---|---|
| `url` | 页面地址 |
| `title` | 页面标题 |
| `markdown` | **干净的正文 Markdown** |
| `chars` | 正文字符数 |
| `truncated` | 是否因超长被截断 |
| `stats` | 质量指标：块数、字符数、段落数、标点数的原始统计 |
| `fetchedAt` | 抓取时间 |

在 Apify 控制台 **Export** 可导出 JSON / CSV。

### 还原能力

- 标题层级 → `#` `##` `###`
- 列表 → `-`（支持缩进层级）
- 代码块 → 三反引号围栏
- 引用 → `>`
- 链接 → `[文字](绝对地址)`（相对地址自动补全）
- 图片 → `![说明](地址)`（可选）

会主动丢弃：`<script>` `<style>` `<nav>` `<footer>` `<aside>` `<form>` `<iframe>` `<button>` 与 HTML 注释、`<head>` 内容。

### 典型用途

- **RAG / 知识库**：先把网页洗成 Markdown 再切片入库，检索质量明显更好
- **AI 摘要与翻译**：去掉噪声再喂模型，省 token 也更准
- **语料清洗**：批量把网页转成训练/微调语料
- **存档**：把网页正文存成可读、可版本管理的纯文本

### 输入参数

| 参数 | 默认 | 说明 |
|---|---|---|
| `urls` | — | 要转换的网页，支持多个 |
| `includeLinks` | 开 | 保留超链接 |
| `includeImages` | 关 | 保留图片 |
| `maxChars` | 100000 | 单页输出上限 |

### 计费

按**成功转换的页面数**计费（`page-converted` 事件）。抓取失败的页面不计费。

### 已知限制（请先读这段再评估是否适用）

**做的好的**：常规文章页、文档页、博客页。实测：

- `docs.python.org` 教程页 → 20415 字符，**开头就是正文**（"In the following examples..."）
- `gnu.org` 哲学文章 → 从 34K 降到 29K，页头/语言选择器/站点导航被切掉
- 合成测试页 → 标题层级、列表、代码块、引用、链接全部正确还原

**做不好的**（诚实说明，避免误用）：

1. **极短页面的目录切不干净。** 当页面正文与导航体量接近时（例如 PEP 页这种一屏短文），
   开头的 "Table of Contents" 可能残留。原因：正文判定靠"文本密度 vs 链接占比"，
   导航和正文一样大时无法可靠区分。
2. **完全依赖前端 JS 渲染的页面（单页应用）** 拿到的 HTML 里没有正文，结果会偏空——
   这类页面需要浏览器渲染型抓取器。
3. **正文内部嵌的复杂组件**（评分卡、相关推荐、评论区）可能残留一部分。
4. 不做 OCR；表格保留为文本行，不还原成 Markdown 表格。

如果你要处理的是大批量**同类**页面，建议先用少量样本跑一遍看效果，再决定是否批量。

# Actor input Schema

## `urls` (type: `array`):

一个或多个网页地址，每行一个。每个页面会输出一条记录。

## `includeLinks` (type: `boolean`):

把页面里的超链接保留为 Markdown 链接 [文字](地址)，相对链接会自动补成绝对地址。

## `includeImages` (type: `boolean`):

把图片保留为 Markdown 图片 ![说明](地址)。

## `maxChars` (type: `integer`):

超长页面截断上限，避免输出过大。

## Actor input object example

```json
{
  "urls": [
    "https://en.wikipedia.org/wiki/Markdown"
  ],
  "includeLinks": true,
  "includeImages": false,
  "maxChars": 100000
}
```

# Actor output Schema

## `overview` (type: `string`):

本次运行的全部结果，可在 Apify 控制台的 Export 里导出为 CSV / Excel / JSON。

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        "https://en.wikipedia.org/wiki/Markdown"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("ordeal/webpage-to-markdown").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "urls": ["https://en.wikipedia.org/wiki/Markdown"] }

# Run the Actor and wait for it to finish
run = client.actor("ordeal/webpage-to-markdown").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    "https://en.wikipedia.org/wiki/Markdown"
  ]
}' |
apify call ordeal/webpage-to-markdown --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,ordeal/webpage-to-markdown"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/4AncibgZ44sdgTfXw/builds/wiaeFo6qGohW4ds6z/openapi.json
