# 网页表格转 Excel / CSV · Table to CSV Extractor (`ordeal/table-to-excel`) Actor

输入网址，自动抽取网页里所有 HTML 表格，第一张表直接转成 CSV / Excel 下载。支持列表与链接抽取，遵守 robots.txt，只抓公开页面。适合查数据、做研究、整理公开统计表。Table to CSV, HTML table scraper, web table extractor.

- **URL**: https://apify.com/ordeal/table-to-excel.md
- **Developed by:** [Li Sun](https://apify.com/ordeal) (community)
- **Categories:** SEO tools, Automation, Developer tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$10.00 / 1,000 抓取一个页面

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## 网页表格转 Excel / CSV

给它一个网址，它把页面里**所有 HTML 表格**抽出来，第一张表直接变成可下载的 **CSV**（也能拿到完整 JSON）。
不写代码、不装插件、不需要登录目标网站。

### 你会拿到什么

每个页面一条记录，包含：

| 字段 | 说明 |
|---|---|
| `url` | 页面地址 |
| `title` / `meta_description` | 页面标题与描述 |
| `tables_count` | 这个页面里找到几个表格 |
| `table_header` | 第一张表的表头 |
| `table_as_records` | **第一张表按行转成的结构化记录**（可直接导出 CSV/Excel） |
| `tables` | 所有表格的原始二维数组 |
| `list_items` | 页面里的列表条目（可选） |
| `links` | 页面里的链接与文字（可选） |

在 Apify 控制台的 **Export** 里可一键导出为 CSV / Excel / JSON。

### 典型用途

- 把维基百科、统计局、政府公开页上的**统计表**转成 Excel 做分析
- 整理**榜单/目录页**（列表 + 链接）
- 把价格表、赛程表、排班表这类**固定格式表格**批量落地
- 做数据研究时，先把网页表格变干净的结构化数据

### 用法

1. 在 **Input** 里填入一个或多个网址（每行一个）
2. 需要列表或链接就勾上对应选项
3. 点 **Start**，运行结束后在 **Export** 里下载 CSV

输入参数：

| 参数 | 默认 | 说明 |
|---|---|---|
| `urls` | — | 要抓的网址，支持多个 |
| `want_tables` | 开 | 抽取表格（主功能） |
| `want_lists` | 关 | 抽取列表条目 |
| `want_links` | 关 | 抽取所有链接 |
| `want_meta` | 开 | 抽取标题与描述 |
| `respect_robots` | 开 | 遵守 robots.txt，禁止抓取的页面会跳过并说明 |
| `delay_seconds` | 1 | 每页间隔，避免给对方服务器压力 |
| `max_pages` | 10 | 单次最多抓多少页 |

### 使用边界（请遵守）

- **只抓公开页面。** 不绕过登录、不绕过验证码。
- **不用于收集个人信息。** 请自行确认你的抓取目标与用途合法合规。
- 默认遵守 `robots.txt`；请勿关闭后用于高频抓取。
- 若目标网站条款明确禁止自动抓取，请不要使用本工具抓取该站点。

### 计费

按**成功解析的页面数**计费（`page-extracted` 事件）。抓取失败的页面不计费。

### 已知限制

- 表格由**静态 HTML** 渲染时效果最好；纯前端 JS 动态生成的表格可能抓不到内容（那些页面返回的 HTML 里没有表格数据）。
- 合并单元格（`rowspan`）会把内容放在首行，`colspan` 会自动补齐空列。
- 不含 OCR：表格里的图片文字识别不了。

# Actor input Schema

## `urls` (type: `array`):

一个或多个网页地址。每行一个。工具会把每个页面里的表格都抽出来。

## `want_tables` (type: `boolean`):

把页面里的 <table> 全抓下来。这是主要功能，默认开。

## `want_lists` (type: `boolean`):

把 <ul>/<ol> 的条目抓下来（适合榜单、目录页）。

## `want_links` (type: `boolean`):

抓取页面里的链接与链接文字（适合做站点地图、找资料）。

## `want_meta` (type: `boolean`):

抓取标题与 meta 描述。

## `respect_robots` (type: `boolean`):

开启后，若目标站 robots.txt 禁止抓取则跳过该页（推荐开启）。

## `delay_seconds` (type: `integer`):

抓多页时的间隔，避免给对方服务器压力。

## `max_pages` (type: `integer`):

本次运行最多抓取的页面数（防止误填超长列表）。

## Actor input object example

```json
{
  "urls": [
    "https://en.wikipedia.org/wiki/List_of_countries_by_GDP_(nominal)"
  ],
  "want_tables": true,
  "want_lists": false,
  "want_links": false,
  "want_meta": true,
  "respect_robots": true,
  "delay_seconds": 1,
  "max_pages": 10
}
```

# Actor output Schema

## `overview` (type: `string`):

本次运行抽到的所有页面与表格，可在 Apify 控制台的 Export 里导出为 CSV / Excel / JSON。

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        "https://en.wikipedia.org/wiki/List_of_countries_by_GDP_(nominal)"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("ordeal/table-to-excel").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "urls": ["https://en.wikipedia.org/wiki/List_of_countries_by_GDP_(nominal)"] }

# Run the Actor and wait for it to finish
run = client.actor("ordeal/table-to-excel").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    "https://en.wikipedia.org/wiki/List_of_countries_by_GDP_(nominal)"
  ]
}' |
apify call ordeal/table-to-excel --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,ordeal/table-to-excel"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/HjMRXJyKmHogyijAO/builds/Vx325wG5viBxmehn7/openapi.json
