# HTML Table Extractor (`bindler/table-extractor`) Actor

Extract every table from any web page into clean rows, JSON and markdown. Correctly handles colspan, rowspan and stacked headers that break other extractors.

- **URL**: https://apify.com/bindler/table-extractor.md
- **Developed by:** [Neil Sangwaiya](https://apify.com/bindler) (community)
- **Categories:** AI, Developer tools, Agents
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $3.00 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## HTML Table Extractor

Pull every data table from any web page into clean rows, as JSON, CSV or markdown.

### Why tables break other extractors

Tables look simple and are not. Four things go wrong, and all four are handled here.

**`colspan` and `rowspan` silently misalign everything.** A cell spanning two rows shifts every later cell one column to the left, so row three onward quietly contains the wrong values. Nothing errors, the data is just wrong. This Actor expands spans into a proper rectangular grid, repeating spanned values so every row stands alone.

**Stacked headers lose their meaning.** Tables with grouped columns have two or three header rows, and taking only the first gives you a bare `Revenue` with no idea which year. This Actor joins them: `2025 / Revenue`.

**Navigation and layout tables come through as data.** Wikipedia navboxes, infoboxes and sidebars are structurally identical to data tables. This Actor filters them by class, by `role="presentation"`, and by link density, since a table that is more than 80% link text is a menu, not data. On one Wikipedia page that reduced six "tables" to the two that were real.

**CSS ends up inside your cells.** A `<style>` tag inside a table is picked up by naive text extraction, so cells fill with `.mw-parser-output .navbar{display:inline...}`. Removed here, along with footnote markers like `[1]`.

### What you get

| Field | Description |
|---|---|
| `url` | Source page |
| `tableIndex` | Position of the table on the page |
| `caption` | Table caption, or the nearest heading above it |
| `headers` | Column names, with stacked headers joined |
| `rows` | Array of objects keyed by column name |
| `rowCount` / `columnCount` | Table dimensions |
| `markdown` | The table rendered as markdown, ready to feed an LLM |
| `scrapedAt` | ISO timestamp |

Set **One record per row** to flatten the output so it exports straight to CSV or a spreadsheet.

### Example input

```json
{
  "urls": ["https://en.wikipedia.org/wiki/List_of_largest_companies_by_revenue"],
  "minRows": 3,
  "rowRecords": false
}
```

### Options

- **Minimum rows** and **Minimum columns** — filter out small layout tables
- **One record per row** — flatten for CSV export
- **Include markdown** — a markdown rendering per table, useful for RAG, since plain-text extraction turns tables into unusable word salad
- **Max tables per page** — cap for pages with very many tables

### Notes

- Reads the HTML the server returns. Tables rendered entirely by JavaScript after load will not appear.
- Nested tables are skipped rather than flattened into their parent cell.

# Actor input Schema

## `urls` (type: `array`):

Pages containing tables.

## `minRows` (type: `integer`):

Ignore tables with fewer rows. Filters out layout tables.

## `minColumns` (type: `integer`):

Ignore tables with fewer columns.

## `rowRecords` (type: `boolean`):

Flatten to one dataset record per table row, ready for CSV export. Off returns one record per table.

## `includeMarkdown` (type: `boolean`):

Add a markdown rendering of each table, useful for feeding to an LLM.

## `maxTablesPerPage` (type: `integer`):

Cap for pages with very many tables.

## Actor input object example

```json
{
  "urls": [
    "https://en.wikipedia.org/wiki/List_of_largest_companies_by_revenue"
  ],
  "minRows": 2,
  "minColumns": 2,
  "rowRecords": false,
  "includeMarkdown": true,
  "maxTablesPerPage": 50
}
```

# Actor output Schema

## `tables` (type: `string`):

Headers, rows and a markdown rendering per table.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        "https://en.wikipedia.org/wiki/List_of_largest_companies_by_revenue"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("bindler/table-extractor").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "urls": ["https://en.wikipedia.org/wiki/List_of_largest_companies_by_revenue"] }

# Run the Actor and wait for it to finish
run = client.actor("bindler/table-extractor").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    "https://en.wikipedia.org/wiki/List_of_largest_companies_by_revenue"
  ]
}' |
apify call bindler/table-extractor --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,bindler/table-extractor"
        }
    }
}

```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/KmhgvAdQXmGsZqf3G/builds/rZa4knxbwFFKCCTPQ/openapi.json
