# Extract HTML Tables From Web Pages as JSON

**Use case:** 

Pulls every table on a page into rows of cells you can load straight into a dataframe. No copy-paste, no merged-header cleanup. $0.003 per page.

## Input

```json
{
  "startUrls": [
    {
      "url": "https://en.wikipedia.org/wiki/List_of_programming_languages"
    }
  ],
  "pageFunction": "async def page_function(page, context, request):\n    rows = await page.eval_on_selector_all(\n        'table tr', 'els => els.map(r => Array.from(\\n            r.querySelectorAll(\\'th,td\\')).map(c => c.innerText.trim()))')\n    return {'url': request.url, 'rowCount': len(rows),\n            'rows': [r for r in rows if r][:500]}",
  "maxRequestsPerCrawl": 1,
  "maxConcurrency": 5,
  "requestTimeoutSecs": 90,
  "waitUntil": "load",
  "respectRobotsTxt": true,
  "proxyConfiguration": {
    "useApifyProxy": true
  },
  "blockResources": true,
  "sessionMode": "auto",
  "maxItems": 1,
  "deliverEmptyRecords": false
}
```

## Output

```json
{
  "url": {
    "label": "URL",
    "format": "link"
  },
  "title": {
    "label": "Title",
    "format": "text"
  },
  "text": {
    "label": "Text",
    "format": "text"
  },
  "textTruncated": {
    "label": "Text cut at 5000",
    "format": "boolean"
  },
  "waitUntilReached": {
    "label": "Fully loaded",
    "format": "boolean"
  }
}
```

## About this Actor

This example demonstrates how to use [Python Web Scraper — Playwright, Any Website](https://apify.com/eszetael_lab/reliable-playwright-scraper) with a specific input configuration. Visit the [Actor detail page](https://apify.com/eszetael_lab/reliable-playwright-scraper) to learn more, explore other use cases, and run it yourself.


## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
This Task's input is already configured above — use it as-is rather than inventing a new one.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For full API examples (JavaScript, Python, CLI, MCP, OpenAPI), see this Task's Actor page: https://apify.com/eszetael_lab/reliable-playwright-scraper.md

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).
