# PDF to Text & Markdown: Extract PDF Content by URL (`accountable_eel/pdf-text-extractor`) Actor

Extracts text or Markdown from PDFs by URL, for a list of PDF URLs from a crawl, a Sheet, or a CRM export. Rebuilds paragraphs and headings from PDF fonts, with best-effort tables, then feeds a RAG pipeline, an LLM prompt, or a row to Google Sheets. No OCR: scanned PDFs are flagged, not charged.

- **URL**: https://apify.com/accountable\_eel/pdf-text-extractor.md
- **Developed by:** [Adrian Voss](https://apify.com/accountable_eel) (community)
- **Categories:** Developer tools, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.52 / 1,000 document processed (up to 20 pages)s

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## PDF to Text & Markdown: Extract PDF Content by URL

Give it a list of PDF URLs and get back clean text or Markdown for each one —
headings and paragraphs rebuilt from the PDF's own font sizes, a best-effort
Markdown table for simple aligned tables, and metadata (title, author, page
count). Built for feeding a RAG pipeline, an LLM prompt, or a row to Google
Sheets — anywhere you'd otherwise hand-copy text out of a PDF. No OCR yet: a
scanned, image-only PDF is flagged (`isScanned`) and never charged.

### Who it's for

Anyone who has a list of PDF URLs — from a web crawl, a Sheet, a CRM export,
or a document management system — and needs a plain PDF to Markdown
conversion, not a raw byte stream. Common uses: preparing PDFs to embed for
RAG or as LLM context, converting a folder of reports from PDF to Markdown
for a wiki, or piping extracted text to Google Sheets or Airtable via an
integration.

### How to use

Run `node scripts/build-schemas.js` once this actor is registered in
`data/portfolio.json` to generate three numbered steps (Console, API curl,
schedule) built from this actor's real input field name and seed input.

### Input

```json
{
  "pdfUrls": [
    "https://www.irs.gov/pub/irs-pdf/fw4.pdf",
    "https://arxiv.org/pdf/1706.03762"
  ],
  "pages": "",
  "outputFormat": "markdown",
  "includePageBreaks": false
}
```

`pdfUrls` is one direct PDF URL per line — for a list of PDF URLs from a
crawl, a Sheet, or a CRM export. A URL pasted without its `https://` (common
when copied from a spreadsheet cell) is filled in automatically.

- **`pages`** (optional) — a page range like `"1-5"`, or a single page like
  `"3"`. Leave empty to extract every page.
- **`outputFormat`** — `"markdown"` (default) rebuilds `#`/`##`/`###`
  headings from font sizes and renders any detected table as a best-effort
  Markdown table; `"text"` keeps the same paragraphs with no Markdown syntax.
- **`includePageBreaks`** — inserts a marker between each page's content (an
  invisible `<!-- page N -->` comment in Markdown, `"--- Page N ---"` in
  plain text) so a downstream chunker can split by page.

### Output

One row per PDF, for example:

Run `node scripts/build-schemas.js` once this actor has a canary task run to
generate a real sample row here. A real row from this build's one live run
(see `docs/eval/builds/pdf-text-extractor.md`), `content` trimmed:

```json
{
  "query": "https://www.irs.gov/pub/irs-pdf/fw4.pdf",
  "found": true,
  "status": "OK",
  "url": "https://www.irs.gov/pub/irs-pdf/fw4.pdf",
  "title": "2026 Form W-4",
  "author": "C:DC:TS:CAR:MP",
  "pageCount": 5,
  "pagesExtracted": 5,
  "outputFormat": "markdown",
  "content": "### Employee's Withholding Certificate OMB No. 1545-0074\n\n## Form W-4\n\nComplete Form W-4 so that your employer can withhold the correct federal income tax from your pay...",
  "wordCount": 5001,
  "isScanned": false,
  "scrapedAt": "2026-09-25T20:29:42.368Z"
}
```

A URL that isn't a PDF, doesn't exist, or is larger than the 50 MB limit
comes back as a row with `"found": false` and is never charged. A scanned,
image-only PDF (`isScanned: true`) is also never charged — see "No OCR in
this version" below.

### How this works

1. **Download** — up to 50 MB per PDF, streamed with a hard cap so one huge
   file can't run away with your run's memory or your bill.
2. **Extract** — [pdfjs-dist](https://www.npmjs.com/package/pdfjs-dist) reads
   the text layer page by page, with each character's font size and
   position.
3. **Reconstruct** — lines are grouped by vertical position, a two-column
   academic layout is detected and read left-column-then-right-column
   instead of line-by-line across the gutter, consecutive lines of similar
   size are merged into paragraphs, and a line noticeably bigger than body
   text becomes a heading (`#`/`##`/`###`, biggest first). A run of 3+
   consecutive lines that each split into well-separated clusters is
   rendered as a best-effort Markdown table.
4. **No OCR in this version** — if a PDF's text layer is almost empty (a
   scanned page with no embedded text), `isScanned` is `true`, `content` is
   `null`, and the row is never charged.

This is a **best-effort** reconstruction, not a layout-perfect one:

- **Two-column detection** works well for the common single-gutter academic
  layout. An exotic multi-region magazine layout, or a page where both
  columns happen to align to the exact same text baseline throughout, can
  read out of order.
- **Table detection** requires at least 3 consecutive lines with clearly
  separated columns. A 1-2 row table, or a table with merged or empty cells
  in irregular places, may render as plain paragraph text instead.
- **Line-wrap dehyphenation** (joining a word split across two lines by a
  hyphen) can't always tell a soft line-break hyphen from a real one in a
  compound word (e.g. "English-to-German" broken exactly at a hyphen can
  come out as "Englishto-German"). This is a known limitation of automatic
  PDF text extraction generally, not specific to this actor.

### Pricing

Pay-per-event. A flat per-run fee covers session/proxy warmup; you're billed
per item only when data is actually found and returned — see
`.actor/pay_per_event.json` for exact prices. A miss is never charged.

One `document` event covers up to 20 extracted pages; a longer document also
bills one `page` event for each page beyond 20, so a 200-page report doesn't
cost the same as a 2-page memo. A scanned PDF with no text layer bills
nothing at all, even though it still returns a row telling you so.

### Use it from Clay, n8n, Make, or an AI agent

Run `node scripts/build-schemas.js` once this actor is priced and registered in
`data/portfolio.json` to generate this section — a run-sync curl example, an n8n
HTTP Request node recipe, a Clay "HTTP API" column recipe, and an MCP line, all
built from this actor's real input field name and seed input.

A common pattern: run this actor on a list of PDF URLs, then send each row's
`content` field to Google Sheets (one row per PDF) as a lightweight PDF-to-
Google-Sheets pipeline, or straight into a RAG pipeline's document loader —
the Markdown output chunks cleanly on its own `#`/`##` headings.

### Data & privacy

This actor only reads the PDF at the URL you give it. Nothing is stored
beyond the run's own dataset; no PDF content is retained after the run ends.

### FAQ

**Does it do OCR on scanned PDFs?** Not in this version. A scanned or
image-only PDF is flagged (`isScanned: true`) and never charged — no
extracted text is fabricated.

**What's the file size limit?** 50 MB per PDF. A larger file comes back as a
free `BAD_FORMAT` miss.

**Can I extract just a few pages from a huge PDF?** Yes — set `pages` to a
range like `"1-10"`. Only the requested pages count toward parsing time and
billing (`pagesExtracted`, not `pageCount`); the whole file still has to be
downloaded first, since PDF pages aren't independently fetchable over HTTP.

# Actor input Schema

## `pdfUrls` (type: `array`):

One direct PDF URL per line, for a list of PDF URLs from a crawl, a Sheet, or a CRM export. Accepted formats: https://www.irs.gov/pub/irs-pdf/fw4.pdf, https://arxiv.org/pdf/1706.03762. You're only charged for the ones we actually find — a miss costs nothing.

## `testRun` (type: `boolean`):

Turn this on to test your input on a small sample before running the full list. Turn it off to process everything.

## `onlyFound` (type: `boolean`):

Only keep rows where something was actually found. Misses are always free, whether or not you show them here.

## `includeKeywords` (type: `array`):

Optional. Only keep results that mention at least one of these words (e.g. a job title, a city, a product name). Leave empty to keep everything.

## `excludeKeywords` (type: `array`):

Optional. Drop any result that mentions one of these words. Leave empty to skip nothing.

## `maxResults` (type: `integer`):

Optional. Stop the run once this many results have been found — useful for a quick, cheap sample. Leave blank for no limit.

## `pages` (type: `string`):

A page range, e.g. "1-5", or a single page, e.g. "3". Leave empty to extract every page. 1-indexed, inclusive.

## `outputFormat` (type: `string`):

"Markdown" (default) rebuilds headings (#/##/###) from font sizes and renders detected tables as a best-effort Markdown table. "Plain text" keeps paragraphs but no Markdown syntax.

## `includePageBreaks` (type: `boolean`):

Insert a marker between each page's content (an HTML comment in Markdown, "--- Page N ---" in plain text) so a downstream chunker can split by page.

## `columns` (type: `array`):

Choose which pieces of information to include in each result row. All are included by default.

## `maxConcurrency` (type: `integer`):

Parallel requests. Keep conservative — this target has no browser fallback, so getting blocked costs more than slow-and-steady.

## `proxyConfiguration` (type: `object`):

Apify Proxy config. Used only for the cheap per-item HEAD check — the actual PDF download deliberately bypasses the proxy (see the actor's README).

## Actor input object example

```json
{
  "pdfUrls": [
    "https://www.irs.gov/pub/irs-pdf/fw4.pdf"
  ],
  "testRun": false,
  "onlyFound": false,
  "includeKeywords": [],
  "excludeKeywords": [],
  "pages": "",
  "outputFormat": "markdown",
  "includePageBreaks": false,
  "columns": [
    "url",
    "title",
    "author",
    "pageCount",
    "pagesExtracted",
    "outputFormat",
    "content",
    "wordCount",
    "isScanned"
  ],
  "maxConcurrency": 5,
  "proxyConfiguration": {
    "useApifyProxy": true
  }
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "pdfUrls": [
        "https://www.irs.gov/pub/irs-pdf/fw4.pdf"
    ],
    "includeKeywords": [],
    "excludeKeywords": []
};

// Run the Actor and wait for it to finish
const run = await client.actor("accountable_eel/pdf-text-extractor").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "pdfUrls": ["https://www.irs.gov/pub/irs-pdf/fw4.pdf"],
    "includeKeywords": [],
    "excludeKeywords": [],
}

# Run the Actor and wait for it to finish
run = client.actor("accountable_eel/pdf-text-extractor").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "pdfUrls": [
    "https://www.irs.gov/pub/irs-pdf/fw4.pdf"
  ],
  "includeKeywords": [],
  "excludeKeywords": []
}' |
apify call accountable_eel/pdf-text-extractor --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,accountable_eel/pdf-text-extractor"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/iGRwSIickdT3uKrUL/builds/2BqKT11mTd5rhY4mp/openapi.json
