# PDF Text Extractor: PDF URL to Text and Markdown per Page (`pistachio_implementation/pdf-text-extractor`) Actor

Convert PDF links to clean text and Markdown in reading order, per document or per page, with title, author, dates and page count. Headings and lists kept for LLM and RAG pipelines. $2 per 1,000 PDFs, failed files free, no OCR needed for digital PDFs.

- **URL**: https://apify.com/pistachio\_implementation/pdf-text-extractor.md
- **Developed by:** [Hay Equipos](https://apify.com/pistachio_implementation) (community)
- **Categories:** AI, Developer tools, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

Pay per event

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## PDF Text Extractor: PDF URL to Text and Markdown per Page

Give this actor links to PDF files and get back clean text and Markdown in reading order, one row per document or one row per page, with the document's title, author, subject, creator, creation and modification dates, page count and file size. Two column layouts (papers, forms, reports) are read column by column, words split across lines are rejoined, and larger fonts become Markdown headings, so the output is ready for search, LLM prompts and RAG chunking.

It downloads each file over plain HTTP and extracts the text layer with pdf.js, the PDF engine used in Firefox. No browser, no OCR service, no API key. You pay per PDF, plus a tiny start fee per run; files that fail cost nothing.

### What you can use it for

- Turn reports, papers, filings, manuals and brochures into text for an LLM or a vector database.
- Chunk long documents by page, with page numbers kept for citations.
- Pull titles, authors and dates from a batch of PDFs into a spreadsheet.
- Monitor PDFs that change (price lists, policies, tenders) by extracting them on a schedule.
- Check which PDFs in a list are scans without a text layer (`likelyScanned`).

### Input

| Field | What it does | Default |
|---|---|---|
| PDF URLs | Direct links to PDF files | required |
| One row per | `document` (whole text in one row) or `pages` (one row per page) | document |
| Include plain text | Text in reading order with paragraph breaks | on |
| Include Markdown | Markdown with headings, bullet lists and a rule between pages | on |
| Maximum pages per PDF | Read at most this many pages from each file | 300 |
| Maximum file size (MB) | Larger files are skipped for free | 50 |
| Maximum text length | Cut text and Markdown per row (0 means no limit) | 0 |
| Respect robots.txt | Skip files a site closes to automated tools (skipped files are free) | on |
| Parallel downloads | PDFs worked on at once; one host always gets one download per second | 2 |

Example input:

```json
{
  "urls": [
    "https://arxiv.org/pdf/1706.03762",
    "https://www.irs.gov/pub/irs-pdf/fw9.pdf"
  ],
  "outputMode": "pages",
  "maxPagesPerPdf": 50
}
```

### Output

Document mode, one row per PDF (text shortened here):

```json
{
  "url": "https://www.irs.gov/pub/irs-pdf/fw9.pdf",
  "finalUrl": "https://www.irs.gov/pub/irs-pdf/fw9.pdf",
  "success": true,
  "fileName": "fw9.pdf",
  "fileSizeBytes": 140815,
  "pageCount": 6,
  "pagesExtracted": 6,
  "title": "Form W-9 (Rev. March 2024)",
  "author": "SE:W:CAR:MP",
  "subject": "Request for Taxpayer Identification Number and Certification",
  "keywords": "Fillable",
  "creator": "Designer 6.5",
  "producer": "Designer 6.5",
  "createdAt": "2024-03-06T13:18:13.000Z",
  "modifiedAt": "2024-03-06T13:18:13.000Z",
  "pdfVersion": "1.7",
  "likelyScanned": false,
  "wordCount": 6277,
  "charCount": 37851,
  "textTruncated": false,
  "text": "FormW-9 Request for Taxpayer Give form to the\n\n(Rev. March 2024) Identification Number and Certification requester. Do not\n\nDepartment of the Treasury send to the IRS...",
  "markdown": "## FormW-9 Request for Taxpayer Give form to the\n\n### (Rev. March 2024) Identification Number and Certification requester. Do not\n\nDepartment of the Treasury send to the IRS...",
  "extractedAt": "2026-09-27T06:23:12.400Z"
}
```

Page mode adds `pageNumber` and gives `wordCount`, `charCount`, `text` and `markdown` for that page only; the document fields repeat on every row so each page stands alone.

Failed files come back as `success: false` with an `error` such as "The server answered HTTP 404", "The URL returned a web page, not a PDF" or "The PDF is password protected".

### Pricing

Pay per event, no subscription, no charge for platform usage on top.

| Event | Price |
|---|---|
| PDF processed | $0.002 ($2 per 1,000 PDFs), the same in document and page mode |
| Actor start | $0.00005 per run (Apify's standard start event) |

Downloads that fail, files that are not PDFs, password protected files, robots.txt skips and files over your size limit are free. You can set a maximum charge per run in Apify and the actor stops cleanly when it is reached.

### Limits

- No OCR. Scanned PDFs without a text layer return little or no text and are flagged with `likelyScanned: true`.
- Tables come out as text lines in reading order, not as structured cells.
- Very complex layouts (three or more columns, text boxes scattered over the page) may read in an imperfect order.
- Links that need a login, a cookie banner click or a captcha before the file downloads return an error row.
- Default memory is 512 MB, enough for typical files up to about 50 MB; give the run more memory for very large files.
- Up to 2,000 PDFs per run.

### FAQ

**Do I need an API key?** No. Paste the links and run.

**Can I process files I have on my computer?** Upload them to an Apify key value store (or any storage with a public link) and pass those links.

**Why is the price per PDF and not per page?** It keeps costs predictable for you and for AI agents. Use "Maximum pages per PDF" to cap very long files.

**What does `likelyScanned` mean?** The file has under about 40 characters of text per page, which usually means scanned images. Those need an OCR tool.

**Is the Markdown exact?** Headings are detected from font size and bullets from bullet characters. It is built for reading and chunking, not for pixel perfect layout.

# Actor input Schema

## `urls` (type: `array`):

Direct links to PDF files. Any public link works, including files you uploaded to an Apify key value store.

## `outputMode` (type: `string`):

"document" returns one row per PDF with the whole text. "pages" returns one row per page (handy for citations and chunking). The price is per PDF either way.

## `includeText` (type: `boolean`):

Plain text in reading order with paragraph breaks.

## `includeMarkdown` (type: `boolean`):

Markdown with headings detected from font size, bullet lists, and pages separated by a rule.

## `maxPagesPerPdf` (type: `integer`):

Read at most this many pages from each file. Rows say when a file was cut.

## `maxFileSizeMb` (type: `integer`):

Files larger than this are skipped for free. For files over 50 MB give the run 1 GB of memory or more.

## `maxTextLength` (type: `integer`):

Cut text and Markdown in each row to this many characters. 0 means no limit.

## `respectRobotsTxt` (type: `boolean`):

Skip files that the site's robots.txt closes to automated tools. Skipped files are free.

## `maxConcurrency` (type: `integer`):

How many PDFs to work on at once. Downloads from one host are always spaced one second apart.

## Actor input object example

```json
{
  "urls": [
    "https://arxiv.org/pdf/1706.03762",
    "https://www.irs.gov/pub/irs-pdf/fw9.pdf"
  ],
  "outputMode": "document",
  "includeText": true,
  "includeMarkdown": true,
  "maxPagesPerPdf": 300,
  "maxFileSizeMb": 50,
  "maxTextLength": 0,
  "respectRobotsTxt": true,
  "maxConcurrency": 2
}
```

# Actor output Schema

## `results` (type: `string`):

All rows the run saved to the default dataset.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        "https://arxiv.org/pdf/1706.03762",
        "https://www.irs.gov/pub/irs-pdf/fw9.pdf"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("pistachio_implementation/pdf-text-extractor").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "urls": [
        "https://arxiv.org/pdf/1706.03762",
        "https://www.irs.gov/pub/irs-pdf/fw9.pdf",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("pistachio_implementation/pdf-text-extractor").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    "https://arxiv.org/pdf/1706.03762",
    "https://www.irs.gov/pub/irs-pdf/fw9.pdf"
  ]
}' |
apify call pistachio_implementation/pdf-text-extractor --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,pistachio_implementation/pdf-text-extractor"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/DMda3Bluc4hoZ3ZBm/builds/WkHW7gJGTXeQcvxDp/openapi.json
