# PDF Text Extractor - Text, Metadata & Page Count from PDF URL (`ninhothedev/pdf-text-extractor`) Actor

$0.5/1K 🔥 PDF text extractor API! Extract full text, metadata & page count from any PDF URL — ready for RAG, LLMs & AI pipelines. No API key. Export JSON, CSV, Excel or API in seconds ⚡

- **URL**: https://apify.com/ninhothedev/pdf-text-extractor.md
- **Developed by:** [ninhothedev](https://apify.com/ninhothedev) (community)
- **Categories:** AI, Developer tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.50 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/platform/actors/running/actors-in-store#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

In JavaScript/TypeScript projects, use official [JavaScript/TypeScript client](https://docs.apify.com/api/client/js/docs.md):

```bash
npm install apify-client
```

In Python projects, use official [Python client library](https://docs.apify.com/api/client/python/docs.md):

```bash
pip install apify-client
```

In shell scripts, use [Apify CLI](https://docs.apify.com/cli/docs.md):

````bash
# MacOS / Linux
curl -fsSL https://apify.com/install-cli.sh | bash
# Windows
irm https://apify.com/install-cli.ps1 | iex
```bash

In AI frameworks, you might use the [Apify MCP server](https://docs.apify.com/integrations/mcp.md).

If your project is in a different language, use the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).


# README

## PDF Text Extractor

Extract clean **text, metadata and page counts from any PDF URL** — built for
RAG pipelines, LLM ingestion and large document datasets. Give it a list of PDF
links and it downloads each file, verifies it is a real PDF, and returns the
full text, optional per-page text, page count, word count and document metadata
(title, author, subject, creator, producer, creation date).

No API key. No login. Datacenter-proxy friendly (it operates on the PDF URLs
**you** provide), so it runs cheaply on the smallest Apify plan.

### Why use this actor

- **RAG / LLM ingestion** — turn PDFs into clean text ready to chunk and embed.
- **Document data extraction** — pull text + metadata from invoices, papers,
  manuals, reports, contracts and public filings at scale.
- **Research** — batch-download and read arXiv, bioRxiv, SSRN or .gov PDFs.
- **Archiving & indexing** — build a searchable text index of a PDF collection.

### Features

- Accepts any list of direct PDF URLs (redirects followed automatically).
- Full concatenated text (capped at 200,000 characters) per document.
- Optional **per-page text** array (`includePages`) for page-accurate chunking.
- Document metadata: title, author, subject, creator, producer, creation date.
- Page count and word count for every PDF.
- Robust: non-PDF URLs are logged and skipped; malformed PDFs never crash a run.
- One dataset item per PDF, ready to export as JSON, CSV, Excel or feed an API.

### Input

| Field | Type | Description |
|-------|------|-------------|
| `mode` | select | Operation mode. Currently `extract`. |
| `urls` | array | Direct links to PDF files to extract. |
| `includePages` | boolean | Add a per-page text array (default `false`). |
| `maxItems` | integer | Max PDFs to process, 1–1000 (default `100`). |

#### Example input

```json
{
  "mode": "extract",
  "urls": [
    "https://arxiv.org/pdf/1706.03762",
    "https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf"
  ],
  "includePages": false,
  "maxItems": 100
}
````

### Output

Each PDF becomes one dataset record:

```json
{
  "url": "https://arxiv.org/pdf/1706.03762",
  "final_url": "https://arxiv.org/pdf/1706.03762",
  "title": null,
  "author": null,
  "subject": null,
  "creator": "LaTeX with hyperref",
  "producer": "pdfTeX-1.40.25",
  "creation_date": "2024-04-10T21:11:43+00:00",
  "page_count": 15,
  "text": "Attention Is All You Need ...",
  "pages": null,
  "word_count": 6123,
  "source": "pdf-extractor",
  "scraped_at": "2026-07-21T16:20:00Z"
}
```

### Pricing

Cheap by design: about **$1 per 1,000 PDFs** plus Apify platform usage. Text
extraction is pure-Python (no browser, no paid proxy), so runs are fast and the
compute cost stays low.

### Related actors

- [Smart Article Extractor](https://apify.com/ninhothedev/smart-article-extractor)
- [Website Content Crawler](https://apify.com/ninhothedev/website-content-crawler)
- [HTML to Markdown Converter](https://apify.com/ninhothedev/html-to-markdown-converter)
- [Wayback Machine Scraper](https://apify.com/ninhothedev/wayback-machine-scraper)

### Keywords

pdf text extractor, extract text from pdf, pdf to text, pdf scraper, pdf parser,
pdf metadata extractor, pdf page count, rag pdf ingestion, llm document
ingestion, document dataset, pdf data extraction, arxiv pdf, ocr alternative,
batch pdf extraction, pdf api.

# Actor input Schema

## `mode` (type: `string`):

Operation mode. 'extract' downloads each PDF URL and extracts its text, page count and document metadata.

## `urls` (type: `array`):

Direct links to PDF files to extract. Each URL should point to a downloadable PDF (redirects are followed). Non-PDF URLs are logged and skipped.

## `includePages` (type: `boolean`):

If enabled, output a 'pages' array with the extracted text of each page separately (each page capped at 20,000 characters). Off by default to keep items small.

## `maxItems` (type: `integer`):

Maximum number of PDFs to process and store in the dataset. Ranges from 1 to 1000.

## Actor input object example

```json
{
  "mode": "extract",
  "urls": [
    "https://arxiv.org/pdf/1706.03762",
    "https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf"
  ],
  "includePages": false,
  "maxItems": 100
}
```

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        "https://arxiv.org/pdf/1706.03762",
        "https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("ninhothedev/pdf-text-extractor").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "urls": [
        "https://arxiv.org/pdf/1706.03762",
        "https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("ninhothedev/pdf-text-extractor").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print("💾 Check your data here: https://console.apify.com/storage/datasets/" + run["defaultDatasetId"])
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    "https://arxiv.org/pdf/1706.03762",
    "https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf"
  ]
}' |
apify call ninhothedev/pdf-text-extractor --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "command": "npx",
            "args": [
                "mcp-remote",
                "https://mcp.apify.com/?tools=ninhothedev/pdf-text-extractor",
                "--header",
                "Authorization: Bearer <YOUR_API_TOKEN>"
            ]
        }
    }
}

```

## OpenAPI specification

```json
{
    "openapi": "3.0.1",
    "info": {
        "title": "PDF Text Extractor - Text, Metadata & Page Count from PDF URL",
        "description": "$0.5/1K 🔥 PDF text extractor API! Extract full text, metadata & page count from any PDF URL — ready for RAG, LLMs & AI pipelines. No API key. Export JSON, CSV, Excel or API in seconds ⚡",
        "version": "0.1",
        "x-build-id": "z4xjc5TQDNplWyVT9"
    },
    "servers": [
        {
            "url": "https://api.apify.com/v2"
        }
    ],
    "paths": {
        "/acts/ninhothedev~pdf-text-extractor/run-sync-get-dataset-items": {
            "post": {
                "operationId": "run-sync-get-dataset-items-ninhothedev-pdf-text-extractor",
                "x-openai-isConsequential": false,
                "summary": "Executes an Actor, waits for its completion, and returns Actor's dataset items in response.",
                "tags": [
                    "Run Actor"
                ],
                "requestBody": {
                    "required": true,
                    "content": {
                        "application/json": {
                            "schema": {
                                "$ref": "#/components/schemas/inputSchema"
                            }
                        }
                    }
                },
                "parameters": [
                    {
                        "name": "token",
                        "in": "query",
                        "required": true,
                        "schema": {
                            "type": "string"
                        },
                        "description": "Enter your Apify token here"
                    }
                ],
                "responses": {
                    "200": {
                        "description": "OK"
                    }
                }
            }
        },
        "/acts/ninhothedev~pdf-text-extractor/runs": {
            "post": {
                "operationId": "runs-sync-ninhothedev-pdf-text-extractor",
                "x-openai-isConsequential": false,
                "summary": "Executes an Actor and returns information about the initiated run in response.",
                "tags": [
                    "Run Actor"
                ],
                "requestBody": {
                    "required": true,
                    "content": {
                        "application/json": {
                            "schema": {
                                "$ref": "#/components/schemas/inputSchema"
                            }
                        }
                    }
                },
                "parameters": [
                    {
                        "name": "token",
                        "in": "query",
                        "required": true,
                        "schema": {
                            "type": "string"
                        },
                        "description": "Enter your Apify token here"
                    }
                ],
                "responses": {
                    "200": {
                        "description": "OK",
                        "content": {
                            "application/json": {
                                "schema": {
                                    "$ref": "#/components/schemas/runsResponseSchema"
                                }
                            }
                        }
                    }
                }
            }
        },
        "/acts/ninhothedev~pdf-text-extractor/run-sync": {
            "post": {
                "operationId": "run-sync-ninhothedev-pdf-text-extractor",
                "x-openai-isConsequential": false,
                "summary": "Executes an Actor, waits for completion, and returns the OUTPUT from Key-value store in response.",
                "tags": [
                    "Run Actor"
                ],
                "requestBody": {
                    "required": true,
                    "content": {
                        "application/json": {
                            "schema": {
                                "$ref": "#/components/schemas/inputSchema"
                            }
                        }
                    }
                },
                "parameters": [
                    {
                        "name": "token",
                        "in": "query",
                        "required": true,
                        "schema": {
                            "type": "string"
                        },
                        "description": "Enter your Apify token here"
                    }
                ],
                "responses": {
                    "200": {
                        "description": "OK"
                    }
                }
            }
        }
    },
    "components": {
        "schemas": {
            "inputSchema": {
                "type": "object",
                "required": [
                    "mode"
                ],
                "properties": {
                    "mode": {
                        "title": "Mode",
                        "enum": [
                            "extract"
                        ],
                        "type": "string",
                        "description": "Operation mode. 'extract' downloads each PDF URL and extracts its text, page count and document metadata.",
                        "default": "extract"
                    },
                    "urls": {
                        "title": "PDF URLs",
                        "type": "array",
                        "description": "Direct links to PDF files to extract. Each URL should point to a downloadable PDF (redirects are followed). Non-PDF URLs are logged and skipped.",
                        "items": {
                            "type": "string"
                        }
                    },
                    "includePages": {
                        "title": "Include per-page text",
                        "type": "boolean",
                        "description": "If enabled, output a 'pages' array with the extracted text of each page separately (each page capped at 20,000 characters). Off by default to keep items small.",
                        "default": false
                    },
                    "maxItems": {
                        "title": "Maximum items",
                        "minimum": 1,
                        "maximum": 1000,
                        "type": "integer",
                        "description": "Maximum number of PDFs to process and store in the dataset. Ranges from 1 to 1000.",
                        "default": 100
                    }
                }
            },
            "runsResponseSchema": {
                "type": "object",
                "properties": {
                    "data": {
                        "type": "object",
                        "properties": {
                            "id": {
                                "type": "string"
                            },
                            "actId": {
                                "type": "string"
                            },
                            "userId": {
                                "type": "string"
                            },
                            "startedAt": {
                                "type": "string",
                                "format": "date-time",
                                "example": "2025-01-08T00:00:00.000Z"
                            },
                            "finishedAt": {
                                "type": "string",
                                "format": "date-time",
                                "example": "2025-01-08T00:00:00.000Z"
                            },
                            "status": {
                                "type": "string",
                                "example": "READY"
                            },
                            "meta": {
                                "type": "object",
                                "properties": {
                                    "origin": {
                                        "type": "string",
                                        "example": "API"
                                    },
                                    "userAgent": {
                                        "type": "string"
                                    }
                                }
                            },
                            "stats": {
                                "type": "object",
                                "properties": {
                                    "inputBodyLen": {
                                        "type": "integer",
                                        "example": 2000
                                    },
                                    "rebootCount": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "restartCount": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "resurrectCount": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "computeUnits": {
                                        "type": "integer",
                                        "example": 0
                                    }
                                }
                            },
                            "options": {
                                "type": "object",
                                "properties": {
                                    "build": {
                                        "type": "string",
                                        "example": "latest"
                                    },
                                    "timeoutSecs": {
                                        "type": "integer",
                                        "example": 300
                                    },
                                    "memoryMbytes": {
                                        "type": "integer",
                                        "example": 1024
                                    },
                                    "diskMbytes": {
                                        "type": "integer",
                                        "example": 2048
                                    }
                                }
                            },
                            "buildId": {
                                "type": "string"
                            },
                            "defaultKeyValueStoreId": {
                                "type": "string"
                            },
                            "defaultDatasetId": {
                                "type": "string"
                            },
                            "defaultRequestQueueId": {
                                "type": "string"
                            },
                            "buildNumber": {
                                "type": "string",
                                "example": "1.0.0"
                            },
                            "containerUrl": {
                                "type": "string"
                            },
                            "usage": {
                                "type": "object",
                                "properties": {
                                    "ACTOR_COMPUTE_UNITS": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "DATASET_READS": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "DATASET_WRITES": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "KEY_VALUE_STORE_READS": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "KEY_VALUE_STORE_WRITES": {
                                        "type": "integer",
                                        "example": 1
                                    },
                                    "KEY_VALUE_STORE_LISTS": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "REQUEST_QUEUE_READS": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "REQUEST_QUEUE_WRITES": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "DATA_TRANSFER_INTERNAL_GBYTES": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "DATA_TRANSFER_EXTERNAL_GBYTES": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "PROXY_RESIDENTIAL_TRANSFER_GBYTES": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "PROXY_SERPS": {
                                        "type": "integer",
                                        "example": 0
                                    }
                                }
                            },
                            "usageTotalUsd": {
                                "type": "number",
                                "example": 0.00005
                            },
                            "usageUsd": {
                                "type": "object",
                                "properties": {
                                    "ACTOR_COMPUTE_UNITS": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "DATASET_READS": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "DATASET_WRITES": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "KEY_VALUE_STORE_READS": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "KEY_VALUE_STORE_WRITES": {
                                        "type": "number",
                                        "example": 0.00005
                                    },
                                    "KEY_VALUE_STORE_LISTS": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "REQUEST_QUEUE_READS": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "REQUEST_QUEUE_WRITES": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "DATA_TRANSFER_INTERNAL_GBYTES": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "DATA_TRANSFER_EXTERNAL_GBYTES": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "PROXY_RESIDENTIAL_TRANSFER_GBYTES": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "PROXY_SERPS": {
                                        "type": "integer",
                                        "example": 0
                                    }
                                }
                            }
                        }
                    }
                }
            }
        }
    }
}
```
