PDF Text Extractor: Text, Markdown & RAG Chunks avatar

PDF Text Extractor: Text, Markdown & RAG Chunks

Pricing

from $3.00 / 1,000 pdf extracteds

Go to Apify Store
PDF Text Extractor: Text, Markdown & RAG Chunks

PDF Text Extractor: Text, Markdown & RAG Chunks

Extract clean text, Markdown with headings, per-page text, metadata and RAG-ready chunks from PDF URLs, or give a web page and every linked PDF is found and extracted. Pay only for successfully extracted PDFs.

Pricing

from $3.00 / 1,000 pdf extracteds

Rating

0.0

(0)

Developer

Digitální produkty pro život

Digitální produkty pro život

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

5 hours ago

Last modified

Share

Turn any PDF into clean, structured text. Paste links to PDF files, or just paste a web page (a reports, downloads or documents page) and every PDF linked from it is found and extracted for you.

You get plain text, Markdown with headings and bullet lists, document metadata, optional text per page and optional RAG-ready chunks for embeddings and vector databases.

What you get for each PDF

FieldExample
url, fileName, sourcePagewhere the PDF came from
title, author, subject, keywordsfrom PDF metadata (title falls back to the biggest heading)
createdAt, modifiedAtISO dates
pageCount, pagesExtracted, wordCount, charCountsize of the document
textclean text with paragraphs
markdownheadings (#, ##, ###) detected from font sizes, bullet lists, joined hyphenated words
pages[{ page, text }] when Text per page is on, great for citations
chunkstext split on paragraph and sentence boundaries when RAG chunk size is set

A run summary with every skipped file and the reason is saved to the SUMMARY record.

Use cases

  • LLM and RAG pipelines: feed reports, manuals and papers into ChatGPT, Claude or a vector database as Markdown or ready-made chunks.
  • Monitor document pages: extract all annual reports, tenders, price lists or policies linked from a company page, and schedule the run.
  • Research: bulk-extract scientific papers (arXiv and others) with titles and page counts.
  • Search and archiving: make PDF libraries full-text searchable.

How to use

  1. Add PDF links or web page links to PDF or page URLs.
  2. Optional: turn on Text per page, set RAG chunk size (e.g. 1000) or limit Max pages per PDF.
  3. Run and download the results as JSON, CSV, Excel or via API.

Example input

{
"urls": ["https://arxiv.org/pdf/1706.03762", "https://www.irs.gov/forms-pubs/about-form-w-9"],
"chunkSize": 1000,
"includePages": true
}

Example output (shortened)

{
"url": "https://arxiv.org/pdf/1706.03762",
"title": "Attention Is All You Need",
"pageCount": 15,
"wordCount": 6225,
"createdAt": "2024-04-10T21:11:43.000Z",
"markdown": "# Attention Is All You Need\n\nAshish Vaswani ...\n\n### Abstract\n\nThe dominant sequence transduction models ...",
"chunks": ["Attention Is All You Need ...", "..."]
}

Pricing

You pay only for successfully extracted PDFs, no matter how many pages they have. Failed downloads, broken files, password-protected files without a password and scanned PDFs without a text layer are never charged. See the Pricing tab for the current price.

Limits and notes

  • Scanned PDFs (images only, no text layer) are skipped, because OCR is not included.
  • Password-protected PDFs work when you enter the password.
  • Files above Max file size (default 50 MB) are skipped.
  • Multi-column layouts follow the reading order stored in the PDF, which is correct for most documents.
  • Only process documents you are allowed to access and use.

Questions or ideas?

Open an issue on the Issues tab. Feature requests are welcome and usually answered within a day or two.

How to use it via API

You can run the Actor from the Apify Console, on a schedule, or from your own code. Get your API token in Apify Console → Settings → Integrations.

Python

from apify_client import ApifyClient
client = ApifyClient("<YOUR_API_TOKEN>")
run = client.actor("digitalni.produkty.pro.zivot/pdf-text-extractor").call(run_input={
"urls": [
"https://arxiv.org/pdf/1706.03762"
],
"includeMarkdown": True,
"chunkSize": 1000
})
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
print(item)

JavaScript / Node.js

import { ApifyClient } from 'apify-client';
const client = new ApifyClient({ token: '<YOUR_API_TOKEN>' });
const run = await client.actor('digitalni.produkty.pro.zivot/pdf-text-extractor').call({
"urls": [
"https://arxiv.org/pdf/1706.03762"
],
"includeMarkdown": true,
"chunkSize": 1000
});
const { items } = await client.dataset(run.defaultDatasetId).listItems();
console.log(items);

Integrations and AI agents

  • Export results as JSON, CSV, Excel, XML or HTML, or open them directly in Google Sheets.
  • Connect to Make, Zapier, n8n, Slack, Google Drive, Airbyte or any webhook to get text, Markdown and RAG chunks from PDFs into your workflow automatically.
  • Use it from AI agents and LLM apps (Claude, ChatGPT, Cursor, LangChain…) through the Apify MCP server: the agent can call this Actor as a tool.
  • Schedule runs (hourly, daily, weekly) to keep data fresh without any code.

FAQ

How much does it cost? $0.003 per successfully extracted PDF, regardless of page count. 1,000 PDFs cost $3. Failed, scanned and skipped files are free.

Does it do OCR on scanned PDFs? No. It extracts the text layer, which is fast and cheap. Image-only scans are detected, reported and not charged.

Can I use it for RAG / LLM pipelines? Yes. Set a RAG chunk size (and overlap) and load the dataset straight into LangChain, LlamaIndex or a vector database through the Apify API.

Can it find PDFs on a web page? Yes. Give a page URL and every linked PDF is discovered and extracted.