PDF Text Extractor: Text, Markdown & RAG Chunks
Pricing
from $3.00 / 1,000 pdf extracteds
PDF Text Extractor: Text, Markdown & RAG Chunks
Extract clean text, Markdown with headings, per-page text, metadata and RAG-ready chunks from PDF URLs, or give a web page and every linked PDF is found and extracted. Pay only for successfully extracted PDFs.
Pricing
from $3.00 / 1,000 pdf extracteds
Rating
0.0
(0)
Developer
Digitální produkty pro život
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
5 hours ago
Last modified
Categories
Share
Turn any PDF into clean, structured text. Paste links to PDF files, or just paste a web page (a reports, downloads or documents page) and every PDF linked from it is found and extracted for you.
You get plain text, Markdown with headings and bullet lists, document metadata, optional text per page and optional RAG-ready chunks for embeddings and vector databases.
What you get for each PDF
| Field | Example |
|---|---|
url, fileName, sourcePage | where the PDF came from |
title, author, subject, keywords | from PDF metadata (title falls back to the biggest heading) |
createdAt, modifiedAt | ISO dates |
pageCount, pagesExtracted, wordCount, charCount | size of the document |
text | clean text with paragraphs |
markdown | headings (#, ##, ###) detected from font sizes, bullet lists, joined hyphenated words |
pages | [{ page, text }] when Text per page is on, great for citations |
chunks | text split on paragraph and sentence boundaries when RAG chunk size is set |
A run summary with every skipped file and the reason is saved to the SUMMARY record.
Use cases
- LLM and RAG pipelines: feed reports, manuals and papers into ChatGPT, Claude or a vector database as Markdown or ready-made chunks.
- Monitor document pages: extract all annual reports, tenders, price lists or policies linked from a company page, and schedule the run.
- Research: bulk-extract scientific papers (arXiv and others) with titles and page counts.
- Search and archiving: make PDF libraries full-text searchable.
How to use
- Add PDF links or web page links to PDF or page URLs.
- Optional: turn on Text per page, set RAG chunk size (e.g.
1000) or limit Max pages per PDF. - Run and download the results as JSON, CSV, Excel or via API.
Example input
{"urls": ["https://arxiv.org/pdf/1706.03762", "https://www.irs.gov/forms-pubs/about-form-w-9"],"chunkSize": 1000,"includePages": true}
Example output (shortened)
{"url": "https://arxiv.org/pdf/1706.03762","title": "Attention Is All You Need","pageCount": 15,"wordCount": 6225,"createdAt": "2024-04-10T21:11:43.000Z","markdown": "# Attention Is All You Need\n\nAshish Vaswani ...\n\n### Abstract\n\nThe dominant sequence transduction models ...","chunks": ["Attention Is All You Need ...", "..."]}
Pricing
You pay only for successfully extracted PDFs, no matter how many pages they have. Failed downloads, broken files, password-protected files without a password and scanned PDFs without a text layer are never charged. See the Pricing tab for the current price.
Limits and notes
- Scanned PDFs (images only, no text layer) are skipped, because OCR is not included.
- Password-protected PDFs work when you enter the password.
- Files above Max file size (default 50 MB) are skipped.
- Multi-column layouts follow the reading order stored in the PDF, which is correct for most documents.
- Only process documents you are allowed to access and use.
Questions or ideas?
Open an issue on the Issues tab. Feature requests are welcome and usually answered within a day or two.
How to use it via API
You can run the Actor from the Apify Console, on a schedule, or from your own code. Get your API token in Apify Console → Settings → Integrations.
Python
from apify_client import ApifyClientclient = ApifyClient("<YOUR_API_TOKEN>")run = client.actor("digitalni.produkty.pro.zivot/pdf-text-extractor").call(run_input={"urls": ["https://arxiv.org/pdf/1706.03762"],"includeMarkdown": True,"chunkSize": 1000})for item in client.dataset(run["defaultDatasetId"]).iterate_items():print(item)
JavaScript / Node.js
import { ApifyClient } from 'apify-client';const client = new ApifyClient({ token: '<YOUR_API_TOKEN>' });const run = await client.actor('digitalni.produkty.pro.zivot/pdf-text-extractor').call({"urls": ["https://arxiv.org/pdf/1706.03762"],"includeMarkdown": true,"chunkSize": 1000});const { items } = await client.dataset(run.defaultDatasetId).listItems();console.log(items);
Integrations and AI agents
- Export results as JSON, CSV, Excel, XML or HTML, or open them directly in Google Sheets.
- Connect to Make, Zapier, n8n, Slack, Google Drive, Airbyte or any webhook to get text, Markdown and RAG chunks from PDFs into your workflow automatically.
- Use it from AI agents and LLM apps (Claude, ChatGPT, Cursor, LangChain…) through the Apify MCP server: the agent can call this Actor as a tool.
- Schedule runs (hourly, daily, weekly) to keep data fresh without any code.
FAQ
How much does it cost? $0.003 per successfully extracted PDF, regardless of page count. 1,000 PDFs cost $3. Failed, scanned and skipped files are free.
Does it do OCR on scanned PDFs? No. It extracts the text layer, which is fast and cheap. Image-only scans are detected, reported and not charged.
Can I use it for RAG / LLM pipelines? Yes. Set a RAG chunk size (and overlap) and load the dataset straight into LangChain, LlamaIndex or a vector database through the Apify API.
Can it find PDFs on a web page? Yes. Give a page URL and every linked PDF is discovered and extracted.