PDF Text Extractor: PDF URL to Text and Markdown per Page
Pricing
Pay per event
PDF Text Extractor: PDF URL to Text and Markdown per Page
Convert PDF links to clean text and Markdown in reading order, per document or per page, with title, author, dates and page count. Headings and lists kept for LLM and RAG pipelines. $2 per 1,000 PDFs, failed files free, no OCR needed for digital PDFs.
Pricing
Pay per event
Rating
0.0
(0)
Developer
Hay Equipos
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
a day ago
Last modified
Categories
Share
Give this actor links to PDF files and get back clean text and Markdown in reading order, one row per document or one row per page, with the document's title, author, subject, creator, creation and modification dates, page count and file size. Two column layouts (papers, forms, reports) are read column by column, words split across lines are rejoined, and larger fonts become Markdown headings, so the output is ready for search, LLM prompts and RAG chunking.
It downloads each file over plain HTTP and extracts the text layer with pdf.js, the PDF engine used in Firefox. No browser, no OCR service, no API key. You pay per PDF, plus a tiny start fee per run; files that fail cost nothing.
What you can use it for
- Turn reports, papers, filings, manuals and brochures into text for an LLM or a vector database.
- Chunk long documents by page, with page numbers kept for citations.
- Pull titles, authors and dates from a batch of PDFs into a spreadsheet.
- Monitor PDFs that change (price lists, policies, tenders) by extracting them on a schedule.
- Check which PDFs in a list are scans without a text layer (
likelyScanned).
Input
| Field | What it does | Default |
|---|---|---|
| PDF URLs | Direct links to PDF files | required |
| One row per | document (whole text in one row) or pages (one row per page) | document |
| Include plain text | Text in reading order with paragraph breaks | on |
| Include Markdown | Markdown with headings, bullet lists and a rule between pages | on |
| Maximum pages per PDF | Read at most this many pages from each file | 300 |
| Maximum file size (MB) | Larger files are skipped for free | 50 |
| Maximum text length | Cut text and Markdown per row (0 means no limit) | 0 |
| Respect robots.txt | Skip files a site closes to automated tools (skipped files are free) | on |
| Parallel downloads | PDFs worked on at once; one host always gets one download per second | 2 |
Example input:
{"urls": ["https://arxiv.org/pdf/1706.03762","https://www.irs.gov/pub/irs-pdf/fw9.pdf"],"outputMode": "pages","maxPagesPerPdf": 50}
Output
Document mode, one row per PDF (text shortened here):
{"url": "https://www.irs.gov/pub/irs-pdf/fw9.pdf","finalUrl": "https://www.irs.gov/pub/irs-pdf/fw9.pdf","success": true,"fileName": "fw9.pdf","fileSizeBytes": 140815,"pageCount": 6,"pagesExtracted": 6,"title": "Form W-9 (Rev. March 2024)","author": "SE:W:CAR:MP","subject": "Request for Taxpayer Identification Number and Certification","keywords": "Fillable","creator": "Designer 6.5","producer": "Designer 6.5","createdAt": "2024-03-06T13:18:13.000Z","modifiedAt": "2024-03-06T13:18:13.000Z","pdfVersion": "1.7","likelyScanned": false,"wordCount": 6277,"charCount": 37851,"textTruncated": false,"text": "FormW-9 Request for Taxpayer Give form to the\n\n(Rev. March 2024) Identification Number and Certification requester. Do not\n\nDepartment of the Treasury send to the IRS...","markdown": "## FormW-9 Request for Taxpayer Give form to the\n\n### (Rev. March 2024) Identification Number and Certification requester. Do not\n\nDepartment of the Treasury send to the IRS...","extractedAt": "2026-09-27T06:23:12.400Z"}
Page mode adds pageNumber and gives wordCount, charCount, text and markdown for that page only; the document fields repeat on every row so each page stands alone.
Failed files come back as success: false with an error such as "The server answered HTTP 404", "The URL returned a web page, not a PDF" or "The PDF is password protected".
Pricing
Pay per event, no subscription, no charge for platform usage on top.
| Event | Price |
|---|---|
| PDF processed | $0.002 ($2 per 1,000 PDFs), the same in document and page mode |
| Actor start | $0.00005 per run (Apify's standard start event) |
Downloads that fail, files that are not PDFs, password protected files, robots.txt skips and files over your size limit are free. You can set a maximum charge per run in Apify and the actor stops cleanly when it is reached.
Limits
- No OCR. Scanned PDFs without a text layer return little or no text and are flagged with
likelyScanned: true. - Tables come out as text lines in reading order, not as structured cells.
- Very complex layouts (three or more columns, text boxes scattered over the page) may read in an imperfect order.
- Links that need a login, a cookie banner click or a captcha before the file downloads return an error row.
- Default memory is 512 MB, enough for typical files up to about 50 MB; give the run more memory for very large files.
- Up to 2,000 PDFs per run.
FAQ
Do I need an API key? No. Paste the links and run.
Can I process files I have on my computer? Upload them to an Apify key value store (or any storage with a public link) and pass those links.
Why is the price per PDF and not per page? It keeps costs predictable for you and for AI agents. Use "Maximum pages per PDF" to cap very long files.
What does likelyScanned mean? The file has under about 40 characters of text per page, which usually means scanned images. Those need an OCR tool.
Is the Markdown exact? Headings are detected from font size and bullets from bullet characters. It is built for reading and chunking, not for pixel perfect layout.