PDF Text Extractor: OCR & Metadata
Pricing
from $2.70 / 1,000 pdf texts
PDF Text Extractor: OCR & Metadata
Extract text from public PDF URLs, including scanned pages with OCR. Choose page ranges and get plain text, Markdown, page details, OCR information, metadata, bookmarks, and page-tagged chunks when available in your dataset.
Pricing
from $2.70 / 1,000 pdf texts
Rating
0.0
(0)
Developer
Maxime Dupré
Maintained by CommunityActor stats
0
Bookmarked
3
Total users
2
Monthly active users
a day ago
Last modified
Categories
Share
📄 Turn PDFs into text with OCR
For developers, researchers, and teams looking for PDF to text OCR, this Actor reads public PDF URLs and returns plain text, Markdown, page details, OCR details, metadata, bookmarks, and text chunks in your dataset. Choose pages when you need part of a document, or use OCR for scanned pages.
- Read a scanned document with Scanned PDF to Text and save its recovered text.
- Extract Text from PDF returns readable text from a public PDF URL.
- Turn a public document into plain text with PDF to Text.
- Read scanned PDF pages with Online OCR and review the returned text.
- Recover text from image-only pages with OCR PDF.
📚 PDF text, pages, and document details
Each saved dataset row describes one submitted PDF. It can include the source URL, file name, media type, file size, page count, extracted text, Markdown, text counts, page-level text, OCR details, PDF metadata, bookmarks, and page-tagged chunks.
▶️ Run PDF text extraction
- Add one or more public PDF URLs.
- Optionally enter page numbers or ranges and choose whether to use OCR for scanned pages.
- Run the Actor and open the PDF results dataset.
⚙️ Input
Input fields
| Field | Type | What it does |
|---|---|---|
pdfSources | array of objects | Required. Adds one or more public PDF sources. |
pdfSources[].url | string (URL) | Public URL of the PDF to read. |
pages | string | Optional page numbers or ranges, separated by commas, such as 1,3-5. Leave empty to read every page. |
useOcr | boolean | Uses OCR for image-only or scanned pages when true. Set it to false to skip OCR and use normal PDF text extraction. Defaults to true. |
Default input
This is the smallest common input from a successful current-beta run:
{"pdfSources": [{"url": "https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf"}],"useOcr": true}
🧾 Output
Dataset fields
| Field | Type | What it does |
|---|---|---|
sourceUrl | string (URL) | Public URL of the PDF read for this row. |
fileName | string | File name found for the PDF, when available. |
mimeType | string | Media type reported for the source file, when available. |
fileSizeBytes | integer | Source PDF size in bytes, when available. |
pageCount | integer | Number of pages in the source PDF. |
text | string | Text read from the requested pages, including OCR text when OCR is used. |
markdown | string | Markdown version of the extracted text with layout cues when available. |
characterCount | integer | Character count for the requested content. |
wordCount | integer | Word count for the requested content. |
pages | array of objects | Text and counts for each requested page. |
pages[].pageNumber | integer | Number of the page in the source PDF. |
pages[].text | string | Text read from that page. |
pages[].characterCount | integer | Character count for that page. |
pages[].wordCount | integer | Word count for that page. |
ocr | object | OCR coverage and signals for the row. |
ocr.pagesProcessed | integer | Number of pages read with OCR. |
ocr.language | string | OCR language, when the source reports it. |
ocr.confidence | number | OCR confidence reported by the source, when available. |
ocr.remainingScannedPages | integer | Scanned pages left without OCR text, when reported. |
metadata | object | Metadata stored in the source PDF, when available. |
metadata.title | string | Title stored in the PDF metadata. |
metadata.author | string | Author stored in the PDF metadata. |
metadata.subject | string | Subject stored in the PDF metadata. |
metadata.keywords | string | Keywords stored in the PDF metadata. |
metadata.creator | string | Application named as the PDF creator, when available. |
metadata.producer | string | Application named as the PDF producer, when available. |
metadata.creationDate | string | Creation date stored in the PDF metadata. |
metadata.modificationDate | string | Modification date stored in the PDF metadata. |
bookmarks | array of objects | Bookmark or table-of-contents entries found in the PDF. |
bookmarks[].title | string | Title of the bookmark entry. |
bookmarks[].pageNumber | integer | Page opened by the bookmark, when it has a destination. |
bookmarks[].level | integer | Nesting level of the bookmark in the PDF outline. |
chunks | array of objects | Page-tagged text chunks for embedding and citation work, when produced. |
chunks[].text | string | Text in the chunk. |
chunks[].pageNumber | integer | Page that the chunk belongs to. |
Standard PDF row
This complete row is from a successful current-beta run:
{"sourceUrl": "https://www.w3.org/WAI/WCAG22/working-examples/pdf-bookmarks/bookmarks.pdf","fileName": "bookmarks.pdf","mimeType": "application/pdf","fileSizeBytes": 106567,"pageCount": 2,"text": "1 Contents Header One ................................................................................................................................................... 2 Header Two for List ................................................................................................................................... 2 Header Three Text ................................................................................................................................ 2 Another Header One ..................................................................................................................................... 2 Another Header Two for List..................................................................................................................... 2","markdown": "1 Contents Header One ................................................................................................................................................... 2 Header Two for List ................................................................................................................................... 2 Header Three Text ................................................................................................................................ 2 Another Header One ..................................................................................................................................... 2 Another Header Two for List..................................................................................................................... 2","characterCount": 776,"wordCount": 28,"pages": [{"pageNumber": 1,"text": "1 Contents Header One ................................................................................................................................................... 2 Header Two for List ................................................................................................................................... 2 Header Three Text ................................................................................................................................ 2 Another Header One ..................................................................................................................................... 2 Another Header Two for List..................................................................................................................... 2","characterCount": 776,"wordCount": 28}],"ocr": {"pagesProcessed": 0},"metadata": {"title": "Working example of creating bookmarks in PDF documents","creator": "Acrobat PDFMaker 9.0 for Word","producer": "Adobe PDF Library 9.0","creationDate": "D:20110202105905-05'00'","modificationDate": "D:20230718095527-07'00'"},"bookmarks": [{"title": "Header One","pageNumber": 2,"level": 0},{"title": "Header Two for List","pageNumber": 2,"level": 1},{"title": "Header Three Text","pageNumber": 2,"level": 2},{"title": "Header Four","pageNumber": 2,"level": 3},{"title": "Another Header One","pageNumber": 2,"level": 0},{"title": "Another Header Two for List","pageNumber": 2,"level": 1},{"title": "Links\r","level": 0},{"title": "WCAG2.0","pageNumber": 2,"level": 1}],"chunks": [{"text": "1 Contents Header One ................................................................................................................................................... 2 Header Two for List ................................................................................................................................... 2 Header Three Text ................................................................................................................................ 2 Another Header One ..................................................................................................................................... 2 Another Header Two for List..................................................................................................................... 2","pageNumber": 1}]}
OCR PDF row
This complete row is from a successful current-beta OCR run:
{"sourceUrl": "https://raw.githubusercontent.com/tinytoolkit-org/pdf-sample-files/4fd733c47ead5446b06372ab910cb78d1c7da55e/out/sample-scanned.pdf","fileName": "sample-scanned.pdf","mimeType": "application/octet-stream","fileSizeBytes": 432160,"pageCount": 3,"text": "Scanned-style sample (image-only pages)\n\nPage 1 of 3\n\nThis document is a machine-generated sample file. It contains no real data: every name, number, and\nfigure on this page is placeholder content produced for software testing.\n\nUse it to exercise PDF tooling — text extraction, page manipulation, rendering, compression — without\nworrying about licensing or privacy. The file is dedicated to the public domain under CCO.\n\nText on this page is real, extractable text set in a standard font, not an image. A text-extraction tool\nshould recover these paragraphs verbatim, in reading order, with no OCR involved.\n\nPage boundaries matter for testing. Each page carries a heading with its own page number so that split,\nextract, reorder, and delete operations can be verified against what the output claims.\n\nIf a tool you are testing reports a different page count, a different order, or drops one of these\nparagraphs, the tool — not this file — is the thing to investigate next.\n\nA reasonable test suite checks the boring cases first: one page, ten pages, a page with nothing on it, a\npage rotated sideways. The files in this collection cover each of those separately.\n\n©CO / public domain — generated sample, no real data","markdown": "Scanned-style sample (image-only pages)\n\nPage 1 of 3\n\nThis document is a machine-generated sample file. It contains no real data: every name, number, and\nfigure on this page is placeholder content produced for software testing.\n\nUse it to exercise PDF tooling — text extraction, page manipulation, rendering, compression — without\nworrying about licensing or privacy. The file is dedicated to the public domain under CCO.\n\nText on this page is real, extractable text set in a standard font, not an image. A text-extraction tool\nshould recover these paragraphs verbatim, in reading order, with no OCR involved.\n\nPage boundaries matter for testing. Each page carries a heading with its own page number so that split,\nextract, reorder, and delete operations can be verified against what the output claims.\n\nIf a tool you are testing reports a different page count, a different order, or drops one of these\nparagraphs, the tool — not this file — is the thing to investigate next.\n\nA reasonable test suite checks the boring cases first: one page, ten pages, a page with nothing on it, a\npage rotated sideways. The files in this collection cover each of those separately.\n\n©CO / public domain — generated sample, no real data","characterCount": 1219,"wordCount": 203,"pages": [{"pageNumber": 1,"text": "Scanned-style sample (image-only pages)\n\nPage 1 of 3\n\nThis document is a machine-generated sample file. It contains no real data: every name, number, and\nfigure on this page is placeholder content produced for software testing.\n\nUse it to exercise PDF tooling — text extraction, page manipulation, rendering, compression — without\nworrying about licensing or privacy. The file is dedicated to the public domain under CCO.\n\nText on this page is real, extractable text set in a standard font, not an image. A text-extraction tool\nshould recover these paragraphs verbatim, in reading order, with no OCR involved.\n\nPage boundaries matter for testing. Each page carries a heading with its own page number so that split,\nextract, reorder, and delete operations can be verified against what the output claims.\n\nIf a tool you are testing reports a different page count, a different order, or drops one of these\nparagraphs, the tool — not this file — is the thing to investigate next.\n\nA reasonable test suite checks the boring cases first: one page, ten pages, a page with nothing on it, a\npage rotated sideways. The files in this collection cover each of those separately.\n\n©CO / public domain — generated sample, no real data","characterCount": 1219,"wordCount": 203}],"ocr": {"pagesProcessed": 1,"language": "eng","confidence": 94,"remainingScannedPages": 0},"metadata": {"title": "Scanned-style sample","author": "tinytoolkit sample files","subject": "Image-only pages — no text layer, use for OCR testing","keywords": "sample test pdf cc0 pdftoolskit.org","creator": "https://pdftoolskit.org/sample-pdfs","producer": "pdf-sample-files generator (pdf-lib)","creationDate": "D:20260719000000Z","modificationDate": "D:20260719000000Z"},"chunks": [{"text": "Scanned-style sample (image-only pages) Page 1 of 3 This document is a machine-generated sample file. It contains no real data: every name, number, and figure on this page is placeholder content produced for software testing. Use it to exercise PDF tooling — text extraction, page manipulation, rendering, compression — without worrying about licensing or privacy. The file is dedicated to the public domain under CCO. Text on this page is real, extractable text set in a standard font, not an image. A text-extraction tool should recover these paragraphs verbatim, in reading order, with no OCR involved. Page boundaries matter for testing. Each page carries a heading with its own page number so that split, extract, reorder, and delete operations can","pageNumber": 1},{"text": "be verified against what the output claims. If a tool you are testing reports a different page count, a different order, or drops one of these paragraphs, the tool — not this file — is the thing to investigate next. A reasonable test suite checks the boring cases first: one page, ten pages, a page with nothing on it, a page rotated sideways. The files in this collection cover each of those separately. ©CO / public domain — generated sample, no real data","pageNumber": 1}]}
💳 Pricing
How charges work
The PDF text event is charged once for each PDF that is processed and returns extracted text. The OCR page event is charged for each scanned or image-only page successfully read with OCR and returned as text. Current tiered rates are shown in the Pricing tab.
🔌 Integrations
Start runs in Apify Console or through the Apify API, then read the PDF results dataset in JSON, CSV, or Excel. Use the dataset in your own script or document workflow.
❓ FAQ
Can this read scanned PDFs?
Yes. Set useOcr to true to read image-only or scanned pages with OCR. The OCR details in each row show the pages processed and any other signals the source reports.
Can I extract only some pages?
Yes. Enter page numbers or ranges in pages, such as 1,3-5. Leave it empty to read every page in each submitted PDF.
Can I submit more than one PDF?
Yes. Add one or more public PDF URL objects to pdfSources. The Actor returns a document result for each submitted source.
Can I upload a local or private PDF?
No. This input accepts public PDF URLs. It does not bypass login pages, paywalls, DRM, or other access restrictions.
Does it return tables or images as separate files?
No. The Actor returns extracted text and Markdown, plus document and page details. It does not extract structured table cells or image files.
📝 Changelog
v0.0 (30-09-2026)
- Initial release.
🆘 Support
For issues, questions, or feature requests, file a ticket and I'll fix or implement it in less than 24h 🫡
🔗 Related Actors
- Webpage Text Extractor - Extract public webpage text and Markdown before or alongside PDF work.
- Markdown to HTML Converter - Turn returned Markdown into HTML for a page or document.
- URL to BibTeX Converter - Create citation data from public paper and article URLs for research workflows.
- PDF Text Extractor — Markdown, OCR & Metadata - Compare another PDF text workflow with Markdown, OCR, and metadata.
- PDF Text Extractor - Multi-Column Layout, OCR & Metadata - Compare a PDF text workflow with multi-column layout, OCR, and metadata.
Made with ❤️ by Maxime Dupré