PDF Text Extractor: OCR & Metadata avatar

PDF Text Extractor: OCR & Metadata

Pricing

from $2.70 / 1,000 pdf texts

Go to Apify Store
PDF Text Extractor: OCR & Metadata

PDF Text Extractor: OCR & Metadata

Extract text from public PDF URLs, including scanned pages with OCR. Choose page ranges and get plain text, Markdown, page details, OCR information, metadata, bookmarks, and page-tagged chunks when available in your dataset.

Pricing

from $2.70 / 1,000 pdf texts

Rating

0.0

(0)

Developer

Maxime Dupré

Maxime Dupré

Maintained by Community

Actor stats

0

Bookmarked

3

Total users

2

Monthly active users

a day ago

Last modified

Share

📄 Turn PDFs into text with OCR

For developers, researchers, and teams looking for PDF to text OCR, this Actor reads public PDF URLs and returns plain text, Markdown, page details, OCR details, metadata, bookmarks, and text chunks in your dataset. Choose pages when you need part of a document, or use OCR for scanned pages.

📚 PDF text, pages, and document details

Each saved dataset row describes one submitted PDF. It can include the source URL, file name, media type, file size, page count, extracted text, Markdown, text counts, page-level text, OCR details, PDF metadata, bookmarks, and page-tagged chunks.

▶️ Run PDF text extraction

  1. Add one or more public PDF URLs.
  2. Optionally enter page numbers or ranges and choose whether to use OCR for scanned pages.
  3. Run the Actor and open the PDF results dataset.

⚙️ Input

Input fields

FieldTypeWhat it does
pdfSourcesarray of objectsRequired. Adds one or more public PDF sources.
pdfSources[].urlstring (URL)Public URL of the PDF to read.
pagesstringOptional page numbers or ranges, separated by commas, such as 1,3-5. Leave empty to read every page.
useOcrbooleanUses OCR for image-only or scanned pages when true. Set it to false to skip OCR and use normal PDF text extraction. Defaults to true.

Default input

This is the smallest common input from a successful current-beta run:

{
"pdfSources": [
{
"url": "https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf"
}
],
"useOcr": true
}

🧾 Output

Dataset fields

FieldTypeWhat it does
sourceUrlstring (URL)Public URL of the PDF read for this row.
fileNamestringFile name found for the PDF, when available.
mimeTypestringMedia type reported for the source file, when available.
fileSizeBytesintegerSource PDF size in bytes, when available.
pageCountintegerNumber of pages in the source PDF.
textstringText read from the requested pages, including OCR text when OCR is used.
markdownstringMarkdown version of the extracted text with layout cues when available.
characterCountintegerCharacter count for the requested content.
wordCountintegerWord count for the requested content.
pagesarray of objectsText and counts for each requested page.
pages[].pageNumberintegerNumber of the page in the source PDF.
pages[].textstringText read from that page.
pages[].characterCountintegerCharacter count for that page.
pages[].wordCountintegerWord count for that page.
ocrobjectOCR coverage and signals for the row.
ocr.pagesProcessedintegerNumber of pages read with OCR.
ocr.languagestringOCR language, when the source reports it.
ocr.confidencenumberOCR confidence reported by the source, when available.
ocr.remainingScannedPagesintegerScanned pages left without OCR text, when reported.
metadataobjectMetadata stored in the source PDF, when available.
metadata.titlestringTitle stored in the PDF metadata.
metadata.authorstringAuthor stored in the PDF metadata.
metadata.subjectstringSubject stored in the PDF metadata.
metadata.keywordsstringKeywords stored in the PDF metadata.
metadata.creatorstringApplication named as the PDF creator, when available.
metadata.producerstringApplication named as the PDF producer, when available.
metadata.creationDatestringCreation date stored in the PDF metadata.
metadata.modificationDatestringModification date stored in the PDF metadata.
bookmarksarray of objectsBookmark or table-of-contents entries found in the PDF.
bookmarks[].titlestringTitle of the bookmark entry.
bookmarks[].pageNumberintegerPage opened by the bookmark, when it has a destination.
bookmarks[].levelintegerNesting level of the bookmark in the PDF outline.
chunksarray of objectsPage-tagged text chunks for embedding and citation work, when produced.
chunks[].textstringText in the chunk.
chunks[].pageNumberintegerPage that the chunk belongs to.

Standard PDF row

This complete row is from a successful current-beta run:

{
"sourceUrl": "https://www.w3.org/WAI/WCAG22/working-examples/pdf-bookmarks/bookmarks.pdf",
"fileName": "bookmarks.pdf",
"mimeType": "application/pdf",
"fileSizeBytes": 106567,
"pageCount": 2,
"text": "1 Contents Header One ................................................................................................................................................... 2 Header Two for List ................................................................................................................................... 2 Header Three Text ................................................................................................................................ 2 Another Header One ..................................................................................................................................... 2 Another Header Two for List..................................................................................................................... 2",
"markdown": "1 Contents Header One ................................................................................................................................................... 2 Header Two for List ................................................................................................................................... 2 Header Three Text ................................................................................................................................ 2 Another Header One ..................................................................................................................................... 2 Another Header Two for List..................................................................................................................... 2",
"characterCount": 776,
"wordCount": 28,
"pages": [
{
"pageNumber": 1,
"text": "1 Contents Header One ................................................................................................................................................... 2 Header Two for List ................................................................................................................................... 2 Header Three Text ................................................................................................................................ 2 Another Header One ..................................................................................................................................... 2 Another Header Two for List..................................................................................................................... 2",
"characterCount": 776,
"wordCount": 28
}
],
"ocr": {
"pagesProcessed": 0
},
"metadata": {
"title": "Working example of creating bookmarks in PDF documents",
"creator": "Acrobat PDFMaker 9.0 for Word",
"producer": "Adobe PDF Library 9.0",
"creationDate": "D:20110202105905-05'00'",
"modificationDate": "D:20230718095527-07'00'"
},
"bookmarks": [
{
"title": "Header One",
"pageNumber": 2,
"level": 0
},
{
"title": "Header Two for List",
"pageNumber": 2,
"level": 1
},
{
"title": "Header Three Text",
"pageNumber": 2,
"level": 2
},
{
"title": "Header Four",
"pageNumber": 2,
"level": 3
},
{
"title": "Another Header One",
"pageNumber": 2,
"level": 0
},
{
"title": "Another Header Two for List",
"pageNumber": 2,
"level": 1
},
{
"title": "Links\r",
"level": 0
},
{
"title": "WCAG2.0",
"pageNumber": 2,
"level": 1
}
],
"chunks": [
{
"text": "1 Contents Header One ................................................................................................................................................... 2 Header Two for List ................................................................................................................................... 2 Header Three Text ................................................................................................................................ 2 Another Header One ..................................................................................................................................... 2 Another Header Two for List..................................................................................................................... 2",
"pageNumber": 1
}
]
}

OCR PDF row

This complete row is from a successful current-beta OCR run:

{
"sourceUrl": "https://raw.githubusercontent.com/tinytoolkit-org/pdf-sample-files/4fd733c47ead5446b06372ab910cb78d1c7da55e/out/sample-scanned.pdf",
"fileName": "sample-scanned.pdf",
"mimeType": "application/octet-stream",
"fileSizeBytes": 432160,
"pageCount": 3,
"text": "Scanned-style sample (image-only pages)\n\nPage 1 of 3\n\nThis document is a machine-generated sample file. It contains no real data: every name, number, and\nfigure on this page is placeholder content produced for software testing.\n\nUse it to exercise PDF tooling — text extraction, page manipulation, rendering, compression — without\nworrying about licensing or privacy. The file is dedicated to the public domain under CCO.\n\nText on this page is real, extractable text set in a standard font, not an image. A text-extraction tool\nshould recover these paragraphs verbatim, in reading order, with no OCR involved.\n\nPage boundaries matter for testing. Each page carries a heading with its own page number so that split,\nextract, reorder, and delete operations can be verified against what the output claims.\n\nIf a tool you are testing reports a different page count, a different order, or drops one of these\nparagraphs, the tool — not this file — is the thing to investigate next.\n\nA reasonable test suite checks the boring cases first: one page, ten pages, a page with nothing on it, a\npage rotated sideways. The files in this collection cover each of those separately.\n\n©CO / public domain — generated sample, no real data",
"markdown": "Scanned-style sample (image-only pages)\n\nPage 1 of 3\n\nThis document is a machine-generated sample file. It contains no real data: every name, number, and\nfigure on this page is placeholder content produced for software testing.\n\nUse it to exercise PDF tooling — text extraction, page manipulation, rendering, compression — without\nworrying about licensing or privacy. The file is dedicated to the public domain under CCO.\n\nText on this page is real, extractable text set in a standard font, not an image. A text-extraction tool\nshould recover these paragraphs verbatim, in reading order, with no OCR involved.\n\nPage boundaries matter for testing. Each page carries a heading with its own page number so that split,\nextract, reorder, and delete operations can be verified against what the output claims.\n\nIf a tool you are testing reports a different page count, a different order, or drops one of these\nparagraphs, the tool — not this file — is the thing to investigate next.\n\nA reasonable test suite checks the boring cases first: one page, ten pages, a page with nothing on it, a\npage rotated sideways. The files in this collection cover each of those separately.\n\n©CO / public domain — generated sample, no real data",
"characterCount": 1219,
"wordCount": 203,
"pages": [
{
"pageNumber": 1,
"text": "Scanned-style sample (image-only pages)\n\nPage 1 of 3\n\nThis document is a machine-generated sample file. It contains no real data: every name, number, and\nfigure on this page is placeholder content produced for software testing.\n\nUse it to exercise PDF tooling — text extraction, page manipulation, rendering, compression — without\nworrying about licensing or privacy. The file is dedicated to the public domain under CCO.\n\nText on this page is real, extractable text set in a standard font, not an image. A text-extraction tool\nshould recover these paragraphs verbatim, in reading order, with no OCR involved.\n\nPage boundaries matter for testing. Each page carries a heading with its own page number so that split,\nextract, reorder, and delete operations can be verified against what the output claims.\n\nIf a tool you are testing reports a different page count, a different order, or drops one of these\nparagraphs, the tool — not this file — is the thing to investigate next.\n\nA reasonable test suite checks the boring cases first: one page, ten pages, a page with nothing on it, a\npage rotated sideways. The files in this collection cover each of those separately.\n\n©CO / public domain — generated sample, no real data",
"characterCount": 1219,
"wordCount": 203
}
],
"ocr": {
"pagesProcessed": 1,
"language": "eng",
"confidence": 94,
"remainingScannedPages": 0
},
"metadata": {
"title": "Scanned-style sample",
"author": "tinytoolkit sample files",
"subject": "Image-only pages — no text layer, use for OCR testing",
"keywords": "sample test pdf cc0 pdftoolskit.org",
"creator": "https://pdftoolskit.org/sample-pdfs",
"producer": "pdf-sample-files generator (pdf-lib)",
"creationDate": "D:20260719000000Z",
"modificationDate": "D:20260719000000Z"
},
"chunks": [
{
"text": "Scanned-style sample (image-only pages) Page 1 of 3 This document is a machine-generated sample file. It contains no real data: every name, number, and figure on this page is placeholder content produced for software testing. Use it to exercise PDF tooling — text extraction, page manipulation, rendering, compression — without worrying about licensing or privacy. The file is dedicated to the public domain under CCO. Text on this page is real, extractable text set in a standard font, not an image. A text-extraction tool should recover these paragraphs verbatim, in reading order, with no OCR involved. Page boundaries matter for testing. Each page carries a heading with its own page number so that split, extract, reorder, and delete operations can",
"pageNumber": 1
},
{
"text": "be verified against what the output claims. If a tool you are testing reports a different page count, a different order, or drops one of these paragraphs, the tool — not this file — is the thing to investigate next. A reasonable test suite checks the boring cases first: one page, ten pages, a page with nothing on it, a page rotated sideways. The files in this collection cover each of those separately. ©CO / public domain — generated sample, no real data",
"pageNumber": 1
}
]
}

💳 Pricing

How charges work

The PDF text event is charged once for each PDF that is processed and returns extracted text. The OCR page event is charged for each scanned or image-only page successfully read with OCR and returned as text. Current tiered rates are shown in the Pricing tab.

🔌 Integrations

Start runs in Apify Console or through the Apify API, then read the PDF results dataset in JSON, CSV, or Excel. Use the dataset in your own script or document workflow.

❓ FAQ

Can this read scanned PDFs?

Yes. Set useOcr to true to read image-only or scanned pages with OCR. The OCR details in each row show the pages processed and any other signals the source reports.

Can I extract only some pages?

Yes. Enter page numbers or ranges in pages, such as 1,3-5. Leave it empty to read every page in each submitted PDF.

Can I submit more than one PDF?

Yes. Add one or more public PDF URL objects to pdfSources. The Actor returns a document result for each submitted source.

Can I upload a local or private PDF?

No. This input accepts public PDF URLs. It does not bypass login pages, paywalls, DRM, or other access restrictions.

Does it return tables or images as separate files?

No. The Actor returns extracted text and Markdown, plus document and page details. It does not extract structured table cells or image files.

📝 Changelog

v0.0 (30-09-2026)

  • Initial release.

🆘 Support

For issues, questions, or feature requests, file a ticket and I'll fix or implement it in less than 24h 🫡

Made with ❤️ by Maxime Dupré