PDF Text Extractor: OCR, Markdown, Page Checks avatar

PDF Text Extractor: OCR, Markdown, Page Checks

Pricing

from $4.00 / 1,000 document returneds

Go to Apify Store
PDF Text Extractor: OCR, Markdown, Page Checks

PDF Text Extractor: OCR, Markdown, Page Checks

Extract text from PDF URLs or your own files, with OCR for scanned pages and a verdict on every page. Each row carries the full text or Markdown, page count, word count, PDF type, and flags for wrong or hidden text layers. Page ranges and passwords work. Failed files are free.

Pricing

from $4.00 / 1,000 document returneds

Rating

0.0

(0)

Developer

Pradio Actors

Pradio Actors

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

8 hours ago

Last modified

Share

What does PDF Text Extractor do?

PDF Text Extractor turns your PDF URLs or uploaded files into text, one row per document, with OCR for scanned pages and a verdict on every page. Each row carries the whole text in one column, a verdict per page and the file's own facts, 27 fields in all. Each page is answered from its embedded text layer and scored on the layer's own statistics; a page with no layer is OCR'd instead. Switch on verifyPages and every page is also rendered and compared with what it shows.

On a 52-document platform run over files it had never seen, 48 returned a data row. Other extractors hand back whatever the embedded layer says. This one flags the layers that look wrong, and can check them against the render: $0.004 per returned document, and a miss costs nothing.

Who uses PDF Text Extractor

WhoWhat they run it for
RAG and search pipeline buildersChecked text into their indexes, so a file with a bad text layer does not silently feed garbage into embeddings
Compliance and legal review teamsA per-document verdict that flags files with suspect or hidden text before a person reads them
Civic-data projectsBulk checks on public government PDFs, such as Federal Register releases and agency technical reports, for scanned or broken files
Document-ingest operatorsAcceptance checks on downloaded PDF batches: which files read clean and which need re-sourcing

Features

  • Layer-first reads. Each page's embedded text layer is pulled directly. Clean digital documents never wait on a render or an OCR pass.
  • OCR only where a page has no layer. A page with nothing embedded is rendered and OCR'd, so scanned files still deliver text. In the platform run, 12 of the 48 files that resolved needed OCR on at least one page.
  • A statistics check on every layer. Each layer page is scored on its own text: improbable word shapes, letter/digit mixes, interior capitals, sparse ragged lines. A page that fails still ships its layer text, flagged suspect_layer. Landscape pages (slides, charts) are not judged on their short lines, and compound words like CubeSat or GitHub are not read as case noise. Dense technical jargon and short portrait pages such as covers can still read suspect, and verifyPages settles it.
  • Optional render verification. verifyPages renders and OCRs every page and compares the layer with what the page shows, catching wrong layers and hidden content the statistics check cannot see. Each verified page is one page-verified event at $0.008.
  • A document-level status. status stays agree unless at least 20% of the pages carry wrong_layer, suspect_layer, hidden_content, no_layer or unreadable; then the dominant flag wins. With verifyPages off, agree means the layers passed the statistics check, not that they were compared with the render.
  • The whole text in one column. text carries every read page joined, so a CSV or Excel export has the full document in one cell. The per-page text stays in page_verdicts.
  • Markdown output. Set outputFormat to markdown or both and markdown carries the document with headings from larger type, paragraphs from line gaps and list items from bullet glyphs and numbering. It is filled only when you ask for it: on the default text output, markdown is null.
  • Batch URL input, or your own files. A plain list of PDF URLs, or files sent inline as base64 in base64Pdfs. One row per document, up to maxItems rows per run.
  • Page ranges. pageRange reads only the pages you name, such as 1-5,8, in every document.
  • Password-protected PDFs. pdfPassword opens them. A missing or wrong password is an uncharged miss row that says which.
  • Misses that name the cause. A failed download says whether the host name did not resolve, the request timed out, the connection was reset or refused, the certificate was rejected or the URL answered an HTTP error, in miss_code and reason. miss_code is filled only on a miss row; on every data row it is null.
  • Per-document facts. Page count, word count, PDF type, the file's own metadata and a 200-character preview on every row.
  • Bounded OCR memory. The OCR engine recycles itself every 400 recognized pages, so memory stays bounded across long batches.

What you can count on

  • You pay only for a document the run answers. An entry it cannot use is pushed as an uncharged ITEM_STATUS row that names the entry and says why.
  • A document is charged only after its row is written to your dataset. A row you cannot see is never billed.
  • A run that finds nothing returns one NO_TEXT_FOUND row that says so, never an empty dataset.
  • A spending limit stops the run cleanly, with a STOPPED_EARLY row saying how many rows were returned and how many were not.
  • Every run writes a RUN_SUMMARY entry with rowsFetched, rowsPushed, rowsCharged, rowsUncharged, duplicatesDropped and stoppedEarly, so a short run and a broken one are told apart.
  • If the run fails in a way no row can answer, it fails with the error in the log. It never returns rows full of nulls and calls it success.
  • No value is invented. On a data row, a field a document does not fill is null, never absent.

Why this one

  • Run on the three documents in the example input, it matched the most-used alternative on this platform row for row, measured 2026-09-22. Its rows also say what a text-only extractor cannot: which pages' text to trust.
  • The most-used alternative bills $0.005 to start a run and $0.005 per document, measured 2026-09-22. This Actor bills $0.004 per returned document, and its start is the platform's $0.00005-per-gigabyte event: lower on both halves of the bill with verifyPages off. With it on, each verified page is billed on top, so a verified document costs more.
  • A miss costs nothing and tells you why. A URL that fails is an uncharged status row naming the entry and the reason.
  • Fill was measured on the platform over files the product had never seen: 48 of 52 documents returned a data row, and agree was the document status on 32 of 52.
  • The source URL, or the name of a file you sent, is on every row, so results join back to your input list without guesswork.
  • The rows keep the field names the category already uses: url, page_count, word_count, pdf_type. A pipeline built on another extractor reads them unchanged.

What data does PDF Text Extractor return?

One data row per document, 27 fields each. An entry the run cannot answer produces an uncharged status row instead. Here is a real row from the example run, in full:

{
"url": "https://www.govinfo.gov/content/pkg/FR-2026-09-21/pdf/2026-19336.pdf",
"file_name": "2026-19336.pdf",
"status": "agree",
"miss_code": null,
"page_count": 2,
"pages_extracted": 2,
"word_count": 741,
"pdf_type": "text_based",
"text": "Presidential Documents \n59979 Federal Register / Vol. 91, No. 181 / Monday, September 21, 2026 / Presidential Documents \nMemorandum of September 16, 2026\nRestoring Reciprocity in Government Procurement\nMemorandum for the Secretary of War[,] the United States Trade\nRepresentative[,] the Director of the Office of Management and Budget[,]\nthe Administrator for Federal Procurement Policy[,] the Administrator of\nGeneral Services[, and] the Administrator of the National Aeronautics\nand Space Administration \nBy the authority vested in me as President by the Constitution and the\nlaws of the United States of America, I hereby direct: \nSection 1 . Purpose and Policy. Canada has unreasonably imposed new bar-\nriers to United States companies seeking to access the Canadian government\nprocurement market by, among other things, establishing preferences for\nCanadian products and Canadian content under its ‘‘Buy Canadian’’ policy.\nCanadian provinces have also limited the access of United States companies\nto their government procurement markets. Meanwhile, Canadian companies\ncontinue to have preferential access to the United States Government procure-\nment system. This includes access to all procurement the United States\nhas agreed to cover at the Federal level under the World Trade Organization\nAgreement on Government Procurement, which amounts to over $280 billion\nannually. My Administration will always act to combat such unreasonable\nor discriminatory practices. \nSec. 2 . Removing Canadian Origin Items From the Federal Procurement\nSystem. (a) The Director of the Office of Management and Budget (Director)\nand the United States Trade Representative (Trade Representative), in coordi-\nnation with the members of the Federal Acquisition Regulatory Council,\nand in consultation with any other senior executive branch official the\nDirector and the Trade Representative deem appropriate, shall, to the extent\nappropriate and consistent with law, identify and take all steps permitted\nby applicable law with respect to Canadian origin items in the Federal\ncivil procurement system that can, where warranted, be removed or made\nnon-available for purchase. Further, the Director, in consultation with any\nsenior executive branch officials he deems appropriate, shall take appropriate\nsteps to notify relevant executive departments and agencies (agencies), as\ndetermined by the Director, of domestic alternatives to Canadian origin\nitems, to the extent permitted by law.\n(b) The Director shall, from time to time, update me on the progress\nof actions taken to implement this memorandum.\n(c) The Trade Representative shall continue to monitor Canada’s treatment\nof United States origin items in the Canadian federal and provincial govern-\nment procurement markets and shall inform me of any circumstances that,\nin the Trade Representative’s opinion, might indicate the need for further\naction. The Trade Representative shall also inform me of any circumstances\nthat, in the Trade Representative’s opinion, might warrant restoring a Cana-\ndian origin item’s availability for Federal civil procurement, such as a change\nin policy by the Canadian government that would end the current treatment\ntoward United States origin items.\n(d) The head of each agency is authorized to and shall take all appropriate\nmeasures within the agency’s authority to implement this memorandum.\nThe head of each agency may, consistent with applicable law, including \nVerDate Sep<11>2014 23:13 Sep 18, 2026 Jkt 268001 PO 00000 Frm 00001 Fmt 4790 Sfmt 4790 E:\\FR\\FM\\21SEO0.SGM 21SEO0\nlotter on DSK8BHNXB4PROD with FR_PREZDOC3\n\n59980 Federal Register / Vol. 91, No. 181 / Monday, September 21, 2026 / Presidential Documents\nsection 301 of title 3, United States Code, redelegate the authority to take\nsuch appropriate measures within the agency. \nSec. 3 . General Provisions. (a) Nothing in this memorandum shall be con-\nstrued to impair or otherwise affect:\n(i) the authority granted by law to an executive department or agency,\nor the head thereof; or\n(ii) the functions of the Director of the Office of Management and Budget\nrelating to budgetary, administrative, or legislative proposals.\n(b) This memorandum shall be implemented consistent with applicable\nlaw and subject to the availability of appropriations.\n(c) This memorandum is not intended to, and does not, create any right\nor benefit, substantive or procedural, enforceable at law or in equity by\nany party against the United States, its departments, agencies, or entities,\nits officers, employees, or agents, or any other person.\n(d) The costs for publication of this memorandum shall be borne by\nthe Office of Management and Budget. \nTHE WHITE HOUSE, \nWashington, September 16, 2026 \n[FR Doc. 2026–19336\nFiled 9–18–26; 11:15 am]\nBilling code 3110–01–P \nVerDate Sep<11>2014 23:13 Sep 18, 2026 Jkt 268001 PO 00000 Frm 00002 Fmt 4790 Sfmt 4790 E:\\FR\\FM\\21SEO0.SGM 21SEO0\nTrump.EPS</GPH>\nlotter on DSK8BHNXB4PROD with FR_PREZDOC3",
"markdown": null,
"page_range": null,
"page_verdicts": [
{
"page": 1,
"verdict": "agree",
"method": "embedded_layer",
"verified": false,
"agreement": null,
"hidden_share": null,
"statistics_score": 0.189,
"text": "Presidential Documents \n59979 Federal Register / Vol. 91, No. 181 / Monday, September 21, 2026 / Presidential Documents \nMemorandum of September 16, 2026\nRestoring Reciprocity in Government Procurement\nMemorandum for the Secretary of War[,] the United States Trade\nRepresentative[,] the Director of the Office of Management and Budget[,]\nthe Administrator for Federal Procurement Policy[,] the Administrator of\nGeneral Services[, and] the Administrator of the National Aeronautics\nand Space Administration \nBy the authority vested in me as President by the Constitution and the\nlaws of the United States of America, I hereby direct: \nSection 1 . Purpose and Policy. Canada has unreasonably imposed new bar-\nriers to United States companies seeking to access the Canadian government\nprocurement market by, among other things, establishing preferences for\nCanadian products and Canadian content under its ‘‘Buy Canadian’’ policy.\nCanadian provinces have also limited the access of United States companies\nto their government procurement markets. Meanwhile, Canadian companies\ncontinue to have preferential access to the United States Government procure-\nment system. This includes access to all procurement the United States\nhas agreed to cover at the Federal level under the World Trade Organization\nAgreement on Government Procurement, which amounts to over $280 billion\nannually. My Administration will always act to combat such unreasonable\nor discriminatory practices. \nSec. 2 . Removing Canadian Origin Items From the Federal Procurement\nSystem. (a) The Director of the Office of Management and Budget (Director)\nand the United States Trade Representative (Trade Representative), in coordi-\nnation with the members of the Federal Acquisition Regulatory Council,\nand in consultation with any other senior executive branch official the\nDirector and the Trade Representative deem appropriate, shall, to the extent\nappropriate and consistent with law, identify and take all steps permitted\nby applicable law with respect to Canadian origin items in the Federal\ncivil procurement system that can, where warranted, be removed or made\nnon-available for purchase. Further, the Director, in consultation with any\nsenior executive branch officials he deems appropriate, shall take appropriate\nsteps to notify relevant executive departments and agencies (agencies), as\ndetermined by the Director, of domestic alternatives to Canadian origin\nitems, to the extent permitted by law.\n(b) The Director shall, from time to time, update me on the progress\nof actions taken to implement this memorandum.\n(c) The Trade Representative shall continue to monitor Canada’s treatment\nof United States origin items in the Canadian federal and provincial govern-\nment procurement markets and shall inform me of any circumstances that,\nin the Trade Representative’s opinion, might indicate the need for further\naction. The Trade Representative shall also inform me of any circumstances\nthat, in the Trade Representative’s opinion, might warrant restoring a Cana-\ndian origin item’s availability for Federal civil procurement, such as a change\nin policy by the Canadian government that would end the current treatment\ntoward United States origin items.\n(d) The head of each agency is authorized to and shall take all appropriate\nmeasures within the agency’s authority to implement this memorandum.\nThe head of each agency may, consistent with applicable law, including \nVerDate Sep<11>2014 23:13 Sep 18, 2026 Jkt 268001 PO 00000 Frm 00001 Fmt 4790 Sfmt 4790 E:\\FR\\FM\\21SEO0.SGM 21SEO0\nlotter on DSK8BHNXB4PROD with FR_PREZDOC3"
},
{
"page": 2,
"verdict": "agree",
"method": "embedded_layer",
"verified": false,
"agreement": null,
"hidden_share": null,
"statistics_score": 0.151,
"text": "59980 Federal Register / Vol. 91, No. 181 / Monday, September 21, 2026 / Presidential Documents\nsection 301 of title 3, United States Code, redelegate the authority to take\nsuch appropriate measures within the agency. \nSec. 3 . General Provisions. (a) Nothing in this memorandum shall be con-\nstrued to impair or otherwise affect:\n(i) the authority granted by law to an executive department or agency,\nor the head thereof; or\n(ii) the functions of the Director of the Office of Management and Budget\nrelating to budgetary, administrative, or legislative proposals.\n(b) This memorandum shall be implemented consistent with applicable\nlaw and subject to the availability of appropriations.\n(c) This memorandum is not intended to, and does not, create any right\nor benefit, substantive or procedural, enforceable at law or in equity by\nany party against the United States, its departments, agencies, or entities,\nits officers, employees, or agents, or any other person.\n(d) The costs for publication of this memorandum shall be borne by\nthe Office of Management and Budget. \nTHE WHITE HOUSE, \nWashington, September 16, 2026 \n[FR Doc. 2026–19336\nFiled 9–18–26; 11:15 am]\nBilling code 3110–01–P \nVerDate Sep<11>2014 23:13 Sep 18, 2026 Jkt 268001 PO 00000 Frm 00002 Fmt 4790 Sfmt 4790 E:\\FR\\FM\\21SEO0.SGM 21SEO0\nTrump.EPS</GPH>\nlotter on DSK8BHNXB4PROD with FR_PREZDOC3"
}
],
"per_page_method": [
"embedded_layer",
"embedded_layer"
],
"layer_ocr_agreement": null,
"wrong_layer_pages": [],
"suspect_layer_pages": [],
"needs_ocr_pages": [],
"unreadable_pages": [],
"verified_pages": [],
"verify_limited": null,
"pages_truncated": false,
"verdict_counts": {
"agree": 2,
"suspect_layer": 0,
"wrong_layer": 0,
"hidden_content": 0,
"no_layer": 0,
"unreadable": 0
},
"text_preview": "Presidential Documents \n59979 Federal Register / Vol. 91, No. 181 / Monday, September 21, 2026 / Presidential Documents \nMemorandum of September 16, 2026\nRestoring Reciprocity in Government Procurem",
"metadata": {
"PDFFormatVersion": "1.7",
"Language": null,
"EncryptFilterName": null,
"IsLinearized": false,
"IsAcroFormPresent": true,
"IsXFAPresent": false,
"IsCollectionPresent": false,
"IsSignaturesPresent": true,
"CreationDate": "D:20260922052149Z",
"Creator": "govinfo, U. S. Government Publishing Office",
"ModDate": "D:20260922012308-04'00'",
"Producer": "iText® Core 9.4.0 (production version) ©2000-2025 Apryse Group NV, Government Publishing Office"
},
"processed_at": "2026-09-25T20:07:21.528Z",
"size_trimmed": null,
"row_type": "ROW"
}

This is a default-path row: outputFormat was text, so text is filled and markdown is null, and each page was answered from its embedded layer, and statistics_score carries how the layer's own text scored on the statistics check. layer_ocr_agreement and verify_limited are null, and wrong_layer_pages and verified_pages are empty, because verifyPages did not run.

Every field on a data row:

FieldWhat it is
urlThe document URL this row answers, echoed from your input. It is also the key duplicate rows are dropped by. Null on a row for a base64Pdfs file.
file_nameThe file name: the last part of the URL path, or the name you gave a base64Pdfs file.
statusThe document's status. agree unless at least 20% of its pages carry wrong_layer, suspect_layer, hidden_content, no_layer or unreadable, in which case the dominant such flag wins. unreadable also when no page could be read at all. On a status row it is the miss word: fetch_failed (the download failed, including a URL that answered an HTTP error such as 403 or 404; miss_code names it), timeout, bad_url, bad_input, failed (the file was fetched but could not be opened, such as a missing password) or aborted.
page_countPages the PDF declares.
word_countWhitespace-separated words across the delivered text of every page.
pdf_typetext_based when every page with text had its own text layer, scanned when every page with text was read by OCR, mixed when some were, unreadable when no page gave any text. A blank page does not decide it: it is listed in unreadable_pages.
pages_extractedPages that delivered any text, by layer or OCR.
textThe whole document text, every read page joined with a blank line between pages. Filled when outputFormat is text (the default) or both; null when it is markdown.
markdownThe whole document as basic Markdown: headings from larger type, paragraphs from line gaps, list items from bullet glyphs and numbering. OCR pages come through as plain paragraphs. Filled when outputFormat is markdown or both; null otherwise.
page_rangeThe pageRange this row was read with, normalised (such as 1-5,8). Null when every page was read.
page_verdictsOne object per processed page: page number, verdict, the side that delivered the text, whether the render verified it, and the delivered text. The verdict is agree, suspect_layer, wrong_layer, hidden_content, no_layer or unreadable. On a page verifyPages did not check, agree means only that the layer passed the statistics check, or was too short to score; nothing was compared with the rendered page. Verified pages carry the layer/OCR agreement score and hidden_share; unverified pages carry the layer's own statistics_score, null on pages too short to score.
per_page_methodPer page, which side delivered its text: embedded_layer, ocr, or none when neither produced any.
layer_ocr_agreementMean of the per-page agreement between the layer's words and the words OCR reads off the rendered page. Filled only when verifyPages ran; null on the default path.
wrong_layer_pagesPage numbers where the rendered page shows text the embedded layer lacks. Filled only when verifyPages ran; those pages deliver OCR text.
suspect_layer_pagesPage numbers whose embedded layer failed the statistics check: impossible word shapes, letter/digit mixes, interior capitals. The default path's flag.
needs_ocr_pagesPage numbers whose delivered text came from OCR rather than the embedded layer.
unreadable_pagesPage numbers that delivered no text from either side: the page had no text layer, and an OCR pass over its render found no words, as on a blank page.
verified_pagesPage numbers the render-and-OCR check actually verified, one page-verified event at $0.008 each. Empty on the default path.
verify_limitedtrue when your charge limit stopped verification before every page; the pages it never reached carry the statistics verdict instead.
pages_truncatedtrue when the per-document page cap (300 by default) stopped the read before every requested page.
verdict_countsTally of page verdicts across the document: agree, suspect_layer, wrong_layer, hidden_content, no_layer, unreadable.
text_previewThe first 200 characters of the delivered text.
metadataThe PDF's own document-info dictionary (title, author, producer, dates) as embedded; null when the file carries none.
processed_atWhen this row was produced, ISO-8601.
size_trimmedNull on almost every row. On a document too large for one dataset row it lists what was shortened to fit: page_verdicts.text first (the per-page text is dropped and the whole text kept), then markdown or text cut at the end.
miss_codeNull on a data row. On an uncharged miss row, the cause in one code (see Troubleshooting).
row_typeROW on every data row; ITEM_STATUS on an uncharged miss; NO_TEXT_FOUND when the run answered and had nothing; STOPPED_EARLY when your charge limit ended the run.

Four more fields appear only on status rows:

FieldWhat it is
reasonWhy the run returned no rows, why it stopped early, or why that entry could not be used.
rowsFetchedHow many rows the run produced before de-duplication and the cap.
rowsReturnedHow many data rows are in the dataset (0 on an empty result; the count so far on STOPPED_EARLY).
rowsRemainingHow many fetched rows were not returned (0 on an empty result).

On a data row, a field a document does not fill is null rather than absent or invented. The dataset's Overview view leads with the verdict columns; the All fields view carries everything, including the verify-only fields that stay empty until verifyPages runs.

How much does it cost?

Pricing is live on Apify's pay-per-event model, and three events can appear on a run:

  • document-returned: $0.004. One document read: its text layer extracted, checked against the text's own statistics, and OCR'd only where a page has no layer. Charged after the row is written to your dataset. This is the primary event.
  • page-verified: $0.008. One page rendered and checked against its embedded text layer. Charged only when verifyPages is on, and only after that page's verdict is computed. On the default path this event never fires.
  • apify-actor-start: $0.00005. The platform's own start charge, once per gigabyte of the run's allocated memory. The Actor's default run allocates 2048 MB.

Nothing else is billed. Status rows and dropped duplicates are never charged: a failed URL, an empty run and a run stopped by your spend cap all cost nothing. You can cap spend on any run; hitting the cap ends it cleanly with a STOPPED_EARLY row saying how many documents came back and how many did not.

PDF URLs in the runRows returned, projecteddocument-returned eventsapify-actor-start
100about 92$0.37$0.00005 per GB allocated
1,000about 923$3.69$0.00005 per GB allocated
10,000about 9,231$36.92$0.00005 per GB allocated

The row counts are projected at the measured 92.3% platform fill. A URL that misses returns a free status row, not a billed one. With verifyPages on, add one page-verified event at $0.008 for each page the render check verifies, up to the per-document page cap (300 by default). Set maxPagesPerDocument or pageRange to bound verification on long files.

How do I use PDF Text Extractor?

  1. Open this Actor's page on Apify and paste your PDF URLs into URLs. The field is already filled in with three public documents, so a first run works immediately.
  2. Leave Verify every page off for the default read, or switch it on to render-check every page at $0.008 per verified page.
  3. Optionally set Output format, Page range, PDF password, OCR Language and Maximum items, or send your own files in PDF files as base64.
  4. Press Start and wait for the run to finish.
  5. Open the Dataset tab. Each document is one row; a URL that failed is an uncharged status row naming it and saying why. The run's counts sit in the key-value store under RUN_SUMMARY.

Example input:

{
"urls": [
"https://www.govinfo.gov/content/pkg/FR-2026-09-21/pdf/2026-19336.pdf",
"https://ntrs.nasa.gov/api/citations/20240009519/downloads/NASA_shielding_SmallSat_collaboration_08062024.pdf",
"https://ntrs.nasa.gov/api/citations/20160001713/downloads/20160001713.pdf"
],
"ocrLanguage": "eng",
"verifyPages": false,
"maxItems": 100
}

The second example, a NASA slide deck, comes back suspect_layer even though its embedded text is clean. Its title slide, a chart slide, a slide of unit abbreviations and its acknowledgements list of surnames and initials read as improbable word shapes to the statistics check. That is the check's known limit on jargon and names, and verifyPages compares those pages with the render to settle it.

Or call it through the Apify API:

curl -X POST "https://api.apify.com/v2/acts/Pradio~pdf-text-verifier/run-sync-get-dataset-items?token=YOUR_APIFY_TOKEN" \
-H "Content-Type: application/json" \
-d '{"urls": ["https://www.govinfo.gov/content/pkg/FR-2026-09-21/pdf/2026-19336.pdf"]}'

Input

InputTypeDefaultWhat it does
urlsarrayfilled in with 3 public PDFsThe PDF document URLs to verify and extract. One data row per document. Give URLs, files sent as base64, or both.
base64PdfsarraynoneYour own PDF files sent inline: a list of {"name": "file.pdf", "data": "<base64>"} objects, each with an optional password. One row per file, named in file_name.
outputFormatstringtextWhich whole-document columns fill: text, markdown or both.
pageRangestringempty (every page)The pages to read in every document, such as 1-5,8 or 10-.
pdfPasswordstring (secret)noneThe password for protected PDFs, used for every document unless an entry carries its own.
ocrLanguagestringengThe language OCR reads where a page's text layer fails or is missing. A Tesseract code such as eng, deu or fra.
verifyPagesbooleanfalseRender each page and check its embedded text against what the page actually shows. Catches wrong and hidden layers the statistics check cannot see. One page-verified event at $0.008 per verified page.
maxItemsinteger100The most rows one run returns, a failed entry's status row included. Rows come back in input order. Entries past the cap are not fetched or read, so nothing is charged for them, not even a verified page, and the run log says the run was capped.
maxPagesPerDocumentinteger300The most pages one document contributes. A longer document is marked pages_truncated. Under verifyPages the cap also bounds page-verified charges per document.

urls

Each entry is an http(s) URL the run fetches with a plain GET. Entries can be plain URL strings or objects with a url key and, optionally, a password for that file. Each document is fetched with a 60-second timeout and an 80 MB size cap. A URL answering a non-PDF body is a bad_input miss. An entry that is not a URL is bad_url. A failed download names its cause in miss_code. Miss rows name the entry and are never charged. It comes filled in with three public government PDFs (a Federal Register document and two NASA technical reports) that need no login.

ocrLanguage

OCR runs in one language per run, chosen for the whole batch. The common two-letter spellings are accepted and mapped for you. en, de, fr, es, it, pt and ar become eng, deu, fra, spa, ita, por and ara. Any other Tesseract code passes through as typed.

base64Pdfs

For files that are not at a public URL. Each entry is an object with the file's name and its bytes as base64 in data (a data:application/pdf;base64, prefix is accepted), plus an optional password. Each file gets one row with url null and its name in file_name. An entry that is not valid base64, or decodes to something other than a PDF, is an uncharged bad_input miss naming it. The whole run input has a size limit, so send large files as URLs.

{
"base64Pdfs": [
{ "name": "invoice-17.pdf", "data": "<base64 of the PDF file>" }
]
}

outputFormat

text (the default) fills the text column. markdown fills markdown instead and leaves text null, and both fills both. The Markdown is inferred from the layout: a line in clearly larger type becomes a heading, a wider gap between lines starts a new paragraph, a line starting with a bullet glyph becomes a list item and a numbered line stays numbered. It is a plain structure for reading and for chunking, not a copy of the page design; tables come through as lines of text. page_verdicts carries the same per-page text in every format.

{
"urls": ["https://ntrs.nasa.gov/api/citations/20160001713/downloads/20160001713.pdf"],
"outputFormat": "markdown"
}

pageRange

Page numbers and ranges separated by commas: 1-5,8 reads pages 1 to 5 and page 8, 10- reads from page 10 to the end. The same range applies to every document in the run, and maxPagesPerDocument still caps how many of those pages are read. A value that is not a page range makes every entry an uncharged bad_input miss with miss_code bad_page_range and the reason, and nothing is fetched. A document the range selects no page of is a miss with miss_code page_range_outside_document.

{
"urls": ["https://ntrs.nasa.gov/api/citations/20160001713/downloads/20160001713.pdf"],
"pageRange": "1-3"
}

pdfPassword

Opens password-protected PDFs. It is used for every document in the run; an entry in urls or base64Pdfs can carry its own password instead. A file that needs a password and gets none is a failed miss with miss_code password_required; a wrong password is a bad_input miss with password_incorrect. Neither is charged. The field is stored as a secret input.

{
"urls": ["https://example.org/reports/protected-report.pdf"],
"pdfPassword": "the file's password"
}

verifyPages

Off by default. When on, every processed page is rendered, OCR'd and compared with its embedded layer. A layer that disagrees with the render is wrong_layer and the page delivers the OCR text. Text the layer carries but the render never shows is hidden_content, and the page keeps the layer text as the evidence. Each verified page is a page-verified charge event at $0.008, charged after the page's verdict is computed. If your charge limit ends verification early, verify_limited is true and the pages never reached keep the statistics verdicts.

Output

A run ends in one of four ways, and the dataset tells them apart:

  • Data rows (row_type: "ROW"): one per document the run answered, each billed $0.004 under document-returned, charged after the row is written.
  • A miss row (row_type: "ITEM_STATUS"): one per entry the run could not use, carrying the entry, the miss word in status, the cause in miss_code and a reason. Pushed so you see the outcome; never charged.
  • Nothing usable (row_type: "NO_TEXT_FOUND"). The input produced no data rows, so the run pushes one status row with rowsFetched, rowsReturned and rowsRemaining. A zero result is an answer, never an empty dataset, and it is not charged.
  • Stopped by your spending cap (row_type: "STOPPED_EARLY"). One uncharged row says how many documents were returned and how many fetched rows were not.

Every run also writes a RUN_SUMMARY entry to the key-value store: rowsFetched, rowsPushed, rowsCharged, rowsUncharged, duplicatesDropped and stoppedEarly. A short run and a broken one are told apart.

What can you do with the data?

Sort documents in a RAG pipeline. A builder splits on status. agree rows index straight away, suspect_layer and wrong_layer rows go to a second look or a verifyPages re-run, and unreadable rows go to a manual queue.

Audit an archive for hidden text. An analyst runs the batch with verifyPages on and filters on verdict_counts.hidden_content. That finds files carrying text the render never shows: white text, invisible runs, injected strings.

Catch broken public documents at scale. A civic-data team runs its folder of releases through the Actor. The rows say which files are scanned images, which carry a suspect layer, and which read clean.

Check a vendor's delivery. An ingest operator points the Actor at the supplier's latest batch and sorts by status and layer_ocr_agreement. The files that fail verification go back.

Worked examples

Use it to put each document's full text in a spreadsheet. The default format fills the text column, so a CSV or Excel export carries one cell per file.

{
"urls": [
"https://www.govinfo.gov/content/pkg/FR-2026-09-21/pdf/2026-19336.pdf",
"https://ntrs.nasa.gov/api/citations/20240009519/downloads/NASA_shielding_SmallSat_collaboration_08062024.pdf"
]
}

Use it to convert PDFs to Markdown for a RAG index. Headings and paragraphs give a chunker natural break points, and status still says which files to hold back.

{
"urls": ["https://ntrs.nasa.gov/api/citations/20160001713/downloads/20160001713.pdf"],
"outputFormat": "both",
"maxPagesPerDocument": 300
}

Use it to read only the first pages of long reports. The summary pages of every file are read and verified, at a lower bill.

{
"urls": ["https://ntrs.nasa.gov/api/citations/20160001713/downloads/20160001713.pdf"],
"pageRange": "1-3",
"verifyPages": true
}

Use it to read your own password-protected files, sent inline. One file here carries its own password.

{
"base64Pdfs": [
{ "name": "statement-march.pdf", "data": "<base64 of the PDF file>" },
{ "name": "statement-april.pdf", "data": "<base64 of the PDF file>", "password": "a different password" }
],
"pdfPassword": "the shared password"
}

Use PDF Text Extractor with AI agents

Paste this line to give a Claude-compatible agent this Actor as an MCP tool:

$claude mcp add --transport http apify "https://mcp.apify.com?tools=Pradio/pdf-text-verifier"

Personal data

This Actor reads only the documents you point it at. It never fetches a site of ours and never goes looking for files.

  • Purpose. The text in your files is read so it can be returned to you, checked page by page. No inference is drawn from the text, and nothing is derived from it beyond the counts and verdicts on the row.
  • Transparency. The delivered text and the file's own metadata may carry personal data or protected text, exactly as they appear in the document. You are the controller of what you upload, and you decide what to run through the Actor.
  • Retention. Nothing is kept beyond the run. The rows live in your dataset; the document bytes and the page renders exist only while the run works.
  • A refused read is respected. A 403 or any other refusal ends that entry as a fetch_failed miss with miss_code http_403 (or the code the host sent), with no retry past it, no login and no wall defeated.

Support

Found a document that reads wrong, or a verdict you disagree with? Open an issue on this Actor's Issues tab with the document's URL. Every issue is read and answered.

Release notes

  • 0.1.62, 2026-09-25: a blank page no longer decides pdf_type. It is still listed in unreadable_pages, and a scanned file with a blank back page now reads scanned instead of mixed.
  • 0.1.52, 2026-09-24: a text column with the whole document; Markdown output behind outputFormat; pageRange; pdfPassword for protected files; your own files as base64 in base64Pdfs; file_name on every row; misses that name their cause in miss_code (DNS, timeout, reset, refused, certificate, HTTP status, password); and the default page cap raised from 60 to 300 pages per document.
  • 0.1.47, 2026-09-24: slide decks and other landscape pages are no longer flagged suspect_layer for their short lines (they can still be flagged for names, web addresses and unit abbreviations), and compound words like CubeSat no longer count as case noise. Measured on the 66-document truth set: the same 18 of 35 generated layers caught, no clean document flagged.
  • 0.1.41, 2026-09-22: first public release. One row per PDF from the public URLs you paste: the document's text, page count, word count and PDF type; a verdict per page with suspect-layer flags computed from the text's own statistics; OCR on the pages that carry no embedded layer; an opt-in per-page render check charged separately; a 60-page cap per document; and a miss row that costs nothing and says why.

Limits

  • 300 pages per document by default. A longer file is processed up to the cap and the row carries pages_truncated: true. The pages that ran are listed in page_verdicts. Raise maxPagesPerDocument for longer files, or name the pages you need with pageRange.
  • Very large documents are shortened to fit one row. A dataset row has a size limit. A document whose text would pass it first loses the per-page text in page_verdicts (the whole text stays in text or markdown), and then that column is cut at the end. size_trimmed says what was shortened.
  • Eighty megabytes and sixty seconds per file. A larger download or a slower answer comes back as a fetch_failed or timeout miss row, not a partial read.
  • Public URLs, or your own files. Each URL must answer a plain GET. A file behind a login or a paywall fails as a miss, as there is no cookie or login input; send such a file in base64Pdfs instead.
  • Protected PDFs need their password. Without pdfPassword, a password-protected file is a failed miss with miss_code password_required.
  • Markdown is inferred, not copied. Headings, paragraphs and lists come from type size, line gaps and bullet glyphs. Tables, columns and footnotes come through as plain lines.
  • The default check reads shape, not meaning. The statistics score catches layers whose text looks generated, but a wrong layer with clean statistics still passes as agree. verifyPages is the check that compares the text to the render.
  • Short pages such as covers can be flagged. A page is scored only when its layer has at least 10 words. statistics_score runs from 0 to 1, and a page scoring 0.5 or more is flagged suspect_layer. One of the things it scores is line length, and a cover, title page or contents page whose lines average three words or fewer scores 1 and is flagged even when its text is clean. On such a page, read the text in page_verdicts before discarding it, or run verifyPages to settle it. A flagged cover changes the document's status only when flagged pages make up at least 20% of the pages read, which a lone cover can do on a short document.
  • A suspect page keeps its layer text. On the default path a flagged page still ships its embedded text, marked suspect_layer; it is not re-OCR'd. Turn on verifyPages to have a wrong layer replaced by the rendered text.
  • Suspect pages are common, and more than a quarter of files are not clean. In the platform run, 36 of the 48 resolved files carried at least one suspect page. 16 of those 48 files ended with a status other than agree. The 20% share keeps a noisy page or two from marking the rest.
  • Pages that deliver nothing happen. 18 of the 48 resolved documents had at least one unreadable page. The row still ships the pages that worked, and unreadable_pages names the rest.
  • One OCR language per run. Multi-language documents are not auto-detected; pick the language that covers most of the batch.
  • Scanned files are OCR quality. A scanned or mixed document's text is whatever OCR reads off the page render. Clean digital documents are more accurate than photographed pages.
  • Fields only verifyPages fills. layer_ocr_agreement and verify_limited stay null on the default path, and wrong_layer_pages and verified_pages stay empty. They fill when verifyPages is on.

Troubleshooting

I pasted 50 URLs but only 47 rows came back. A failed URL is not the cause, because every failed entry gets its own ITEM_STATUS row. Two things make fewer rows. A URL you pasted twice comes back once, and the repeat is counted in RUN_SUMMARY.duplicatesDropped. And maxItems caps how many rows a run pushes.

A miss row says what went wrong in miss_code. dns_failure: the host name does not resolve. timeout: the host did not answer or the download stalled. connection_reset or connection_refused: the host closed or refused the connection. host_unreachable, tls_error (a rejected certificate) and redirect_loop say the same for the network. http_ and a number (such as http_404 or http_403): the URL answered that HTTP status. not_pdf, corrupt_pdf and too_large are about the file. password_required and password_incorrect are about pdfPassword. bad_page_range and page_range_outside_document are about pageRange. bad_base64 is a base64Pdfs entry that does not decode. bad_url is an entry that is not a URL.

The document's status is agree, but suspect_layer_pages lists a page. That is the materiality rule working. A file is marked only when at least 20% of its pages carry a flag, so one noisy page does not mark the document. The page's text still shipped, and page_verdicts lists every page either way.

layer_ocr_agreement is null and verified_pages is empty. verifyPages was off, so the run never rendered a page and there was nothing to agree against. Re-run with verifyPages on for the agreement score and the wrong_layer and hidden_content verdicts.

verify_limited is true on my row. Your charge limit stopped render verification before every page. The pages it never reached carry the statistics verdict instead. Raise the limit and re-run to verify the rest.

pages_truncated is true on my row. The document has more pages than maxPagesPerDocument (300 by default) and only that many were processed. page_verdicts shows which pages ran; raise the input to read further.

Still stuck? Open an issue on the Issues tab of this listing.

FAQ

Can I use integrations with PDF Text Extractor? Yes. Runs can be scheduled on Apify, and the dataset exports to JSON, CSV or Excel. The platform's integrations (Make, Zapier, n8n and webhooks) can pick up the rows when a run finishes. Any tool that can build the input JSON can drive the Actor through the API.

Can I use PDF Text Extractor with the Apify API? Yes. Call the run-sync endpoint with the same input JSON shown above. The dataset items come back in the response. Or start an asynchronous run and read the dataset when it finishes. The apify-client packages for JavaScript and Python wrap both paths.

Can I use PDF Text Extractor through an MCP server? Yes. The one line in the section above registers this Actor with Apify's MCP server, so an agent can call it as a tool.

Is it legal to extract text from PDFs? The Actor reads only documents you supply or public URLs you name; it never goes looking for files on its own. The rights in the text are yours or your licensor's, and you are the controller of what you upload. For public documents such as the government PDFs in the example input, extraction is generally lawful. For licensed or confidential files, your own license governs. A refused or blocked read is respected: the run never logs in and never defeats a wall.

How fast is PDF Text Extractor? In the platform run (run GWPmIEvhbbs7LFsQo), 52 public PDFs took 461 seconds at the default memory, about 9 seconds a file. Files are read one after another, so the time grows with the list. A file whose pages all carry a text layer is the fast case. Each page that needs OCR is rendered and recognised first, so scanned files take longer, and verifyPages renders every page, which makes it the slowest setting.

How many PDFs can one run take? As many as maxItems allows: 100 rows by default, and you can raise it. Each file is held to 80 MB, sixty seconds to download and 300 pages unless you raise maxPagesPerDocument. For a very long list, split it across runs, or set a spending limit and the run stops cleanly at it with a STOPPED_EARLY row.

See also

Not affiliated

PDF Text Extractor is an independent tool. It is not affiliated with, endorsed by or sponsored by any document host or standards body. That includes the hosts of the example documents (govinfo.gov and ntrs.nasa.gov), the PDF format's steward, and the Tesseract OCR project. It reads only the files you point it at.