PDF Text, Metadata & Tables Extractor
Pricing
from $1.00 / 1,000 text pages
PDF Text, Metadata & Tables Extractor
Extract page text, optional tables and PDF metadata from public digital PDF URLs. Get one page per row, document hashes, blank-page markers and coverage summaries.
Pricing
from $1.00 / 1,000 text pages
Rating
0.0
(0)
Developer
Akshay Aggarwal
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
3 days ago
Last modified
Categories
Share
Turn public digital PDFs into one dataset row per page for document research, indexing, and downstream automation. Each row has extracted text, optional tables, document metadata, a SHA-256 hash, and a clear marker for pages without a digital text layer. The OUTPUT record reports coverage and errors for every document.
Quick start
{"pdfUrls": ["https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf"],"maxPagesPerDocument": 20,"maxFileSizeMb": 10,"includeTables": false}
For a compact output example, a three-page sample PDF containing a ruled table and one blank page produces these rows:
[{"page_number": 1, "text": "Alpha report\nName Value\nBeta 7", "tables": [[ ["Name", "Value"], ["Beta", "7"] ]], "text_layer_present": true, "page_status": "text"},{"page_number": 2, "text": "", "tables": [], "text_layer_present": false, "page_status": "no_text_layer"}]
The full rows also include document_url, document_hash, metadata and truncation flags. This example uses a generated sample PDF; the quick-start URL above is a public W3C PDF.
Output and billing
The default dataset contains one saved page per row: document_url (final redirect URL), document_hash (SHA-256 of source bytes), page_number (1-based), text, tables, metadata, text_layer_present, page_status, text_truncated, and tables_truncated. Text includes text inside detected tables. Tables are arrays of rows and cell strings; null means the parser found an empty cell. Table detection works best on ruled digital tables and can miss complex layouts. When includeTables is false, tables is an empty array.
OUTPUT is a key-value record, separate from the result dataset. It includes totals for URLs requested, documents fetched/skipped/with errors, pages processed/saved, text and blank pages saved, at least how many pages were truncated, and document-level status, errors, and limit reasons. A page limit stops after the chosen number; the Actor detects at least one extra page but does not count all omitted pages. Failed PDFs have no paid result rows. A page parse error is recorded for that page and later pages are attempted.
The billing unit is one saved page with a digital text layer (text-page event). Blank/image-only page rows have no event, and there is no separate document fee.
Scope and limits
- Public direct HTTPS PDF URLs only, 1–10 per run. No login, private/local address, IP literal, custom port, or HTTP link. Redirects are checked before connection. Downloads require a PDF content type and
%PDF-signature, and are streamed with a byte cap. maxPagesPerDocument: 1–200, default 20.maxFileSizeMb: 1–20 MiB, default 10.includeTables: default false.- Extracted text is capped at 100,000 characters per page. Tables are capped at 20 per page, 100 rows per table, 20 columns per row, and 2,000 characters per cell; truncation flags are set in the row. Table discovery itself can be costly on complex pages, so use a small page limit first.
- No OCR, LLM calls, scanned-page transcription, encrypted-PDF unlocking, form filling, JavaScript execution, or external enrichment. A blank/image-only page is emitted with empty text and
no_text_layerstatus. Embedded images and attachments are omitted. - Source hosts may block anonymous cloud requests or change their files. Check
OUTPUTfor partial results and failed URLs; saved dataset rows remain available.