PDF to JSON Converter — Tables & Markdown
Pricing
from $2.80 / 1,000 pdf processeds
PDF to JSON Converter — Tables & Markdown
PDF parser and converter for structured JSON, CSV-ready rows and Markdown. Extract reading-order text, page-level Markdown, heuristic tables and document metadata with bounded file size, page, timeout and retry controls.
Pricing
from $2.80 / 1,000 pdf processeds
Rating
0.0
(0)
Developer
Rosario Vitale
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
0
Monthly active users
19 hours ago
Last modified
Categories
Share
PDF to JSON/CSV API — Text, Tables & Structured Data
Why use this Actor?
Convert public PDF files into structured JSON or CSV-ready data. Extract page text, reconstructed reading-order lines, optional table rows and document metadata with page, file-size, timeout and retry limits for reliable document-processing pipelines.
Features
- PDF URLs — Direct links to the PDF files you want to convert into structured data. Each URL must point straight to a .pdf file.
- Extract tables — Try to detect tables and return them as rows of cells (heuristic, based on text positions).
- Extract metadata — Include document metadata (title, author, producer, creation date) when available.
- Max pages per PDF — Limit how many pages to read per PDF. Use 0 to read all pages.
- Max PDF size (MB) — Reject PDF files larger than this limit before parsing.
- Download timeout (seconds) — Maximum time allowed for each PDF download attempt.
- Download retries — Retries for transient network errors, rate limits, and server errors.
Use cases
- Pdf-to-json pipelines.
- Document text extraction.
- Table reconstruction.
- Document metadata processing.
Example input
{"pdfUrls": ["https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf"],"extractTables": false,"extractMetadata": true,"maxPages": 0,"maxPdfSizeMb": 50,"requestTimeoutSecs": 30}
Pricing & cost control
Use the bounded input limits and filters to keep runs predictable. Pay-per-result Actors only charge primary result rows; summary, status and monitoring metadata are designed to add context without inflating result volume.
FAQ
What is this Actor for?
It is designed for PDF-to-JSON pipelines, document text extraction, table reconstruction.
Can I run it on a schedule?
Yes. You can schedule Actor runs on Apify and send the resulting dataset into automations, webhooks, storage, or downstream APIs.
How do I control cost and run size?
Use the input limits and filters shown in the Actor input form. The Actor applies bounded defaults and hard caps so large jobs remain predictable.
Search keywords
pdf to json, pdf to json converter, pdf to json converter free, pdf to json file, pdf to json free, pdf to json python, pdf to json converter free online, pdf to json converter i love pdf, pdf to json i2pdf, pdf to json converter online, pdf table extractor, pdf table extractor python, pdf table extractor online, pdf table extractor npm
Turn any PDF into clean, structured data your code can actually use. Give the Actor one or more PDF URLs and get back text per page, reconstructed lines in reading order, optional table detection, and document metadata — as JSON or CSV.
No more copy-pasting from PDFs by hand or fighting with brittle local libraries. Send a URL, get structured output.
What it does
- 📄 Text extraction — full text of every page, in natural reading order.
- 📐 Line reconstruction — text items are grouped by position into real lines, not a jumbled blob.
- 📊 Table detection (optional) — heuristically splits rows into cells so you can rebuild tables.
- 🏷️ Metadata (optional) — title, author, producer and creation date when present.
- 🔁 Batch — pass many PDF URLs in a single run.
Input
| Field | Type | Description |
|---|---|---|
pdfUrls | array of strings | Direct links to the PDF files (required). |
extractTables | boolean | Detect tables and return rows of cells. Default false. |
extractMetadata | boolean | Include document metadata. Default true. |
maxPages | integer | Max pages to read per PDF. 0 = all. Default 0. |
Example input
{"pdfUrls": ["https://raw.githubusercontent.com/mozilla/pdf.js/master/web/compressed.tracemonkey-pldi-09.pdf"],"extractTables": false,"extractMetadata": true,"maxPages": 0}
Output
One dataset item per PDF:
{"url": "https://.../document.pdf","success": true,"numPages": 14,"pagesExtracted": 14,"metadata": { "Producer": "pdfeTeX-1.21a", "Creator": "TeX", "CreationDate": "..." },"pages": [{"pageNumber": 1,"text": "Trace-based Just-in-Time Type Specialization ...","lines": ["Trace-based Just-in-Time Type Specialization ...", "Languages"],"tables": [["Cell A", "Cell B"], ["1", "2"]]}],"fullText": "Trace-based Just-in-Time Type Specialization ..."}
Export the dataset as JSON, CSV, Excel, or HTML straight from the run, or pull it through the Apify API.
Common use cases
- Extract data from invoices, receipts, price lists, and bank statements.
- Feed PDF text into search, RAG pipelines, or LLMs.
- Turn reports and catalogs into spreadsheets.
- Archive and index document text at scale.
Notes & limits
- Works on text-based PDFs. Scanned/image-only PDFs contain no selectable text, so they need OCR (not included in this version).
- Table detection is a position-based heuristic — great for clean, grid-like tables, approximate for complex layouts.
pdfUrlsmust be direct links to the PDF file (not a viewer page).
Pricing
Pay-per-result: you are billed per PDF successfully processed. Failed downloads/parses are returned with success: false and are not charged.
How to use
Add one or more direct PDF URLs, enable metadata or heuristic table extraction when needed, optionally limit pages and download resources, and run the Actor. Results are available as a structured dataset and through the Apify API.
Reliability controls
Downloads use bounded timeouts and retries. The Actor checks HTTP status, content type, declared and actual file size, and the PDF signature before parsing. Oversized or non-PDF responses fail cleanly instead of consuming unbounded memory.
FAQ and support
Table extraction is heuristic because PDF files do not contain a universal table structure. Image-only scanned PDFs require OCR. For reproducible parsing problems, include the PDF URL when shareable, the run ID, and expected page in the Actor Issues tab.