PDF to JSON Converter — Tables & Markdown avatar

PDF to JSON Converter — Tables & Markdown

Pricing

from $2.80 / 1,000 pdf processeds

Go to Apify Store
PDF to JSON Converter — Tables & Markdown

PDF to JSON Converter — Tables & Markdown

PDF parser and converter for structured JSON, CSV-ready rows and Markdown. Extract reading-order text, page-level Markdown, heuristic tables and document metadata with bounded file size, page, timeout and retry controls.

Pricing

from $2.80 / 1,000 pdf processeds

Rating

0.0

(0)

Developer

Rosario Vitale

Rosario Vitale

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

0

Monthly active users

19 hours ago

Last modified

Share

PDF to JSON/CSV API — Text, Tables & Structured Data

Why use this Actor?

Convert public PDF files into structured JSON or CSV-ready data. Extract page text, reconstructed reading-order lines, optional table rows and document metadata with page, file-size, timeout and retry limits for reliable document-processing pipelines.

Features

  • PDF URLs — Direct links to the PDF files you want to convert into structured data. Each URL must point straight to a .pdf file.
  • Extract tables — Try to detect tables and return them as rows of cells (heuristic, based on text positions).
  • Extract metadata — Include document metadata (title, author, producer, creation date) when available.
  • Max pages per PDF — Limit how many pages to read per PDF. Use 0 to read all pages.
  • Max PDF size (MB) — Reject PDF files larger than this limit before parsing.
  • Download timeout (seconds) — Maximum time allowed for each PDF download attempt.
  • Download retries — Retries for transient network errors, rate limits, and server errors.

Use cases

  • Pdf-to-json pipelines.
  • Document text extraction.
  • Table reconstruction.
  • Document metadata processing.

Example input

{
"pdfUrls": [
"https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf"
],
"extractTables": false,
"extractMetadata": true,
"maxPages": 0,
"maxPdfSizeMb": 50,
"requestTimeoutSecs": 30
}

Pricing & cost control

Use the bounded input limits and filters to keep runs predictable. Pay-per-result Actors only charge primary result rows; summary, status and monitoring metadata are designed to add context without inflating result volume.

FAQ

What is this Actor for?
It is designed for PDF-to-JSON pipelines, document text extraction, table reconstruction.

Can I run it on a schedule?
Yes. You can schedule Actor runs on Apify and send the resulting dataset into automations, webhooks, storage, or downstream APIs.

How do I control cost and run size?
Use the input limits and filters shown in the Actor input form. The Actor applies bounded defaults and hard caps so large jobs remain predictable.

Search keywords

pdf to json, pdf to json converter, pdf to json converter free, pdf to json file, pdf to json free, pdf to json python, pdf to json converter free online, pdf to json converter i love pdf, pdf to json i2pdf, pdf to json converter online, pdf table extractor, pdf table extractor python, pdf table extractor online, pdf table extractor npm

Turn any PDF into clean, structured data your code can actually use. Give the Actor one or more PDF URLs and get back text per page, reconstructed lines in reading order, optional table detection, and document metadata — as JSON or CSV.

No more copy-pasting from PDFs by hand or fighting with brittle local libraries. Send a URL, get structured output.

What it does

  • 📄 Text extraction — full text of every page, in natural reading order.
  • 📐 Line reconstruction — text items are grouped by position into real lines, not a jumbled blob.
  • 📊 Table detection (optional) — heuristically splits rows into cells so you can rebuild tables.
  • 🏷️ Metadata (optional) — title, author, producer and creation date when present.
  • 🔁 Batch — pass many PDF URLs in a single run.

Input

FieldTypeDescription
pdfUrlsarray of stringsDirect links to the PDF files (required).
extractTablesbooleanDetect tables and return rows of cells. Default false.
extractMetadatabooleanInclude document metadata. Default true.
maxPagesintegerMax pages to read per PDF. 0 = all. Default 0.

Example input

{
"pdfUrls": [
"https://raw.githubusercontent.com/mozilla/pdf.js/master/web/compressed.tracemonkey-pldi-09.pdf"
],
"extractTables": false,
"extractMetadata": true,
"maxPages": 0
}

Output

One dataset item per PDF:

{
"url": "https://.../document.pdf",
"success": true,
"numPages": 14,
"pagesExtracted": 14,
"metadata": { "Producer": "pdfeTeX-1.21a", "Creator": "TeX", "CreationDate": "..." },
"pages": [
{
"pageNumber": 1,
"text": "Trace-based Just-in-Time Type Specialization ...",
"lines": ["Trace-based Just-in-Time Type Specialization ...", "Languages"],
"tables": [["Cell A", "Cell B"], ["1", "2"]]
}
],
"fullText": "Trace-based Just-in-Time Type Specialization ..."
}

Export the dataset as JSON, CSV, Excel, or HTML straight from the run, or pull it through the Apify API.

Common use cases

  • Extract data from invoices, receipts, price lists, and bank statements.
  • Feed PDF text into search, RAG pipelines, or LLMs.
  • Turn reports and catalogs into spreadsheets.
  • Archive and index document text at scale.

Notes & limits

  • Works on text-based PDFs. Scanned/image-only PDFs contain no selectable text, so they need OCR (not included in this version).
  • Table detection is a position-based heuristic — great for clean, grid-like tables, approximate for complex layouts.
  • pdfUrls must be direct links to the PDF file (not a viewer page).

Pricing

Pay-per-result: you are billed per PDF successfully processed. Failed downloads/parses are returned with success: false and are not charged.

How to use

Add one or more direct PDF URLs, enable metadata or heuristic table extraction when needed, optionally limit pages and download resources, and run the Actor. Results are available as a structured dataset and through the Apify API.

Reliability controls

Downloads use bounded timeouts and retries. The Actor checks HTTP status, content type, declared and actual file size, and the PDF signature before parsing. Oversized or non-PDF responses fail cleanly instead of consuming unbounded memory.

FAQ and support

Table extraction is heuristic because PDF files do not contain a universal table structure. Image-only scanned PDFs require OCR. For reproducible parsing problems, include the PDF URL when shareable, the run ID, and expected page in the Actor Issues tab.