PDF to Markdown & Table Extractor — CSV, JSON · $0.03/page avatar

PDF to Markdown & Table Extractor — CSV, JSON · $0.03/page

Pricing

Pay per event

Go to Apify Store
PDF to Markdown & Table Extractor — CSV, JSON · $0.03/page

PDF to Markdown & Table Extractor — CSV, JSON · $0.03/page

Extract text, tables and fields from PDF and DOCX files. Returns Markdown with headings and lists, tables as CSV rows, and named fields such as invoice number, date and total found by label. Rule-based, repeatable results, no subscription. Image-only scans are flagged. Failed files are not charged.

Pricing

Pay per event

Rating

0.0

(0)

Developer

Ambolt

Ambolt

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

4 hours ago

Last modified

Share

Extract text, tables and fields from PDF and DOCX files. Returns Markdown with headings and lists, tables as CSV rows, and named fields such as invoice number, date and total found by label. Scanned pages are read with OCR at $0.06 a page. No subscription. Failed files are not charged.

Turn PDF and DOCX files into clean Markdown with headings, lists and tables, every table also as CSV rows, and optionally named fields (invoice number, total, dates) found by label. Layout rules, no AI model: the same file gives the same result. Per page pricing, files fetched from a URL, up to 20 files per run. Scanned pages are read with OCR (Tesseract) and billed separately.

What you get

Structured JSON for each run, ready for spreadsheets, alerts and AI agents. Also available as an MCP tool (doc_to_json) and as a pay-per-call HTTP endpoint.

Input

FieldTypeDefaultDescription
urlsarrayPublic URLs of PDF or DOCX files (up to 20).
maxPagesPerFileinteger50Stop after this many pages per PDF. Each processed page is one billable result.
fieldHintsarrayOptional fields to find, one per line as name
ocrstring (auto, off)autoauto reads pages without a text layer (scans) with OCR, billed as ocr-page; off skips them.
ocrLanguagesstringengTesseract language codes joined by +, for example eng+dan+deu.
includeTablesCsvbooleantrueAdd every detected table as CSV text next to the rows.

Example input

{
"urls": [
"https://www.rfc-editor.org/rfc/rfc9110.pdf"
],
"maxPagesPerFile": 3
}

Example output

{
"files": 1,
"extracted": 1,
"pagesProcessed": 4,
"billableResults": 4,
"datasetItems": [
{
"url": "https://www.rfc-editor.org/rfc/rfc9110.pdf",
"ok": true,
"kind": "pdf",
"pages": 4,
"characters": 5134,
…

Pricing

EventPrice
page (each successful result)$0.030

Failed runs are not charged. Free trial credits from Apify cover your first runs.

FAQ

Is a result guaranteed to be complete? Each result carries the date it was read. If the upstream service is down the run fails with a clear message and you are not charged.

Can I call it from code or an AI agent? Yes: run it through the Apify API, or use the MCP tool and pay-per-call endpoint listed at https://ambolt.dev.

Is this affiliated with the data source? No. Source names are used descriptively; each result keeps its source attribution and licence line.

Where do I report a problem? Open an issue on this Actor's page; replies are written by the Ambolt service team, usually within a day.

Review

If this Actor saved you time, a short review on its Store page helps others find it and tells us what to improve.

Keywords

PDF to Markdown, PDF table extraction, PDF to CSV, PDF to JSON, DOCX to Markdown, invoice data extraction, document parsing, RAG, data entry automation, document to JSON

Notes

Data comes from the source named in each result and keeps its licence and attribution. Informational only, not financial, legal or tax advice. Open an issue on the Actor page for questions; replies come under the product name.